Claude Shannon gave the world a single, elegant equation in 1948 that turned the fuzzy idea of “information” into something measurable. His entropy formula, H(X) = −∑ p(x) log p(x), tells us the average uncertainty in a set of outcomes and remains the backbone of information theory. But the basic formula assumes we already know the probabilities of every outcome, and that the only thing worth measuring is statistical uncertainty. Real information problems are messier. Sometimes we gather information by asking questions one at a time. Sometimes we want to weight rare events differently. Sometimes the uncertainty is not about chance at all, but about vagueness. To handle these situations, researchers built a family of Shannon-type measures that keep the spirit of his work while stretching it in new directions. This post looks at three of them: Picard’s decision tree approach, Rényi information, and fuzzy measures.
Table of Contents
- Why Shannon’s formula needed company
- Picard’s decision tree approach
- How a questionnaire becomes a decision tree
- Measuring information through questions
- Why this matters for information science
- Rényi information
- The generalized formula
- Special cases hidden inside one equation
- Where Rényi measures are useful
- Fuzzy measures
- From crisp sets to degrees of membership
- De Luca and Termini’s fuzzy entropy
- Applications in information work
- Bringing the three together
Why Shannon’s formula needed company
Shannon’s entropy works beautifully when we have a clean probability distribution and a clear communication channel. It quantifies the average level of uncertainty tied to a random variable and sets the absolute limit on how far data can be compressed without loss. The trouble is that this model treats information as a one-shot affair: the source emits symbols, and we measure the uncertainty across all of them at once.
In practice, information is often acquired step by step, and not every uncertain situation is governed by probability. A reference librarian narrowing down a patron’s vague request, a survey designer choosing which question to ask first, or a system trying to classify a document that only “partly” belongs to a category, all face problems the original formula does not directly address. The three measures below were developed precisely to fill these gaps.
Picard’s decision tree approach
The French mathematician C.-F. Picard built a bridge between information theory and the practical act of asking questions. His work, published as Graphes et questionnaires in 1972 and later translated as Graphs and Questionnaires, models the process of gathering information as a structured sequence of questions rather than a single measurement.
How a questionnaire becomes a decision tree
Picard’s central idea is that a questionnaire is really a decision tree. Each question is a node. Each possible answer is a branch leading to the next question or to a final outcome. Imagine you want to identify which of eight equally likely outcomes is true. If you ask well-designed yes/no questions, each answer cuts the set of possibilities roughly in half, and you can reach the answer in about three questions. That number, three, is not a coincidence: it equals log₂ 8, the Shannon entropy of eight equally likely events. The structure of the tree and the information content of the source are deeply linked.
Measuring information through questions
In this framework, the value of a question is judged by how much uncertainty it removes. Before a question is asked, the situation has a certain entropy. After the answer arrives, the remaining uncertainty is usually smaller. The difference is the information the question delivered. This is the same logic that later powered information gain in machine learning, where decision tree algorithms choose the attribute that reduces entropy the most at each split. Picard’s questionnaire theory anticipated and formalised this connection between tree structure and entropy long before it became routine in data science.
His approach also separates the idea of uncertainty from the act of measurement. Mathematicians such as B. Forte extended this thinking, exploring entropies with and without probabilities and applying them directly to questionnaire theory. This matters because not every question has neat probabilities attached, yet we still want to know which question is “worth more.”
Why this matters for information science
For anyone working with users, this is more than abstract theory. When a librarian conducts a reference interview, each question is a node in an implicit decision tree, and a skilled professional instinctively asks the question that narrows the search most. The same principle guides good survey design: ordering questions so that each one extracts maximum information keeps questionnaires short and respondents engaged. Picard’s model gives a mathematical reason for what experienced practitioners already do by intuition.
Rényi information
The second extension comes from the Hungarian mathematician Alfréd Rényi, who in 1961 asked a sharp question: is Shannon entropy the only sensible way to measure information, or just one member of a larger family? He answered by deriving a generalized measure now called Rényi entropy, built to be the most general form that still preserves additivity for independent events.
The generalized formula
Rényi entropy of order α is written as H_α(X) = [1 / (1 − α)] log ∑ pᵢ^α, where α is a positive parameter that tunes the measure’s sensitivity. The clever part is that this single equation is not one measure but a whole spectrum of them. By turning the dial on α, you shift how much weight the formula gives to common versus rare outcomes. A low α treats all outcomes more evenly, while a high α lets the most probable events dominate.
Special cases hidden inside one equation
What makes Rényi’s family so useful is that several well-known measures sit inside it as special cases. When α approaches 1, the formula collapses neatly back into ordinary Shannon entropy, which is why Rényi entropy is rightly called a generalization rather than a replacement. At α = 0 it becomes the Hartley or max-entropy, which counts the number of possible outcomes regardless of their probabilities. At α = 2 it gives collision entropy, linked to the chance that two independent draws yield the same value. As α grows very large it tends toward min-entropy, which focuses entirely on the single most likely outcome and is widely used in cryptography to gauge worst-case unpredictability.
Where Rényi measures are useful
Because one parameter controls so much, Rényi information is valuable wherever a single average is too blunt. In ecology and statistics it underpins diversity indices that describe how varied a population is. In signal processing it helps analyse complex time-frequency patterns. For informetrics and scientometrics, this flexibility is appealing because citation distributions, keyword frequencies, and journal usage are often highly skewed. A measure that can be tuned to emphasise either the dominant journals or the long tail of rarely cited works gives analysts a richer picture than a single fixed formula would.
Fuzzy measures
The third extension changes the very nature of the uncertainty being measured. Shannon and Rényi both deal with randomness, where outcomes are uncertain because of chance. Fuzzy measures deal with vagueness, where the uncertainty comes from imprecise boundaries rather than probability.
From crisp sets to degrees of membership
The foundation is Lotfi Zadeh’s theory of fuzzy sets. In ordinary “crisp” set theory, an element either belongs to a set or it does not, with no middle ground. A fuzzy set replaces this hard yes/no with a degree of membership between 0 and 1. A document might belong to the category “highly relevant” with a membership of 0.8 and to “moderately relevant” with 0.3. This captures the partial, overlapping way real concepts behave, which is far closer to how people actually judge relevance and meaning.
De Luca and Termini’s fuzzy entropy
In 1972, A. De Luca and S. Termini gave this idea a formal measure with their definition of a non-probabilistic entropy for fuzzy sets. Their formula mirrors Shannon’s but uses membership values in place of probabilities. The key insight is where the measure peaks. Fuzzy entropy is at its maximum when a membership value is 0.5, the point of greatest vagueness, because the element sits exactly between belonging and not belonging. It falls to zero when membership is 0 or 1, where the set behaves like a crisp set and there is no fuzziness at all. The measure quantifies how “fuzzy” the set is, not how random it is.
Later researchers added their own versions. Bart Kosko proposed a geometric fuzzy entropy based on how close a fuzzy set sits to its nearest crisp corner, and several axiomatic refinements have since extended the idea to more complex fuzzy set types. The common thread is measuring uncertainty that has nothing to do with probability.
Applications in information work
Fuzzy measures fit information science naturally because so much of it involves graded judgments. Relevance in information retrieval is rarely all-or-nothing; a search result can be partly on target. Subject indexing, document classification, and recommender systems all deal with concepts whose boundaries blur. Fuzzy entropy gives a way to quantify and compare that vagueness, helping designers build systems that handle “sort of relevant” results gracefully instead of forcing every item into a rigid yes or no.
Bringing the three together
These three measures are not rivals competing to replace Shannon’s formula. They are companions, each extending it along a different axis. Picard’s decision trees stretch information measurement across a sequence of questions, showing how uncertainty falls step by step. Rényi information generalizes the shape of the measure, packing many entropies into one tunable family. Fuzzy measures change the kind of uncertainty being captured, moving from randomness to vagueness. Together they show that Shannon’s 1948 insight was not a finished destination but a starting point, flexible enough to grow into tools for surveys, skewed data, and imprecise concepts alike.
What do you think? If you were designing a reference interview or a user survey as a decision tree, which single question would remove the most uncertainty about what your user actually needs? And when you judge a search result as “partly relevant,” are you really dealing with probability, or with the kind of vagueness that fuzzy measures were built to describe?
References
- https://en.wikipedia.org/wiki/Entropy_(information_theory)
- https://link.springer.com/chapter/10.1007/3-540-18579-8_35
- https://www.sciencedirect.com/science/article/abs/pii/0306457384900700
- https://en.wikipedia.org/wiki/R%C3%A9nyi_entropy
- https://www.sciencedirect.com/science/article/abs/pii/0165011479900204
- https://doi.org/10.3390/axioms13080556

Leave a Reply