Claude Shannon gave the world a single, elegant equation in 1948 that turned the fuzzy idea of “information” into something measurable. His entropy formula, H(X) = −∑ p(x) log p(x), tells us the average uncertainty in a set of outcomes and remains the backbone of information theory. But the basic formula assumes we already know the probabilities of every outcome, and that the only thing worth measuring is statistical uncertainty. Real information problems are messier. Sometimes we gather information by asking questions one at a time. Sometimes we want to weight rare events differently. Sometimes the uncertainty is not about chance at all, but about vagueness. To handle these situations, researchers built a family of Shannon-type measures that keep the spirit of his work while stretching it in new directions. This post looks at three of them: Picard’s decision tree approach, Rényi information, and fuzzy measures.

Table of Contents

Why Shannon’s formula needed company

Shannon’s entropy works beautifully when we have a clean probability distribution and a clear communication channel. It quantifies the average level of uncertainty tied to a random variable and sets the absolute limit on how far data can be compressed without loss. The trouble is that this model treats information as a one-shot affair: the source emits symbols, and we measure the uncertainty across all of them at once.

In practice, information is often acquired step by step, and not every uncertain situation is governed by probability. A reference librarian narrowing down a patron’s vague request, a survey designer choosing which question to ask first, or a system trying to classify a document that only “partly” belongs to a category, all face problems the original formula does not directly address. The three measures below were developed precisely to fill these gaps.

Picard’s decision tree approach

The French mathematician C.-F. Picard built a bridge between information theory and the practical act of asking questions. His work, published as Graphes et questionnaires in 1972 and later translated as Graphs and Questionnaires, models the process of gathering information as a structured sequence of questions rather than a single measurement.

How a questionnaire becomes a decision tree

Picard’s central idea is that a questionnaire is really a decision tree. Each question is a node. Each possible answer is a branch leading to the next question or to a final outcome. Imagine you want to identify which of eight equally likely outcomes is true. If you ask well-designed yes/no questions, each answer cuts the set of possibilities roughly in half, and you can reach the answer in about three questions. That number, three, is not a coincidence: it equals log₂ 8, the Shannon entropy of eight equally likely events. The structure of the tree and the information content of the source are deeply linked.

Measuring information through questions

In this framework, the value of a question is judged by how much uncertainty it removes. Before a question is asked, the situation has a certain entropy. After the answer arrives, the remaining uncertainty is usually smaller. The difference is the information the question delivered. This is the same logic that later powered information gain in machine learning, where decision tree algorithms choose the attribute that reduces entropy the most at each split. Picard’s questionnaire theory anticipated and formalised this connection between tree structure and entropy long before it became routine in data science.

His approach also separates the idea of uncertainty from the act of measurement. Mathematicians such as B. Forte extended this thinking, exploring entropies with and without probabilities and applying them directly to questionnaire theory. This matters because not every question has neat probabilities attached, yet we still want to know which question is “worth more.”

Why this matters for information science

For anyone working with users, this is more than abstract theory. When a librarian conducts a reference interview, each question is a node in an implicit decision tree, and a skilled professional instinctively asks the question that narrows the search most. The same principle guides good survey design: ordering questions so that each one extracts maximum information keeps questionnaires short and respondents engaged. Picard’s model gives a mathematical reason for what experienced practitioners already do by intuition.

Rényi information

The second extension comes from the Hungarian mathematician Alfréd Rényi, who in 1961 asked a sharp question: is Shannon entropy the only sensible way to measure information, or just one member of a larger family? He answered by deriving a generalized measure now called Rényi entropy, built to be the most general form that still preserves additivity for independent events.

The generalized formula

Rényi entropy of order α is written as H_α(X) = [1 / (1 − α)] log ∑ pᵢ^α, where α is a positive parameter that tunes the measure’s sensitivity. The clever part is that this single equation is not one measure but a whole spectrum of them. By turning the dial on α, you shift how much weight the formula gives to common versus rare outcomes. A low α treats all outcomes more evenly, while a high α lets the most probable events dominate.

Special cases hidden inside one equation

What makes Rényi’s family so useful is that several well-known measures sit inside it as special cases. When α approaches 1, the formula collapses neatly back into ordinary Shannon entropy, which is why Rényi entropy is rightly called a generalization rather than a replacement. At α = 0 it becomes the Hartley or max-entropy, which counts the number of possible outcomes regardless of their probabilities. At α = 2 it gives collision entropy, linked to the chance that two independent draws yield the same value. As α grows very large it tends toward min-entropy, which focuses entirely on the single most likely outcome and is widely used in cryptography to gauge worst-case unpredictability.

Where Rényi measures are useful

Because one parameter controls so much, Rényi information is valuable wherever a single average is too blunt. In ecology and statistics it underpins diversity indices that describe how varied a population is. In signal processing it helps analyse complex time-frequency patterns. For informetrics and scientometrics, this flexibility is appealing because citation distributions, keyword frequencies, and journal usage are often highly skewed. A measure that can be tuned to emphasise either the dominant journals or the long tail of rarely cited works gives analysts a richer picture than a single fixed formula would.

Fuzzy measures

The third extension changes the very nature of the uncertainty being measured. Shannon and Rényi both deal with randomness, where outcomes are uncertain because of chance. Fuzzy measures deal with vagueness, where the uncertainty comes from imprecise boundaries rather than probability.

From crisp sets to degrees of membership

The foundation is Lotfi Zadeh’s theory of fuzzy sets. In ordinary “crisp” set theory, an element either belongs to a set or it does not, with no middle ground. A fuzzy set replaces this hard yes/no with a degree of membership between 0 and 1. A document might belong to the category “highly relevant” with a membership of 0.8 and to “moderately relevant” with 0.3. This captures the partial, overlapping way real concepts behave, which is far closer to how people actually judge relevance and meaning.

De Luca and Termini’s fuzzy entropy

In 1972, A. De Luca and S. Termini gave this idea a formal measure with their definition of a non-probabilistic entropy for fuzzy sets. Their formula mirrors Shannon’s but uses membership values in place of probabilities. The key insight is where the measure peaks. Fuzzy entropy is at its maximum when a membership value is 0.5, the point of greatest vagueness, because the element sits exactly between belonging and not belonging. It falls to zero when membership is 0 or 1, where the set behaves like a crisp set and there is no fuzziness at all. The measure quantifies how “fuzzy” the set is, not how random it is.

Later researchers added their own versions. Bart Kosko proposed a geometric fuzzy entropy based on how close a fuzzy set sits to its nearest crisp corner, and several axiomatic refinements have since extended the idea to more complex fuzzy set types. The common thread is measuring uncertainty that has nothing to do with probability.

Applications in information work

Fuzzy measures fit information science naturally because so much of it involves graded judgments. Relevance in information retrieval is rarely all-or-nothing; a search result can be partly on target. Subject indexing, document classification, and recommender systems all deal with concepts whose boundaries blur. Fuzzy entropy gives a way to quantify and compare that vagueness, helping designers build systems that handle “sort of relevant” results gracefully instead of forcing every item into a rigid yes or no.

Bringing the three together

These three measures are not rivals competing to replace Shannon’s formula. They are companions, each extending it along a different axis. Picard’s decision trees stretch information measurement across a sequence of questions, showing how uncertainty falls step by step. Rényi information generalizes the shape of the measure, packing many entropies into one tunable family. Fuzzy measures change the kind of uncertainty being captured, moving from randomness to vagueness. Together they show that Shannon’s 1948 insight was not a finished destination but a starting point, flexible enough to grow into tools for surveys, skewed data, and imprecise concepts alike.

What do you think? If you were designing a reference interview or a user survey as a decision tree, which single question would remove the most uncertainty about what your user actually needs? And when you judge a search result as “partly relevant,” are you really dealing with probability, or with the kind of vagueness that fuzzy measures were built to describe?

How useful was this post?

Click on a star to rate it!

Average rating 0 / 5. Vote count: 0

No votes so far! Be the first to rate this post.

We are sorry that this post was not useful for you!

Let us improve this post!

Tell us how we can improve this post?

References
  1. https://en.wikipedia.org/wiki/Entropy_(information_theory)
  2. https://link.springer.com/chapter/10.1007/3-540-18579-8_35
  3. https://www.sciencedirect.com/science/article/abs/pii/0306457384900700
  4. https://en.wikipedia.org/wiki/R%C3%A9nyi_entropy
  5. https://www.sciencedirect.com/science/article/abs/pii/0165011479900204
  6. https://doi.org/10.3390/axioms13080556

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

Informetrics & Scientometrics

1 Information and Measurement

  1. Information Revisited
  2. Framework for Information Exchange
  3. Measurement Techniques
  4. Informativeness
  5. Standardization of Measurement

2 Measure of Information

  1. Information and Entropy
  2. Shannon Information
  3. Probabilistic Information
  4. Properties of Shannon Information
  5. Derivation of Shannon Information Formula
  6. Normalization Condition
  7. Relating Semantic Value to Shannon Type Measures
  8. Other Shannon Type Measures of Information
  9. Semantic Information
  10. Fuzzy Information Measure
  11. Other Information Measures

3 Informetrics – Definition, Scope and Evolution

  1. Definitions
  2. Scope
  3. Evolution
  4. Summary

4 Sociology of Science and Scientometrics

  1. Sociology of Science
  2. Growth of Scientific Knowledge
  3. Social Organization in Research Areas
  4. Approaches of Scientometrics to Sociology of Science
  5. Models of Growth of Knowledge

5 Organizations Engaged in Scientometrics and Informetrics Studies

  1. Organizations Engaged in or Supporting Scientometrics/Informetrics Studies
  2. Websites
  3. Research Groups/Discussion Groups
  4. Periodical Publications
  5. Conferences/Seminars/Workshops/Congresses
  6. Individuals Engaged in the Study and Research in Scientometrics/Informetrics

6 Law of Scattering and its Applications

  1. Introduction
  2. Historical Account
  3. Bradford’s Law
  4. Verbal Form of Bradford’s Law
  5. Applications of Bradford’s Law
  6. Graphical Representation of Bradford’s Law
  7. Conditions for Bradford’s Law
  8. Falling Tail of Bradford Curve: The Groos Droop
  9. Ambiguity in Bradford’s Law
  10. Fitting Bibliographic Data to Bradford’s Law

7 Rank and Size Frequency Models

  1. Representations and Organization of Numerical Data
  2. Size – Frequency Approach
  3. Rank – Frequency Approach
  4. Size – Frequency Models
  5. Rank – Frequency Cumulative (Fractional) Models
  6. Rank – Frequency Cumulative (Non-Fractional) Models
  7. Rank – Frequency Non – Cumulative Models

8 Informetrics Phenomena

  1. Terminology and Historical Development
  2. Selected Laws of Bibliometrics and Informetrics
  3. Informetrics Phenomena in Science
  4. Practical Applications of Informetrics

9 Analysis of Library Related Data

  1. Necessity for Analytical Studies in Libraries
  2. Citation Counting: A Versatile Tool for Journal Selection
  3. An Alternative Method of Citation Analysis
  4. Selection of New Source Journals to Eliminate Bias Due to Country, and Language
  5. Weightage Formula to Correct Citation for Post-War Periodicals
  6. Three New Bibliometric Parameters to Re-Rank Scientific Periodicals
  7. Garfield’s Methods for Cito-Analytical Studies
  8. Librametric Analysis
  9. Bibliometric Analysis
  10. Informetrics
  11. Scientometrics: Its Genesis, Scope, Definition, and Applications

10 User Studies

  1. User Studies
  2. Questionnaire Method
  3. Interview Method
  4. Diary Method
  5. Observation Method
  6. Planning a Survey
  7. Classification and Tabulation of Data
  8. Analysis of Data
  9. Presentation of Results
  10. Important User Studies
  11. Application of User Studies

11 Laws of Scientific Productivity

  1. Scientific Productivity – Influencing Factors
  2. Scientific Productivity – Problems in Measurement
  3. Scientific Productivity – Distribution Characteristics
  4. Lotka’s Law
  5. Statistical Distributions or Models
  6. Application of Lotka’s Law
  7. Goodness-of-Fit Test

12 Growth and Obsolescence of Literature

  1. Growth of Literature
  2. Obsolescence of Literature
  3. Growth Vs Obsolescence of Literature

13 Science Indicators

  1. Indicators
  2. Towards Science Indicators
  3. Historical Aspects
  4. Functions of Science Indicators
  5. S&T Indicators for the Developing Countries
  6. Types of Indicators
  7. Validity and Reliability of Indicators
  8. Building S&T Indicators
  9. Literature Based Indicators
  10. Patent Indicators

14 Mapping of Science

  1. Cognitive Mapping
  2. Journal-to-journal Citation Maps
  3. Co-citation Maps
  4. Co-word Maps
  5. Co-classification Maps
  6. Descriptive Mapping

15 Elements of Statistics

  1. Data and Its Measurement
  2. Graphical Representation
  3. Measures of Central Tendency
  4. Measure of Variability
  5. Correlation and Regression

16 Probability Distributions and their Applications

  1. Probability – Definition
  2. Random Variables
  3. Joint Probability Distribution
  4. Conditional Probability Distribution
  5. Some Special Distributions
  6. Applications of Probability

17 Regression Analysis

  1. Simple Linear Regression
  2. Multiple Regression
  3. Stepwise Regression
  4. Regression with Qualitative Explanatory Variables

18 Cluster Analysis and Factor Analysis

  1. Introduction
  2. Cluster Analysis
  3. Factor Analysis
  4. Examples of Cluster and Factor Analysis