Every research field has a handful of names that appear again and again, and a long tail of authors who publish just once or twice. The same lopsided pattern shows up in journals: a few titles carry most of the important papers, while hundreds carry only a stray article. Size-frequency models are the mathematical tools that turn this everyday observation into something measurable and predictable. They sit at the heart of bibliometrics and informetrics, helping researchers describe how scholarly output is shared among authors, journals, and even words. This post explains what these models are, starting with the most famous one, Lotka’s law, and then moving to the contributions of Kendall and Fairthorne.

Table of Contents

What size-frequency models actually measure

To understand size-frequency models, you first need two ideas: sources and items. Sources are the producers, such as authors or journals. Items are what they produce, such as papers or articles. A size-frequency model asks a simple question: how many sources produce exactly one item, how many produce two, how many produce three, and so on?

This is different from a rank-frequency model. In a rank-frequency approach, sources are first ranked in decreasing order of productivity, and you study how output falls as rank increases. Zipf’s law of word frequencies and Bradford’s law of scattering are rank-frequency models, while Lotka’s law is the classic size-frequency model that describes how items are distributed over sources. The mathematician Leo Egghe formalised this so completely that he named the whole framework Lotkaian informetrics, built on the size-frequency function f(n) = C/n^a.

The key feature of every one of these models is that they are not symmetrical, bell-shaped curves. They are highly skewed. A small number of sources account for a large share of output, and a very large number of sources account for very little. This concentration is what makes the models so useful for prediction.

Lotka’s law: the foundation of author productivity

In 1926, the statistician Alfred J. Lotka examined author productivity in chemistry and physics. He counted how many authors had written one paper, two papers, three papers, and so on, using data from Chemical Abstracts and a history of physics. What he found became one of the founding laws of the field.

The inverse square relationship

Lotka observed that the number of authors producing a given number of papers drops off sharply in a regular way. In its most quoted form, the law states that the number of authors producing n papers is about 1/n² of those producing a single paper. So if 100 authors write one paper each, roughly 25 write two, about 11 write three, and so on. The pattern also implies that around 60 percent of all contributors in a field publish only once. This inverse square relationship is why Lotka’s law is often called the inverse square law of scientific productivity.

The intuition behind this is the idea that “success breeds success.” Authors who have already published find it easier to publish again, attracting collaborators, funding, and recognition. Derek de Solla Price captured this formally as a cumulative advantage process, where the rich get richer in publication terms.

The generalized form

The exponent of 2 was Lotka’s specific finding, but his original work was based on a more general inverse power relationship. The generalized form is usually written as x^n · y = C, where x is the number of papers, y is the number of authors producing that many papers, and n and C are constants estimated from the actual data. When n equals 2, you get the classic inverse square law. In practice, the exponent varies from field to field, so researchers calculate it for the dataset in front of them rather than assuming it is always 2.

Testing whether the law holds

A model is only useful if you can check it against reality. Bibliometricians test Lotka’s law by comparing the observed distribution of authors against the distribution the formula predicts. Two statistical tests are commonly used: the Kolmogorov-Smirnov (K-S) test and the Chi-square test. Both measure whether the gap between the observed and expected values is small enough to be due to chance.

Indian research output features heavily in this kind of testing. A study published in the DRDO’s DESIDOC Journal of Library and Information Technology examined Gastritis research and applied the K-S and Chi-square tests to check how well Lotka’s law described author productivity. Studies like this show that the law does not fit every dataset perfectly; the exponent and constant must be tuned, and sometimes the law fails the goodness-of-fit test entirely. That is a normal and healthy part of how the model is used.

Beyond Lotka: Kendall, Fairthorne and journal productivity

Lotka focused on authors, but the same logic applies to journals. A few journals publish a large share of the papers on any topic, while a long tail of journals publish only a handful. Two figures helped extend size-frequency thinking to journal productivity and to prediction.

Kendall and the scatter of journal literature

The statistician Maurice Kendall studied how the literature of a subject is spread across journals. His 1960 paper on the bibliography of operational research, published in the Operational Research Quarterly, looked at how articles in a defined field cluster in a small group of core journals and scatter thinly across many others. This connects directly to Bradford’s law of scattering, and it gave bibliographers a way to estimate how many journals contribute a given number of articles to a field. For a librarian building a subject collection, that prediction is gold: it suggests which journals are essential and which are marginal.

Fairthorne’s hyperbolic distributions

Robert A. Fairthorne pulled these threads together in a landmark 1969 paper titled Empirical Hyperbolic Distributions for Bibliometric Description and Prediction, published in the Journal of Documentation. The same article is historically important for another reason: it sits in the very issue where Alan Pritchard introduced the word “bibliometrics” to replace the older phrase “statistical bibliography.”

Fairthorne’s central insight was that the bibliometric distributions of Bradford, Zipf, and Mandelbrot are all members of one family of hyperbolic distributions. Because they share the same mathematical shape, they can be used not just to describe data after the fact but to predict future patterns, such as how many journals will contribute a fixed number of articles as a literature grows. Later scholars confirmed this unity. Work on Bradford’s law notes that Fairthorne, Price, and Bookstein all argued that the various bibliometric distributions are deeply consistent, behaving as special cases of a single hyperbolic form.

How bibliometricians put these models to work

These models are not just academic curiosities. They shape practical decisions in libraries, research administration, and publishing.

Collection development: By predicting which journals carry the bulk of articles in a subject, a library can prioritise subscriptions and weed peripheral titles without losing access to the core literature. This is especially valuable when budgets are tight and serials are expensive.

Identifying core contributors: Size-frequency analysis quickly separates the small group of highly productive authors from the large group of occasional contributors. Funding bodies, departments, and editors use this to spot the active core of a field.

Research evaluation: When a department or country measures its research performance, knowing the expected shape of author productivity helps set realistic benchmarks. An institution can ask whether its productivity distribution matches or deviates from the pattern Lotka predicts.

Forecasting growth: Because the distributions are hyperbolic and stable, they let analysts estimate how a literature will expand, how scatter will increase, and how many new sources are likely to appear. Springer’s work on bibliometric models for journal productivity reviews many such models built for exactly this purpose.

It is worth remembering that these are empirical laws, derived from observation rather than from first principles. They describe strong tendencies, not iron rules. A dataset may follow Lotka’s law closely, loosely, or not at all, and the honest researcher reports the goodness-of-fit results either way. The value of the model lies in giving a clear expectation against which real data can be compared.

Why the models agree with each other

One of the most satisfying ideas in this area is that Lotka’s, Bradford’s, and Zipf’s laws are not three separate discoveries but three views of the same underlying phenomenon. Lotka counts sources by their output size. Zipf ranks sources and watches output decline. Bradford accumulates journals into zones of equal yield. As analysis of Zipf’s law has shown, these distributions are mathematically equivalent and connect to the wider family that includes Pareto and Price. Fairthorne’s hyperbolic framework, and later Egghe’s Lotkaian informetrics, made this unity explicit. Understanding one of these laws well gives you a foothold on all of them.

What do you think? If a small fraction of authors really do produce most of the research in any field, should research funding follow that concentration or deliberately work against it? And when a real dataset fails the goodness-of-fit test for Lotka’s law, does that weaken the law, or does it reveal something interesting about that particular field?

How useful was this post?

Click on a star to rate it!

Average rating 0 / 5. Vote count: 0

No votes so far! Be the first to rate this post.

We are sorry that this post was not useful for you!

Let us improve this post!

Tell us how we can improve this post?

References
  1. https://www.sciencedirect.com/science/article/pii/S2694610625000104
  2. https://link.springer.com/article/10.1023/A:1017919924342
  3. https://en.wikipedia.org/wiki/Lotka%27s_law
  4. https://asistdl.onlinelibrary.wiley.com/doi/10.1002/asi.4630270505
  5. https://publications.drdo.gov.in/ojs/index.php/djlit/article/view/21019
  6. https://asistdl.onlinelibrary.wiley.com/doi/abs/10.1002/asi.4630240207
  7. https://www.emerald.com/jd/article-abstract/25/4/319/206708/Empirical-Hyperbolic-Distributions-Bradford-Zipf
  8. https://www.researchgate.net/publication/280218558_Bradford's_Empirical_Law
  9. https://link.springer.com/article/10.1007/BF00353144
  10. https://asistdl.onlinelibrary.wiley.com/doi/abs/10.1002/asi.4630330507

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

Informetrics & Scientometrics

1 Information and Measurement

  1. Information Revisited
  2. Framework for Information Exchange
  3. Measurement Techniques
  4. Informativeness
  5. Standardization of Measurement

2 Measure of Information

  1. Information and Entropy
  2. Shannon Information
  3. Probabilistic Information
  4. Properties of Shannon Information
  5. Derivation of Shannon Information Formula
  6. Normalization Condition
  7. Relating Semantic Value to Shannon Type Measures
  8. Other Shannon Type Measures of Information
  9. Semantic Information
  10. Fuzzy Information Measure
  11. Other Information Measures

3 Informetrics – Definition, Scope and Evolution

  1. Definitions
  2. Scope
  3. Evolution
  4. Summary

4 Sociology of Science and Scientometrics

  1. Sociology of Science
  2. Growth of Scientific Knowledge
  3. Social Organization in Research Areas
  4. Approaches of Scientometrics to Sociology of Science
  5. Models of Growth of Knowledge

5 Organizations Engaged in Scientometrics and Informetrics Studies

  1. Organizations Engaged in or Supporting Scientometrics/Informetrics Studies
  2. Websites
  3. Research Groups/Discussion Groups
  4. Periodical Publications
  5. Conferences/Seminars/Workshops/Congresses
  6. Individuals Engaged in the Study and Research in Scientometrics/Informetrics

6 Law of Scattering and its Applications

  1. Introduction
  2. Historical Account
  3. Bradford’s Law
  4. Verbal Form of Bradford’s Law
  5. Applications of Bradford’s Law
  6. Graphical Representation of Bradford’s Law
  7. Conditions for Bradford’s Law
  8. Falling Tail of Bradford Curve: The Groos Droop
  9. Ambiguity in Bradford’s Law
  10. Fitting Bibliographic Data to Bradford’s Law

7 Rank and Size Frequency Models

  1. Representations and Organization of Numerical Data
  2. Size – Frequency Approach
  3. Rank – Frequency Approach
  4. Size – Frequency Models
  5. Rank – Frequency Cumulative (Fractional) Models
  6. Rank – Frequency Cumulative (Non-Fractional) Models
  7. Rank – Frequency Non – Cumulative Models

8 Informetrics Phenomena

  1. Terminology and Historical Development
  2. Selected Laws of Bibliometrics and Informetrics
  3. Informetrics Phenomena in Science
  4. Practical Applications of Informetrics

9 Analysis of Library Related Data

  1. Necessity for Analytical Studies in Libraries
  2. Citation Counting: A Versatile Tool for Journal Selection
  3. An Alternative Method of Citation Analysis
  4. Selection of New Source Journals to Eliminate Bias Due to Country, and Language
  5. Weightage Formula to Correct Citation for Post-War Periodicals
  6. Three New Bibliometric Parameters to Re-Rank Scientific Periodicals
  7. Garfield’s Methods for Cito-Analytical Studies
  8. Librametric Analysis
  9. Bibliometric Analysis
  10. Informetrics
  11. Scientometrics: Its Genesis, Scope, Definition, and Applications

10 User Studies

  1. User Studies
  2. Questionnaire Method
  3. Interview Method
  4. Diary Method
  5. Observation Method
  6. Planning a Survey
  7. Classification and Tabulation of Data
  8. Analysis of Data
  9. Presentation of Results
  10. Important User Studies
  11. Application of User Studies

11 Laws of Scientific Productivity

  1. Scientific Productivity – Influencing Factors
  2. Scientific Productivity – Problems in Measurement
  3. Scientific Productivity – Distribution Characteristics
  4. Lotka’s Law
  5. Statistical Distributions or Models
  6. Application of Lotka’s Law
  7. Goodness-of-Fit Test

12 Growth and Obsolescence of Literature

  1. Growth of Literature
  2. Obsolescence of Literature
  3. Growth Vs Obsolescence of Literature

13 Science Indicators

  1. Indicators
  2. Towards Science Indicators
  3. Historical Aspects
  4. Functions of Science Indicators
  5. S&T Indicators for the Developing Countries
  6. Types of Indicators
  7. Validity and Reliability of Indicators
  8. Building S&T Indicators
  9. Literature Based Indicators
  10. Patent Indicators

14 Mapping of Science

  1. Cognitive Mapping
  2. Journal-to-journal Citation Maps
  3. Co-citation Maps
  4. Co-word Maps
  5. Co-classification Maps
  6. Descriptive Mapping

15 Elements of Statistics

  1. Data and Its Measurement
  2. Graphical Representation
  3. Measures of Central Tendency
  4. Measure of Variability
  5. Correlation and Regression

16 Probability Distributions and their Applications

  1. Probability – Definition
  2. Random Variables
  3. Joint Probability Distribution
  4. Conditional Probability Distribution
  5. Some Special Distributions
  6. Applications of Probability

17 Regression Analysis

  1. Simple Linear Regression
  2. Multiple Regression
  3. Stepwise Regression
  4. Regression with Qualitative Explanatory Variables

18 Cluster Analysis and Factor Analysis

  1. Introduction
  2. Cluster Analysis
  3. Factor Analysis
  4. Examples of Cluster and Factor Analysis