Some researchers publish prolifically across decades, while most contribute just a single paper before disappearing from the record. This is not random chance. Back in 1926, a scientist named Alfred J. Lotka noticed a striking regularity in how scientific output spreads across people, and he captured it in a simple mathematical relationship. That relationship, now known as Lotka’s Law, became the first quantitative law of what we today call scientometrics and informetrics. It tells us something profound about the unequal nature of scholarly productivity, and it still shapes how we study research output nearly a century later.

Table of Contents

What is Lotka’s Law?

Lotka’s Law describes the frequency distribution of scientific productivity among authors. In plain terms, it predicts how many authors will write one paper, how many will write two, how many will write three, and so on. The law was introduced by Alfred James Lotka, an American demographer and statistician, in a paper titled “The frequency distribution of scientific productivity,” published in the Journal of the Washington Academy of Sciences. It holds the distinction of being the earliest of the three classical bibliometric laws, the other two being Bradford’s Law of scattering and Zipf’s Law of word frequency.

The core idea is that a small number of authors produce a large share of the literature, while the majority publish very little. Productivity, in other words, is heavily skewed. A handful of highly active researchers dominate output in any given field, and they sit atop a long tail of one-time contributors.

The mathematical formulation

Lotka expressed the relationship he observed through a compact equation. The general form states that the number of authors producing a given number of papers is inversely proportional to the square of that number of papers. Written out, the relation is:

yx = C / xn

Here, yx is the number of authors who have published x papers, while C and n are two constants estimated from the actual data of a specific field. In Lotka’s original analysis, the value of C came out to 0.6079 when n was set to 2, and 0.5669 when n was 1.888. When the exponent n equals 2, the formula becomes the famous inverse square relationship, which is why the law is popularly called the inverse square law of scientific productivity.

The practical meaning of the inverse square version is easy to grasp through a series of fractions. About 60 percent of authors in a field contribute just one publication, around 15 percent contribute two, roughly 7 percent contribute three, and only about 6 percent produce more than ten articles. Each step down the productivity ladder is governed by the square of the paper count. The number of authors writing two papers is one-fourth (1/2²) of those writing one, the number writing three is one-ninth (1/3²), and so on.

A concrete illustration of the proportions

Working with a round figure makes the pattern vivid. Suppose a field has 1,000 authors. According to the inverse square law, about 600 of them produce a single paper, 250 produce two papers, 111 produce three papers, and 63 produce four papers. Notice how steeply the numbers fall. The drop from one paper to two is dramatic, and by the time we reach authors with four publications, we are dealing with a tiny minority. This concentration of output among few producers is the signature feature of the law.

Empirical findings behind the law

Lotka did not arrive at his law through abstract reasoning alone. He built it on careful counting of real data. To derive the relationship, he tallied the number of personal names appearing in the 1907 to 1916 decennial index of Chemical Abstracts, covering only entries under the letters A and B of the alphabet. Against each name, he recorded how many entries appeared, giving him a distribution of authors by their output.

To check whether the pattern extended beyond chemistry, he applied the same counting method to a second source. He used the name index in Felix Auerbach’s work on the history of physics, which covered the entire span of physics through the year 1900. By drawing on two distinct disciplines, Lotka hoped to show that the distribution was not peculiar to a single subject.

The role of the exponent and the slope

When Lotka fitted his data, he did not get a clean value of exactly 2 for the exponent. Instead, using the method of least squares, he found the slope of the curve to be 1.888 for the Chemical Abstracts data and 2.021 for the Auerbach physics data. The chemistry figure sat slightly below 2, while the physics figure sat just above it. Lotka treated 2 as a convenient rounded approximation, and this is how the inverse square label entered common usage.

It is worth pausing on a methodological detail here. Lotka identified his power-law behaviour by fitting a regression line on a logarithmic plot and judging the quality of fit. Modern statisticians have raised concerns about this approach. A recent analysis points out that Lotka heavily truncated the right-hand tail of his data and relied on a regression-line fitting method that current theory considers unreliable for power-law distributions. This does not invalidate the law as a useful approximation, but it reminds us that the original derivation rested on the statistical conventions of its time.

How well does the law hold across fields?

Since 1926, hundreds of studies have tested whether Lotka’s Law fits the literature of specific disciplines. The results are genuinely mixed, which is itself an important finding. The law works as a broad approximation rather than a rigid rule, and the exponent often shifts away from 2 depending on the field.

For example, research on information science literature found that an exponent of around 3.5 gave a better fit than the standard value of 2, while studies of library science and computer science also produced exponents that diverged from the inverse square form. Some humanities and map librarianship datasets matched the classic law closely, whereas medicine and certain other fields did not conform well at all. The general lesson is that author productivity is always skewed, but the exact steepness of the skew varies from one discipline to another.

Lotka’s Law in the Indian research context

Indian researchers have actively tested the applicability of this law to local scientific output, and the findings echo the global pattern of partial fit. A study examining research productivity in the Council of Scientific and Industrial Research found that the distribution did not follow the strict inverse square form, with the author productivity pattern in CSIR samples diverging from Lotka’s prediction. One explanation offered was the longer period of research participation among scientists in such institutions, which inflates the output of senior contributors.

More recent scientometric work continues this tradition. A study on gastritis research published in a Defence Research and Development Organisation journal used statistical tools like the Kolmogorov-Smirnov test and the Chi-Square test, and it found notable discrepancies between the observed publication distribution and what Lotka’s Law predicted. Such testing procedures, especially the goodness-of-fit checks, have become standard practice when applying the law to any new body of literature.

Why Lotka’s Law still matters

The enduring value of Lotka’s Law lies in what it reveals about the structure of knowledge production. It belongs to a family of three foundational laws that together describe how scholarly information is distributed. Bradford’s Law deals with how journal literature on a subject scatters across publications, Zipf’s Law studies the frequency of words, and Lotka’s Law focuses specifically on author productivity. Each law captures a different dimension of the same underlying phenomenon, namely the highly uneven concentration of output.

From a single law to Lotkaian informetrics

Lotka’s modest 1926 paper eventually grew into an entire theoretical framework. The mathematician Leo Egghe developed what he termed “Lotkaian informetrics,” presenting informetric results from the viewpoint of size-frequency functions within a broader model of information production processes. This body of work shows that the relationship Lotka observed connects mathematically to other distributions, including applications linked to Zipf’s Law across many fields. What began as a count of names in a chemistry index became a generalised power-law approach to studying productivity.

The same skewed pattern that Lotka identified has also turned out to be remarkably universal. The inverse power form crops up not only in author productivity but in the frequency of words, the productivity of journals, and even in economic and demographic phenomena, suggesting that a common underlying mechanism may drive the concentration of activity across very different domains. This is part of why the law continues to attract attention from researchers in computer science, linguistics, and beyond, not just library and information science.

Practical uses today

For students and practitioners, the law offers a useful diagnostic tool. By applying it to a body of literature, one can quickly identify the small group of highly productive authors who anchor a field and distinguish them from the large pool of occasional contributors. This helps in mapping research communities, allocating attention to core authors, and understanding the dynamics of collaboration. With modern computing, applying these laws to enormous datasets has become far easier than it was in Lotka’s era of manual counting, opening the door to large-scale analyses across global research databases.

What do you think? If author productivity follows such a predictable and skewed pattern across so many fields, what does this tell us about how scientific talent and opportunity are actually distributed in the research community? And should funding and recognition systems be designed around the few prolific authors at the top, or around supporting the vast majority who contribute only a single paper?

How useful was this post?

Click on a star to rate it!

Average rating 5 / 5. Vote count: 1

No votes so far! Be the first to rate this post.

We are sorry that this post was not useful for you!

Let us improve this post!

Tell us how we can improve this post?

References
  1. https://files.eric.ed.gov/fulltext/EJ1115017.pdf
  2. https://www.sciencedirect.com/science/article/abs/pii/S0306457398000272
  3. https://arxiv.org/pdf/2103.03738
  4. https://arxiv.org/pdf/1411.0928
  5. https://arxiv.org/pdf/1601.04950
  6. https://dl.acm.org/doi/abs/10.1002/asi.23785
  7. https://arxiv.org/pdf/2102.09182
  8. https://publications.drdo.gov.in/ojs/index.php/djlit/article/view/21019
  9. https://ebooks.inflibnet.ac.in/liscp10/chapter/classical-law-of-bibliometrics/
  10. https://link.springer.com/article/10.1007/s11192-005-0211-5

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

Informetrics & Scientometrics

1 Information and Measurement

  1. Information Revisited
  2. Framework for Information Exchange
  3. Measurement Techniques
  4. Informativeness
  5. Standardization of Measurement

2 Measure of Information

  1. Information and Entropy
  2. Shannon Information
  3. Probabilistic Information
  4. Properties of Shannon Information
  5. Derivation of Shannon Information Formula
  6. Normalization Condition
  7. Relating Semantic Value to Shannon Type Measures
  8. Other Shannon Type Measures of Information
  9. Semantic Information
  10. Fuzzy Information Measure
  11. Other Information Measures

3 Informetrics – Definition, Scope and Evolution

  1. Definitions
  2. Scope
  3. Evolution
  4. Summary

4 Sociology of Science and Scientometrics

  1. Sociology of Science
  2. Growth of Scientific Knowledge
  3. Social Organization in Research Areas
  4. Approaches of Scientometrics to Sociology of Science
  5. Models of Growth of Knowledge

5 Organizations Engaged in Scientometrics and Informetrics Studies

  1. Organizations Engaged in or Supporting Scientometrics/Informetrics Studies
  2. Websites
  3. Research Groups/Discussion Groups
  4. Periodical Publications
  5. Conferences/Seminars/Workshops/Congresses
  6. Individuals Engaged in the Study and Research in Scientometrics/Informetrics

6 Law of Scattering and its Applications

  1. Introduction
  2. Historical Account
  3. Bradford’s Law
  4. Verbal Form of Bradford’s Law
  5. Applications of Bradford’s Law
  6. Graphical Representation of Bradford’s Law
  7. Conditions for Bradford’s Law
  8. Falling Tail of Bradford Curve: The Groos Droop
  9. Ambiguity in Bradford’s Law
  10. Fitting Bibliographic Data to Bradford’s Law

7 Rank and Size Frequency Models

  1. Representations and Organization of Numerical Data
  2. Size – Frequency Approach
  3. Rank – Frequency Approach
  4. Size – Frequency Models
  5. Rank – Frequency Cumulative (Fractional) Models
  6. Rank – Frequency Cumulative (Non-Fractional) Models
  7. Rank – Frequency Non – Cumulative Models

8 Informetrics Phenomena

  1. Terminology and Historical Development
  2. Selected Laws of Bibliometrics and Informetrics
  3. Informetrics Phenomena in Science
  4. Practical Applications of Informetrics

9 Analysis of Library Related Data

  1. Necessity for Analytical Studies in Libraries
  2. Citation Counting: A Versatile Tool for Journal Selection
  3. An Alternative Method of Citation Analysis
  4. Selection of New Source Journals to Eliminate Bias Due to Country, and Language
  5. Weightage Formula to Correct Citation for Post-War Periodicals
  6. Three New Bibliometric Parameters to Re-Rank Scientific Periodicals
  7. Garfield’s Methods for Cito-Analytical Studies
  8. Librametric Analysis
  9. Bibliometric Analysis
  10. Informetrics
  11. Scientometrics: Its Genesis, Scope, Definition, and Applications

10 User Studies

  1. User Studies
  2. Questionnaire Method
  3. Interview Method
  4. Diary Method
  5. Observation Method
  6. Planning a Survey
  7. Classification and Tabulation of Data
  8. Analysis of Data
  9. Presentation of Results
  10. Important User Studies
  11. Application of User Studies

11 Laws of Scientific Productivity

  1. Scientific Productivity – Influencing Factors
  2. Scientific Productivity – Problems in Measurement
  3. Scientific Productivity – Distribution Characteristics
  4. Lotka’s Law
  5. Statistical Distributions or Models
  6. Application of Lotka’s Law
  7. Goodness-of-Fit Test

12 Growth and Obsolescence of Literature

  1. Growth of Literature
  2. Obsolescence of Literature
  3. Growth Vs Obsolescence of Literature

13 Science Indicators

  1. Indicators
  2. Towards Science Indicators
  3. Historical Aspects
  4. Functions of Science Indicators
  5. S&T Indicators for the Developing Countries
  6. Types of Indicators
  7. Validity and Reliability of Indicators
  8. Building S&T Indicators
  9. Literature Based Indicators
  10. Patent Indicators

14 Mapping of Science

  1. Cognitive Mapping
  2. Journal-to-journal Citation Maps
  3. Co-citation Maps
  4. Co-word Maps
  5. Co-classification Maps
  6. Descriptive Mapping

15 Elements of Statistics

  1. Data and Its Measurement
  2. Graphical Representation
  3. Measures of Central Tendency
  4. Measure of Variability
  5. Correlation and Regression

16 Probability Distributions and their Applications

  1. Probability – Definition
  2. Random Variables
  3. Joint Probability Distribution
  4. Conditional Probability Distribution
  5. Some Special Distributions
  6. Applications of Probability

17 Regression Analysis

  1. Simple Linear Regression
  2. Multiple Regression
  3. Stepwise Regression
  4. Regression with Qualitative Explanatory Variables

18 Cluster Analysis and Factor Analysis

  1. Introduction
  2. Cluster Analysis
  3. Factor Analysis
  4. Examples of Cluster and Factor Analysis