Some researchers publish prolifically across decades, while most contribute just a single paper before disappearing from the record. This is not random chance. Back in 1926, a scientist named Alfred J. Lotka noticed a striking regularity in how scientific output spreads across people, and he captured it in a simple mathematical relationship. That relationship, now known as Lotka’s Law, became the first quantitative law of what we today call scientometrics and informetrics. It tells us something profound about the unequal nature of scholarly productivity, and it still shapes how we study research output nearly a century later.
Table of Contents
- What is Lotka’s Law?
- The mathematical formulation
- A concrete illustration of the proportions
- Empirical findings behind the law
- The role of the exponent and the slope
- How well does the law hold across fields?
- Lotka’s Law in the Indian research context
- Why Lotka’s Law still matters
- From a single law to Lotkaian informetrics
- Practical uses today
What is Lotka’s Law?
Lotka’s Law describes the frequency distribution of scientific productivity among authors. In plain terms, it predicts how many authors will write one paper, how many will write two, how many will write three, and so on. The law was introduced by Alfred James Lotka, an American demographer and statistician, in a paper titled “The frequency distribution of scientific productivity,” published in the Journal of the Washington Academy of Sciences. It holds the distinction of being the earliest of the three classical bibliometric laws, the other two being Bradford’s Law of scattering and Zipf’s Law of word frequency.
The core idea is that a small number of authors produce a large share of the literature, while the majority publish very little. Productivity, in other words, is heavily skewed. A handful of highly active researchers dominate output in any given field, and they sit atop a long tail of one-time contributors.
The mathematical formulation
Lotka expressed the relationship he observed through a compact equation. The general form states that the number of authors producing a given number of papers is inversely proportional to the square of that number of papers. Written out, the relation is:
yx = C / xn
Here, yx is the number of authors who have published x papers, while C and n are two constants estimated from the actual data of a specific field. In Lotka’s original analysis, the value of C came out to 0.6079 when n was set to 2, and 0.5669 when n was 1.888. When the exponent n equals 2, the formula becomes the famous inverse square relationship, which is why the law is popularly called the inverse square law of scientific productivity.
The practical meaning of the inverse square version is easy to grasp through a series of fractions. About 60 percent of authors in a field contribute just one publication, around 15 percent contribute two, roughly 7 percent contribute three, and only about 6 percent produce more than ten articles. Each step down the productivity ladder is governed by the square of the paper count. The number of authors writing two papers is one-fourth (1/2²) of those writing one, the number writing three is one-ninth (1/3²), and so on.
A concrete illustration of the proportions
Working with a round figure makes the pattern vivid. Suppose a field has 1,000 authors. According to the inverse square law, about 600 of them produce a single paper, 250 produce two papers, 111 produce three papers, and 63 produce four papers. Notice how steeply the numbers fall. The drop from one paper to two is dramatic, and by the time we reach authors with four publications, we are dealing with a tiny minority. This concentration of output among few producers is the signature feature of the law.
Empirical findings behind the law
Lotka did not arrive at his law through abstract reasoning alone. He built it on careful counting of real data. To derive the relationship, he tallied the number of personal names appearing in the 1907 to 1916 decennial index of Chemical Abstracts, covering only entries under the letters A and B of the alphabet. Against each name, he recorded how many entries appeared, giving him a distribution of authors by their output.
To check whether the pattern extended beyond chemistry, he applied the same counting method to a second source. He used the name index in Felix Auerbach’s work on the history of physics, which covered the entire span of physics through the year 1900. By drawing on two distinct disciplines, Lotka hoped to show that the distribution was not peculiar to a single subject.
The role of the exponent and the slope
When Lotka fitted his data, he did not get a clean value of exactly 2 for the exponent. Instead, using the method of least squares, he found the slope of the curve to be 1.888 for the Chemical Abstracts data and 2.021 for the Auerbach physics data. The chemistry figure sat slightly below 2, while the physics figure sat just above it. Lotka treated 2 as a convenient rounded approximation, and this is how the inverse square label entered common usage.
It is worth pausing on a methodological detail here. Lotka identified his power-law behaviour by fitting a regression line on a logarithmic plot and judging the quality of fit. Modern statisticians have raised concerns about this approach. A recent analysis points out that Lotka heavily truncated the right-hand tail of his data and relied on a regression-line fitting method that current theory considers unreliable for power-law distributions. This does not invalidate the law as a useful approximation, but it reminds us that the original derivation rested on the statistical conventions of its time.
How well does the law hold across fields?
Since 1926, hundreds of studies have tested whether Lotka’s Law fits the literature of specific disciplines. The results are genuinely mixed, which is itself an important finding. The law works as a broad approximation rather than a rigid rule, and the exponent often shifts away from 2 depending on the field.
For example, research on information science literature found that an exponent of around 3.5 gave a better fit than the standard value of 2, while studies of library science and computer science also produced exponents that diverged from the inverse square form. Some humanities and map librarianship datasets matched the classic law closely, whereas medicine and certain other fields did not conform well at all. The general lesson is that author productivity is always skewed, but the exact steepness of the skew varies from one discipline to another.
Lotka’s Law in the Indian research context
Indian researchers have actively tested the applicability of this law to local scientific output, and the findings echo the global pattern of partial fit. A study examining research productivity in the Council of Scientific and Industrial Research found that the distribution did not follow the strict inverse square form, with the author productivity pattern in CSIR samples diverging from Lotka’s prediction. One explanation offered was the longer period of research participation among scientists in such institutions, which inflates the output of senior contributors.
More recent scientometric work continues this tradition. A study on gastritis research published in a Defence Research and Development Organisation journal used statistical tools like the Kolmogorov-Smirnov test and the Chi-Square test, and it found notable discrepancies between the observed publication distribution and what Lotka’s Law predicted. Such testing procedures, especially the goodness-of-fit checks, have become standard practice when applying the law to any new body of literature.
Why Lotka’s Law still matters
The enduring value of Lotka’s Law lies in what it reveals about the structure of knowledge production. It belongs to a family of three foundational laws that together describe how scholarly information is distributed. Bradford’s Law deals with how journal literature on a subject scatters across publications, Zipf’s Law studies the frequency of words, and Lotka’s Law focuses specifically on author productivity. Each law captures a different dimension of the same underlying phenomenon, namely the highly uneven concentration of output.
From a single law to Lotkaian informetrics
Lotka’s modest 1926 paper eventually grew into an entire theoretical framework. The mathematician Leo Egghe developed what he termed “Lotkaian informetrics,” presenting informetric results from the viewpoint of size-frequency functions within a broader model of information production processes. This body of work shows that the relationship Lotka observed connects mathematically to other distributions, including applications linked to Zipf’s Law across many fields. What began as a count of names in a chemistry index became a generalised power-law approach to studying productivity.
The same skewed pattern that Lotka identified has also turned out to be remarkably universal. The inverse power form crops up not only in author productivity but in the frequency of words, the productivity of journals, and even in economic and demographic phenomena, suggesting that a common underlying mechanism may drive the concentration of activity across very different domains. This is part of why the law continues to attract attention from researchers in computer science, linguistics, and beyond, not just library and information science.
Practical uses today
For students and practitioners, the law offers a useful diagnostic tool. By applying it to a body of literature, one can quickly identify the small group of highly productive authors who anchor a field and distinguish them from the large pool of occasional contributors. This helps in mapping research communities, allocating attention to core authors, and understanding the dynamics of collaboration. With modern computing, applying these laws to enormous datasets has become far easier than it was in Lotka’s era of manual counting, opening the door to large-scale analyses across global research databases.
What do you think? If author productivity follows such a predictable and skewed pattern across so many fields, what does this tell us about how scientific talent and opportunity are actually distributed in the research community? And should funding and recognition systems be designed around the few prolific authors at the top, or around supporting the vast majority who contribute only a single paper?
References
- https://files.eric.ed.gov/fulltext/EJ1115017.pdf
- https://www.sciencedirect.com/science/article/abs/pii/S0306457398000272
- https://arxiv.org/pdf/2103.03738
- https://arxiv.org/pdf/1411.0928
- https://arxiv.org/pdf/1601.04950
- https://dl.acm.org/doi/abs/10.1002/asi.23785
- https://arxiv.org/pdf/2102.09182
- https://publications.drdo.gov.in/ojs/index.php/djlit/article/view/21019
- https://ebooks.inflibnet.ac.in/liscp10/chapter/classical-law-of-bibliometrics/
- https://link.springer.com/article/10.1007/s11192-005-0211-5

Leave a Reply