Every field of research has a hidden mathematical rhythm. A handful of authors publish most of the papers, a small cluster of journals carries the bulk of the important articles, and a few words appear far more often than the rest. These are not coincidences. They are predictable patterns first observed nearly a century ago, and they form the foundation of bibliometrics and informetrics. Understanding these laws helps librarians build better collections, helps researchers find the most relevant literature faster, and gives information scientists a framework for measuring how knowledge spreads.
Table of Contents
- Classical bibliometric laws
- Lotka’s law: the pattern of author productivity
- Bradford’s law: the scattering of journal literature
- Zipf’s law: the frequency of words
- Stochastic models in bibliometrics
- What stochastic models do
- Key ideas behind the models
- Applications of bibliometric laws
- Optimising journal subscriptions and periodical management
- Guiding literature searches and research strategy
- Mapping fields and measuring productivity
- Limitations to keep in mind
Classical bibliometric laws
Bibliometrics applies mathematical and statistical methods to bibliographic data such as books, articles, authors, and citations. Three classical laws sit at the heart of the discipline: Lotka’s Law, Bradford’s Law, and Zipf’s Law. Each describes a different kind of skewed distribution, where a small portion of the population accounts for a large portion of the output. Together they explain how scientific productivity and information are distributed across people, sources, and words.
Lotka’s law: the pattern of author productivity
Formulated by the statistician Alfred J. Lotka in 1926, this law describes the frequency of publication by authors in any given field. Lotka observed that a small number of researchers produce a large share of the literature, while most authors publish only once or twice.
The law states that the number of authors who make n contributions is approximately 1/n² of those who make a single contribution. In practical terms, if 100 authors publish exactly one paper on a topic, then roughly 25 authors (100 divided by 2², or 4) will publish two papers, about 11 authors (100 divided by 3², or 9) will publish three papers, and so on. The result is a steep decline. A productive core of scholars dominates output while a long tail of occasional contributors makes up the majority of names. Researchers across disciplines, from educational sciences to artificial intelligence, continue to test whether their fields conform to this pattern.
Bradford’s law: the scattering of journal literature
Samuel C. Bradford, a British librarian, proposed his law of scattering in 1934. It describes how articles on a particular subject are distributed across journals. Bradford noticed that while a few journals publish many articles on a topic, the rest are scattered thinly across a very large number of less relevant periodicals.
To explain this, Bradford divided journals into zones of roughly equal article output. The first zone, called the nucleus, contains a small number of highly productive core journals. Each successive zone needs progressively more journals to contribute the same number of articles. If the nucleus holds a few core titles, the next zone might need several times more journals to match it, and the third zone many times more again. This multiplier creates the characteristic scattering. A study on conservation and preservation literature found that just five journals in the nucleus accounted for nearly 200 articles, illustrating how concentrated core publications can be.
Bradford’s Law is closely related to the well-known 80/20 rule, or Pareto distribution, where roughly 80% of the useful articles come from about 20% of the journals. Of the three classical laws, Bradford’s has found the widest practical application in libraries.
Zipf’s law: the frequency of words
The linguist George Kingsley Zipf published his law in 1949 after studying how often words appear in texts. He found that the frequency of any word is inversely proportional to its rank in a frequency table. The most common word appears about twice as often as the second most common, three times as often as the third, and so on.
Although Zipf’s Law began in linguistics, it applies neatly to bibliometrics. When analysts examine the keywords or index terms in a body of research, a few terms dominate while most appear rarely. This helps map the structure of a subject and reveals which concepts define a field. Studies with Zipf’s Law are more time-consuming than the other two because they involve counting thousands of words, which is why this work depends heavily on computers today.
These three laws are connected. Researchers have shown mathematical relationships linking Lotka, Bradford, and Zipf, suggesting they are different expressions of the same underlying tendency toward concentration in information systems.
Stochastic models in bibliometrics
The classical laws are powerful, but they share a limitation. They describe static snapshots based on empirical observation. They tell us how literature is distributed at one moment, but they do not explain how those distributions build up or change over time. Academic publishing is dynamic, and researchers need more flexible tools to capture that movement. This is where stochastic models come in.
What stochastic models do
A stochastic model is a mathematical model that incorporates randomness and uncertainty. Instead of assuming a fixed, predictable outcome, it treats bibliometric events such as the writing of a paper or the receipt of a citation as processes shaped by chance. These models predict how bibliometric phenomena evolve over time by accounting for the variable, non-deterministic nature of scholarly activity.
The shift from static laws to dynamic models gave rise to what is often called dynamic scientometrics, which builds sophisticated models of scientific growth, obsolescence, and citation processes. These models are valuable not only in theory but also for evaluation and prediction.
Key ideas behind the models
One important principle is cumulative advantage, sometimes called the “success breeds success” effect. An author who has already published is more likely to publish again, and a paper that has been cited often is more likely to attract further citations. This feedback loop helps explain why the distributions become so skewed and why Lotka’s steep curve appears in the first place.
Another major application is the study of obsolescence or ageing, which examines how the use and citation of literature declines as it gets older. Researchers have built stochastic models of the citation process using mixtures of Poisson processes to describe how citations accumulate while older work gradually ages. Similar models have been applied to library loans and circulation, predicting how often a book will be borrowed in future years and helping libraries decide when to move material to storage.
The work of Leo Egghe on Lotkaian informetrics extended these ideas into a general framework called Information Production Processes, treating authors producing papers, journals producing articles, and texts producing words as variations of one underlying process. This unifying view shows how the classical laws and the newer dynamic models belong to the same theoretical family.
Applications of bibliometric laws
These laws are not just academic curiosities. They shape practical decisions in libraries and information centres every day, especially where budgets are limited and choices must be deliberate.
Optimising journal subscriptions and periodical management
Bradford’s Law is the most directly useful for collection management. Because a small nucleus of core journals carries most of the important articles on any subject, a library can identify those core titles and prioritise them for subscription. This is especially valuable when funds are tight. Rather than spreading a budget thinly across hundreds of peripheral journals, a library can secure the core that serves the majority of its users’ needs and treat the rest through inter-library loan or document delivery.
The same logic guides cancellation decisions. When subscription costs rise, librarians can use Bradford analysis and usage data to identify low-yield journals in the outer zones that can be dropped with minimal impact on patrons. Bradford’s work even sets a lower bound on the number of journals a collection must hold to adequately cover a given subject, giving collection managers an evidence-based target instead of guesswork.
Guiding literature searches and research strategy
For researchers and students, the laws offer a smarter way to search. Bradford’s Law suggests starting a literature review with the core journals of a field, where the densest and most relevant material lives. Once those are exhausted, the searcher can move outward to mid-zone and peripheral journals for specialised or unusual perspectives. This strategy saves time and ensures the most influential studies are not missed.
Mapping fields and measuring productivity
Lotka’s Law and Zipf’s Law support research evaluation and subject analysis. By studying author productivity through Lotka’s Law, institutions and funding bodies can understand how research output is concentrated and identify the most active contributors in a discipline. Zipf’s Law, applied to keywords and index terms, helps map the conceptual structure of a field and supports indexing, automatic abstracting, and information retrieval system design. Together these tools feed into national science policy, helping countries analyse their science output and set research goals.
Limitations to keep in mind
No law is perfect. The classical laws were developed in an era of print journals and limited collaboration, and they do not always fit modern data cleanly. Empirical values often deviate from the ideal formulas, which is why researchers have proposed many refined versions. Relevance judgements can change the shape of a Bradford analysis, and a core identified through citations may differ from one identified through usage. These laws are best treated as reliable rules of thumb rather than exact predictions, and they work most powerfully when combined with current data and the dynamic insight of stochastic models.
What do you think? If a small core of journals truly carries most of the valuable research in your field, should libraries invest almost entirely in that core, or does the long tail of peripheral journals still deserve protection? And as research increasingly moves to open-access platforms and preprint servers, do you think these century-old laws will still hold, or will new digital patterns of publishing reshape them?
References
- https://en.wikipedia.org/wiki/Bradford's_law
- https://arxiv.org/pdf/2102.09182
- https://files.eric.ed.gov/fulltext/EJ1115017.pdf
- https://link.springer.com/article/10.1023/A:1005669410627
- https://www.tandfonline.com/doi/abs/10.1300/J105v08n01_06
- https://www.tandfonline.com/doi/full/10.1080/03615260801970774
- https://pmc.ncbi.nlm.nih.gov/articles/PMC12476466/
- https://www.researchgate.net/publication/227835133_The_Theoretical_Foundation_of_Zipf's_Law_and_Its_Application_to_the_Bibliographic_Database_Environment
- https://scispace.com/pdf/bibliometrics-as-a-research-field-a-course-on-theory-and-3x1alqfsgw.pdf
- https://link.springer.com/article/10.1023/A:1012751509975
- https://ebooks.inflibnet.ac.in/liscp10/chapter/classical-law-of-bibliometrics/
- https://www.researchgate.net/publication/280218558_Bradford's_Empirical_Law

Leave a Reply