Every year, millions of research papers pour into journals worldwide, and Indian researchers contribute a steadily growing share of them. But raw numbers alone tell us little. How do we measure whether a field is flourishing? Who are the most productive authors? How fast does a new idea travel from one laboratory to the next? These are not questions of intuition. They are questions of measurement, and an entire branch of library and information science exists to answer them using statistics. This is the world of informetrics, and when applied to science itself, it reveals striking patterns hidden inside the chaos of scholarly publishing.
Table of Contents
- Scientometrics and informetrics: two lenses on the same data
- Why this matters for measuring science
- Lotka’s law and its application to scientific productivity
- What the law predicts in plain terms
- Does the law always hold?
- Why prolific authors cluster at the top
- Diffusion of information in science
- Why the Poisson distribution fits the spread of ideas
- From Poisson processes to epidemic models
- What diffusion patterns reveal
- Bringing the three ideas together
Scientometrics and informetrics: two lenses on the same data
The terms informetrics and scientometrics are often used together, and they overlap heavily, but they are not identical. Informetrics is the broad study of the quantitative aspects of information in any form, drawing on mathematics and statistics to examine how information is produced, recorded, and used. It sits at the top of a family of related fields and grew out of two older traditions: bibliometrics and scientometrics.
Scientometrics is narrower and more focused. It concentrates specifically on scientific publications and is sometimes called the “science of science,” with the goal of guiding research policy and decision-making. Bibliometrics, the oldest of the three, applies statistical methods to books, articles, and other recorded communication. The term itself was introduced in 1969 to describe the application of mathematical and statistical methods to documents.
So how do they relate? Researchers who have compared the three fields find that they differ mainly in their subject background but share the same theories, methods, technologies, and applications. Bibliometrics is rooted in library science, scientometrics in the sociology and history of science, and informetrics acts as the umbrella that contains both. Because of this overlap, the same statistical tool, a citation count or an authorship distribution, can be called bibliometric, scientometric, or informetric depending on who is using it and why.
Why this matters for measuring science
The practical importance of these fields lies in measurement. Universities use these methods to evaluate departments. Funding agencies use them to decide where money should go. Librarians use them to manage collections and identify the core journals of a discipline. In India, this work has a well-established home. The CSIR-National Institute of Science, Technology and Development Studies (NISTADS) has long produced bibliometric assessments of the country’s research output, and an analysis of papers published by Indian scholars showed a steep rise in scientometric studies after 1995 compared with earlier decades, reflecting growing institutional interest in measuring science systematically.
Lotka’s law and its application to scientific productivity
One of the oldest and most famous patterns in this field describes a simple but surprising truth: most scientific work is done by a small number of people. This is captured by Lotka’s law, formulated by Alfred J. Lotka in 1926.
The law states that the number of authors producing a given number of papers is inversely proportional to the square of that number of papers. In its standard form, the relationship is written as yx = C / xn, where yx is the number of authors who have published x papers, while C and n are constants estimated for a particular dataset. Lotka originally found values close to n = 2, which is why the law is popularly called the inverse square law of scientific productivity.
What the law predicts in plain terms
If 100 authors each publish exactly one paper, then about 25 authors (100 divided by 2², or 4) will publish two papers, about 11 authors (100 divided by 3², or 9) will publish three, and so on. The result is a steep curve: a handful of highly prolific authors at the top, and a long tail of researchers who publish only once or twice. It expresses the share of authors publishing a certain number of articles as a fixed ratio to those publishing just one. This concentration of productivity is one of the most consistently observed structures in scholarly literature.
Does the law always hold?
Here is where measurement becomes interesting, because Lotka’s law does not fit every field perfectly. Since 1926, researchers have tested it across dozens of disciplines with mixed results. Studies have found that an exponent of 3.5 fits information science better, while library science and computer science required still different values, meaning the strict inverse square form does not universally apply. The law works best as a generalised inverse power law, with the exponent adjusted to suit the data, rather than a rigid rule fixed at the value of two.
Indian research confirms this caution. A study of author productivity in library and information science open access journals found that the observed values differed significantly from those predicted by Lotka’s law under chi-square and Kolmogorov-Smirnov goodness-of-fit tests, so that particular literature did not adhere to the inverse square pattern. Similarly, studies based on the Council of Scientific and Industrial Research (CSIR) productivity data have repeatedly failed to follow the classic formulation. The lesson is that Lotka’s law is a powerful starting hypothesis, not an iron law. Testing whether it holds, and finding the exponent that fits, is itself a valuable scientometric exercise.
Why prolific authors cluster at the top
Interestingly, Lotka’s distribution has a mirror image. When authors are ranked by their number of publications, the most prolific author at rank one tends to dominate, and the output falls off rapidly down the ranking. This ranking form resembles patterns seen elsewhere in nature and society, such as the distribution of city sizes or word frequencies. The recurrence of this shape suggests that scientific productivity is governed by a kind of cumulative advantage: success tends to breed further success, as established researchers attract more resources, collaborators, and opportunities to publish.
Diffusion of information in science
Productivity tells us who creates knowledge. The next question is how that knowledge spreads. The diffusion of information within scientific communities is the study of how ideas, methods, and findings travel from their origin through citations, collaborations, and communication networks until they reach the wider field.
This spread is rarely instant or uniform. A new finding might sit unnoticed for years, then suddenly be cited heavily once the field is ready for it. To capture this irregular, event-based pattern, informetricians turn to probability models, and one of the most useful is the Poisson distribution.
Why the Poisson distribution fits the spread of ideas
The Poisson distribution describes the probability of a given number of independent events happening within a fixed interval, when those events occur at some average rate but at unpredictable moments. Citations to a paper fit this description well. Each citation is, in a sense, a discrete event arriving over time, and the rate at which they arrive changes as a paper ages.
Researchers have built citation models directly on this idea. One influential approach uses mixtures of non-homogeneous Poisson processes to describe the citation process while accounting for the ageing and obsolescence of literature. In such models, the “non-homogeneous” part is crucial: the rate of citation is not constant but rises and then falls as a paper grows older and is eventually superseded. These models can explain why some papers are never cited at all and why others have a long delay before their first citation.
From Poisson processes to epidemic models
The spread of ideas resembles the spread of a contagion, and this analogy has shaped a whole class of diffusion models. Just as an infection passes from person to person, a scientific concept passes from one researcher to another through reading and citing. Researchers have asked directly whether epidemic models can describe the diffusion of topics across disciplines, treating the adoption of a new idea much like the transmission of a disease through a susceptible population. Related work has used stochastic models to describe the evolution and ageing of scientific disciplines over time.
Network-based approaches take this further. By combining citation analysis with social network analysis, scholars trace knowledge as it moves through the connections between papers and authors. One such study introduced a citation-based directed network model with a time dimension to capture how scientific ideas spread from a network point of view. These maps reveal the pathways along which influence flows, identifying the key papers that act as bridges between research communities.
What diffusion patterns reveal
Studying diffusion is not merely academic. It shows policymakers which fields are emerging and which are stagnating. It helps identify the institutions and journals that act as hubs, accelerating the spread of new work. For a country building its research capacity, understanding how knowledge diffuses can inform where to invest in collaboration, open access, and communication infrastructure. When information moves freely and quickly through a scientific community, the whole system becomes more productive, which closes the loop back to the productivity patterns Lotka described a century ago.
Bringing the three ideas together
These three strands form a connected picture of how science works as a measurable system. Scientometrics and informetrics provide the toolkit and vocabulary for quantitative analysis. Lotka’s law describes the structure of who produces knowledge, revealing the heavy concentration of output among a few prolific authors. Diffusion models, built on Poisson processes and epidemic analogies, describe how that knowledge then travels through the community over time.
Together they transform scholarly communication from something mysterious into something that can be charted, predicted, and improved. For students of library and information science, mastering these phenomena means being able to read the hidden statistical signature of any research field, whether it is cancer research, solar cell development, or library science itself.
What do you think? If Lotka’s law shows that a small group of researchers produces most of the published work, should evaluation systems reward sheer volume of output, or find better ways to recognise quality and influence? And if ideas diffuse through science much like an epidemic, what could institutions do to help good research spread faster and reach the people who need it?
References
- https://en.wikipedia.org/wiki/Informetrics
- https://taylorandfrancis.com/knowledge/Engineering_and_technology/Computer_science/Informetrics/
- https://www.researchgate.net/publication/318940072_Are_Scientometrics_Informetrics_and_Bibliometrics_different
- https://www.researchgate.net/publication/317026767_Bibliometrics_and_scientometrics_in_india_An_overview_of_studies_during_1995-2014_Part_I_Indian_publication_output_and_its_citation_impact
- https://www.sciencedirect.com/science/article/abs/pii/S0306457398000272
- https://digitalcommons.usf.edu/si_facpub/135/
- https://arxiv.org/pdf/2102.09182
- https://www.emerald.com/insight/content/doi/10.1108/DLP-10-2020-0103/full/html
- https://link.springer.com/article/10.1023/A:1012751509975
- https://arxiv.org/pdf/1201.0676
- https://link.springer.com/article/10.1007/s11192-011-0554-z

Leave a Reply