When you search a database for articles on a topic, you quickly notice something strange: a handful of journals supply most of the relevant papers, while the rest are scattered thinly across dozens or even hundreds of titles. Cumulative rank-frequency models are the mathematical tools that turn this everyday observation into something measurable and predictable. They sit at the heart of bibliometrics, helping librarians and researchers map how scholarly literature concentrates and spreads. In this post, we will unpack fractional (cumulative) rank-frequency models, look closely at Cole’s and Leimkuhler’s contributions, examine the logarithmic equations developed by scholars like Egghe and Asai, and see how all of this plays out in real research.
Table of Contents
- What cumulative rank-frequency models actually describe
- Why the cumulative view is so useful
- Cole’s reference scattering model
- The constants that shape the curve
- Leimkuhler’s model and the logarithmic shape
- The Leimkuhler curve and its links to economics
- Logarithmic functions from Egghe and Asai
- Asai’s graph-oriented contribution
- Related logarithmic formulations
- Real-world applications across journals
- From information science to specialised fields
- Why this matters for libraries and researchers
What cumulative rank-frequency models actually describe
A rank-frequency distribution arranges items by size in decreasing order, then pairs each item with its rank. In bibliometrics, the “items” are usually journals, and their “size” is the number of articles or citations they contribute on a subject. A rank-size or rank-frequency distribution like this often follows a power law, where a few sources dominate and a long tail trails behind.
The word cumulative changes the picture. Instead of asking how many articles the journal at rank 5 produced, a cumulative model asks how many articles the top 5 journals produced together. As you move down the ranks, the running total keeps climbing until it reaches the full count. When you plot this cumulative total against the rank, you get a smooth rising curve rather than a jagged scatter of points.
These models are also called fractional models because they are frequently expressed in proportions. Rather than counting raw articles, you work with the fraction of total productivity captured by a given fraction of the most productive sources. This makes datasets of very different sizes directly comparable, which is exactly what you want when comparing, say, a small specialist field against a sprawling discipline like medicine.
Why the cumulative view is so useful
The cumulative curve directly exposes concentration. A curve that shoots up steeply at the start tells you a small core of journals carries most of the literature. This is the core-and-scatter pattern that Samuel Bradford first described, where articles on a subject distribute themselves unevenly across a small nucleus of journals and a much larger periphery. Cumulative models give that intuition a precise shape, letting you estimate how many “core” journals you need to subscribe to or index in order to capture a target percentage of the literature.
Cole’s reference scattering model
P. F. Cole was among the early researchers who experimented with Bradford-type data and tried to capture its shape mathematically. Cole studied reference scattering and worked with cumulative journal data, treating the slope of the cumulative curve as a meaningful quantity in its own right.
In Cole’s formulation, F(x) stands for the cumulative number of papers contained in the x most productive journals, and the slope of the curve was named the reference scattering coefficient. Cole argued that this coefficient might be a characteristic feature of the subject field itself, meaning different disciplines could have their own typical scattering values. Working with petroleum literature, Cole found a value of roughly 0.43 for this slope parameter. The takeaway is conceptual as much as numerical: the steepness of the cumulative curve is not random noise but a property that reflects how a field’s literature behaves.
The constants that shape the curve
Cumulative models are typically written with two constants, often labelled a and b. The constant b controls how quickly the curve climbs, which in turn reflects how steeply article frequency drops as journal rank increases. A larger b means the literature concentrates faster in the top journals. Because these constants can be estimated from observed data, the model lets you describe a whole dataset with just a couple of numbers, which is enormously convenient for comparison and prediction.
Leimkuhler’s model and the logarithmic shape
Ferdinand F. Leimkuhler, a professor of industrial engineering at Purdue University, analysed published Bradford data and derived what has become one of the most widely used cumulative formulations in the field. His cumulative function expresses R(r), the cumulative number of articles contributed by journals ranked 1 through r, as a logarithmic function of the rank, of the general form R(r) = a + b log r.
This logarithmic structure is the mathematical signature of Bradford-style scattering. Because the rank sits inside a logarithm, equal jumps in cumulative output require ever-larger jumps in the number of journals you add. In plain terms, the first few journals deliver a lot, but you need increasingly many additional journals to keep adding the same number of articles. That is precisely the diminishing-returns behaviour librarians observe when building collections.
The Leimkuhler curve and its links to economics
When you plot the cumulative proportion of total productivity against the cumulative proportion of sources, ranked from most to least productive, you produce what is now called the Leimkuhler curve. This curve is a direct cousin of the Lorenz curve used in economics to depict income inequality. The difference is mainly in the ordering of sources, but the underlying idea of measuring concentration is shared. Researchers have repeatedly connected the Leimkuhler curve to the Gini index, borrowing inequality measures from economics to quantify how unevenly citations or articles are distributed across journals.
Logarithmic functions from Egghe and Asai
The logarithmic family of cumulative models was substantially deepened by Leo Egghe, a leading figure in informetrics. In a landmark 1990 paper, Egghe developed the theory of Bradford’s law and showed how to calculate Leimkuhler’s law from it, deriving a theoretical formula for the Bradford multiplier and for the number of items produced by the most productive source in each Bradford group. Crucially, he extended the approach to incomplete bibliographies, such as citation tables truncated before the so-called Groos droop, and showed how properties of the full, unknown bibliography could be estimated from a partial one.
Egghe’s version of the Leimkuhler model has become a practical workhorse. Applying it begins with choosing the number of Bradford groups, p, which can be selected freely, and then determining the Bradford multiplier (k) and the number of journals in the nucleus. As a study of artificial intelligence research notes, the calculation uses Euler’s constant (approximately 1.781) within the formula for the multiplier, and the choice of p depends on the size of the bibliography being analysed. This procedure lets researchers estimate the zone-wise distribution of journals using clean exponential and logarithmic relationships rather than rough verbal counting.
Asai’s graph-oriented contribution
I. Asai contributed a general formulation of Bradford’s distribution through what is described as a graph-oriented approach. This work, referenced in Egghe’s own bibliography, sits within the same effort to give Bradford-style scattering a rigorous mathematical and logarithmic underpinning. Together, the contributions of Egghe and Asai moved the field from describing scattering verbally toward modelling it with equations that can be fitted, tested, and extended.
Related logarithmic formulations
Leimkuhler was not alone in proposing a logarithmic cumulative form. B. C. Brookes derived a closely related expression for journal productivity, and the yield formula developed by Haspers expressed cumulative yield as a logarithm of rank plus a constant, requiring three parameters instead of two for a closer fit. Eugene Garfield’s overview of these statistical patterns catalogues several such formulations side by side, including Leimkuhler’s and Brookes’s, underscoring that the logarithmic shape is a recurring theme rather than a single isolated equation.
Real-world applications across journals
The real value of these models shows up when they meet actual data. Consider research output during the COVID-19 pandemic. A bibliometric study indexed in Web of Science examined a dataset of nearly 21,000 articles spread across roughly 3,300 journals. The researchers applied Bradford’s zoning to split journals into three zones, each contributing about a third of the articles, and then validated the distribution using the Leimkuhler logarithmic model to construct a Bradford-Leimkuhler curve. Egghe’s theoretical formulation was used to interpret the characteristics of the three zones. This combination identified the core journals driving pandemic research and showed how the literature concentrated.
From information science to specialised fields
The same toolkit has been applied widely. A study of information science literature drawn from Scopus applied both the verbal formulation and the Egghe-expanded Leimkuhler model, finding that the simple verbal form produced large errors while the Leimkuhler model gave a far better fit after computing the multiplier and nucleus size. In artificial intelligence research, scholars used the Leimkuhler equations to estimate the average number of articles per journal and map the growth of the field. These cumulative models have also been pressed into service for fluid mechanics, production engineering, and public health literature, consistently helping analysts pinpoint a small set of high-yield journals.
Why this matters for libraries and researchers
For a librarian managing a limited budget, these models answer a concrete question: which journals must we hold to cover most of the literature our users need? For a researcher entering a new field, the cumulative curve quickly reveals the core venues worth reading and targeting. Recent work continues to refine the approach, with a 2024 study introducing mixture distributions to improve the fit of Leimkuhler curves where the standard form falls short, showing that this is an active rather than a settled area. The fractional, proportion-based nature of these models is what makes them portable across fields and dataset sizes.
What do you think? If a single logarithmic curve can describe how articles scatter across journals in fields as different as petroleum engineering and pandemic medicine, what does that say about the way scholarly knowledge organises itself? And as databases grow ever larger, do you think classic models like Leimkuhler’s will keep holding up, or will newer mixture-based approaches eventually replace them?
References
- https://en.wikipedia.org/wiki/Rank%E2%80%93size_distribution
- https://link.springer.com/article/10.1007/BF00353144
- https://ebooks.inflibnet.ac.in/liscp10/chapter/bradford-distributions-an-overview/
- https://www.sciencedirect.com/science/article/abs/pii/S1751157708000138
- https://link.springer.com/article/10.1007/s13171-024-00363-9
- https://asistdl.onlinelibrary.wiley.com/doi/abs/10.1002/(SICI)1097-4571(199010)41:7%3C469::AID-ASI1%3E3.0.CO;2-P
- https://jscires.org/full-text/6556/
- https://asistdl.onlinelibrary.wiley.com/doi/abs/10.1002/asi.4630270503
- https://garfield.library.upenn.edu/essays/v4p476y1979-80.pdf
- https://scientifictemper.com/index.php/tst/article/view/2366
- https://arxiv.org/abs/2401.07052

Leave a Reply