When you sort journals, authors, or keywords from the most productive to the least and then start adding them up one by one, you are working with a cumulative rank-frequency distribution. The “non-fractional” version of this approach is one of the oldest and most practical tools in bibliometrics, used to answer a deceptively simple question: if I know an item’s rank, how many publications should I expect it to account for? This post unpacks how these models work, how their equations are built, and where they actually get used in research analysis.
Table of Contents
- What “rank-frequency cumulative” actually means
- Non-fractional distributions: the Bradford family
- Wilkinson’s model
- Haspers’s model
- The unifying picture
- Mathematical formulations: how the equations differ
- The logarithmic backbone
- Where the models diverge
- The connection to Zipf and Mandelbrot
- Applications: predicting publications from rank
- Estimating how many journals cover a subject
- Library collection and budgeting decisions
- Tracking the growth and spread of a research field
- Beyond journals
- Why non-fractional models still matter
What “rank-frequency cumulative” actually means
Start with the basics. In a rank-frequency distribution, you arrange sources in decreasing order of output. The most productive journal gets rank 1, the next gets rank 2, and so on. This idea comes from linguistics, where Zipf’s law describes the number of items in a source when sources are ranked in decreasing order of frequency. Applied to scholarly literature, the most prolific journal in a field publishes far more papers than the one ranked second, which in turn beats the third, and the gap keeps narrowing down the list.
The word cumulative changes the picture. Instead of asking “how many papers does the journal at rank 10 publish?”, a cumulative model asks “how many papers do the top 10 journals publish together?” You are summing output as you move down the ranked list. This running total is what makes the model so useful for planning and prediction, because real decisions are usually about groups of sources, not single ones.
Finally, non-fractional tells you how the counting is done. In non-fractional counting, every publication is counted as a whole unit and assigned fully to its source. A paper co-authored across three journals’ editorial scope, or credited to multiple authors, is not split into thirds. This contrasts with fractional models, where credit is divided. Fractional counting was developed to handle authorship division, but the non-fractional approach keeps the arithmetic simpler and is the natural fit when you are counting how many articles a journal contains.
Non-fractional distributions: the Bradford family
The cumulative non-fractional approach is most closely associated with Bradford’s law of scattering. Bradford observed that if you sort journals by how many articles they publish on a given subject, you can divide them into zones each containing roughly the same number of articles, but with the number of journals per zone growing in a fixed ratio. Mathematically, Bradford-type distributions are cumulative versions of Lotka’s distribution, which is exactly why they belong in a discussion of cumulative rank-frequency models.
The trouble is that Bradford’s original formulation was ambiguous. It could be expressed verbally or graphically, and the two did not always agree, especially for the most productive “core” journals. This gap is what motivated a series of refinements, and it is here that the models named in this topic enter the story.
Wilkinson’s model
E. A. Wilkinson tackled the inconsistency directly in a paper whose title says it all: “The Ambiguity of Bradford’s Law,” published in 1972. Wilkinson examined the mismatch between the verbal statement of Bradford’s law and its graphical form, and proposed a way to resolve it so that the cumulative distribution of articles against the logarithm of journal rank behaves consistently. The practical value of this is that it sharpened how the linear, log-based portion of the Bradford curve should be interpreted, particularly for the long tail of less productive journals.
Haspers’s model
J. H. Haspers approached the same problem from a different angle in “The Yield Formula and Bradford’s Law,” published in 1976. The yield formula focuses on the relationship between the cumulative number of articles (the “yield”) and the cumulative number of journals contributing them. Rather than treating the Bradford curve as a single rigid line, Haspers’s formulation gives a way to estimate the yield you can expect from adding more sources, which is directly a cumulative prediction. Both Wilkinson and Haspers, then, were trying to make the same underlying distribution behave predictably across its whole range.
The unifying picture
These were not isolated efforts. By 1981, a careful review of eight previously published mathematical models of Bradford’s distribution showed they could be reduced to a single general formulation of the form y = a log(x + c) + b, where y is the cumulative ratio of articles and x relates to journal rank. Wilkinson’s and Haspers’s models both fall inside this family, differing in how they set the constants and what part of the curve they emphasize. The broader lesson, demonstrated mathematically by Egghe, is that the distributions of Lotka, Bradford, and Zipf are all equivalent expressions of the same phenomenon viewed from different angles.
Mathematical formulations: how the equations differ
The reason there are several models for one phenomenon is that a cumulative rank-frequency curve has different behaviour in different regions, and no single simple equation captures all of it perfectly. Understanding the equations helps you see why.
The logarithmic backbone
The signature feature of non-fractional Bradford-type models is the logarithm. When you plot cumulative articles against journal rank, a raw plot is a steep curve that is hard to read. Take the logarithm of the rank, and a large stretch of the curve straightens into a line. This is why logarithmic functions sit at the heart of these models: they linearize data that follows a power-law pattern, making it far easier to fit a straight line and read off predictions. The general y = a log(x + c) + b form is exactly this idea, with the constant c shifting the curve so that the steep early part involving the core journals also fits the line.
Where the models diverge
The differences between formulations come down to how each handles the two ends of the distribution. The earliest plots tended to bend away from the straight line near the core (the “Groos droop” effect that many later studies wrestled with). Wilkinson’s resolution of the ambiguity addressed how to read this consistently, while Haspers’s yield formula reframed the relationship so that cumulative yield could be projected. The constants a, b, and c are not just curve-fitting noise; a controls the steepness, b sets the starting level, and c absorbs the curvature near rank 1. Choosing different values, or a slightly different functional form, is what separates one named model from another.
The connection to Zipf and Mandelbrot
The same logarithmic structure links these models to the wider power-law literature. The Zipf-Mandelbrot law, proposed by Benoit Mandelbrot, generalizes Zipf’s law by adding a shift parameter to the rank, which is conceptually the same role played by the constant c in the Bradford formulation. In other words, the non-fractional cumulative models are part of a family of empirical hyperbolic distributions described as Bradford-Zipf-Mandelbrot for bibliometric description and prediction. Recognising this kinship means a technique learned for journals transfers cleanly to word frequencies, citation counts, or city sizes.
Applications: predicting publications from rank
The real payoff of these models is prediction. Once you have fitted the equation to a sample of your data, you can extend it to estimate quantities you have not directly counted.
Estimating how many journals cover a subject
The classic use is completing a bibliography. If you have surveyed the most productive journals in a field and counted their relevant articles, the cumulative model lets you estimate how many additional, less productive journals you would need to scan to capture a target percentage of the total literature. Leimkuhler turned the Bradford distribution into a workable formula precisely for this, and the Leimkuhler model has been used to produce quantitative estimates of research output, including a recent study of artificial intelligence publishing indexed in Web of Science from 2011 to 2020.
Library collection and budgeting decisions
For a librarian, the cumulative curve answers a budget question directly. If the top zone of journals delivers, say, a third of all relevant articles, and subscribing to the next zone delivers another third but requires three times as many subscriptions, the model quantifies the diminishing return. An early demonstration of this used the Bradford-Zipf distribution to estimate efficiency values for a journal circulation system. The principle scales from a single library to national resource-sharing decisions.
Tracking the growth and spread of a research field
Because the parameters of the fitted curve change as a field evolves, the models double as a diagnostic tool. A flattening curve means output is spreading across more sources; a steepening one means it is concentrating in a core. In the AI study mentioned above, comparing productivity distributions across years showed that the field grew not only in volume of publications but also in the breadth of subject areas it reached. The same method can map any emerging discipline.
Beyond journals
Nothing in the mathematics restricts these models to journals. Because the distribution form is shared across phenomena, the same cumulative rank-frequency fit applies to authors ranked by output, keywords ranked by frequency, or citations ranked by count. Researchers have used the framework to model the distributions of citations and impact factors, treating these as among the most relevant problems in current informetric research. Learn the technique once and it generalizes widely.
Why non-fractional models still matter
It is fair to ask why a cluster of models from the 1970s deserves attention now. The answer is that they are transparent and robust. Non-fractional counting avoids the assumptions and disputes that come with splitting credit, and the logarithmic form is easy to fit with a simple regression. For a great many practical questions about journals and bibliographies, this simplicity is a feature, not a limitation. The more elaborate fractional and continuous models matter when you need precision about authorship or fine-grained citation analysis, but for predicting how much literature a set of ranked sources will yield, the non-fractional cumulative approach remains a dependable first tool.
What do you think? If you were building a reading list for an unfamiliar research topic, would you trust a cumulative rank-frequency estimate to tell you how many journals to scan? And where do you think whole-counting (non-fractional) gives a clearer answer than splitting credit across sources?
References
- https://arxiv.org/pdf/1011.1533
- https://link.springer.com/chapter/10.1007/978-3-319-41631-1_4
- https://asistdl.onlinelibrary.wiley.com/doi/abs/10.1002/asi.4630320206
- https://en.wikipedia.org/wiki/Zipf's_law
- https://link.springer.com/article/10.1007/BF02016844
- https://jscires.org/full-text/6556/
- https://www.sciencedirect.com/science/article/abs/pii/S1751157711000824

Leave a Reply