When we line up journals, authors, or keywords from the most productive to the least and then look at how many items each one contributes, a striking pattern appears. A handful of sources dominate, and a long tail of smaller contributors stretches out behind them. Rank-frequency non-cumulative models are the mathematical tools that describe this pattern directly. Instead of adding up totals as we move down the list, these models tell us exactly how much the source at each individual rank contributes. This makes them powerful for both describing existing data and predicting how a collection of publications is likely to be distributed.
Table of Contents
- What “rank-frequency” and “non-cumulative” actually mean
- The foundation: Zipf’s law as a non-cumulative model
- Why the basic model needs refinement
- Non-cumulative rank models: Fairthorne and the hyperbolic family
- Mandelbrot’s extension
- Chen and Leimkuhler: unifying the laws
- Use in informetrics: predicting publication distribution
- Fitting and testing the model
- Example in action: analysing journal contributions
- A note of caution
- Why these models matter for information science
What “rank-frequency” and “non-cumulative” actually mean
The rank-frequency approach starts with a simple step: arrange all the sources in a dataset in decreasing order of their output. The source that produces the most gets rank 1, the next gets rank 2, and so on. Once everything is ranked, we record the frequency, that is, the number of items each source produces.
The term non-cumulative is the key distinction here. In a cumulative model, the value at any rank includes everything that came before it, so you read off running totals such as “the top five journals together published 200 articles.” In a non-cumulative model, each rank stands on its own. You read the frequency of that single source only, for example “the journal at rank 5 published 12 articles.” Non-cumulative models therefore preserve the individual data points rather than rolling them into a growing sum.
This difference matters because non-cumulative distributions keep the raw shape of the data visible. As researchers analysing ranking techniques have pointed out, rank distributions retain empirical information that cumulative forms tend to smooth away, which gives them a theoretical advantage when you want to study the underlying structure rather than just the broad totals.
The foundation: Zipf’s law as a non-cumulative model
The most famous non-cumulative rank-frequency model is Zipf’s law. Originally developed by the linguist George Kingsley Zipf to describe word frequencies in text, it states that the frequency of an item is inversely proportional to its rank. When values are sorted in decreasing order, the n-th entry tends to be roughly inversely proportional to n.
In practical terms, this means the second-ranked source produces about half as much as the first, the third produces about a third as much, and so on. The relationship is usually written as:
f × r = C
Here f is the frequency of a source, r is its rank, and C is a constant for the dataset. Because the product of frequency and rank stays roughly constant, you can estimate the output of a source at any rank once you know the constant. This is what makes Zipf’s law a predictive tool and not just a descriptive one.
Zipf’s law is recognised as one of the three classical empirical laws of bibliometrics, alongside Lotka’s law of author productivity and Bradford’s law of scattering. While Lotka and Bradford are often expressed in size-frequency or cumulative terms, Zipf’s law in its basic form is non-cumulative because it speaks to the frequency of each ranked item individually.
Why the basic model needs refinement
The simple Zipf relationship works well in the middle of a distribution but tends to break down at the extremes. The very top-ranked items often contribute more than the formula predicts, and the long tail of low-frequency items behaves differently too. Modern analyses of large text corpora confirm this: when huge datasets are plotted on a log-log scale, the tail bends downward instead of following the straight line that pure Zipf’s law would produce. This limitation is exactly what prompted later researchers to develop more flexible non-cumulative models.
Non-cumulative rank models: Fairthorne and the hyperbolic family
One of the most influential contributions to this area came from R. A. Fairthorne in 1969. In his paper on empirical hyperbolic distributions for bibliometric description and prediction, Fairthorne brought together the work of Bradford, Zipf, and Mandelbrot under a single umbrella. He argued that the various bibliometric laws are all manifestations of the same underlying hyperbolic pattern.
A hyperbolic relationship is one where the frequency multiplied by the rank, raised to some power, stays constant. The basic Zipf form is the simplest version of this. Fairthorne’s key insight was that these distributions are not isolated curiosities but a connected family, and that a hyperbolic model can be used to both describe a body of literature and predict its future shape. The very title of his work, pairing “description” with “prediction,” signals that these models are meant to forecast distributions, not merely catalogue them.
Mandelbrot’s extension
The mathematician Benoit Mandelbrot refined Zipf’s law to handle the deviations seen at the top of distributions. The resulting Zipf-Mandelbrot law adds an extra constant inside the formula so that the curve can bend to fit real data more closely. It is often written in the form where frequency is proportional to a constant divided by (rank + a shift parameter), raised to a power.
The shift parameter is what gives this model its flexibility. When that parameter is zero, the formula collapses back into ordinary Zipf’s law, which means Zipf’s law is simply a special case of Mandelbrot’s more general model. This extension is particularly useful when the highest-ranked sources do not behave the way pure Zipf’s law expects, which is common in real publication data.
Chen and Leimkuhler: unifying the laws
Another important step came from Y. S. Chen and F. F. Leimkuhler. In 1986 they published a study establishing a relationship between Lotka’s law, Bradford’s law, and Zipf’s law, showing that these three pillars of bibliometrics are mathematically connected rather than separate phenomena.
This unification is significant for anyone working with non-cumulative models. It means that if you understand the rank-frequency behaviour described by Zipf, you also hold a key to the productivity patterns described by Lotka and the scattering patterns described by Bradford. Chen and Leimkuhler demonstrated that these laws produce similar probability distributions, with the differences lying mainly in the kind of data being analysed, whether that data is words in a text, authors and their papers, or articles scattered across journals.
For the analyst, this consistency is reassuring. It suggests that the long-tail pattern we see in scholarly output is not an accident of one particular dataset but a structural feature of how knowledge production works.
Use in informetrics: predicting publication distribution
The practical value of non-cumulative rank-frequency models lies in prediction. Once you have ranked your sources and confirmed that the data follows a Zipfian or hyperbolic pattern, you can estimate values you have not directly measured.
Consider a few concrete uses. A librarian building a journal collection can predict how many articles on a topic the lower-ranked journals are likely to publish, helping decide where subscription money is best spent. A research administrator can forecast how output will concentrate among a small group of prolific authors. A subject specialist can identify the “core” sources that carry most of the literature and separate them from the marginal ones.
The prediction works because the constant in the Zipf formula is stable for a given collection. If you measure the frequency of the top few sources, you can calculate the constant and then project the expected frequency at any other rank. This is why these models are valued for assessing the growth and structure of a subject by arranging publication data with statistical techniques.
Fitting and testing the model
Before trusting a prediction, analysts check how well the data actually fits the model. A common method is to plot frequency against rank on a log-log graph. If the data conforms to Zipf’s law, the points fall close to a straight line whose slope corresponds to the exponent of the distribution. A clean straight line suggests the pattern is genuine and not the product of chance; a bending or broken line suggests a Mandelbrot-style correction may fit better.
Example in action: analysing journal contributions
Suppose a student is studying which journals publish research on a particular subject and gathers the following counts of articles. The journals have already been arranged from the most productive to the least.
Rank 1: 60 articles
Rank 2: 31 articles
Rank 3: 19 articles
Rank 4: 16 articles
Rank 5: 12 articles
To test whether this follows Zipf’s law, we multiply each frequency by its rank and look at the products:
Rank 1: 60 × 1 = 60
Rank 2: 31 × 2 = 62
Rank 3: 19 × 3 = 57
Rank 4: 16 × 4 = 64
Rank 5: 12 × 5 = 60
The products cluster tightly around 60, which tells us the constant C for this collection is roughly 60. This is the non-cumulative pattern in action. Each rank is examined on its own, and the frequency-times-rank product stays nearly constant across the list.
Now we can predict. If we want to estimate how many articles the journal at rank 6 is likely to publish, we divide the constant by the rank: 60 ÷ 6 = 10 articles. For rank 10, the estimate would be 60 ÷ 10 = 6 articles. Without having to count the lower-ranked journals individually, we have a reasonable forecast of their output.
This is exactly the kind of forecasting that makes the model useful in collection development and research evaluation. Notice how different this is from a cumulative approach, which would instead report running totals such as “the top three journals published 110 articles between them.” The non-cumulative view keeps each journal’s contribution distinct, which is what lets us model the rank-by-rank pattern and extend it.
A note of caution
Real data is rarely as tidy as a textbook example. The products will not be perfectly constant, and the top-ranked and bottom-ranked sources are the most likely to drift away from the pattern. This is precisely why the Zipf-Mandelbrot extension exists, and why goodness-of-fit testing matters. Pure hyperbolic decay produces the classic straight-line Zipf plot, but other decay rates create noticeable differences at small and large ranks. Treating the prediction as an estimate rather than an exact value is the sensible approach.
Why these models matter for information science
Non-cumulative rank-frequency models give information professionals a quantitative grip on something that would otherwise feel vague: the uneven way knowledge concentrates in a few productive sources. They turn a messy list of journals or authors into a predictable curve. From that curve flow practical decisions about which journals to subscribe to, which sources form the core of a field, and how a body of literature is likely to grow.
The lineage from Zipf through Fairthorne, Mandelbrot, and Chen and Leimkuhler also shows how the discipline matured. What began as a description of word frequencies became a unified framework linking the major bibliometric laws, capable of both explaining the past distribution of publications and forecasting the future one.
What do you think? If a small number of journals consistently dominate the literature in your field, should libraries concentrate their limited budgets on those core titles, or deliberately support the long tail of smaller journals? And when a real dataset only loosely follows Zipf’s law, how much should we trust the predictions we draw from it?
References
- https://www.sciencedirect.com/science/article/abs/pii/0306457384900384
- https://en.wikipedia.org/wiki/Zipf's_law
- https://files.eric.ed.gov/fulltext/EJ1115017.pdf
- https://www.sciencedirect.com/science/article/abs/pii/S1751157707000296
- https://www.emerald.com/jd/article/25/4/319/206708/Empirical-Hyperbolic-Distributions-Bradford-Zipf
- https://link.springer.com/chapter/10.1007/978-3-319-41631-1_4
- https://link.springer.com/article/10.1007/s11192-020-03476-8
- https://tefkos.comminfo.rutgers.edu/Courses/e530/Readings/Jayroe%20Bibliometrics%20for%20Dummies%202008.pdf
- https://www.sciencedirect.com/science/article/pii/S2694610625000104
- https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5628998/

Leave a Reply