Some journals publish hundreds of papers on a topic. Many publish only one or two. When you line up every source from the most productive to the least, a striking pattern emerges that holds across word counts, city populations, website visits, and scientific output. This pattern is captured by the rank-frequency approach, one of the foundational ideas in informetrics and scientometrics. It lets analysts move from a messy pile of numbers to a clear, predictable structure that explains why a few contributors dominate while a long tail does the routine work.
Table of Contents
What the rank-frequency approach actually means
The rank-frequency approach is a way of organising data by listing every source in decreasing order of how much it contributes, then studying the relationship between a source’s position on that list and the size of its contribution. A “source” can be a journal, an author, a keyword, or a website. The “contribution” is whatever you are counting, such as the number of papers, citations, or word occurrences.
The mechanics are simple. You rank the most productive source as 1, the next as 2, and so on. Then you record the frequency, which is the actual count attached to each rank. The result is a rank-frequency distribution that describes the number of items in a source when sources are ranked in decreasing order of frequency. Plotted out, this distribution almost always slopes steeply downward, because the top-ranked source contributes far more than anything below it.
This is different from a closely related idea called the size-frequency distribution, and the distinction matters. A size-frequency function counts how many sources have a given number of items, for example how many authors published exactly two papers. A rank-frequency function instead counts how many items belong to the source sitting at a particular rank. In essence, the roles of sources and items are interchanged between the two functions, which is why they are described as dual descriptions of the same two-dimensional reality. Lotka’s law of author productivity is a size-frequency law, while Zipf’s and Pareto’s laws are rank-frequency laws.
Where the idea came from
The rank-frequency approach grew out of linguistics. George Kingsley Zipf proposed that the frequency of a word is inversely proportional to its rank, after studying how often different words appear in English texts. The most common word, “the”, appears roughly twice as often as the second most common word, three times as often as the third, and so on. Beyond about rank 1,000 the relationship breaks down, but across the useful middle range it is remarkably stable.
Zipf’s own verification was textual. His rank-frequency word distribution was found to approximate the simple equation of an equilateral hyperbola, written as r × f = C, where r is rank, f is frequency, and C is a constant. In plain terms, if you multiply a word’s rank by its frequency, you keep getting roughly the same number. That constant product is the signature of a rank-frequency relationship and the quickest test of whether Zipf’s law fits a dataset.
Why rank matters in bibliometrics
Library and information science adopted the rank-frequency idea because scholarly output behaves much like word frequencies. A handful of journals, authors, and institutions carry the bulk of the published research in any field, while a large number of others contribute marginally. Ranking makes this concentration visible and measurable rather than just intuitive.
This is why the three classical laws are taught together. The three most common worldwide are Lotka’s inverse square law of author productivity, Bradford’s law of scattering, and Zipf’s law of word frequencies, and each is a different lens on the same uneven distribution. Lotka looks at how many authors produce how many papers. Bradford ranks journals to find the core set that publishes most of the literature on a subject. Zipf ranks words by frequency. All three describe the same underlying tendency for output to cluster at the top.
Finding the core contributors
The most practical use of ranking is identifying who or what sits at the top. When journals are ranked by the number of relevant articles they publish, a small core stands out clearly. In one study of educational sciences, the six most productive journals accounted for nearly half of all the articles in the field. For a librarian deciding which journals to subscribe to, or a researcher choosing where to submit, that core list is far more useful than an undifferentiated catalogue of thousands of titles.
Ranking also exposes effects that shape scientific careers. Productivity tends to split researchers into a small, highly prolific group and a large group with limited output, a pattern linked to the Matthew effect, where highly ranked researchers attract disproportionate attention. Rank-frequency data is what makes these social dynamics quantifiable instead of anecdotal.
The power-law backbone
Mathematically, the rank-frequency relationship is a power law. The frequency of an item of rank r is proportional to r raised to a negative exponent, written f(r) ∝ r⁻ᵅ, with α typically around 1. The single most important visual consequence is that a power law becomes a straight line when you plot it on log-log axes. So if you take the logarithm of rank and the logarithm of frequency and the points fall on a downward-sloping straight line, you are almost certainly looking at a Zipf-type rank-frequency distribution.
This straight-line test is useful because it is hard to fake. When the number of visitors to websites is plotted against rank, the relationship is nearly linear on a log-log plot with a slope of about minus one, which is what makes it Zipfian. The same diagnostic applies to citations, downloads, and paper counts, which is why the rank-frequency approach travels so easily between fields.
Building a rank-frequency table and making predictions
The real payoff of the approach is prediction. Once you know the constant in the rank-frequency relationship, you can estimate the contribution of a source at any rank, even ranks you have not directly measured. This turns a descriptive table into a forecasting tool.
Consider a simple worked example. Suppose you survey a research field and find that the most prolific journal, ranked 1, published 120 papers on a topic last year. Using Zipf’s idealised relationship, the rank-2 journal would publish roughly half that, the rank-3 journal roughly a third, and so on. The table fills out like this:
Rank 1: 120 papers (the observed top contributor).
Rank 2: about 60 papers (120 ÷ 2).
Rank 3: about 40 papers (120 ÷ 3).
Rank 4: about 30 papers (120 ÷ 4).
Rank 5: about 24 papers (120 ÷ 5).
Each estimate comes from dividing the top journal’s output by the rank number, which is the direct consequence of the constant product rule. The same logic runs in reverse. If you spot a journal you believe is the fifth most productive but cannot easily count its full output, the model predicts roughly 24 papers, giving you a sensible benchmark to check against. Zipf’s law also implies that a term’s occurrence frequency can be estimated from its rank by dividing the Zipf constant by the rank number, which is exactly the operation performed in the journal table above.
Knowing the limits of the prediction
Rank-frequency predictions are estimates, not guarantees, and treating them as exact is a common mistake. Real data rarely matches the idealised “half, then a third” pattern perfectly. The fit is usually best in the middle of the ranking and weaker at the extremes. In one detailed text analysis, ranks 8 to 600 showed a spread of no more than about 25 percent in the Zipf product, tightening to around 10 percent for ranks 40 to 300. The very top ranks and the long tail tend to deviate, which is why the original Zipf formula was later refined.
That refinement came from Benoît Mandelbrot, who adjusted the formula to correct the poor fit at very high and very low frequencies that weakened the original word-counting model. For most LIS coursework and routine bibliometric analysis, the plain Zipf version is enough to grasp the structure. But it helps to remember that the rank-frequency table is a model of a tendency, and that empirical counts will scatter around the predicted line rather than sit exactly on it.
One distribution, many disciplines
What makes the rank-frequency approach worth learning is its reach beyond libraries. The same downward-sloping power law that describes journal output also describes the frequency of words in natural languages, the sizes of urban agglomerations, firm sizes, and internet traffic. A method developed to count words in a novel now helps map scientific collaboration, rank research keywords, and identify core journals. Learning to read a rank-frequency table in bibliometrics therefore gives you a tool that transfers directly to economics, linguistics, web analytics, and beyond.
What do you think? If a research field’s top journal suddenly doubled its output, how would that change the rest of the predicted rank-frequency table, and would the core set of journals stay the same? And when real publication data deviates from Zipf’s neat pattern, is that a failure of the model or a signal that something interesting is happening in the field?
References
- https://arxiv.org/pdf/1011.1533
- https://arxiv.org/pdf/1411.0928
- https://www.britannica.com/topic/Zipfs-law
- https://ebooks.inflibnet.ac.in/liscp10/chapter/classical-law-of-bibliometrics/
- https://www.sciencedirect.com/science/article/pii/S2694610625000104
- https://files.eric.ed.gov/fulltext/EJ1115017.pdf
- https://link.springer.com/chapter/10.1007/978-3-319-41631-1_4
- https://www.researchgate.net/publication/264873465_Zipf's_Law_What_and_Why
- https://www.hpl.hp.com/research/idl/papers/ranking/ranking.html
- https://image-ppubs.uspto.gov/dirsearch-public/print/downloadPdf/5293552
- https://arxiv.org/pdf/cs/0610091
- https://arxiv.org/pdf/0908.0501

Leave a Reply