When we line up journals, authors, or keywords from the most productive to the least and then look at how many items each one contributes, a striking pattern appears. A handful of sources dominate, and a long tail of smaller contributors stretches out behind them. Rank-frequency non-cumulative models are the mathematical tools that describe this pattern directly. Instead of adding up totals as we move down the list, these models tell us exactly how much the source at each individual rank contributes. This makes them powerful for both describing existing data and predicting how a collection of publications is likely to be distributed.

Table of Contents

What “rank-frequency” and “non-cumulative” actually mean

The rank-frequency approach starts with a simple step: arrange all the sources in a dataset in decreasing order of their output. The source that produces the most gets rank 1, the next gets rank 2, and so on. Once everything is ranked, we record the frequency, that is, the number of items each source produces.

The term non-cumulative is the key distinction here. In a cumulative model, the value at any rank includes everything that came before it, so you read off running totals such as “the top five journals together published 200 articles.” In a non-cumulative model, each rank stands on its own. You read the frequency of that single source only, for example “the journal at rank 5 published 12 articles.” Non-cumulative models therefore preserve the individual data points rather than rolling them into a growing sum.

This difference matters because non-cumulative distributions keep the raw shape of the data visible. As researchers analysing ranking techniques have pointed out, rank distributions retain empirical information that cumulative forms tend to smooth away, which gives them a theoretical advantage when you want to study the underlying structure rather than just the broad totals.

The foundation: Zipf’s law as a non-cumulative model

The most famous non-cumulative rank-frequency model is Zipf’s law. Originally developed by the linguist George Kingsley Zipf to describe word frequencies in text, it states that the frequency of an item is inversely proportional to its rank. When values are sorted in decreasing order, the n-th entry tends to be roughly inversely proportional to n.

In practical terms, this means the second-ranked source produces about half as much as the first, the third produces about a third as much, and so on. The relationship is usually written as:

f × r = C

Here f is the frequency of a source, r is its rank, and C is a constant for the dataset. Because the product of frequency and rank stays roughly constant, you can estimate the output of a source at any rank once you know the constant. This is what makes Zipf’s law a predictive tool and not just a descriptive one.

Zipf’s law is recognised as one of the three classical empirical laws of bibliometrics, alongside Lotka’s law of author productivity and Bradford’s law of scattering. While Lotka and Bradford are often expressed in size-frequency or cumulative terms, Zipf’s law in its basic form is non-cumulative because it speaks to the frequency of each ranked item individually.

Why the basic model needs refinement

The simple Zipf relationship works well in the middle of a distribution but tends to break down at the extremes. The very top-ranked items often contribute more than the formula predicts, and the long tail of low-frequency items behaves differently too. Modern analyses of large text corpora confirm this: when huge datasets are plotted on a log-log scale, the tail bends downward instead of following the straight line that pure Zipf’s law would produce. This limitation is exactly what prompted later researchers to develop more flexible non-cumulative models.

Non-cumulative rank models: Fairthorne and the hyperbolic family

One of the most influential contributions to this area came from R. A. Fairthorne in 1969. In his paper on empirical hyperbolic distributions for bibliometric description and prediction, Fairthorne brought together the work of Bradford, Zipf, and Mandelbrot under a single umbrella. He argued that the various bibliometric laws are all manifestations of the same underlying hyperbolic pattern.

A hyperbolic relationship is one where the frequency multiplied by the rank, raised to some power, stays constant. The basic Zipf form is the simplest version of this. Fairthorne’s key insight was that these distributions are not isolated curiosities but a connected family, and that a hyperbolic model can be used to both describe a body of literature and predict its future shape. The very title of his work, pairing “description” with “prediction,” signals that these models are meant to forecast distributions, not merely catalogue them.

Mandelbrot’s extension

The mathematician Benoit Mandelbrot refined Zipf’s law to handle the deviations seen at the top of distributions. The resulting Zipf-Mandelbrot law adds an extra constant inside the formula so that the curve can bend to fit real data more closely. It is often written in the form where frequency is proportional to a constant divided by (rank + a shift parameter), raised to a power.

The shift parameter is what gives this model its flexibility. When that parameter is zero, the formula collapses back into ordinary Zipf’s law, which means Zipf’s law is simply a special case of Mandelbrot’s more general model. This extension is particularly useful when the highest-ranked sources do not behave the way pure Zipf’s law expects, which is common in real publication data.

Chen and Leimkuhler: unifying the laws

Another important step came from Y. S. Chen and F. F. Leimkuhler. In 1986 they published a study establishing a relationship between Lotka’s law, Bradford’s law, and Zipf’s law, showing that these three pillars of bibliometrics are mathematically connected rather than separate phenomena.

This unification is significant for anyone working with non-cumulative models. It means that if you understand the rank-frequency behaviour described by Zipf, you also hold a key to the productivity patterns described by Lotka and the scattering patterns described by Bradford. Chen and Leimkuhler demonstrated that these laws produce similar probability distributions, with the differences lying mainly in the kind of data being analysed, whether that data is words in a text, authors and their papers, or articles scattered across journals.

For the analyst, this consistency is reassuring. It suggests that the long-tail pattern we see in scholarly output is not an accident of one particular dataset but a structural feature of how knowledge production works.

Use in informetrics: predicting publication distribution

The practical value of non-cumulative rank-frequency models lies in prediction. Once you have ranked your sources and confirmed that the data follows a Zipfian or hyperbolic pattern, you can estimate values you have not directly measured.

Consider a few concrete uses. A librarian building a journal collection can predict how many articles on a topic the lower-ranked journals are likely to publish, helping decide where subscription money is best spent. A research administrator can forecast how output will concentrate among a small group of prolific authors. A subject specialist can identify the “core” sources that carry most of the literature and separate them from the marginal ones.

The prediction works because the constant in the Zipf formula is stable for a given collection. If you measure the frequency of the top few sources, you can calculate the constant and then project the expected frequency at any other rank. This is why these models are valued for assessing the growth and structure of a subject by arranging publication data with statistical techniques.

Fitting and testing the model

Before trusting a prediction, analysts check how well the data actually fits the model. A common method is to plot frequency against rank on a log-log graph. If the data conforms to Zipf’s law, the points fall close to a straight line whose slope corresponds to the exponent of the distribution. A clean straight line suggests the pattern is genuine and not the product of chance; a bending or broken line suggests a Mandelbrot-style correction may fit better.

Example in action: analysing journal contributions

Suppose a student is studying which journals publish research on a particular subject and gathers the following counts of articles. The journals have already been arranged from the most productive to the least.

Rank 1: 60 articles
Rank 2: 31 articles
Rank 3: 19 articles
Rank 4: 16 articles
Rank 5: 12 articles

To test whether this follows Zipf’s law, we multiply each frequency by its rank and look at the products:

Rank 1: 60 × 1 = 60
Rank 2: 31 × 2 = 62
Rank 3: 19 × 3 = 57
Rank 4: 16 × 4 = 64
Rank 5: 12 × 5 = 60

The products cluster tightly around 60, which tells us the constant C for this collection is roughly 60. This is the non-cumulative pattern in action. Each rank is examined on its own, and the frequency-times-rank product stays nearly constant across the list.

Now we can predict. If we want to estimate how many articles the journal at rank 6 is likely to publish, we divide the constant by the rank: 60 ÷ 6 = 10 articles. For rank 10, the estimate would be 60 ÷ 10 = 6 articles. Without having to count the lower-ranked journals individually, we have a reasonable forecast of their output.

This is exactly the kind of forecasting that makes the model useful in collection development and research evaluation. Notice how different this is from a cumulative approach, which would instead report running totals such as “the top three journals published 110 articles between them.” The non-cumulative view keeps each journal’s contribution distinct, which is what lets us model the rank-by-rank pattern and extend it.

A note of caution

Real data is rarely as tidy as a textbook example. The products will not be perfectly constant, and the top-ranked and bottom-ranked sources are the most likely to drift away from the pattern. This is precisely why the Zipf-Mandelbrot extension exists, and why goodness-of-fit testing matters. Pure hyperbolic decay produces the classic straight-line Zipf plot, but other decay rates create noticeable differences at small and large ranks. Treating the prediction as an estimate rather than an exact value is the sensible approach.

Why these models matter for information science

Non-cumulative rank-frequency models give information professionals a quantitative grip on something that would otherwise feel vague: the uneven way knowledge concentrates in a few productive sources. They turn a messy list of journals or authors into a predictable curve. From that curve flow practical decisions about which journals to subscribe to, which sources form the core of a field, and how a body of literature is likely to grow.

The lineage from Zipf through Fairthorne, Mandelbrot, and Chen and Leimkuhler also shows how the discipline matured. What began as a description of word frequencies became a unified framework linking the major bibliometric laws, capable of both explaining the past distribution of publications and forecasting the future one.

What do you think? If a small number of journals consistently dominate the literature in your field, should libraries concentrate their limited budgets on those core titles, or deliberately support the long tail of smaller journals? And when a real dataset only loosely follows Zipf’s law, how much should we trust the predictions we draw from it?

How useful was this post?

Click on a star to rate it!

Average rating 0 / 5. Vote count: 0

No votes so far! Be the first to rate this post.

We are sorry that this post was not useful for you!

Let us improve this post!

Tell us how we can improve this post?

References
  1. https://www.sciencedirect.com/science/article/abs/pii/0306457384900384
  2. https://en.wikipedia.org/wiki/Zipf's_law
  3. https://files.eric.ed.gov/fulltext/EJ1115017.pdf
  4. https://www.sciencedirect.com/science/article/abs/pii/S1751157707000296
  5. https://www.emerald.com/jd/article/25/4/319/206708/Empirical-Hyperbolic-Distributions-Bradford-Zipf
  6. https://link.springer.com/chapter/10.1007/978-3-319-41631-1_4
  7. https://link.springer.com/article/10.1007/s11192-020-03476-8
  8. https://tefkos.comminfo.rutgers.edu/Courses/e530/Readings/Jayroe%20Bibliometrics%20for%20Dummies%202008.pdf
  9. https://www.sciencedirect.com/science/article/pii/S2694610625000104
  10. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5628998/

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

Informetrics & Scientometrics

1 Information and Measurement

  1. Information Revisited
  2. Framework for Information Exchange
  3. Measurement Techniques
  4. Informativeness
  5. Standardization of Measurement

2 Measure of Information

  1. Information and Entropy
  2. Shannon Information
  3. Probabilistic Information
  4. Properties of Shannon Information
  5. Derivation of Shannon Information Formula
  6. Normalization Condition
  7. Relating Semantic Value to Shannon Type Measures
  8. Other Shannon Type Measures of Information
  9. Semantic Information
  10. Fuzzy Information Measure
  11. Other Information Measures

3 Informetrics – Definition, Scope and Evolution

  1. Definitions
  2. Scope
  3. Evolution
  4. Summary

4 Sociology of Science and Scientometrics

  1. Sociology of Science
  2. Growth of Scientific Knowledge
  3. Social Organization in Research Areas
  4. Approaches of Scientometrics to Sociology of Science
  5. Models of Growth of Knowledge

5 Organizations Engaged in Scientometrics and Informetrics Studies

  1. Organizations Engaged in or Supporting Scientometrics/Informetrics Studies
  2. Websites
  3. Research Groups/Discussion Groups
  4. Periodical Publications
  5. Conferences/Seminars/Workshops/Congresses
  6. Individuals Engaged in the Study and Research in Scientometrics/Informetrics

6 Law of Scattering and its Applications

  1. Introduction
  2. Historical Account
  3. Bradford’s Law
  4. Verbal Form of Bradford’s Law
  5. Applications of Bradford’s Law
  6. Graphical Representation of Bradford’s Law
  7. Conditions for Bradford’s Law
  8. Falling Tail of Bradford Curve: The Groos Droop
  9. Ambiguity in Bradford’s Law
  10. Fitting Bibliographic Data to Bradford’s Law

7 Rank and Size Frequency Models

  1. Representations and Organization of Numerical Data
  2. Size – Frequency Approach
  3. Rank – Frequency Approach
  4. Size – Frequency Models
  5. Rank – Frequency Cumulative (Fractional) Models
  6. Rank – Frequency Cumulative (Non-Fractional) Models
  7. Rank – Frequency Non – Cumulative Models

8 Informetrics Phenomena

  1. Terminology and Historical Development
  2. Selected Laws of Bibliometrics and Informetrics
  3. Informetrics Phenomena in Science
  4. Practical Applications of Informetrics

9 Analysis of Library Related Data

  1. Necessity for Analytical Studies in Libraries
  2. Citation Counting: A Versatile Tool for Journal Selection
  3. An Alternative Method of Citation Analysis
  4. Selection of New Source Journals to Eliminate Bias Due to Country, and Language
  5. Weightage Formula to Correct Citation for Post-War Periodicals
  6. Three New Bibliometric Parameters to Re-Rank Scientific Periodicals
  7. Garfield’s Methods for Cito-Analytical Studies
  8. Librametric Analysis
  9. Bibliometric Analysis
  10. Informetrics
  11. Scientometrics: Its Genesis, Scope, Definition, and Applications

10 User Studies

  1. User Studies
  2. Questionnaire Method
  3. Interview Method
  4. Diary Method
  5. Observation Method
  6. Planning a Survey
  7. Classification and Tabulation of Data
  8. Analysis of Data
  9. Presentation of Results
  10. Important User Studies
  11. Application of User Studies

11 Laws of Scientific Productivity

  1. Scientific Productivity – Influencing Factors
  2. Scientific Productivity – Problems in Measurement
  3. Scientific Productivity – Distribution Characteristics
  4. Lotka’s Law
  5. Statistical Distributions or Models
  6. Application of Lotka’s Law
  7. Goodness-of-Fit Test

12 Growth and Obsolescence of Literature

  1. Growth of Literature
  2. Obsolescence of Literature
  3. Growth Vs Obsolescence of Literature

13 Science Indicators

  1. Indicators
  2. Towards Science Indicators
  3. Historical Aspects
  4. Functions of Science Indicators
  5. S&T Indicators for the Developing Countries
  6. Types of Indicators
  7. Validity and Reliability of Indicators
  8. Building S&T Indicators
  9. Literature Based Indicators
  10. Patent Indicators

14 Mapping of Science

  1. Cognitive Mapping
  2. Journal-to-journal Citation Maps
  3. Co-citation Maps
  4. Co-word Maps
  5. Co-classification Maps
  6. Descriptive Mapping

15 Elements of Statistics

  1. Data and Its Measurement
  2. Graphical Representation
  3. Measures of Central Tendency
  4. Measure of Variability
  5. Correlation and Regression

16 Probability Distributions and their Applications

  1. Probability – Definition
  2. Random Variables
  3. Joint Probability Distribution
  4. Conditional Probability Distribution
  5. Some Special Distributions
  6. Applications of Probability

17 Regression Analysis

  1. Simple Linear Regression
  2. Multiple Regression
  3. Stepwise Regression
  4. Regression with Qualitative Explanatory Variables

18 Cluster Analysis and Factor Analysis

  1. Introduction
  2. Cluster Analysis
  3. Factor Analysis
  4. Examples of Cluster and Factor Analysis