When a librarian decides which journals to subscribe to, or when a researcher judges where to publish, both rely on the idea that some journals matter more than others. But how do you measure that importance objectively? The answer that has shaped library science for nearly a century is citation analysis – counting how often a journal is cited by others. The method began with a simple count proposed by Gross and Gross in 1927, yet that simplicity hid serious flaws. This post explains why straight citation counting can mislead us, how the Indian scientometrician I. N. Sengupta devised a clever correction, and what we learn by comparing different ranking methods side by side.

Table of Contents

The origin of citation counting

The story starts in 1927, when P. L. K. Gross and E. M. Gross published a study in the journal Science titled “College libraries and chemical education.” They examined every reference in one year’s issues of the Journal of the American Chemical Society and tallied which periodicals were cited most often. Their goal was practical: to help small college libraries decide which chemistry journals were essential when budgets were tight.

This was the first citation analysis ever conducted, and the principle behind it became foundational. Gross and Gross treated the raw count of citations a journal received as a direct measure of its importance. More citations meant a more valuable journal. This idea seeded everything that followed in bibliometrics, including Eugene Garfield’s later work on citation indexing and the impact factor.

How the method works

The Gross method is appealingly straightforward. You select a set of source journals in a field, you go through their reference lists, and you count how many times each cited journal appears. Rank the journals from most-cited to least-cited, and you have a list of the “core” periodicals for that subject. Gross and Gross even noticed a pattern that later matched Zipf’s Law – a small number of journals were cited very frequently, while a large number were cited only once or twice.

For libraries with limited funds, this was gold. It gave an objective, replicable, quantitative basis for collection development decisions, replacing guesswork and personal preference. That objectivity remains the strongest argument in favour of citation analysis even today.

Problems with country and language

The trouble is that a raw citation count measures more than just quality. It measures visibility, accessibility, and the citing habits of researchers – and these are heavily skewed.

Consider language bias. Papers published in English consistently receive more citations than equally good papers in other languages, simply because English is more widely read. A study of the Science Citation Index found that English-language articles yield far higher citation impacts than non-English ones, and that citation databases provide uneven coverage of foreign-language journals to begin with. For a country like India, where some valuable research appears in regional or specialised outlets, this means genuinely important work can look unimportant in a raw count.

Then there is country bias. Researchers tend to cite work from their own nation more readily, partly because they read national journals and partly because of professional networks. Studies have shown that scientists are more likely to cite papers published in national languages and that authors from large, well-funded research communities receive more citations than those from smaller ones. The result is that the citation totals of journals from the United States and Western Europe are inflated relative to journals from elsewhere, regardless of the actual research merit involved.

A raw count also captures self-citation and the sheer size of a journal. A large journal that publishes hundreds of papers a year will naturally collect more citations than a small specialist journal, even if the small one is more influential per article. Garfield’s later insight – to divide citations by the number of citable items – was a response to exactly this size problem.

Sengupta’s weightage formula

One bias that the Gross method completely ignores is chronology. This is the problem that the Indian scientometrician I. N. Sengupta set out to solve, and his correction is a notable contribution to the field from India.

The post-war journal problem

Sengupta observed that ranking lists are built by counting citations with equal weightage given to every citation, regardless of when the cited journal began publishing. This sounds fair, but it quietly punishes newer journals. As he explained in his work, the position occupied by post-war journals may not be an accurate index of their real demand or usefulness.

Why? Because a journal that started in, say, the 1960s has had far fewer years to accumulate citations than a journal founded before the Second World War. Many post-war journals had been in existence for less than twenty years when these ranking studies were done, so their citation counts covered a much shorter window than those of older, pre-war journals. A young journal of excellent quality simply has not had time to build up a citation record comparable to an established one. Raw counting therefore discriminates against new journals through no fault of their content.

How the formula corrects the bias

To level the playing field, Sengupta proposed an off-setting weightage formula. Instead of treating all citations equally, the formula applies a corrective weight that accounts for the shorter period during which newer journals could have been cited. In effect, it boosts the credit given to citations of younger journals so that their scores reflect performance per unit of available time rather than raw accumulation.

This adjustment changes the rankings meaningfully. When the weightage is applied, several post-war journals rise to positions that better reflect their current usefulness to working researchers, while the artificial advantage of older journals is reduced. The correction does not throw out citation counting; it refines it so that the age of a journal stops distorting the result.

Application to microbiology and biochemistry

Sengupta did not leave the formula as theory. He applied it to real disciplines. In one study he used the weightage formula to rerank periodicals in the field of microbiology, producing a revised list that he recommended in preference to an earlier ranking he himself had compiled without the correction. He carried out similar work on the literature of biochemistry, examining how the growth of the literature changed the ranking of periodicals over time.

Sengupta was a major figure in Indian library and information science more broadly. He is also widely credited with helping define the scope of bibliometrics and related fields like scientometrics and informetrics, which makes his methodological work all the more significant for students in this country.

Comparing citation methods

Sengupta’s formula is one of several attempts to improve on raw counting. Comparing these methods reveals an important lesson: no single number tells the whole story.

Different methods, different insights

Garfield’s impact factor, introduced in the 1960s, normalises citations by dividing them by the number of citable articles, which addresses the size problem that Gross ignored. Later indicators went further. Metrics such as SNIP (Source Normalised Impact per Paper), SJR (SCImago Journal Rank), the Eigenfactor, and the h-index each correct for a different weakness – field differences, the prestige of the citing journal, or the longevity bias in older papers.

What is striking is that these indicators do not agree with one another. A detailed review of business and management journals found that although the various metrics appear highly correlated, in practice they lead to large differences in journal rankings. A journal can rank near the top by one measure and slip considerably by another. Each method, in other words, illuminates a different facet of what “importance” means.

Beyond simple counting

This is why modern bibliometrics treats citation counts as evidence rather than verdicts. A raw count from the Gross method tells you about gross visibility. Sengupta’s weighted count tells you about usefulness adjusted for age. The impact factor tells you about average citations per paper. A wise librarian or researcher reads several of these together, alongside qualitative judgement about the field.

The deeper point connects back to the biases we began with. Citation metrics do not capture the reasons a work is cited – whether the citation is praise, criticism, or routine acknowledgement. They should never be the sole criterion for assessing research merit, and they must always be interpreted in light of the discipline’s publication practices. Sengupta’s contribution matters precisely because it reminds us that even an objective-looking number carries hidden assumptions worth questioning.

What do you think? If a brilliant new journal and an old established one received the same number of citations, which one would you consider more important, and why? And in the Indian context, how should we account for valuable research that appears in regional or non-English journals when we rank scholarly work?

How useful was this post?

Click on a star to rate it!

Average rating 0 / 5. Vote count: 0

No votes so far! Be the first to rate this post.

We are sorry that this post was not useful for you!

Let us improve this post!

Tell us how we can improve this post?

References
  1. https://www.science.org/doi/10.1126/science.66.1713.385
  2. https://www.tandfonline.com/doi/full/10.1080/0194262X.2023.2238013
  3. https://scholarworks.calstate.edu/downloads/6w924c50c
  4. https://arxiv.org/pdf/astro-ph/0401228
  5. https://pmc.ncbi.nlm.nih.gov/articles/PMC3826058/
  6. https://link.springer.com/article/10.1007/BF02026414
  7. https://link.springer.com/article/10.1007/BF00353144
  8. https://www.frontiersin.org/journals/research-metrics-and-analytics/articles/10.3389/frma.2021.742311/full
  9. https://arxiv.org/pdf/1604.06685
  10. https://www.frontiersin.org/articles/202382

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

Informetrics & Scientometrics

1 Information and Measurement

  1. Information Revisited
  2. Framework for Information Exchange
  3. Measurement Techniques
  4. Informativeness
  5. Standardization of Measurement

2 Measure of Information

  1. Information and Entropy
  2. Shannon Information
  3. Probabilistic Information
  4. Properties of Shannon Information
  5. Derivation of Shannon Information Formula
  6. Normalization Condition
  7. Relating Semantic Value to Shannon Type Measures
  8. Other Shannon Type Measures of Information
  9. Semantic Information
  10. Fuzzy Information Measure
  11. Other Information Measures

3 Informetrics – Definition, Scope and Evolution

  1. Definitions
  2. Scope
  3. Evolution
  4. Summary

4 Sociology of Science and Scientometrics

  1. Sociology of Science
  2. Growth of Scientific Knowledge
  3. Social Organization in Research Areas
  4. Approaches of Scientometrics to Sociology of Science
  5. Models of Growth of Knowledge

5 Organizations Engaged in Scientometrics and Informetrics Studies

  1. Organizations Engaged in or Supporting Scientometrics/Informetrics Studies
  2. Websites
  3. Research Groups/Discussion Groups
  4. Periodical Publications
  5. Conferences/Seminars/Workshops/Congresses
  6. Individuals Engaged in the Study and Research in Scientometrics/Informetrics

6 Law of Scattering and its Applications

  1. Introduction
  2. Historical Account
  3. Bradford’s Law
  4. Verbal Form of Bradford’s Law
  5. Applications of Bradford’s Law
  6. Graphical Representation of Bradford’s Law
  7. Conditions for Bradford’s Law
  8. Falling Tail of Bradford Curve: The Groos Droop
  9. Ambiguity in Bradford’s Law
  10. Fitting Bibliographic Data to Bradford’s Law

7 Rank and Size Frequency Models

  1. Representations and Organization of Numerical Data
  2. Size – Frequency Approach
  3. Rank – Frequency Approach
  4. Size – Frequency Models
  5. Rank – Frequency Cumulative (Fractional) Models
  6. Rank – Frequency Cumulative (Non-Fractional) Models
  7. Rank – Frequency Non – Cumulative Models

8 Informetrics Phenomena

  1. Terminology and Historical Development
  2. Selected Laws of Bibliometrics and Informetrics
  3. Informetrics Phenomena in Science
  4. Practical Applications of Informetrics

9 Analysis of Library Related Data

  1. Necessity for Analytical Studies in Libraries
  2. Citation Counting: A Versatile Tool for Journal Selection
  3. An Alternative Method of Citation Analysis
  4. Selection of New Source Journals to Eliminate Bias Due to Country, and Language
  5. Weightage Formula to Correct Citation for Post-War Periodicals
  6. Three New Bibliometric Parameters to Re-Rank Scientific Periodicals
  7. Garfield’s Methods for Cito-Analytical Studies
  8. Librametric Analysis
  9. Bibliometric Analysis
  10. Informetrics
  11. Scientometrics: Its Genesis, Scope, Definition, and Applications

10 User Studies

  1. User Studies
  2. Questionnaire Method
  3. Interview Method
  4. Diary Method
  5. Observation Method
  6. Planning a Survey
  7. Classification and Tabulation of Data
  8. Analysis of Data
  9. Presentation of Results
  10. Important User Studies
  11. Application of User Studies

11 Laws of Scientific Productivity

  1. Scientific Productivity – Influencing Factors
  2. Scientific Productivity – Problems in Measurement
  3. Scientific Productivity – Distribution Characteristics
  4. Lotka’s Law
  5. Statistical Distributions or Models
  6. Application of Lotka’s Law
  7. Goodness-of-Fit Test

12 Growth and Obsolescence of Literature

  1. Growth of Literature
  2. Obsolescence of Literature
  3. Growth Vs Obsolescence of Literature

13 Science Indicators

  1. Indicators
  2. Towards Science Indicators
  3. Historical Aspects
  4. Functions of Science Indicators
  5. S&T Indicators for the Developing Countries
  6. Types of Indicators
  7. Validity and Reliability of Indicators
  8. Building S&T Indicators
  9. Literature Based Indicators
  10. Patent Indicators

14 Mapping of Science

  1. Cognitive Mapping
  2. Journal-to-journal Citation Maps
  3. Co-citation Maps
  4. Co-word Maps
  5. Co-classification Maps
  6. Descriptive Mapping

15 Elements of Statistics

  1. Data and Its Measurement
  2. Graphical Representation
  3. Measures of Central Tendency
  4. Measure of Variability
  5. Correlation and Regression

16 Probability Distributions and their Applications

  1. Probability – Definition
  2. Random Variables
  3. Joint Probability Distribution
  4. Conditional Probability Distribution
  5. Some Special Distributions
  6. Applications of Probability

17 Regression Analysis

  1. Simple Linear Regression
  2. Multiple Regression
  3. Stepwise Regression
  4. Regression with Qualitative Explanatory Variables

18 Cluster Analysis and Factor Analysis

  1. Introduction
  2. Cluster Analysis
  3. Factor Analysis
  4. Examples of Cluster and Factor Analysis