When you cite two earlier papers together in your own research, you are making a quiet statement: these works belong in the same conversation. Now imagine thousands of researchers doing this across millions of papers. Patterns begin to emerge. Certain documents keep appearing side by side in reference lists, signalling a hidden intellectual relationship. This is the foundation of co-citation analysis, one of the most powerful techniques in scientometrics for revealing the structure of knowledge itself.

Table of Contents

What co-citation actually means

Co-citation is a specific relationship between two documents. Two documents are said to be co-cited when a later, third document references both of them together. The more frequently this pairing occurs across the literature, the stronger their co-citation relationship.

The key insight is that this relationship is created by the citing authors, not by the original authors. When many researchers independently cite Paper A and Paper B in the same reference list, they are collectively signalling that these two works are intellectually connected. Co-citation frequency therefore becomes a measurable indicator of subject similarity between documents.

This idea was introduced in 1973 by Henry Small, then working at the Institute for Scientific Information. In his foundational paper, he defined co-citation as a new measure of the relationship between two documents and demonstrated it using examples from particle physics literature. The Russian researcher Irina Marshakova independently developed the same concept in the same year, although her work received less attention because it was published in Russian. Both are now credited as co-originators.

Interestingly, the idea did not come purely from information science. Small later explained that the concept emerged from his earlier work documenting the history of nuclear physics, an origin that the references in his own paper did not fully reveal.

How it differs from bibliographic coupling

Co-citation is often confused with bibliographic coupling, an earlier concept proposed by M. M. Kessler in 1963. The two are related but point in opposite directions. Bibliographic coupling links two documents when they share common references in their own bibliographies. Co-citation links two documents when they are later cited together by others.

This difference has an important consequence. Bibliographic coupling is fixed at the moment of publication, because a paper’s reference list never changes. Co-citation, on the other hand, is dynamic. Because future citations depend on how a field evolves, the co-citation frequency between two papers can keep changing over time. This makes co-citation a forward-looking measure, while bibliographic coupling is retrospective.

Small himself tested these methods. When he compared direct citation, co-citation, and bibliographic coupling using a set of physics papers, he found that co-citation patterns agreed much more closely with direct citation patterns than bibliographic coupling did. This is one reason co-citation became the preferred tool for mapping science.

From pairs to maps

A single co-citation pair is just one data point. The real power appears when these pairs are aggregated. When thousands of co-citation relationships are combined, they form a network. In this network, each document is a node, and the strength of the co-citation relationship determines how closely two nodes are placed together.

Documents that are frequently co-cited cluster together, while unrelated documents drift apart. The result is a visual map of a scientific field, where tight clusters represent specialties or topics, and the gaps between them represent disciplinary boundaries.

This clustering process is largely automated. Software identifies groups of documents that are strongly co-cited and then labels each cluster, often using significant words drawn from the titles or abstracts of the papers that cite the cluster. The outcome is a self-organising classification of science that emerges from citation behaviour rather than from any predetermined subject scheme.

Research fronts and the intellectual base

One of the most useful distinctions in this field is between the intellectual base and the research front. These two concepts describe different layers of a scientific specialty, and co-citation analysis is central to identifying both.

The intellectual base consists of the foundational, frequently co-cited works that give a field its shape. The intellectual base is the essential knowledge required to undertake research in a domain, and it includes the significant or classic works that have defined a broad area of enquiry. Because these are the documents being cited, co-citation analysis identifies them directly.

The research front, by contrast, refers to the current, active problems that researchers are working on right now. A research front consists of a cluster of co-cited core papers together with the group of current source papers that cite one or more of these core papers. In other words, the cited members of a cluster form its intellectual base, while the papers citing the cluster form its research front.

This relationship was explored in detail by Olle Persson, who applied these methods to information science literature. His study showed that a co-citation map of the most co-cited authors closely resembled maps of the field produced by other methods, and that research fronts could be defined as clusters of articles drawing on similar parts of the intellectual base.

Variants of co-citation analysis

The original method analysed individual documents, but researchers soon realised the same logic could be applied to other units. This gave rise to several variants, each revealing a different aspect of scientific structure.

Document co-citation analysis

This is the original form, in which individual documents are the units of analysis. It is excellent for identifying specific influential papers and the topics that connect them. It works best when you want to know which exact works form the backbone of a field.

Author co-citation analysis

Here the unit shifts from documents to authors. Two authors are co-cited when a later paper cites the work of both. This approach was developed by Howard White and Belver Griffith in 1981 and later refined in a well-known study by White and McCain, who in 1998 mapped the field of information science using author co-citation analysis. Author co-citation analysis is valuable for understanding the intellectual structure of a discipline in terms of its key thinkers rather than its individual publications.

Journal co-citation analysis

In this variant, journals become the unit of analysis. Mapping which journals are co-cited together reveals how disciplines and subfields relate to one another at a broader level. It is particularly useful for understanding the boundaries and overlaps between entire research areas.

The tools that make mapping possible

Co-citation analysis on a large scale is impossible by hand. Modern science mapping relies on specialised software that extracts citation data, calculates co-citation frequencies, builds networks, and visualises clusters. Two tools dominate this space.

VOSviewer, developed by Nees Jan van Eck and Ludo Waltman, is widely used for constructing and visualising bibliometric networks. It builds networks of journals, researchers, and articles based on co-citation, bibliographic coupling, or co-authorship relationships, and provides graphical representations of the scientific landscape. Its clean, intuitive maps make it a popular starting point for researchers.

CiteSpace, developed by Chaomei Chen, takes a more analytical approach. It pinpoints research focal points by analysing clusters of publications and is especially good at detecting emerging trends and intellectual turning points. Many studies use the two tools together, relying on VOSviewer for initial visual exploration and CiteSpace for deeper analysis of specific trends and clusters.

These tools typically draw their data from large citation databases such as the Web of Science, which provides the structured citation records needed for the calculations. The combination of rich data and powerful visualisation has made co-citation mapping accessible to researchers across every discipline.

Where co-citation analysis is used today

The applications extend far beyond information science. Researchers use co-citation maps to understand the development of fields as varied as blockchain technology, medicine, and management studies. The method has even been extended into the digital humanities, where it has been applied to large corpora of early modern letters to explore how historical figures and ideas connect.

For students and early researchers, co-citation maps offer a practical benefit. Instead of reading hundreds of papers blindly, you can use a co-citation map to identify the foundational works of a field, see how its subtopics relate, and spot where the active research is happening. It turns an overwhelming body of literature into a navigable landscape. This is one reason research support based on such analysis has become a valued service in academic libraries.

The limitations to keep in mind

Co-citation analysis is powerful but not perfect. Because it depends on accumulated citations, very recent papers cannot be mapped well, as they have not yet had time to be co-cited. This introduces a bias towards older, established works in the intellectual base.

The method also reflects citation behaviour, which carries its own biases. Highly cited papers attract more citations regardless of merit, and citation practices vary between disciplines. A map produced from co-citation data shows what the citing community treats as connected, which is not always the same as objective intellectual similarity. Used thoughtfully, however, co-citation analysis remains one of the clearest windows we have into the structure and evolution of scientific knowledge.

What do you think? If co-citation maps reveal the intellectual structure of a field through the collective behaviour of citing authors, how much should we trust them to define what is truly central to a discipline? And as artificial intelligence begins to assist researchers in finding and citing sources, how might that reshape the co-citation maps of the future?

How useful was this post?

Click on a star to rate it!

Average rating 5 / 5. Vote count: 1

No votes so far! Be the first to rate this post.

We are sorry that this post was not useful for you!

Let us improve this post!

Tell us how we can improve this post?

References
  1. https://asistdl.onlinelibrary.wiley.com/doi/10.1002/asi.4630240406
  2. https://en.wikipedia.org/wiki/Co-citation
  3. https://arxiv.org/pdf/1806.00224
  4. https://arxiv.org/pdf/1511.05078
  5. https://intarch.ac.uk/journal/issue42/8/5.html
  6. https://clarivate.com/academia-government/essays/research-fronts/
  7. https://asistdl.onlinelibrary.wiley.com/doi/abs/10.1002/(SICI)1097-4571(199401)45:1%3C31::AID-ASI4%3E3.0.CO;2-G
  8. https://www.vosviewer.com/download/f-x2.pdf
  9. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11408372/
  10. https://academic.oup.com/dsh/article/39/1/321/7512118

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

Informetrics & Scientometrics

1 Information and Measurement

  1. Information Revisited
  2. Framework for Information Exchange
  3. Measurement Techniques
  4. Informativeness
  5. Standardization of Measurement

2 Measure of Information

  1. Information and Entropy
  2. Shannon Information
  3. Probabilistic Information
  4. Properties of Shannon Information
  5. Derivation of Shannon Information Formula
  6. Normalization Condition
  7. Relating Semantic Value to Shannon Type Measures
  8. Other Shannon Type Measures of Information
  9. Semantic Information
  10. Fuzzy Information Measure
  11. Other Information Measures

3 Informetrics – Definition, Scope and Evolution

  1. Definitions
  2. Scope
  3. Evolution
  4. Summary

4 Sociology of Science and Scientometrics

  1. Sociology of Science
  2. Growth of Scientific Knowledge
  3. Social Organization in Research Areas
  4. Approaches of Scientometrics to Sociology of Science
  5. Models of Growth of Knowledge

5 Organizations Engaged in Scientometrics and Informetrics Studies

  1. Organizations Engaged in or Supporting Scientometrics/Informetrics Studies
  2. Websites
  3. Research Groups/Discussion Groups
  4. Periodical Publications
  5. Conferences/Seminars/Workshops/Congresses
  6. Individuals Engaged in the Study and Research in Scientometrics/Informetrics

6 Law of Scattering and its Applications

  1. Introduction
  2. Historical Account
  3. Bradford’s Law
  4. Verbal Form of Bradford’s Law
  5. Applications of Bradford’s Law
  6. Graphical Representation of Bradford’s Law
  7. Conditions for Bradford’s Law
  8. Falling Tail of Bradford Curve: The Groos Droop
  9. Ambiguity in Bradford’s Law
  10. Fitting Bibliographic Data to Bradford’s Law

7 Rank and Size Frequency Models

  1. Representations and Organization of Numerical Data
  2. Size – Frequency Approach
  3. Rank – Frequency Approach
  4. Size – Frequency Models
  5. Rank – Frequency Cumulative (Fractional) Models
  6. Rank – Frequency Cumulative (Non-Fractional) Models
  7. Rank – Frequency Non – Cumulative Models

8 Informetrics Phenomena

  1. Terminology and Historical Development
  2. Selected Laws of Bibliometrics and Informetrics
  3. Informetrics Phenomena in Science
  4. Practical Applications of Informetrics

9 Analysis of Library Related Data

  1. Necessity for Analytical Studies in Libraries
  2. Citation Counting: A Versatile Tool for Journal Selection
  3. An Alternative Method of Citation Analysis
  4. Selection of New Source Journals to Eliminate Bias Due to Country, and Language
  5. Weightage Formula to Correct Citation for Post-War Periodicals
  6. Three New Bibliometric Parameters to Re-Rank Scientific Periodicals
  7. Garfield’s Methods for Cito-Analytical Studies
  8. Librametric Analysis
  9. Bibliometric Analysis
  10. Informetrics
  11. Scientometrics: Its Genesis, Scope, Definition, and Applications

10 User Studies

  1. User Studies
  2. Questionnaire Method
  3. Interview Method
  4. Diary Method
  5. Observation Method
  6. Planning a Survey
  7. Classification and Tabulation of Data
  8. Analysis of Data
  9. Presentation of Results
  10. Important User Studies
  11. Application of User Studies

11 Laws of Scientific Productivity

  1. Scientific Productivity – Influencing Factors
  2. Scientific Productivity – Problems in Measurement
  3. Scientific Productivity – Distribution Characteristics
  4. Lotka’s Law
  5. Statistical Distributions or Models
  6. Application of Lotka’s Law
  7. Goodness-of-Fit Test

12 Growth and Obsolescence of Literature

  1. Growth of Literature
  2. Obsolescence of Literature
  3. Growth Vs Obsolescence of Literature

13 Science Indicators

  1. Indicators
  2. Towards Science Indicators
  3. Historical Aspects
  4. Functions of Science Indicators
  5. S&T Indicators for the Developing Countries
  6. Types of Indicators
  7. Validity and Reliability of Indicators
  8. Building S&T Indicators
  9. Literature Based Indicators
  10. Patent Indicators

14 Mapping of Science

  1. Cognitive Mapping
  2. Journal-to-journal Citation Maps
  3. Co-citation Maps
  4. Co-word Maps
  5. Co-classification Maps
  6. Descriptive Mapping

15 Elements of Statistics

  1. Data and Its Measurement
  2. Graphical Representation
  3. Measures of Central Tendency
  4. Measure of Variability
  5. Correlation and Regression

16 Probability Distributions and their Applications

  1. Probability – Definition
  2. Random Variables
  3. Joint Probability Distribution
  4. Conditional Probability Distribution
  5. Some Special Distributions
  6. Applications of Probability

17 Regression Analysis

  1. Simple Linear Regression
  2. Multiple Regression
  3. Stepwise Regression
  4. Regression with Qualitative Explanatory Variables

18 Cluster Analysis and Factor Analysis

  1. Introduction
  2. Cluster Analysis
  3. Factor Analysis
  4. Examples of Cluster and Factor Analysis