When researchers analyze thousands of articles, authors, or journals, the raw data is often too vast and tangled to interpret directly. Cluster analysis is the technique that brings order to this chaos. It is a data reduction method that groups objects so that items within a group are similar to one another and different from items in other groups. In informetrics and scientometrics, this is the backbone of techniques like co-citation analysis, bibliographic coupling, and the mapping of research fronts. This guide breaks down the two major families of clustering, the role of distance measurement, and the practical steps involved in running a clustering exercise.

Table of Contents

What cluster analysis actually does

Cluster analysis is an unsupervised learning technique, which means it finds structure in data without being told in advance what the groups should be. The algorithm itself discovers the natural groupings. The core principle is simple: similar objects lie close to one another in a data space, while dissimilar objects lie far apart.

In scientometric work, the “objects” being clustered could be documents, authors, keywords, or journals. The “features” used to measure similarity might be shared citations, co-occurring terms, or shared references. By grouping these objects, a researcher can reveal the thematic structure of a discipline, identify emerging subfields, or detect clusters of closely related research papers. This is why clustering is treated as a data reduction tool: it compresses a huge, messy dataset into a smaller number of interpretable groups.

There are two broad approaches to clustering, and understanding the difference between them is the foundation of everything that follows.

Hierarchical cluster analysis

Hierarchical clustering builds a hierarchy of clusters, producing a tree-like structure rather than a single flat set of groups. A major advantage of this approach is that you do not need to specify the number of clusters in advance. Instead, the method reveals groupings at many levels of granularity, and you decide afterwards where to cut the tree.

There are two opposite directions in which this hierarchy can be built: agglomerative and divisive.

Agglomerative clustering: the bottom-up approach

Agglomerative clustering works in a “bottom-up” manner. Each object starts out as its own single-member cluster. At each step, the two clusters that are most similar are merged into one larger cluster. This continues until every object belongs to a single, all-encompassing cluster.

The general procedure can be broken into clear stages:

Step 1 – Start with singletons. Treat each data point as its own cluster. If you have 500 papers, you begin with 500 clusters.

Step 2 – Compute the distance matrix. Calculate the dissimilarity between every pair of clusters or points using a chosen distance measure such as Euclidean distance.

Step 3 – Merge the closest pair. Find the two clusters with the smallest dissimilarity and combine them into a new cluster.

Step 4 – Update and repeat. Recalculate the distances between the new cluster and all remaining clusters, then repeat the merging step until only one cluster remains.

This is the more commonly used of the two hierarchical approaches. One caution is that it can be computationally expensive for large datasets, because the dissimilarity matrix must be updated at every iteration.

Divisive clustering: the top-down approach

Divisive clustering does exactly the opposite. It begins with all objects in a single cluster at the root and then iteratively splits clusters into smaller ones until each object stands alone. This method is also known as DIANA (Divisive Analysis).

Divisive methods are less common in practice. A useful rule of thumb is that agglomerative clustering is better at identifying small clusters, while divisive clustering is better at identifying large clusters. Because the divisive approach considers the global structure of the data before splitting, it can be useful when the goal is to identify a few large, distinct groups first.

How linkage methods decide what to merge

When clusters contain more than one point, you need a rule to define the distance between two clusters. This rule is called the linkage criterion, and the choice strongly affects the final result. The four most widely used methods are single, complete, average, and Ward’s linkage.

Single linkage (nearest neighbour) defines the distance between two clusters as the shortest distance between any two of their points. It tends to produce long, chain-like clusters and is sensitive to noise.

Complete linkage (farthest neighbour) uses the longest distance between any two points in the two clusters. It produces more compact, well-separated clusters and avoids the chaining effect.

Average linkage uses the average distance between all pairs of points across the two clusters. It strikes a balance between the single and complete methods and is less sensitive to outliers.

Ward’s method is different from the others. Instead of measuring distance between points directly, it merges the pair of clusters that produces the smallest increase in total within-cluster variance. This tends to create compact clusters of roughly similar size. In scientometric studies, Ward’s method is often the preferred choice. In one comparison of bibliographically coupled information-science papers, complete and single linkage failed to give acceptable results, while Ward clustering produced a stable and clear thematic structure.

Reading a dendrogram

The result of hierarchical clustering is visualised as a dendrogram, a tree diagram where the root represents the entire dataset and the leaves represent individual objects. The height at which two branches join shows the dissimilarity at which those clusters were merged. The longer the branch, the less similar the joined clusters are.

To extract a specific number of groups, you cut the dendrogram with a horizontal line at a chosen height. Cutting at a lower height produces more groups with greater internal similarity, while cutting higher produces fewer, broader groups. The practical goal is to find a cut that yields the most interpretable partition. A common technique is to draw the horizontal line where there is the largest vertical gap with no merging, and count the branches it crosses to determine the optimal number of clusters.

Non-hierarchical cluster analysis

Non-hierarchical methods, also called partitioning methods, take a different route. They group objects into a number of clusters that is fixed in advance by the researcher. There is no tree and no hierarchy; the dataset is divided directly into non-overlapping groups.

A key practical advantage is efficiency. Because these methods do not need to calculate and store a full dissimilarity matrix at every step, they apply quickly to very large datasets. This makes partitioning methods attractive when working with the enormous citation and text datasets common in scientometrics. The most widely used partitioning method is K-means clustering.

K-means clustering step by step

K-means, a method credited to MacQueen in 1967, assigns each item to the cluster with the nearest centroid (mean). A centroid is simply the centre point of a cluster, calculated as the mean of all points within it. The algorithm proceeds as follows:

Step 1 – Choose K. Decide how many clusters you want and initialise K centroids, often by random selection.

Step 2 – Assign points. Assign each data point to the nearest centroid based on a distance measure.

Step 3 – Recompute centroids. Recalculate each centroid as the mean of all points currently assigned to that cluster.

Step 4 – Repeat. Repeat the assignment and recomputation steps until no points change clusters, a state known as convergence.

The overall aim is to minimise the total distance between points and their assigned centroid, which is equivalent to minimising the within-cluster variance. Unlike hierarchical clustering, K-means allows reallocation, meaning an object can move from one cluster to another in later iterations.

K-means is popular because it is simple, fast, and scalable, with a running time that grows linearly with the number of points. However, it has a notable weakness: it is non-deterministic. Because it starts from random centroids, running it twice on the same data can give different results. The choice of initial centres strongly influences the final clusters, so analysts often run the algorithm several times and compare the outcomes. Related partitioning methods such as K-medoids and Fuzzy C-means address some of these limitations.

Choosing between hierarchical and partitioning methods

The two families serve different needs. Hierarchical methods produce richer, more detailed information suitable for browsing and exploratory analysis, but they struggle with very large datasets. K-means and its variants are far more efficient and provide sufficient structure for most large-scale tasks, but they require you to fix the number of clusters beforehand. Many scientometric workflows use both: a partitioning method to handle scale, and a hierarchical view to interpret the relationships between the resulting groups.

Key steps and distance measurement

Regardless of which algorithm you choose, every clustering exercise rests on a measure of how far apart two objects are. The choice of distance measure is one of the most consequential decisions in the whole process, because it has a strong influence on the clustering results.

Euclidean distance

Euclidean distance is the classical and most common measure, and it is the default in most clustering software. It is simply the straight-line distance between two points, the familiar “as the crow flies” measurement. For two vectors, it is the square root of the sum of the squared differences across all features. Because it squares the differences, Euclidean distance gives more weight to large gaps, and observations with high feature values tend to be grouped together.

Non-Euclidean measures

Several alternatives exist for situations where straight-line distance is not the best fit. Manhattan distance, also called city-block distance, sums the absolute differences across coordinates, like a taxi navigating a grid of streets. Minkowski distance is a generalisation that includes both Euclidean and Manhattan as special cases through a tunable parameter.

In scientometrics, correlation-based and cosine-style measures are especially important. Correlation-based distance considers two objects similar if their feature profiles are highly correlated, even if their raw values differ widely. This matters when comparing documents or authors whose citation patterns have the same shape but different magnitudes. The cosine measure, closely related, focuses on the angle between two vectors and is heavily used in document clustering where term-frequency vectors of different lengths must be compared.

Experimental comparisons show there is no universally superior metric. Some studies find that Euclidean distance converges in fewer iterations, while others report better-quality results from Manhattan distance on particular datasets. The right choice depends on the nature of your data and the questions you are asking. One technical note worth remembering: centroid-based and Ward’s methods are correctly defined only when Euclidean distance is used.

The practical workflow

Pulling these ideas together, a typical cluster analysis follows a clear sequence. First, select and prepare the variables, often standardising them so that no single feature dominates simply because of its scale. Second, choose an appropriate distance measure for the data. Third, select a clustering method, whether hierarchical with a chosen linkage rule or a partitioning method such as K-means. Fourth, run the algorithm and determine the number of clusters, either by cutting a dendrogram or by setting K in advance. Finally, validate and interpret the clusters using domain knowledge, because a statistically valid grouping is only useful if it makes sense in the context of the research field.

This last stage is where scientometric expertise meets statistical technique. A cluster of papers is only a “research front” or “thematic group” if a knowledgeable analyst can recognise it as one. The algorithm reduces the data; the researcher gives it meaning.

What do you think? If you were mapping the intellectual structure of your own discipline, would you trust an efficient K-means partition with a pre-set number of themes, or would you prefer the exploratory freedom of a dendrogram even at the cost of speed? And how much should the choice of distance measure depend on whether you are clustering citations, keywords, or full documents?

How useful was this post?

Click on a star to rate it!

Average rating 0 / 5. Vote count: 0

No votes so far! Be the first to rate this post.

We are sorry that this post was not useful for you!

Let us improve this post!

Tell us how we can improve this post?

References
  1. https://www.geeksforgeeks.org/machine-learning/hierarchical-clustering/
  2. https://www.geeksforgeeks.org/difference-between-hierarchical-and-non-hierarchical-clustering/
  3. https://www.datanovia.com/en/lessons/agglomerative-hierarchical-clustering/
  4. https://www.linkedin.com/pulse/comprehensive-overview-hierarchical-clustering-divisive-nandini-verma-2kxff
  5. https://medium.com/@iqra.bismi/different-linkage-methods-used-in-hierarchical-clustering-627bde3787e8
  6. https://codesignal.com/learn/courses/hierarchical-clustering-deep-dive/lessons/understanding-linkage-criteria-in-hierarchical-clustering
  7. https://arxiv.org/pdf/1107.5266
  8. https://en.wikipedia.org/wiki/Hierarchical_clustering
  9. https://arxiv.org/pdf/2112.01372
  10. https://arxiv.org/pdf/1908.07390
  11. http://sfb649.wiwi.hu-berlin.de/fedc_homepage/xplore/tutorials/xaghtmlnode54.html
  12. https://www.ibm.com/think/topics/k-means-clustering
  13. https://en.wikipedia.org/wiki/Document_clustering
  14. https://www.datanovia.com/en/lessons/clustering-distance-measures/
  15. https://arxiv.org/pdf/1112.5593
  16. https://iopscience.iop.org/article/10.1088/1742-6596/1566/1/012058/pdf
  17. https://docs.scipy.org/doc/scipy/reference/generated/scipy.cluster.hierarchy.linkage.html

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

Informetrics & Scientometrics

1 Information and Measurement

  1. Information Revisited
  2. Framework for Information Exchange
  3. Measurement Techniques
  4. Informativeness
  5. Standardization of Measurement

2 Measure of Information

  1. Information and Entropy
  2. Shannon Information
  3. Probabilistic Information
  4. Properties of Shannon Information
  5. Derivation of Shannon Information Formula
  6. Normalization Condition
  7. Relating Semantic Value to Shannon Type Measures
  8. Other Shannon Type Measures of Information
  9. Semantic Information
  10. Fuzzy Information Measure
  11. Other Information Measures

3 Informetrics – Definition, Scope and Evolution

  1. Definitions
  2. Scope
  3. Evolution
  4. Summary

4 Sociology of Science and Scientometrics

  1. Sociology of Science
  2. Growth of Scientific Knowledge
  3. Social Organization in Research Areas
  4. Approaches of Scientometrics to Sociology of Science
  5. Models of Growth of Knowledge

5 Organizations Engaged in Scientometrics and Informetrics Studies

  1. Organizations Engaged in or Supporting Scientometrics/Informetrics Studies
  2. Websites
  3. Research Groups/Discussion Groups
  4. Periodical Publications
  5. Conferences/Seminars/Workshops/Congresses
  6. Individuals Engaged in the Study and Research in Scientometrics/Informetrics

6 Law of Scattering and its Applications

  1. Introduction
  2. Historical Account
  3. Bradford’s Law
  4. Verbal Form of Bradford’s Law
  5. Applications of Bradford’s Law
  6. Graphical Representation of Bradford’s Law
  7. Conditions for Bradford’s Law
  8. Falling Tail of Bradford Curve: The Groos Droop
  9. Ambiguity in Bradford’s Law
  10. Fitting Bibliographic Data to Bradford’s Law

7 Rank and Size Frequency Models

  1. Representations and Organization of Numerical Data
  2. Size – Frequency Approach
  3. Rank – Frequency Approach
  4. Size – Frequency Models
  5. Rank – Frequency Cumulative (Fractional) Models
  6. Rank – Frequency Cumulative (Non-Fractional) Models
  7. Rank – Frequency Non – Cumulative Models

8 Informetrics Phenomena

  1. Terminology and Historical Development
  2. Selected Laws of Bibliometrics and Informetrics
  3. Informetrics Phenomena in Science
  4. Practical Applications of Informetrics

9 Analysis of Library Related Data

  1. Necessity for Analytical Studies in Libraries
  2. Citation Counting: A Versatile Tool for Journal Selection
  3. An Alternative Method of Citation Analysis
  4. Selection of New Source Journals to Eliminate Bias Due to Country, and Language
  5. Weightage Formula to Correct Citation for Post-War Periodicals
  6. Three New Bibliometric Parameters to Re-Rank Scientific Periodicals
  7. Garfield’s Methods for Cito-Analytical Studies
  8. Librametric Analysis
  9. Bibliometric Analysis
  10. Informetrics
  11. Scientometrics: Its Genesis, Scope, Definition, and Applications

10 User Studies

  1. User Studies
  2. Questionnaire Method
  3. Interview Method
  4. Diary Method
  5. Observation Method
  6. Planning a Survey
  7. Classification and Tabulation of Data
  8. Analysis of Data
  9. Presentation of Results
  10. Important User Studies
  11. Application of User Studies

11 Laws of Scientific Productivity

  1. Scientific Productivity – Influencing Factors
  2. Scientific Productivity – Problems in Measurement
  3. Scientific Productivity – Distribution Characteristics
  4. Lotka’s Law
  5. Statistical Distributions or Models
  6. Application of Lotka’s Law
  7. Goodness-of-Fit Test

12 Growth and Obsolescence of Literature

  1. Growth of Literature
  2. Obsolescence of Literature
  3. Growth Vs Obsolescence of Literature

13 Science Indicators

  1. Indicators
  2. Towards Science Indicators
  3. Historical Aspects
  4. Functions of Science Indicators
  5. S&T Indicators for the Developing Countries
  6. Types of Indicators
  7. Validity and Reliability of Indicators
  8. Building S&T Indicators
  9. Literature Based Indicators
  10. Patent Indicators

14 Mapping of Science

  1. Cognitive Mapping
  2. Journal-to-journal Citation Maps
  3. Co-citation Maps
  4. Co-word Maps
  5. Co-classification Maps
  6. Descriptive Mapping

15 Elements of Statistics

  1. Data and Its Measurement
  2. Graphical Representation
  3. Measures of Central Tendency
  4. Measure of Variability
  5. Correlation and Regression

16 Probability Distributions and their Applications

  1. Probability – Definition
  2. Random Variables
  3. Joint Probability Distribution
  4. Conditional Probability Distribution
  5. Some Special Distributions
  6. Applications of Probability

17 Regression Analysis

  1. Simple Linear Regression
  2. Multiple Regression
  3. Stepwise Regression
  4. Regression with Qualitative Explanatory Variables

18 Cluster Analysis and Factor Analysis

  1. Introduction
  2. Cluster Analysis
  3. Factor Analysis
  4. Examples of Cluster and Factor Analysis