When researchers analyze thousands of articles, authors, or journals, the raw data is often too vast and tangled to interpret directly. Cluster analysis is the technique that brings order to this chaos. It is a data reduction method that groups objects so that items within a group are similar to one another and different from items in other groups. In informetrics and scientometrics, this is the backbone of techniques like co-citation analysis, bibliographic coupling, and the mapping of research fronts. This guide breaks down the two major families of clustering, the role of distance measurement, and the practical steps involved in running a clustering exercise.
Table of Contents
- What cluster analysis actually does
- Hierarchical cluster analysis
- Agglomerative clustering: the bottom-up approach
- Divisive clustering: the top-down approach
- How linkage methods decide what to merge
- Reading a dendrogram
- Non-hierarchical cluster analysis
- K-means clustering step by step
- Choosing between hierarchical and partitioning methods
- Key steps and distance measurement
- Euclidean distance
- Non-Euclidean measures
- The practical workflow
What cluster analysis actually does
Cluster analysis is an unsupervised learning technique, which means it finds structure in data without being told in advance what the groups should be. The algorithm itself discovers the natural groupings. The core principle is simple: similar objects lie close to one another in a data space, while dissimilar objects lie far apart.
In scientometric work, the “objects” being clustered could be documents, authors, keywords, or journals. The “features” used to measure similarity might be shared citations, co-occurring terms, or shared references. By grouping these objects, a researcher can reveal the thematic structure of a discipline, identify emerging subfields, or detect clusters of closely related research papers. This is why clustering is treated as a data reduction tool: it compresses a huge, messy dataset into a smaller number of interpretable groups.
There are two broad approaches to clustering, and understanding the difference between them is the foundation of everything that follows.
Hierarchical cluster analysis
Hierarchical clustering builds a hierarchy of clusters, producing a tree-like structure rather than a single flat set of groups. A major advantage of this approach is that you do not need to specify the number of clusters in advance. Instead, the method reveals groupings at many levels of granularity, and you decide afterwards where to cut the tree.
There are two opposite directions in which this hierarchy can be built: agglomerative and divisive.
Agglomerative clustering: the bottom-up approach
Agglomerative clustering works in a “bottom-up” manner. Each object starts out as its own single-member cluster. At each step, the two clusters that are most similar are merged into one larger cluster. This continues until every object belongs to a single, all-encompassing cluster.
The general procedure can be broken into clear stages:
Step 1 – Start with singletons. Treat each data point as its own cluster. If you have 500 papers, you begin with 500 clusters.
Step 2 – Compute the distance matrix. Calculate the dissimilarity between every pair of clusters or points using a chosen distance measure such as Euclidean distance.
Step 3 – Merge the closest pair. Find the two clusters with the smallest dissimilarity and combine them into a new cluster.
Step 4 – Update and repeat. Recalculate the distances between the new cluster and all remaining clusters, then repeat the merging step until only one cluster remains.
This is the more commonly used of the two hierarchical approaches. One caution is that it can be computationally expensive for large datasets, because the dissimilarity matrix must be updated at every iteration.
Divisive clustering: the top-down approach
Divisive clustering does exactly the opposite. It begins with all objects in a single cluster at the root and then iteratively splits clusters into smaller ones until each object stands alone. This method is also known as DIANA (Divisive Analysis).
Divisive methods are less common in practice. A useful rule of thumb is that agglomerative clustering is better at identifying small clusters, while divisive clustering is better at identifying large clusters. Because the divisive approach considers the global structure of the data before splitting, it can be useful when the goal is to identify a few large, distinct groups first.
How linkage methods decide what to merge
When clusters contain more than one point, you need a rule to define the distance between two clusters. This rule is called the linkage criterion, and the choice strongly affects the final result. The four most widely used methods are single, complete, average, and Ward’s linkage.
Single linkage (nearest neighbour) defines the distance between two clusters as the shortest distance between any two of their points. It tends to produce long, chain-like clusters and is sensitive to noise.
Complete linkage (farthest neighbour) uses the longest distance between any two points in the two clusters. It produces more compact, well-separated clusters and avoids the chaining effect.
Average linkage uses the average distance between all pairs of points across the two clusters. It strikes a balance between the single and complete methods and is less sensitive to outliers.
Ward’s method is different from the others. Instead of measuring distance between points directly, it merges the pair of clusters that produces the smallest increase in total within-cluster variance. This tends to create compact clusters of roughly similar size. In scientometric studies, Ward’s method is often the preferred choice. In one comparison of bibliographically coupled information-science papers, complete and single linkage failed to give acceptable results, while Ward clustering produced a stable and clear thematic structure.
Reading a dendrogram
The result of hierarchical clustering is visualised as a dendrogram, a tree diagram where the root represents the entire dataset and the leaves represent individual objects. The height at which two branches join shows the dissimilarity at which those clusters were merged. The longer the branch, the less similar the joined clusters are.
To extract a specific number of groups, you cut the dendrogram with a horizontal line at a chosen height. Cutting at a lower height produces more groups with greater internal similarity, while cutting higher produces fewer, broader groups. The practical goal is to find a cut that yields the most interpretable partition. A common technique is to draw the horizontal line where there is the largest vertical gap with no merging, and count the branches it crosses to determine the optimal number of clusters.
Non-hierarchical cluster analysis
Non-hierarchical methods, also called partitioning methods, take a different route. They group objects into a number of clusters that is fixed in advance by the researcher. There is no tree and no hierarchy; the dataset is divided directly into non-overlapping groups.
A key practical advantage is efficiency. Because these methods do not need to calculate and store a full dissimilarity matrix at every step, they apply quickly to very large datasets. This makes partitioning methods attractive when working with the enormous citation and text datasets common in scientometrics. The most widely used partitioning method is K-means clustering.
K-means clustering step by step
K-means, a method credited to MacQueen in 1967, assigns each item to the cluster with the nearest centroid (mean). A centroid is simply the centre point of a cluster, calculated as the mean of all points within it. The algorithm proceeds as follows:
Step 1 – Choose K. Decide how many clusters you want and initialise K centroids, often by random selection.
Step 2 – Assign points. Assign each data point to the nearest centroid based on a distance measure.
Step 3 – Recompute centroids. Recalculate each centroid as the mean of all points currently assigned to that cluster.
Step 4 – Repeat. Repeat the assignment and recomputation steps until no points change clusters, a state known as convergence.
The overall aim is to minimise the total distance between points and their assigned centroid, which is equivalent to minimising the within-cluster variance. Unlike hierarchical clustering, K-means allows reallocation, meaning an object can move from one cluster to another in later iterations.
K-means is popular because it is simple, fast, and scalable, with a running time that grows linearly with the number of points. However, it has a notable weakness: it is non-deterministic. Because it starts from random centroids, running it twice on the same data can give different results. The choice of initial centres strongly influences the final clusters, so analysts often run the algorithm several times and compare the outcomes. Related partitioning methods such as K-medoids and Fuzzy C-means address some of these limitations.
Choosing between hierarchical and partitioning methods
The two families serve different needs. Hierarchical methods produce richer, more detailed information suitable for browsing and exploratory analysis, but they struggle with very large datasets. K-means and its variants are far more efficient and provide sufficient structure for most large-scale tasks, but they require you to fix the number of clusters beforehand. Many scientometric workflows use both: a partitioning method to handle scale, and a hierarchical view to interpret the relationships between the resulting groups.
Key steps and distance measurement
Regardless of which algorithm you choose, every clustering exercise rests on a measure of how far apart two objects are. The choice of distance measure is one of the most consequential decisions in the whole process, because it has a strong influence on the clustering results.
Euclidean distance
Euclidean distance is the classical and most common measure, and it is the default in most clustering software. It is simply the straight-line distance between two points, the familiar “as the crow flies” measurement. For two vectors, it is the square root of the sum of the squared differences across all features. Because it squares the differences, Euclidean distance gives more weight to large gaps, and observations with high feature values tend to be grouped together.
Non-Euclidean measures
Several alternatives exist for situations where straight-line distance is not the best fit. Manhattan distance, also called city-block distance, sums the absolute differences across coordinates, like a taxi navigating a grid of streets. Minkowski distance is a generalisation that includes both Euclidean and Manhattan as special cases through a tunable parameter.
In scientometrics, correlation-based and cosine-style measures are especially important. Correlation-based distance considers two objects similar if their feature profiles are highly correlated, even if their raw values differ widely. This matters when comparing documents or authors whose citation patterns have the same shape but different magnitudes. The cosine measure, closely related, focuses on the angle between two vectors and is heavily used in document clustering where term-frequency vectors of different lengths must be compared.
Experimental comparisons show there is no universally superior metric. Some studies find that Euclidean distance converges in fewer iterations, while others report better-quality results from Manhattan distance on particular datasets. The right choice depends on the nature of your data and the questions you are asking. One technical note worth remembering: centroid-based and Ward’s methods are correctly defined only when Euclidean distance is used.
The practical workflow
Pulling these ideas together, a typical cluster analysis follows a clear sequence. First, select and prepare the variables, often standardising them so that no single feature dominates simply because of its scale. Second, choose an appropriate distance measure for the data. Third, select a clustering method, whether hierarchical with a chosen linkage rule or a partitioning method such as K-means. Fourth, run the algorithm and determine the number of clusters, either by cutting a dendrogram or by setting K in advance. Finally, validate and interpret the clusters using domain knowledge, because a statistically valid grouping is only useful if it makes sense in the context of the research field.
This last stage is where scientometric expertise meets statistical technique. A cluster of papers is only a “research front” or “thematic group” if a knowledgeable analyst can recognise it as one. The algorithm reduces the data; the researcher gives it meaning.
What do you think? If you were mapping the intellectual structure of your own discipline, would you trust an efficient K-means partition with a pre-set number of themes, or would you prefer the exploratory freedom of a dendrogram even at the cost of speed? And how much should the choice of distance measure depend on whether you are clustering citations, keywords, or full documents?
References
- https://www.geeksforgeeks.org/machine-learning/hierarchical-clustering/
- https://www.geeksforgeeks.org/difference-between-hierarchical-and-non-hierarchical-clustering/
- https://www.datanovia.com/en/lessons/agglomerative-hierarchical-clustering/
- https://www.linkedin.com/pulse/comprehensive-overview-hierarchical-clustering-divisive-nandini-verma-2kxff
- https://medium.com/@iqra.bismi/different-linkage-methods-used-in-hierarchical-clustering-627bde3787e8
- https://codesignal.com/learn/courses/hierarchical-clustering-deep-dive/lessons/understanding-linkage-criteria-in-hierarchical-clustering
- https://arxiv.org/pdf/1107.5266
- https://en.wikipedia.org/wiki/Hierarchical_clustering
- https://arxiv.org/pdf/2112.01372
- https://arxiv.org/pdf/1908.07390
- http://sfb649.wiwi.hu-berlin.de/fedc_homepage/xplore/tutorials/xaghtmlnode54.html
- https://www.ibm.com/think/topics/k-means-clustering
- https://en.wikipedia.org/wiki/Document_clustering
- https://www.datanovia.com/en/lessons/clustering-distance-measures/
- https://arxiv.org/pdf/1112.5593
- https://iopscience.iop.org/article/10.1088/1742-6596/1566/1/012058/pdf
- https://docs.scipy.org/doc/scipy/reference/generated/scipy.cluster.hierarchy.linkage.html

Leave a Reply