Informetric data is often messy and high-dimensional. A single dataset might track how dozens of countries publish across hundreds of subfields, or how thousands of papers cite one another. Raw numbers like these are nearly impossible to interpret by eye. This is where cluster analysis and factor analysis step in. Both are multivariate techniques that reduce complexity and reveal hidden structure. Cluster analysis groups similar entities together, while factor analysis uncovers the underlying dimensions that explain why variables move together. To see how these methods actually work, it helps to walk through two concrete studies: one that clustered countries by their strengths in chemistry, and one that mapped the intellectual base of artificial intelligence research.
Table of Contents
- Cluster analysis: Grouping countries by chemistry specialization
- Why raw counts are not enough
- How the data gets clustered
- The role of the dendrogram
- Factor analysis: Mapping the intellectual base of AI
- From citations to a correlation structure
- How variables get grouped into factors
- Interpreting the factors
- Where clustering and factoring meet
- Why these techniques matter
Cluster analysis: Grouping countries by chemistry specialization
A classic application of cluster analysis in scientometrics comes from a cross-national study of research specialization in chemistry. The researchers compared the specialization profiles of eleven countries across nine subfields of chemistry, looking at both publication output and citation impact. The goal was simple to state but hard to answer without statistics: which countries resemble each other in what they are good at?
Why raw counts are not enough
The first problem the study had to solve was size. A large country naturally publishes more papers than a small one, and some subfields are simply bigger than others. Comparing raw publication counts would tell you mostly about country size, not about specialization. To get around this, the researchers used relative indicators called the activity index and the attractivity index. These normalize the data so that a country’s strength in, say, organic chemistry is measured relative to its own total output and relative to global activity in that subfield. Only after this normalization does the data describe genuine specialization rather than sheer volume.
How the data gets clustered
Once each country has a profile of relative strengths across the nine subfields, you can measure how similar any two countries are. Two countries with high values in the same subfields are “close”; two with opposite strengths are “far apart.” Cluster analysis turns these pairwise similarities into groups.
The study used hierarchical clustering, which builds groups step by step. Most hierarchical methods work in an agglomerative, bottom-up way: every country starts as its own cluster, and at each step the two closest clusters are merged, until everything ends up in a single group. The definition of “closest” depends on a choice called the linkage method. Single linkage uses the distance between the two nearest members of each cluster and tends to produce long, chain-like groups. Complete linkage uses the farthest members and produces compact groups. Ward’s method merges whichever pair produces the smallest increase in overall variance, which is why it often creates evenly sized, well-separated clusters and remains one of the most widely used choices in practice.
The distance metric matters too. Euclidean distance measures straight-line distance between profiles, while correlation-based distances capture whether two countries rise and fall together across subfields. Different combinations of metric and linkage can produce noticeably different groupings from the same data, so these choices are never trivial.
The role of the dendrogram
The output of hierarchical clustering is best understood through a dendrogram, a tree-like diagram that records the entire merging process.
Reading a dendrogram is straightforward once you know what the parts mean. Each country sits at the bottom as a “leaf.” As you move up, branches join together. The two legs of each U-shaped link show which clusters were merged, and the height of the link represents the distance between them. Countries that join low down are very similar; countries that only connect near the top are quite different. To actually decide on a set of clusters, you draw a horizontal line across the diagram, and everything joined below that line forms one cluster. Cut low and you get many small, tight groups; cut high and you get a few broad ones.
One honest limitation is worth noting. A dendrogram is a summary of the full distance matrix, and like any summary it loses some information, which is why it is most reliable at the bottom where it shows which items are genuinely close. In the chemistry study, the resulting tree grouped countries with similar research priorities together, and the authors interpreted these groupings as a reflection of how national science policies had shaped each country’s chemistry profile. A related study on European research institutions found that eight clusters were the optimal solution for classifying institutions by their publication profiles, showing how the same logic scales from countries down to individual organizations.
Factor analysis: Mapping the intellectual base of AI
Cluster analysis tells you which entities group together. Factor analysis answers a different question: what are the hidden dimensions underlying a set of variables? In scientometrics, it is the workhorse behind co-citation analysis, which is used to map the intellectual structure of a research field, including a fast-moving one like artificial intelligence.
From citations to a correlation structure
The core idea of co-citation analysis is that when two authors or two documents are repeatedly cited together, they probably belong to the same line of thought. By counting how often each pair is co-cited, you build a matrix that captures the relationships across an entire field. The goal is to identify the intellectual structure of a knowledge domain through the groupings formed by accumulated co-citation trails in the literature.
That co-citation matrix becomes the input for factor analysis. Factor analysis is a multivariate technique whose overall objective is data summarization and data reduction, identifying a small set of underlying dimensions, called factors, that explain the correlations among many variables. In an AI mapping study, the “variables” are the cited authors or key documents, and each resulting factor represents a research specialty or sub-theme within AI.
How variables get grouped into factors
Factor analysis works through the correlations in the matrix to produce eigenvalues and eigenvectors, with the aim of finding a few factors that explain most of the variation. A few technical terms make the process clearer:
Eigenvalues indicate how much variance each factor explains. Kaiser’s criterion keeps only factors with an eigenvalue greater than one, on the logic that a useful factor should explain more than a single variable’s worth of variance. A scree plot offers a visual alternative: you examine the graph of eigenvalues and stop factoring where the line levels off, the point where additional factors stop adding meaningful structure.
Factor loadings tell you how strongly each variable belongs to each factor. An author “loads” on a factor when their citation pattern is highly correlated with that factor. In a rotated component matrix, the loadings are by definition equal to the correlation of the factor with the variable, so a loading near 1.0 means a near-perfect fit. Rotation is an important final step. It adjusts the axes so that each variable loads cleanly on one factor rather than being spread thinly across several, which makes the factors far easier to interpret.
Interpreting the factors
The numbers alone do not name anything. The researcher has to read each factor by looking at which authors or documents load most heavily on it, then assign a meaningful label. A landmark author co-citation analysis of information retrieval research did exactly this, producing an interpretation of the field in terms of distinct factors, where authors loading on the first factor were identified as information-retrieval theory researchers focused on topics like probability of relevance and Boolean retrieval. Applied to AI, the same approach would surface factors corresponding to recognizable sub-communities, since scientometric studies of AI consistently find that genetic algorithms, neural networks, fuzzy logic, and machine learning are among the most widely used techniques in the literature. Each of these could emerge as a separate factor, and a variable that loads on more than one factor reveals an author or topic that bridges specialties.
Where clustering and factoring meet
These two techniques are often used side by side rather than in competition. Modern bibliometric software makes this routine. Tools built on the R bibliometrix package, for instance, offer factorial analysis methods such as correspondence analysis, alongside dendrogram charts that display the hierarchical relationships among topics. A typical AI mapping study might use factor analysis to identify the major research themes and a dendrogram to show how the keywords within those themes nest inside one another. The factor solution gives you the big dimensions; the dendrogram gives you the fine-grained tree.
Why these techniques matter
Both methods turn unmanageable tables of numbers into something a human can reason about. Cluster analysis is the right tool when you want to sort entities into groups, whether those entities are countries, institutions, or documents, and the dendrogram lets you choose how fine or coarse those groups should be. Factor analysis is the right tool when you want to discover the latent dimensions behind many correlated variables and use them to describe the intellectual makeup of a field. The chemistry example shows the grouping logic at the level of nations; the AI example shows the dimension-finding logic at the level of ideas. Together they form a large part of the analytical toolkit that scientometrics uses to make sense of the ever-growing volume of scientific literature.
It also helps to remember what each technique cannot do. Hierarchical clustering will not tell you the correct number of clusters; that decision rests on where you cut the tree and on your judgment about the field. Factor analysis will not name your factors; interpretation is always a human act of reading the loadings in light of subject knowledge. The mathematics organizes the evidence, but the meaning still comes from the analyst.
What do you think? If you were mapping a field you know well, would you trust a dendrogram’s grouping more than a factor solution, or would you want to see both before drawing conclusions? And how much should the choice of linkage method or rotation technique change how confident you feel about the patterns these analyses reveal?
References
- https://link.springer.com/article/10.1007/BF02016551
- https://scienceinsights.org/how-to-interpret-a-dendrogram-in-hierarchical-clustering/
- https://docs.scipy.org/doc/scipy/reference/generated/scipy.cluster.hierarchy.dendrogram.html
- https://www.displayr.com/what-is-dendrogram/
- https://link.springer.com/article/10.1007/s11192-009-0425-z
- https://arxiv.org/pdf/1002.1985
- https://www.sciencedirect.com/topics/medicine-and-dentistry/factor-analysis
- https://support.sas.com/resources/papers/proceedings/proceedings/sugi31/200-31.pdf
- https://arxiv.org/pdf/0911.3416
- https://www.researchgate.net/publication/238835213_Mapping_the_intellectual_structure_of_information_retrieval_studies_An_author_co-citation_analysis_1987-1997
- https://www.researchgate.net/publication/360273144_Scientometric_Analysis_of_the_Research_Paper_Output_on_Artificial_Intelligence_A_Study
- https://www.researchgate.net/publication/397525749_SCIENTOMETRIC_PROFILING_OF_LITERATURE_ON_'ARTIFICIAL_INTELLIGENCE_AND_LIBRARIES'_WITH_REFERENCE_TO_SCOPUS_BIBLIOGRAPHICAL_DATABASE

Leave a Reply