Bibliometric and scientometric studies generate enormous amounts of numerical information. A single research evaluation might track hundreds of institutions, dozens of subject fields, thousands of journals, and millions of citations. Making sense of all this requires a structured way to organize the numbers and a set of statistical tools to find patterns hidden inside them. This is where multidimensional data, along with techniques like cluster analysis and factor analysis, becomes central to the work of any information scientist. Understanding how raw bibliometric data is arranged into a matrix, and how that matrix is then analysed, is the foundation for almost every advanced study in scientometrics.

Table of Contents

What multidimensional data means in scientometrics

In everyday language, “dimension” refers to a measurement like length or width. In statistics, a dimension is simply a variable, a single attribute we measure. When we measure only one attribute, such as the number of papers published by a university, we have one-dimensional data. The moment we measure several attributes at once, like publications, citations, collaborations, and journals, the data becomes multidimensional.

Scientometric data is almost always multidimensional. Consider a study comparing universities. For each university we might record its output in physics, chemistry, biology, mathematics, and engineering. Each subject is a separate dimension. A single university is now described not by one number but by a set of numbers. When we line up many universities, each described by many subjects, the dataset becomes too complex to read by eye. We need a systematic structure to hold it. That structure is the matrix.

Multidimensional analysis is the branch of statistics that handles data where many variables are observed together. According to the discussion of multivariate methods at Pennsylvania State University, the underlying structure of real-world data is almost always simpler than the number of variables suggests, which is exactly why reduction techniques are so useful. The goal is to simplify without losing the meaningful signal.

Matrix representation of bibliometric data

A matrix is a rectangular arrangement of numbers in rows and columns. It is the standard way to store multidimensional data because it keeps every value in a fixed, addressable position. In bibliometrics, the matrix is the starting point for nearly all quantitative analysis.

Imagine a table where each row stands for an institution and each column stands for a research field. The cell where a particular row meets a particular column holds a value, for example, the number of papers that institution published in that field. This simple table is a data matrix. A large bibliometric study might have a matrix with 50 institutions as rows and 10 subject fields as columns, giving 500 individual data points arranged in a single organized grid.

The power of this arrangement is that it can represent many different kinds of relationships. The rows and columns do not always have to be institutions and fields. A matrix can be built where rows are documents and columns are journals, or rows are authors and columns are keywords. The bibliometrix R-tool documentation shows how the same matrix logic underlies co-citation, co-word, collaboration, and coupling networks. The structure stays the same even when the meaning of the rows and columns changes.

Row vectors and column vectors

Once data is in a matrix, we can read it in two directions, and each direction tells a different story.

A row vector is a single horizontal line of the matrix. It describes one object completely across all variables. If a row represents one university, the row vector is the full profile of that university across every research field, its publication count in physics, then chemistry, then biology, and so on. To compare the overall research personality of two universities, you compare their row vectors.

A column vector is a single vertical line of the matrix. It describes one variable across all objects. If a column represents the field of chemistry, the column vector lists how every university in the study performs in chemistry. To understand which subject is dominant across the whole sector, you study the column vectors.

This dual reading is important. The same matrix supports two different analytical questions. Analysing rows groups similar institutions together. Analysing columns groups similar fields together. In bibliometric software, a matrix is often described as a bipartite structure where, as explained in the bibliometrix analysis guide, the generic element takes a value of 1 when a connection exists between a row item and a column item and 0 when it does not. The column sum then tells you how frequently each variable appears, and sorting those sums reveals the most influential sources or fields.

The proximity matrix

There is a second kind of matrix that appears constantly in this work. Instead of holding raw counts, it holds measures of how close or how similar two items are. This is called a proximity matrix. A proximity matrix is usually square, with the same items listed on both the rows and the columns, and each cell records the similarity or distance between a pair of items.

As described in Springer’s chapter on cluster analysis, many analyses begin not with the raw data matrix but with this proximity matrix. The entries can represent either dissimilarity, which is a measure of distance, or similarity, which is a measure of association. The choice of how to measure nearness depends on the subject, the scale of measurement, and the type of variables involved. The proximity matrix is the bridge between raw data and the grouping methods that follow.

Analysing multidimensional data

Once the matrix is built, the real work begins. A large matrix is too dense to interpret directly, so we apply statistical methods that either group its contents or shrink its size. Three families of methods dominate scientometric practice.

Multivariate regression analysis

Regression studies the relationship between variables, specifically how one or more independent variables influence a dependent variable. Ordinary regression handles one outcome at a time. Multivariate regression extends this to situations where several outcomes are modelled together, which suits the multidimensional nature of bibliometric data.

In a scientometric setting, a researcher might ask whether the size of an institution, its funding, and its number of collaborations jointly predict its citation impact. Regression supplies a way to quantify these relationships and to test which inputs matter most. One caution noted in the Pennsylvania State University materials is that ordinary regression can struggle when the input variables are themselves correlated, a common problem when many bibliometric indicators move together. This is one reason the next two techniques are so valuable, because they can produce uncorrelated variables that make regression cleaner.

Cluster analysis

Cluster analysis is concerned with group identification. Its goal is to partition a set of observations into a number of unknown groups, called clusters, so that observations within a group are similar to one another while observations in different groups are not. Crucially, the number and nature of these groups are not known in advance. The method discovers them from the data itself.

In bibliometric work, cluster analysis is used to group institutions with similar research profiles, to identify clusters of related journals, or to reveal thematic clusters of keywords that define the structure of a field. The technique works on the proximity matrix, using the distances between items to decide which belong together. A widely used way to display the result is the dendrogram, a tree-like diagram that, as explained by EDUCBA’s comparison of clustering and factor analysis, shows how clusters merge at each step until they form a single group. Two of the most common algorithms are hierarchical clustering and k-means partitioning, the latter assigning points to a set number of clusters by minimising their distance from cluster centres.

It is worth noting what cluster analysis does not do. It classifies objects into groups, but it does not place them in a low-dimensional space for visual mapping. When the dimensionality of the data itself needs to be explored visually, a related technique called multidimensional scaling is used instead.

Data reduction techniques

The third family addresses a different problem. A matrix may have so many columns, so many variables, that interpretation becomes impossible. Data reduction techniques shrink the number of variables while keeping most of the information. The two most important methods here are principal component analysis and factor analysis.

Principal Component Analysis (PCA) transforms a set of correlated variables into a smaller set of uncorrelated variables called principal components. As described in ScienceDirect’s chapter on data reduction, these components are arranged so that the first few explain most of the variation in the original data. If twenty bibliometric indicators can be summarised by three components that together capture most of the variance, the analysis becomes far more manageable, and those components can even be used as inputs for later regression.

Factor analysis serves a closely related purpose but with a different emphasis. According to the documentation on factor analysis as a data reduction method, factor analysis assumes that observed variables are influenced by a smaller number of hidden, underlying variables called factors. Where PCA simply captures total variance, factor analysis tries to capture the shared variance, the common ground that links several variables together. In scientometrics, factor analysis is often applied to co-citation matrices to uncover the intellectual structure of a field, revealing the hidden research themes that connect groups of papers.

The mechanics of factor analysis involve a few standard steps. Eigenvalues are computed to decide how many factors to keep, and a common rule, summarised by GeeksforGeeks, is to retain factors with an eigenvalue greater than one. A scree plot helps confirm this choice by showing where the curve levels off. Finally, a rotation method such as Varimax is applied to make the factors easier to interpret.

How the techniques fit together

These methods are not rivals; they form a workflow. The matrix holds the data. A proximity matrix measures nearness. Cluster analysis sorts items into groups. Data reduction techniques like PCA and factor analysis cut down the number of variables and expose hidden structure. Regression then tests relationships, often using the cleaner, uncorrelated outputs that reduction produces. Together they turn a wall of raw numbers into an interpretable map of how science is organized, who collaborates with whom, and which themes hold a discipline together.

What do you think? If you were building a matrix to compare research institutions in your own state, would you choose research fields or journals as your columns, and how would that choice change the patterns you could detect? When a cluster analysis groups two institutions together, does that grouping reflect a real intellectual connection or just a statistical artefact of how the data was measured?

How useful was this post?

Click on a star to rate it!

Average rating 0 / 5. Vote count: 0

No votes so far! Be the first to rate this post.

We are sorry that this post was not useful for you!

Let us improve this post!

Tell us how we can improve this post?

References
  1. https://online.stat.psu.edu/stat505/book/export/html/670
  2. https://massimoaria.github.io/bibliometrix/
  3. https://warin.ca/shiny/bibliometrix/
  4. https://link.springer.com/content/pdf/10.1007%2F978-0-387-22771-9_9.pdf
  5. https://www.educba.com/cluster-analysis-vs-factor-analysis/
  6. https://www.sciencedirect.com/science/article/pii/S0922348708702053
  7. https://docs.tibco.com/data-science/GUID-DFFC90D7-F292-4474-A842-2D2CEAA47BE2.html
  8. https://www.geeksforgeeks.org/machine-learning/introduction-to-factor-analytics/

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

Informetrics & Scientometrics

1 Information and Measurement

  1. Information Revisited
  2. Framework for Information Exchange
  3. Measurement Techniques
  4. Informativeness
  5. Standardization of Measurement

2 Measure of Information

  1. Information and Entropy
  2. Shannon Information
  3. Probabilistic Information
  4. Properties of Shannon Information
  5. Derivation of Shannon Information Formula
  6. Normalization Condition
  7. Relating Semantic Value to Shannon Type Measures
  8. Other Shannon Type Measures of Information
  9. Semantic Information
  10. Fuzzy Information Measure
  11. Other Information Measures

3 Informetrics – Definition, Scope and Evolution

  1. Definitions
  2. Scope
  3. Evolution
  4. Summary

4 Sociology of Science and Scientometrics

  1. Sociology of Science
  2. Growth of Scientific Knowledge
  3. Social Organization in Research Areas
  4. Approaches of Scientometrics to Sociology of Science
  5. Models of Growth of Knowledge

5 Organizations Engaged in Scientometrics and Informetrics Studies

  1. Organizations Engaged in or Supporting Scientometrics/Informetrics Studies
  2. Websites
  3. Research Groups/Discussion Groups
  4. Periodical Publications
  5. Conferences/Seminars/Workshops/Congresses
  6. Individuals Engaged in the Study and Research in Scientometrics/Informetrics

6 Law of Scattering and its Applications

  1. Introduction
  2. Historical Account
  3. Bradford’s Law
  4. Verbal Form of Bradford’s Law
  5. Applications of Bradford’s Law
  6. Graphical Representation of Bradford’s Law
  7. Conditions for Bradford’s Law
  8. Falling Tail of Bradford Curve: The Groos Droop
  9. Ambiguity in Bradford’s Law
  10. Fitting Bibliographic Data to Bradford’s Law

7 Rank and Size Frequency Models

  1. Representations and Organization of Numerical Data
  2. Size – Frequency Approach
  3. Rank – Frequency Approach
  4. Size – Frequency Models
  5. Rank – Frequency Cumulative (Fractional) Models
  6. Rank – Frequency Cumulative (Non-Fractional) Models
  7. Rank – Frequency Non – Cumulative Models

8 Informetrics Phenomena

  1. Terminology and Historical Development
  2. Selected Laws of Bibliometrics and Informetrics
  3. Informetrics Phenomena in Science
  4. Practical Applications of Informetrics

9 Analysis of Library Related Data

  1. Necessity for Analytical Studies in Libraries
  2. Citation Counting: A Versatile Tool for Journal Selection
  3. An Alternative Method of Citation Analysis
  4. Selection of New Source Journals to Eliminate Bias Due to Country, and Language
  5. Weightage Formula to Correct Citation for Post-War Periodicals
  6. Three New Bibliometric Parameters to Re-Rank Scientific Periodicals
  7. Garfield’s Methods for Cito-Analytical Studies
  8. Librametric Analysis
  9. Bibliometric Analysis
  10. Informetrics
  11. Scientometrics: Its Genesis, Scope, Definition, and Applications

10 User Studies

  1. User Studies
  2. Questionnaire Method
  3. Interview Method
  4. Diary Method
  5. Observation Method
  6. Planning a Survey
  7. Classification and Tabulation of Data
  8. Analysis of Data
  9. Presentation of Results
  10. Important User Studies
  11. Application of User Studies

11 Laws of Scientific Productivity

  1. Scientific Productivity – Influencing Factors
  2. Scientific Productivity – Problems in Measurement
  3. Scientific Productivity – Distribution Characteristics
  4. Lotka’s Law
  5. Statistical Distributions or Models
  6. Application of Lotka’s Law
  7. Goodness-of-Fit Test

12 Growth and Obsolescence of Literature

  1. Growth of Literature
  2. Obsolescence of Literature
  3. Growth Vs Obsolescence of Literature

13 Science Indicators

  1. Indicators
  2. Towards Science Indicators
  3. Historical Aspects
  4. Functions of Science Indicators
  5. S&T Indicators for the Developing Countries
  6. Types of Indicators
  7. Validity and Reliability of Indicators
  8. Building S&T Indicators
  9. Literature Based Indicators
  10. Patent Indicators

14 Mapping of Science

  1. Cognitive Mapping
  2. Journal-to-journal Citation Maps
  3. Co-citation Maps
  4. Co-word Maps
  5. Co-classification Maps
  6. Descriptive Mapping

15 Elements of Statistics

  1. Data and Its Measurement
  2. Graphical Representation
  3. Measures of Central Tendency
  4. Measure of Variability
  5. Correlation and Regression

16 Probability Distributions and their Applications

  1. Probability – Definition
  2. Random Variables
  3. Joint Probability Distribution
  4. Conditional Probability Distribution
  5. Some Special Distributions
  6. Applications of Probability

17 Regression Analysis

  1. Simple Linear Regression
  2. Multiple Regression
  3. Stepwise Regression
  4. Regression with Qualitative Explanatory Variables

18 Cluster Analysis and Factor Analysis

  1. Introduction
  2. Cluster Analysis
  3. Factor Analysis
  4. Examples of Cluster and Factor Analysis