Every time you read that India ranks among the top three or four nations in scientific publications, or that gross R&D expenditure has tripled over a decade, you are looking at the output of carefully constructed science and technology indicators. These numbers do not appear by themselves. They are built, deliberately and methodically, from raw data gathered across thousands of laboratories, universities, and companies. Understanding how S&T indicators are built tells you a great deal about how research itself is monitored, funded, and steered. This post walks through what actually goes into constructing reliable indicators, the global standards that govern the process, and why a deep understanding of the science and technology system has to come first.
Table of Contents
- What science and technology indicators actually measure
- Why patents sit at the interface
- The Frascati Manual and the global standards behind the numbers
- The five core criteria for identifying R&D
- The Frascati family and its companion manuals
- How indicators are built step by step
- Start with a theoretical framework
- Select data, then normalise and aggregate
- Validate, test, and document everything
- Why understanding the system has to come first
- The role of classification systems
- Choosing the right primary indicator
- How indicators are built in the Indian context
- From survey to published indicator
- What the finished indicators reveal
- Common pitfalls when building indicators
What science and technology indicators actually measure
An S&T indicator is a statistical measure that captures some aspect of scientific and technological activity. Broadly, the field has settled on two dominant kinds of data. The first is input data, such as money spent on research and the number of people employed in it. The second is output data, mainly patents, publications, and citations. A widely cited analysis of S&T indicators notes that resource or science inputs (R&D spending measured in money or time) and technology or science outputs (patents, innovations, and bibliometric data like articles and citations) are the two types of data that dominate the field.
This input-output framing matters because it shapes how research is monitored. Input indicators tell you how much a country, sector, or institution is investing. Output indicators tell you what that investment produces. When the two are read together, policymakers can ask sharper questions about efficiency and return. Bibliometrics, in particular, has become a major source of output indicators. A foundational OECD report explains that bibliometrics rests on counting and statistically analysing scientific output in the form of articles, publications, citations, and patents to evaluate research activities, laboratories, scientists, and the performance of countries.
Why patents sit at the interface
Patents occupy an interesting position. They are sometimes treated as an output of science and sometimes as an input to the economy. Researchers have long noted that patents are indicators of inventions whose main function is the legal protection of intellectual property, but which can also serve as input indicators to the economy or output indicators of science depending on the model being used. This dual nature is exactly why building indicators requires care. The same piece of data can mean different things depending on the conceptual framework you place around it.
The Frascati Manual and the global standards behind the numbers
You cannot build reliable indicators without agreed definitions. If one country counts a software trial as research and another does not, their numbers cannot be compared. This is the problem the Frascati Manual was created to solve. In June 1963, OECD experts met with the National Experts on Science and Technology Indicators (NESTI) group in Frascati, Italy, and drafted the first version of what became officially titled The Proposed Standard Practice for Surveys of Research and Experimental Development.
The manual has grown into the international reference for the entire field. The OECD describes it as the internationally recognised methodology for collecting and using R&D statistics, providing definitions of key concepts, classifications, and data collection guidelines for statistics on the human and financial resources devoted to R&D. The current seventh edition was published in 2015 and was the product of more than 120 experts from nearly 40 countries and international organisations.
The five core criteria for identifying R&D
The Frascati Manual does not let anyone simply declare an activity to be research. To count as R&D, an activity must satisfy five criteria: it must be novel, creative, uncertain in its outcome, systematic, and transferable or reproducible. These criteria force a consistent boundary around what gets measured. The manual also defines R&D itself in famous terms, describing it as creative work undertaken on a systematic basis to increase the stock of knowledge and to devise new applications from it. Getting this definition right is the first building block of any indicator, because everything counted afterwards depends on it.
The Frascati family and its companion manuals
The Frascati Manual does not stand alone. Over the decades the OECD has produced a set of related methodological guides known as the “Frascati family.” Among the most important is the Oslo Manual, which deals specifically with innovation. The Frascati Manual itself was first published in 1963 under the title “Proposed standard for research and experimental development surveys” and has since been promoted by the OECD’s NESTI working group, integrating more than 50 countries. For anyone constructing indicators, knowing which manual governs which type of measurement is essential. R&D resources follow Frascati; innovation activity follows Oslo; bibliometric output draws on separate scientometric conventions.
How indicators are built step by step
Raw data and good definitions are necessary but not sufficient. The numbers must then be assembled into indicators that are valid, comparable, and transparent. When several individual measures are combined into a single summary figure, the result is called a composite indicator. Innovation rankings and competitiveness indices are common examples. The construction of these measures has its own carefully documented methodology.
The OECD’s guide to composite indicators is candid about the nature of the task. It states that composite indicators are much like mathematical or computational models, and their construction owes more to the craftsmanship of the modeller than to universally accepted scientific rules. In other words, there is no single mechanical recipe. The justification for an indicator lies in its fitness for the intended purpose and in its acceptance by peers.
Start with a theoretical framework
The first and most important step is conceptual, not statistical. The OECD handbook is blunt about this, warning that what is badly defined is likely to be badly measured, and that a sound theoretical framework is the starting point in constructing composite indicators. The framework must define the phenomenon being measured, break it into sub-components, and select individual indicators and weights that reflect their relative importance. Skipping this step produces numbers that look precise but mean very little.
Select data, then normalise and aggregate
Once the framework is fixed, the builder selects individual indicators, handles missing data, normalises values so that different units can be combined, and then aggregates them. Researchers reviewing the field note that variations can exist in the indicators chosen, in the normalisation and aggregation methods used, and in how indicators and results are validated, but that systematic and defensible approaches common to many tools do exist. Each of these choices changes the final number, which is why documenting them is not optional.
Validate, test, and document everything
A good indicator can withstand scrutiny. Much of the public scepticism toward indicators comes from poor transparency, so the OECD handbook places special emphasis on documentation and metadata, recommending that relevant documentation be prepared at the end of each phase to ensure the coherence of the whole process. Validation also includes testing how sensitive the final ranking is to the modelling choices made along the way. An indicator whose results swing wildly when a weight is changed is not yet ready for use.
Why understanding the system has to come first
This is the part that is easy to overlook. You cannot build accurate metrics for a system you do not understand. The intellectual organisation of science is not arbitrary. It evolved historically, and indicators have to respect that structure. One detailed account of the evolution of science indicators traces this back to wartime science, observing that the idea that scientific knowledge could be deliberately organised and controlled from a mission perspective was a result of World War II, after which a new science and technology policy had to be formulated under peacetime conditions.
Understanding the system means knowing how its parts fit together. The Frascati Manual organises measurement around the sectors that actually perform research, dealing primarily with measuring the expenditure and personnel resources devoted to R&D in the sectors performing it, namely higher education, government, business, and private non-profit organisations. If your survey design does not map onto these real performing sectors, the data you collect will not aggregate sensibly.
The role of classification systems
A second dimension of system understanding is classification. To compare research across fields, you need an agreed scheme for sorting it. The manual provides exactly this through its Fields of Science classification, which groups research into natural sciences, engineering and technology, medical and health sciences, agricultural sciences, social sciences, and humanities. Without such a classification, a publication count is just a heap of papers with no analytical structure. With it, you can track which disciplines are growing and where a country’s strengths lie.
Choosing the right primary indicator
System understanding also tells you which single number best captures a given dimension. For national R&D effort, the standard headline measure is Gross Domestic Expenditure on R&D, or GERD. Training material on the Frascati Manual confirms that GERD is the primary indicator for the R&D activities of a country and should be based on performer reports rather than on information from the source of R&D funds. That subtle rule, measuring at the performer rather than the funder, only makes sense once you understand how money flows through the research system. Build the indicator without that understanding and you risk double counting.
How indicators are built in the Indian context
India offers a clear example of this whole machinery in action. The body responsible is the National Science and Technology Management Information System (NSTMIS), a division of the Department of Science and Technology. It has been entrusted with building the information base on a continuous basis on resources devoted to scientific and technological activities, covering collection, collation, analysis, and dissemination.
Crucially, the construction does not happen in isolation from global standards. NSTMIS reports that it has been generating a database for the S&T sector since 1973 and, for the sake of international comparability, has adopted UNESCO and OECD guidelines on standards, concepts, and definitions for the collection of science statistics. This is the link between the global Frascati framework and a working national system. The definitions agreed in Italy decades ago directly shape the questionnaire sent to Indian laboratories today.
From survey to published indicator
The process runs through periodic national surveys. The national S&T survey that fed into a recent statistics report captured information from around 5,000 R&D organisations spanning the public sector, private sector, multinational companies, higher education, scientific and industrial research organisations, and NGOs, using a structured questionnaire built on international standardisation of S&T resources. Once the data is collected and validated, it becomes the basis for the published indicators that policymakers rely on. The output is explicitly designed to serve evidence-based decisions, functioning as a source book on S&T used by planners, researchers, and policymakers both within the country and internationally.
What the finished indicators reveal
The end result is a set of comparable figures that monitor research at the national level. These include GERD as a share of GDP, the size of the R&D workforce, patent filings, and publication output. International comparison is built in. One Indian statistics summary notes, drawing on OECD and UNESCO sources, that India spent 0.64% of its GDP on R&D in 2020-21, compared with figures such as 2.4% for China and 1.3% for Brazil among other developing economies. A number like that is only meaningful because every country in the comparison built its measure using the same underlying Frascati definitions. That is the whole point of the standards: they make the numbers speak to one another.
Common pitfalls when building indicators
Even with good standards, several traps recur. The first is treating an output count as a measure of quality. Bibliometric researchers caution that the quality of patents is not uniform, not all patents carry the same technical or economic significance, and it is therefore ill advised to compare patent applications across diverse technologies or industries. A raw count flattens these differences.
The second pitfall is opacity. As the OECD handbook stresses, a composite indicator built without transparent documentation invites distrust, however sophisticated its mathematics. The third is over-reliance. Even strong indicators are simplifications of a complex reality, and reading too much into a single number can mislead more than it informs. The discipline of building indicators is, in large part, the discipline of staying honest about what they can and cannot tell you.
What do you think? If you were designing a new indicator to capture the real-world impact of research rather than just its volume, what would you choose to measure, and which part of the science system would you need to understand most deeply to get it right?
References
- https://www.sciencedirect.com/science/article/abs/pii/S004873339800050X
- https://www.oecd.org/content/dam/oecd/en/publications/reports/1997/01/bibliometric-indicators-and-analysis-of-research-systems_g17a152e/208277770603.pdf
- https://arxiv.org/pdf/1203.1006
- https://en.wikipedia.org/wiki/Frascati_Manual
- https://www.oecd.org/en/about/projects/frascati-manual-development.html
- https://en.wikipedia.org/wiki/Research_and_development_intensity
- https://www.ovtt.org/en/resources/frascati-manual/
- https://www.oecd.org/content/dam/oecd/en/publications/reports/2008/08/handbook-on-constructing-composite-indicators-methodology-and-user-guide_g1gh9301/9789264043466-en.pdf
- https://www.oecd.org/content/dam/oecd/en/publications/reports/2005/08/handbook-on-constructing-composite-indicators_g17a16e3/533411815016.pdf
- https://www.nationalacademies.org/read/27317/chapter/5
- https://arxiv.org/pdf/0911.4298
- https://sesricdiag.blob.core.windows.net/oicstatcom/STI_Understandin_the_Frascati_Manual_EN.pdf
- https://dst.gov.in/scientific-programmes/scientific-engineering-research/national-science-technology-management-information-system-nstmis
- https://dst.gov.in/sites/default/files/Updated%20RD%20Statistics%20at%20a%20Glance%202022-23.pdf
- https://dst.gov.in/sites/default/files/Research%20and%20Deveopment%20Statistics%202019-20_0.pdf
- https://www.nstmis-dst.org/Pdfs/R&D%20Statistics%20at%20a%20Glance,%202022-23.pdf

Leave a Reply