An indexing system is only as good as the results it produces. You can build the most elaborate classification scheme or the most detailed subject headings, but if a searcher cannot find what they need, the effort is wasted. This is why evaluation sits at the heart of information work. Measuring the effectiveness of an indexing system tells us whether documents are described well enough to be retrieved when they matter and ignored when they do not. This post walks through why evaluation matters, the metrics that quantify performance, the landmark experiments that shaped the field, and the practical techniques that make indexing better.

Table of Contents

Why evaluate an indexing system?

Indexing exists to serve retrieval. When a user submits a query, the indexing system decides which documents surface and which stay hidden. If indexing is sloppy, relevant material gets buried and irrelevant material clutters the results. Evaluation is the disciplined way of checking whether the system is doing its real job: connecting people with the information they need.

The components needed for a controlled evaluation are now standard. You need a collection of documents, a set of queries or requests, and a set of relevance judgments that say which documents should be considered relevant to each query. Together, these three components form what is called a test collection, and they let us compare different indexing approaches on equal footing.

Evaluation also exposes hidden weaknesses. An indexing scheme might work beautifully for general topics but fail on specialised ones. It might depend heavily on the skill of individual indexers, producing inconsistent results. Without measurement, these problems remain invisible until users complain. Systematic evaluation turns vague dissatisfaction into specific, fixable failures.

Key performance metrics

Two measures dominate the evaluation of indexing and retrieval: recall and precision. They were first used by Allen Kent and colleagues in the 1950s and have remained central ever since. Both depend on a clear notion of which documents are relevant to a query.

Recall

Recall measures completeness. It is the proportion of all relevant documents in the collection that the system actually retrieves. If a database holds 100 documents relevant to a query and the system returns 70 of them, recall is 70%. High recall means the system is good at not missing things. This matters most in exhaustive searches, such as legal discovery or systematic literature reviews, where missing a single relevant item can be costly.

Precision

Precision measures accuracy. It is the proportion of retrieved documents that are actually relevant. If a search returns 100 documents and 70 of them are relevant, precision is 70%. High precision means the user is not forced to wade through irrelevant material. This matters most in quick reference work and decision support, where users want focused answers without information overload.

The recall-precision trade-off

Recall and precision usually pull in opposite directions. Improving one tends to harm the other. To raise recall, a system broadens its search and pulls in more documents, which inevitably includes irrelevant ones and lowers precision. To raise precision, a system tightens its criteria, which risks discarding relevant documents and lowering recall. The Cranfield researchers found this relationship to be effectively inescapable: within a system’s normal operating range, a 1% gain in precision tended to come with a 3% drop in recall.

This is why there is no single “best” setting. The right balance depends on the user’s need. A patent examiner cannot afford to miss prior art and will tolerate lower precision for higher recall. A busy clinician wants a handful of trustworthy results and will accept lower recall for higher precision. Good evaluation reports both numbers rather than collapsing them into a misleading single score.

The role of relevance

Every metric depends on relevance, and relevance is harder to pin down than it looks. It is subjective, context-specific, and can vary between assessors and even for the same person at different times. Some evaluations therefore use graded relevance, recording whether a document is of major, minor, or no value rather than a simple yes or no. Because relevance judgments anchor every other measurement, the care taken in producing them largely determines how trustworthy an evaluation is.

Major retrieval experiments

The metrics above did not appear in a vacuum. They were forged through a series of large experiments that defined how the field tests its systems. Three stand out.

The Cranfield experiments

The Cranfield experiments, conducted by Cyril W. Cleverdon at the College of Aeronautics between 1958 and 1966, are widely regarded as the birth of formal information retrieval evaluation. They were designed to test the efficiency of indexing systems in a controlled, laboratory-like setting.

The first phase, Cranfield I, compared four different indexing methods, including the Universal Decimal Classification, an alphabetical subject catalogue, a faceted scheme, and the Uniterm system of coordinate indexing. A striking result emerged: all four methods performed at roughly the same level of recall. This suggested that the choice of index language mattered less than people had assumed, and that the intellectual stage of analysing a document’s subject mattered more.

The second phase, Cranfield II, dug into the components of indexing that actually drive performance, especially exhaustivity and specificity. The experiments established several enduring conclusions: there is an optimum level of indexing exhaustivity beyond which recall barely improves while precision suffers badly, and the inverse relationship between recall and precision is unavoidable. The lasting contribution, though, was the Cranfield paradigm itself, the test-collection methodology of documents, queries, and relevance judgments that still underpins evaluation today.

The MEDLARS evaluation

Where Cranfield was a laboratory experiment, the MEDLARS evaluation tested a real, working system. MEDLARS, the Medical Literature Analysis and Retrieval System, was run by the U.S. National Library of Medicine. Between 1966 and 1967, F. W. Lancaster carried out one of the earliest evaluations of a computer-based retrieval system, and it was the first application of recall and precision to a large, operational database.

Lancaster analysed around 300 actual search requests. The results showed the system operating, on average, at about 58% recall and 50% precision. More valuable than the numbers was the detailed failure analysis. By examining where searches went wrong, the study traced problems to specific causes in indexing, searching, and the way users expressed their needs. This led to concrete reforms, including a redesigned search request form to better capture the real information need behind a query, and an expanded, more accessible vocabulary. MEDLARS proved that evaluation could directly drive improvement in a live service.

The TREC experiments

By the early 1990s, test collections had grown stale and small relative to the volumes of digital text appearing. In 1990, NIST was asked to build a very large test collection as part of the DARPA TIPSTER program, and from this grew the Text REtrieval Conference (TREC), launched in 1992 and co-sponsored by NIST and a U.S. Department of Defense agency.

TREC scaled the Cranfield approach to roughly a million documents and, for the first time, let many research groups compare their systems on the same data using the same scoring methods. To make relevance judging feasible at this scale, TREC introduced pooling: the top results from many systems are combined into a pool, and human assessors judge only that pool rather than the entire collection. The conference is organised into tracks covering tasks such as web search, question answering, and more specialised domains. Its impact has been enormous; one study estimated that without TREC, American internet users would have spent billions of additional hours searching the web over a decade. TREC is, in effect, the Cranfield tradition carried into the era of large-scale digital search.

Improving indexing efficiency

Evaluation is only useful if it leads to better indexing. Several levers, many of them identified in the experiments above, let us tune performance toward the recall or precision a given audience needs.

Adjusting exhaustivity and specificity

Exhaustivity is how many of a document’s topics are captured by index terms. Specificity is how precisely those terms describe the concepts. These two qualities map almost directly onto our metrics: high exhaustivity tends to raise recall while high specificity tends to raise precision. An indexer who assigns many terms makes a document findable through more routes but also pulls it into searches where it only marginally fits. An indexer who uses very specific terms ensures that the documents retrieved are tightly on target. Tuning these two qualities is the most direct way to shift a system’s balance.

Controlled vocabulary and thesauri

Natural language is full of synonyms and ambiguous words. Without control, the same concept scatters across an index under different labels, and searchers miss material simply because they used a different word. A controlled vocabulary, supported by a thesaurus that records broader, narrower, and related terms, brings these variants together. Linking “aam” to “mango,” or grouping a concept with its broader and narrower terms, helps recall by uniting synonyms and helps precision by disambiguating terms. This structure is one of the most reliable ways to raise both metrics at once.

Improving indexing consistency

If two indexers describe the same document differently, retrieval becomes unpredictable. Indexing consistency, the degree to which indexers agree, is therefore a quality factor in its own right. Encouragingly, the Cranfield work found that trained indexers could index consistently and well even without deep subject expertise, and that beyond roughly four minutes of indexing time per document, extra effort produced little improvement. Clear indexing policies, good training, and well-defined rules all push consistency upward and make the system’s behaviour dependable.

Failure analysis and iteration

The lesson of MEDLARS is that you learn most from the searches that fail. Examining why a relevant document was missed, or why an irrelevant one was retrieved, points to the exact cause, whether a misinterpreted subject, a missing vocabulary term, or a poorly expressed query. Beyond recall and precision, supplementary measures such as fallout, which looks at how well a system rejects irrelevant material, can pinpoint specific weaknesses like poor synonym handling. Treating evaluation as a continuous loop, where results feed changes and changes are re-measured, is what turns a static index into an improving one.

Matching the system to the user

Finally, efficiency is not absolute; it is relative to purpose. The most efficient indexing system is the one tuned to its audience. An academic database may prioritise recall and discovery. A consumer search tool may prioritise precision and speed. A teaching collection may aim for a balance that gives comprehensive yet manageable results. Defining the target before optimising prevents the common mistake of chasing a number that does not serve the actual user.

What do you think? If you were designing an indexing system for your own library or database, would you lean toward higher recall or higher precision, and what would tip that decision? And given how subjective relevance can be, how much should we trust a single recall or precision figure as a measure of quality?

How useful was this post?

Click on a star to rate it!

Average rating 0 / 5. Vote count: 0

No votes so far! Be the first to rate this post.

We are sorry that this post was not useful for you!

Let us improve this post!

Tell us how we can improve this post?

References
  1. https://en.wikipedia.org/wiki/Cyril_Cleverdon
  2. https://nlp.stanford.edu/IR-book/html/htmledition/evaluation-of-ranked-retrieval-results-1.html
  3. https://www-nlpir.nist.gov/projects/irlib/pubs/cranv1p1/cranv1p1_text/01_002.txt
  4. https://en.wikipedia.org/wiki/Cranfield_experiments
  5. https://en.wikipedia.org/wiki/Frederick_Wilfrid_Lancaster
  6. https://eric.ed.gov/?id=ED022494
  7. https://trec.nist.gov/about.html

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

Organising and Managing Information

1 Basic Concepts

  1. Meanings of Classification
  2. Classification and Organisation
  3. Uses of Classification
  4. Scope of Classification
  5. Process of Classification
  6. Genus-Species Relation
  7. Nature of Classification
  8. Classification as a Tool
  9. Knowledge Classification
  10. Library Classification
  11. Modern Library Classification
  12. Uses of Classification in a Library
  13. Limitations of Classification

2 Type of classification

  1. Fixed and Relative Location Systems
  2. By Design Methodology
  3. Knowledge Classification and Library Classification
  4. Web Classifications: Ontologies
  5. By Areas of Applications
  6. By Form of Literature
  7. Print and Electronic Versions

3 Postulational Approach

  1. Postulational Approach
  2. Idea Plane
  3. Canons of Characteristics
  4. Canons for Succession of Characteristics
  5. Canons for Arrays
  6. Canons for Chain of Classes
  7. Verbal Plane
  8. Notational Plane
  9. Canons of Notation
  10. Hospitality in Array
  11. Hospitality in Chain
  12. Problems of Notation

4 Comparative Study of Schemes of classification

  1. Comparative Librarianship
  2. Introduction to the Major Schemes of Classification
  3. Discipline and Main Class
  4. Notation
  5. Extent of Use and Popularity
  6. Historical Contribution

5 Basic Concepts

  1. Library Catalogue
  2. Laws of Library Science and Library Catalogue
  3. Library Catalogue vis-a-vis Other Library Records
  4. Cataloguing and the Role of Technology
  5. Symbiosis

6 Types and forms of catalogues

  1. Author Catalogue
  2. Name Catalogue
  3. Title Catalogue
  4. Alphabetical Subject Catalogue
  5. Dictionary Catalogue
  6. Classified Catalogue
  7. Comparison of Dictionary and Classified Catalogue
  8. Alphabetico-Classed Catalogue
  9. Outer/Physical Forms of a Catalogue
  10. Bound Register Form
  11. Printed Book Form
  12. Sheaf Form
  13. Card Form
  14. Computer-Produced Book Form
  15. Microform Catalogue
  16. MARC and Online Catalogue
  17. CD-ROM Catalogue
  18. Comparative Study of Physical Forms of Catalogues

7 Formats and standards

  1. Bibliographic Record Formats
  2. Types of Formats
  3. Exchange Formats: Structure and Content
  4. ISBD (International Standard Bibliographic Description)
  5. ISO 2709
  6. MARC and MARC 21
  7. USMARC
  8. UK MARC
  9. UNIMARC
  10. CCF (Common Communication Format)
  11. Indian Standards

8 Cataloguing of non-book material

  1. Non-Book Material
  2. Problems of Cataloguing Non-Book Material
  3. Cataloguing Non-Book Material
  4. Bibliographic Description of Non-Book Material (AACR-2 Rev.Ed.)
  5. Changes in AACR 2R and Amendments 2002
  6. Resources Description and Access (RDA)

9 Basics of Subject Indexing

  1. Subject Indexing: Origin and Development
  2. Meaning and Purpose
  3. Cataloguing Versus Indexing
  4. Indexing Principles and Process
  5. Evaluation of Indexing

10 Indexing languages

  1. Meaning and Scope
  2. Natural Language vs. Indexing Language
  3. Structure of Indexing Language
  4. Attributes of an Indexing Language
  5. Vocabulary Control
  6. Types of Indexing Languages
  7. Library of Congress Subject Headings
  8. Sears List of Subject Headings

11 Indexing Techniques

  1. Derivative Indexing and Assignment Indexing
  2. Pre-Coordinate Indexing System
  3. Cutter’s Contribution
  4. Kaiser’s Contribution
  5. Chain Indexing
  6. PRECIS (Preserved Context Index System)
  7. POPSI (Postulate Based Permuted Subject Indexing)
  8. Post-Coordinate Indexing
  9. Uniterm Indexing
  10. Keyword Indexing
  11. Computerised Indexing
  12. Indexing Internet Resources

12 Conceptual Changes- Impact of Technology

  1. Knowledge Hierarchy
  2. Knowledge Organisation: Concept
  3. Knowledge Organisation in the Pre-Digital Age
  4. Knowledge Organisation Systems: Types
  5. Planning Knowledge Organisation Systems
  6. Linking Interrelated Digital Resources
  7. Universal Access to Heterogeneous Networked Resources
  8. Future of Knowledge Organisation Systems on the Web

13 Online Catalogues- Design and Services

  1. Physical Catalogue to OPAC: Changing Perspectives
  2. Descriptive Catalogue
  3. Standards
  4. Electronic Catalogue
  5. Online Catalogue
  6. Next-Generation Catalogue
  7. MARC Compliant Database
  8. Machine-Readable Cataloguing: Structural Design
  9. Metadata Tools for Cataloguing Networked Resources
  10. OPAC – Online Catalogue Interface
  11. Online Cataloguing Utility Services

14 Overview of Web Indexing, Metadata, Interoperability and Ontologies

  1. Web Indexing
  2. Metadata
  3. Ontology
  4. Interoperability