Every time you search a library catalogue or an academic database and instantly find the right document, an indexing system is quietly doing its job behind the scenes. Indexing is the process of representing the subject content of a document through terms that help users locate it later. But not all indexing works the same way. Some methods pull terms straight out of the document itself, while others rely on an indexer who assigns terms from a fixed list of approved vocabulary. These two broad approaches are known as derivative indexing and assignment indexing. Understanding how they differ is fundamental to grasping how information is organised, stored, and retrieved.
Table of Contents
- What is derivative indexing?
- How derivative indexing extracts terms
- What is assignment indexing?
- The role of conceptual analysis
- Derivative vs. assignment indexing: key differences
- Examples and applications of derivative indexing
- Keyword indexing: KWIC, KWOC, and KWAC
- Citation indexing
- Examples and applications of assignment indexing
- Chain indexing
- PRECIS and classification schemes
- Choosing the right method
What is derivative indexing?
Derivative indexing, also called extractive indexing or derived indexing, creates index entries by extracting terms directly from the document. The indexer (or a computer program) picks out significant words from the title, abstract, or full text and uses them as index terms without any modification. As described in the literature on derived indexing, this method relies solely on the information that is already present in the document. It does not add anything from the indexer’s own knowledge or from external sources.
The key feature here is the use of natural language. The vocabulary comes from how the author actually wrote the document. If an author titled a paper “The Cat in the Hat,” derivative indexing would generate entries under “Cat” and “Hat” because those are the words present in the text. This makes the process largely mechanical and fast, which is why it works well with automated systems and large databases.
How derivative indexing extracts terms
Because derivative indexing depends on the document’s own words, it is sometimes called automatic or natural language indexing. A computer can scan thousands of titles in seconds, identify significant words, and build an index without human intervention. This speed is a major advantage. However, it also creates problems. Authors do not always choose the most representative words for their titles. A paper titled “Expert System” might actually be about library automation, but a purely title-based index would never reveal that. This is the well-known limitation of natural language: synonyms get scattered, related concepts are separated, and false retrieval becomes common.
What is assignment indexing?
Assignment indexing, also known as assigned indexing or concept indexing, takes the opposite approach. Here, a human indexer reads the document, analyses its subject content, and then selects one or more subject headings or descriptors from a controlled vocabulary to represent it. According to the description of assigned indexing, the index terms do not need to appear anywhere in the title or text of the document. The indexer assigns them based on what the document is about, not what words it happens to use.
The controlled vocabulary is a prescribed list of authorised terms. Common examples include the Library of Congress Subject Headings, the Library of Congress Classification, and the Dewey Decimal Classification. By translating the document’s concepts into these standardised terms, assignment indexing brings consistency. The same concept is always represented by the same term, regardless of how different authors phrased it.
The role of conceptual analysis
Assignment indexing involves genuine intellectual effort. The indexer performs a conceptual analysis of the document, identifies the concepts expressed in it, and then assigns terms for those concepts from the controlled vocabulary according to set rules and procedures. As explained in an overview of indexing systems and techniques, this process also displays the syntactic and semantic relationships between terms. Systems like Chain Indexing, PRECIS, POPSI, and classification schemes all fall under this category. This is why assignment indexing is considered far more powerful for representing complex subjects, even though it is slower and more labour-intensive.
Derivative vs. assignment indexing: key differences
The single most important difference between the two methods is the use of controlled vocabulary. Derivative indexing does not rely on any external vocabulary, while assignment indexing depends entirely on it. From this one distinction, several other differences flow.
Source of terms: In derivative indexing, terms come from the document’s own text. In assignment indexing, terms come from a predetermined list, and they may not appear in the document at all.
Process: Derivative indexing is mechanical and can be automated. Assignment indexing is intellectual and usually requires a trained human indexer.
Consistency: Assignment indexing ensures uniformity because the same concept always maps to the same authorised term. Derivative indexing inherits the inconsistencies of natural language, where one idea may be expressed in many different ways.
Speed and cost: Derivative indexing is fast and cheap, making it suitable for huge volumes of material. Assignment indexing is slower and more expensive but offers greater precision and depth.
Examples and applications of derivative indexing
Derivative indexing is best understood through its most common real-world applications. These are the systems that build indexes straight from a document’s own words.
Keyword indexing: KWIC, KWOC, and KWAC
Keyword indexing is the classic example of derivative indexing. The KWIC (Keyword in Context) system is built on the principle that the title of a document represents its contents, almost like a one-line abstract. As outlined in the study material on keyword indexing, a KWIC index makes an entry under each significant word in the title, keeping the rest of the title alongside it to preserve context. The method was popularised by Hans Peter Luhn at IBM in the late 1950s, and it became a landmark in automated information retrieval.
Several variations followed. KWOC (Keyword Out of Context) places the keyword separately, usually at the beginning of the entry, rather than embedding it within the title line. KWAC (Keyword and Context) goes a step further by enriching the title keywords with additional significant words drawn from the abstract or contents, which helps when a title does not fully express the subject. All of these are derivative because they extract terms directly from the document.
Citation indexing
Citation indexing is another form of derivative indexing. Instead of extracting subject keywords, it tracks how documents reference or cite one another. The index is built from the citations listed within a document. This creates a network of scholarly connections, where you can move from one paper to the others it cites or to later papers that cite it. Academic databases use this principle to measure relevance and impact in the research community. Because the linking terms (the citations) are taken directly from the documents themselves, citation indexing belongs firmly in the derivative category.
Examples and applications of assignment indexing
Assignment indexing is represented by some of the most sophisticated systems developed in Library and Information Science. These systems rely on controlled vocabulary and on the analysis of subject relationships.
Chain indexing
Chain indexing, developed by Dr. S.R. Ranganathan, is a semi-mechanical method for deriving subject index entries from the class number of a document. He first described it in his book on the theory of library cataloguing. According to the explanation of chain indexing using DDC, each class number is analysed as a series of links, that is, steps of division running from the main class down to the specific subject. The “sought links” become subject index entries, prepared starting from the most specific level upward. Because the indexer assigns subject headings based on an analysis of the classification number rather than the document’s own words, chain indexing is treated as a form of assignment indexing. It was used in the British National Bibliography from the 1950s until the early 1970s.
PRECIS and classification schemes
PRECIS (Preserved Context Indexing System) was designed by Derek Austin to replace chain indexing in the British National Bibliography and to suit computerised production. As noted in the discussion of subject cataloguing systems, PRECIS uses a machine-held thesaurus to regulate semantic relationships, generating “see” and “see also” references automatically. POPSI (Postulate-based Permuted Subject Indexing), created by G. Bhattacharyya, applies Ranganathan’s postulates and is regarded as an improved version of the chain procedure because it does not depend directly on the class number.
Classification schemes such as the Dewey Decimal Classification and the Universal Decimal Classification are also part of the assignment family. They categorise documents into subject areas using a controlled, hierarchical structure. In every one of these systems, the defining feature is the same: a controlled vocabulary and an indexer who assigns terms based on subject analysis.
Choosing the right method
Neither method is universally better. Derivative indexing suits situations where speed, low cost, and large volume matter most, such as big digital databases and search engines. Assignment indexing suits situations that demand consistency, precision, and the ability to represent complex subjects accurately, such as library catalogues and specialised bibliographies. Many modern information systems combine both, using automated keyword extraction alongside human-assigned controlled terms to get the best of each approach.
What do you think? If you were building an index for a large research database, would you lean towards the speed of derivative indexing or the precision of assignment indexing? And as automated systems grow more capable, do you think controlled vocabularies will still be needed in the future?
References
- https://www.librarianshipstudies.com/2016/08/derived-indexing.html
- https://www.librarianshipstudies.com/2017/02/assigned-indexing.html
- https://mlsu.ac.in/econtents/413_Indexing%20techniques%20and%20process.pdf
- https://egyankosh.ac.in/bitstream/123456789/38415/5/Unit-13.pdf
- https://egyankosh.ac.in/bitstream/123456789/38416/5/Unit-14.pdf
- https://ebooks.inflibnet.ac.in/lisp3/chapter/subject-cataloguing-chain-procedure-popsi-and-precis/

Leave a Reply