When you scan an index at the back of a book, you are using one of the oldest tools in information retrieval. Keyword indexing takes this idea further by building searchable entries directly from the words inside a document’s title. But for such an index to actually work, each entry must follow a precise structure. Without it, a list of keywords would just be a jumble of words with no way to point a reader back to the actual document. This post breaks down the exact structure and format that makes keyword indexing function, looking at its three essential parts and how they fit together.

Table of Contents

Why structure matters in keyword indexing

Keyword indexing rests on a simple assumption: the title of a document acts as a one-line abstract of its contents. The significant words in that title reveal what the document is about. The system was pioneered by Hans Peter Luhn, a computer scientist at IBM, who introduced his Keyword in Context (KWIC) technique in November 1958 at the International Conference on Scientific Information. The underlying concept, however, is even older. Librarian Andrea Crestadoro had proposed and manually implemented a “Keywords in Titles” approach for the Manchester Public Library in the mid-1800s.

The genius of the method is that it generates index entries automatically from existing titles, without the need for a trained indexer to assign subject terms from a controlled vocabulary. But automation only works if every entry obeys a fixed format. The structure is what turns raw title words into a usable finding aid. Each entry must tell the reader three things: which word they searched under, what surrounds that word to give it meaning, and where to physically locate the full document. These three demands map directly onto the three parts of every keyword index entry.

The three-part structure of an entry

Every entry in a keyword index is built from three components. Understanding these parts is the key to reading and constructing such an index correctly.

1. The keyword

The keyword is the significant word from the title that serves as the access point or approach term. This is the word under which the entry is filed and the word a user would look up. In a single document, several entries are generated because each significant word in the title becomes a keyword in turn. So a title with four important words produces four separate entries, each led by a different keyword. The keyword is what makes the entry findable in an alphabetical sequence.

2. The context

The context is made up of the remaining words of the title, presented alongside the keyword. These words specify the document’s context and prevent the keyword from being misread. The whole point of keeping these surrounding words is to preserve the meaning that the title originally carried. A keyword sitting alone tells you very little; the same keyword shown with the rest of its title tells you exactly how it is being used in that document.

3. The identification code

The identification or location code is a code, usually a serial number assigned to the entry, that points to where the full bibliographic details of the document can be found. It sits at the extreme right of the entry. This code is the bridge between the index and the actual document. Without it, a reader could identify a relevant title but would have no way to retrieve the item itself. The use of a numeric code to act as a reliable address reflects Luhn’s broader interest in using numbers to validate and locate records – he is also remembered for the Luhn algorithm used to check identification numbers such as credit card numbers.

How an entry is laid out on the page

The format of a keyword index follows a strict visual logic so that the three parts are always identifiable at a glance. The keyword is positioned so it stands out, while the context fills in around it and the identification code anchors the right-hand edge.

The role of the slash symbol

A forward slash “/” is used within an entry to mark the end of the title. Because the words of a title are rotated to bring different keywords to the front, the slash signals where the original title would have ended and shows how the remaining words wrap around. This small punctuation mark helps the reader mentally reassemble the full title from a rotated entry.

Keyword on the left versus keyword in the centre

There are two common formatting choices for where the keyword appears. In one layout, the keyword is shifted to the extreme left-hand side of the entry, with the rest of the title following it. In the other, the keyword appears in the centre of the line, with parts of the title on either side. Both layouts carry the same three components; they only differ in how the keyword is visually highlighted. The centred format is the classic look most people associate with a KWIC index.

A worked example

Consider a document with the title Classification of Books in a University Library, assigned the identification code 1279. Watch how the structure and format come together step by step. This worked sequence follows the approach described in standard accounts of KWIC indexing structure.

Step one: selecting the keywords

First, the significant words are picked out of the title. Words like “of,” “in,” and “a” are dropped because they carry no subject meaning. What remains are the four keywords: Classification, Books, University, and Library. Each of these will become the access point for one entry.

Step two: generating the entries

Next, the title is rotated so that each keyword in turn moves to the lead position, while the rest of the title trails behind it to supply context. The identification code 1279 stays fixed at the right of every entry:

CLASSIFICATION of Books in a University Library  1279
Books in a University Library/Classification of  1279
UNIVERSITY Library/Classification of Books in  1279
LIBRARY/Classification of Books in University  1279

Notice that each line is the same title rotated to a different starting word, the slash marks where the title wraps, and the code 1279 anchors every entry. The keyword, the context, and the identification code are all present in each line.

Step three: filing the entries

Finally, all the generated entries are merged into a single alphabetical sequence with entries from other documents. Once filed, a reader looking up “Library” will land on the relevant entry, see the full title for context, and use the code 1279 to retrieve the document.

Selecting keywords: separating the significant from the trivial

The quality of a keyword index depends almost entirely on which words are chosen as keywords. The goal is to keep words that carry subject meaning and discard those that do not.

The stop list

Non-significant words are removed using a stop list – a stored list of articles, prepositions, conjunctions, and other common words that should never become keywords. When the index is prepared by computer, the machine compares each title word against the stop list and rejects any match. Words like “the,” “of,” “in,” “and,” and “a” are typical stop-list entries. This is the same principle behind the way a KWIC index makes every non-stop word in a title searchable.

Editorial intervention

Selection can also be done by a human editor who marks the significant words before the title is processed. This editorial step matters because automatic selection is not perfect. A purely mechanical system may keep a word that is grammatically significant but useless for retrieval, or it may struggle with titles that do not clearly express their subject. Where a title is vague, the words alone may not represent the true content of the document, which is a known limitation of title-based indexing.

Contextual integrity: keeping meaning intact

One reason the context component exists is to protect against ambiguity. A single keyword pulled out on its own can be badly misleading. The word “Library” could refer to a building, a software library, or a collection. By displaying the keyword alongside the rest of its title, the index preserves the sense in which the word was originally used.

This is precisely what distinguishes the in-context approach from its variant, KWOC (Keyword Out of Context). In a KWOC arrangement, the keyword is lifted out and placed separately, often as a heading, with the title shown beside or below it. The trade-off is real: removing the keyword from its natural place in the title can make entries quicker to scan but weakens the immediate sense of context. The in-context format keeps the keyword embedded in its title precisely so that contextual integrity is never lost.

Identification codes and document retrieval

The identification code deserves a closer look because it is the part that makes the whole system practically useful. Finding a relevant title is only half the job; the reader still needs to get to the document.

What the code points to

The serial number in each entry is not random. It links to a separate master list, sometimes called a bibliography or document register, where the full citation lives – author, full title, publisher, year, and shelf location. The index itself stays compact because it does not repeat all those details in every entry. Instead, the short code does the work of pointing to them. This separation keeps the index lean and easy to scan while still guaranteeing that every entry leads back to a complete record.

Why serial numbers work well

Serial numbers are ideal as location codes because they are short, unique, and easy to sort. A number takes up little space at the right edge of an entry, and it can be matched instantly against the master register. Because the same code is repeated across all the entries generated from one title, a reader who finds the document under any of its keywords arrives at the same identification number and therefore the same document. This consistency is what allows a user to approach a single document from several different access points and still retrieve it reliably.

Putting the structure and format together

The structure and format of a keyword index are not arbitrary conventions; each element solves a specific problem. The keyword answers “what did I search for,” the context answers “what does this actually mean,” and the identification code answers “where do I find it.” The format – keyword positioned left or centre, the slash marking the title’s end, and the code pinned to the right – exists to make those three answers instantly readable. When all three parts are present and correctly laid out, a simple list of title words becomes a powerful, self-generating retrieval tool that pointed the way toward the automated search systems we rely on today.

What do you think? If automatic keyword selection depends so heavily on the words an author happens to choose for a title, how reliable can a title-derived index really be for documents with vague or clever titles? And in an age of full-text search engines, does the disciplined three-part structure of keyword indexing still have lessons worth preserving?

How useful was this post?

Click on a star to rate it!

Average rating 0 / 5. Vote count: 0

No votes so far! Be the first to rate this post.

We are sorry that this post was not useful for you!

Let us improve this post!

Tell us how we can improve this post?

References
  1. https://www.historyofinformation.com/detail.php?id=2065
  2. https://en.wikipedia.org/wiki/Luhn_algorithm
  3. https://www.librarianshipstudies.com/2017/02/keyword-in-context-kwic-indexing.html
  4. https://en.wikipedia.org/wiki/Key_Word_in_Context

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

Document Processing – Practice

1 Introduction, Structure and Organisation

  1. 19th Edition of DDC
  2. Notation
  3. Properties of Decimal Fractions
  4. Basic Plan and Convention of a Minimum of Three Digits
  5. Volume 1: Introduction and Tables
  6. Volume 2: Schedules
  7. Volume 3: Relative Index
  8. Transcription of a Class Number
  9. Numbers in Square Brackets, Obsolete Entries

2 Definitions, Notes and Instructions

  1. Definitions, Explanation and Scope
  2. Different Types of Notes
  3. Inclusion Notes
  4. Class Hereโ€™ Notes
  5. Class Elsewhereโ€™ Notes
  6. Class … in…โ€™ Notes
  7. For… see…โ€™ Notes
  8. Class comprehensive works in…โ€™ Notes
  9. Notes Appearing under Discontinued (Bracketed) Entries
  10. Importance of Various Notes
  11. Centred Heading/Centred Entry
  12. Number Building Notes

3 Introduction to Three Summaries and Steps in Classifying Documents

  1. Ten Main Classes
  2. Second Summary of 100 Divisions
  3. Third Summary of 1000 Sections
  4. Multi-Level Summaries
  5. Steps in Classifying Documents
  6. Subject Analysis
  7. Steps in Practical Classification

4 Relative Index and Its Use

  1. Need and Importance
  2. Nomenclature: Relative Index
  3. Scope of the Index
  4. Organisation of the Index
  5. Numbers Given Against Entries

5 Study of Tables and Schedules

  1. Tables
  2. Schedules

6 Auxiliary Tables and Devices

  1. Number-Building with Tables
  2. Use of Table 1: Standard Subdivisions
  3. Use of Table 2: Areas
  4. Table 3: Subdivisions of Individual Literatures
  5. Use of Table 4: Subdivisions of Individual Languages
  6. Use of Table 6: Languages
  7. Use of Table 5: Racial, Ethnic, National Groups
  8. Use of Table 7: Persons

7 Practical Classification

  1. Simple Synthesis
  2. Multiple Synthesis
  3. Order of Precedence
  4. Other Means for Fixing Priority of Numbers
  5. Table of Precedence for Standard Subdivisions

8 AACR-2R Preliminaries

  1. Structure of AACR-2R
  2. Levels of Description
  3. Style of Writing
  4. Types of Entries
  5. Items in the Catalogue Entry
  6. Skeleton Card
  7. Added Entries and Tracing Section
  8. Subject Headings

9 Choice and Rendering of Headings and Statement of Responsibility

  1. Personal Author
  2. Heading for Personal Author
  3. Western Names
  4. Indian Names
  5. Cataloguing Practice
  6. Single Personal Author
  7. Shared Responsibility
  8. Books under Editorial Direction
  9. Pseudonymous Authors
  10. Corporate Bodies

10 Cataloguing Multivolumes, Serial Publications and Nonprint Media

  1. Multi-volume Books
  2. Serial Publications
  3. Cataloguing of Non-Print Media

11 Marc-21 Cataloguing

  1. MARC Standards
  2. MARC 21 Structure
  3. MARC 21 Tags and Subfields

12 Structure of Sears List of Subject Headings (18th Edition)

  1. History of the Sears List
  2. Features of the 18th Edition
  3. Principles of Vocabulary Control
  4. Principles of the Sears List
  5. Structure of the Sears List
  6. Grammar of the Subject Headings
  7. Key Headings
  8. Adding Subdivisions
  9. Categories of Subject Headings Omitted
  10. Limitations

13 Keyword Indexing

  1. Keyword Indexing โ€“ Concept
  2. Structure and Format of Keyword Indexing
  3. Indexing Process
  4. Variants of Keyword Indexing
  5. Advantages and Disadvantages of Keyword Indexing

14 Chain Indexing (DDC โ€“ 19th Edition)

  1. Classified Catalogue
  2. Problems of Subject Cataloguing
  3. Advent of Chain Indexing
  4. Mechanism of Chain Indexing
  5. Step by Step Method
  6. Working with the DDC
  7. Preparation of Class Index Entries (CIEs)
  8. Chain Indexing for a Dictionary Catalogue
  9. Advantages of Chain Indexing
  10. Limitations and Problems of Chain Indexing