When you scan an index at the back of a book, you are using one of the oldest tools in information retrieval. Keyword indexing takes this idea further by building searchable entries directly from the words inside a document’s title. But for such an index to actually work, each entry must follow a precise structure. Without it, a list of keywords would just be a jumble of words with no way to point a reader back to the actual document. This post breaks down the exact structure and format that makes keyword indexing function, looking at its three essential parts and how they fit together.
Table of Contents
- Why structure matters in keyword indexing
- The three-part structure of an entry
- 1. The keyword
- 2. The context
- 3. The identification code
- How an entry is laid out on the page
- The role of the slash symbol
- Keyword on the left versus keyword in the centre
- A worked example
- Step one: selecting the keywords
- Step two: generating the entries
- Step three: filing the entries
- Selecting keywords: separating the significant from the trivial
- The stop list
- Editorial intervention
- Contextual integrity: keeping meaning intact
- Identification codes and document retrieval
- What the code points to
- Why serial numbers work well
- Putting the structure and format together
Why structure matters in keyword indexing
Keyword indexing rests on a simple assumption: the title of a document acts as a one-line abstract of its contents. The significant words in that title reveal what the document is about. The system was pioneered by Hans Peter Luhn, a computer scientist at IBM, who introduced his Keyword in Context (KWIC) technique in November 1958 at the International Conference on Scientific Information. The underlying concept, however, is even older. Librarian Andrea Crestadoro had proposed and manually implemented a “Keywords in Titles” approach for the Manchester Public Library in the mid-1800s.
The genius of the method is that it generates index entries automatically from existing titles, without the need for a trained indexer to assign subject terms from a controlled vocabulary. But automation only works if every entry obeys a fixed format. The structure is what turns raw title words into a usable finding aid. Each entry must tell the reader three things: which word they searched under, what surrounds that word to give it meaning, and where to physically locate the full document. These three demands map directly onto the three parts of every keyword index entry.
The three-part structure of an entry
Every entry in a keyword index is built from three components. Understanding these parts is the key to reading and constructing such an index correctly.
1. The keyword
The keyword is the significant word from the title that serves as the access point or approach term. This is the word under which the entry is filed and the word a user would look up. In a single document, several entries are generated because each significant word in the title becomes a keyword in turn. So a title with four important words produces four separate entries, each led by a different keyword. The keyword is what makes the entry findable in an alphabetical sequence.
2. The context
The context is made up of the remaining words of the title, presented alongside the keyword. These words specify the document’s context and prevent the keyword from being misread. The whole point of keeping these surrounding words is to preserve the meaning that the title originally carried. A keyword sitting alone tells you very little; the same keyword shown with the rest of its title tells you exactly how it is being used in that document.
3. The identification code
The identification or location code is a code, usually a serial number assigned to the entry, that points to where the full bibliographic details of the document can be found. It sits at the extreme right of the entry. This code is the bridge between the index and the actual document. Without it, a reader could identify a relevant title but would have no way to retrieve the item itself. The use of a numeric code to act as a reliable address reflects Luhn’s broader interest in using numbers to validate and locate records – he is also remembered for the Luhn algorithm used to check identification numbers such as credit card numbers.
How an entry is laid out on the page
The format of a keyword index follows a strict visual logic so that the three parts are always identifiable at a glance. The keyword is positioned so it stands out, while the context fills in around it and the identification code anchors the right-hand edge.
The role of the slash symbol
A forward slash “/” is used within an entry to mark the end of the title. Because the words of a title are rotated to bring different keywords to the front, the slash signals where the original title would have ended and shows how the remaining words wrap around. This small punctuation mark helps the reader mentally reassemble the full title from a rotated entry.
Keyword on the left versus keyword in the centre
There are two common formatting choices for where the keyword appears. In one layout, the keyword is shifted to the extreme left-hand side of the entry, with the rest of the title following it. In the other, the keyword appears in the centre of the line, with parts of the title on either side. Both layouts carry the same three components; they only differ in how the keyword is visually highlighted. The centred format is the classic look most people associate with a KWIC index.
A worked example
Consider a document with the title Classification of Books in a University Library, assigned the identification code 1279. Watch how the structure and format come together step by step. This worked sequence follows the approach described in standard accounts of KWIC indexing structure.
Step one: selecting the keywords
First, the significant words are picked out of the title. Words like “of,” “in,” and “a” are dropped because they carry no subject meaning. What remains are the four keywords: Classification, Books, University, and Library. Each of these will become the access point for one entry.
Step two: generating the entries
Next, the title is rotated so that each keyword in turn moves to the lead position, while the rest of the title trails behind it to supply context. The identification code 1279 stays fixed at the right of every entry:
CLASSIFICATION of Books in a University Library 1279
Books in a University Library/Classification of 1279
UNIVERSITY Library/Classification of Books in 1279
LIBRARY/Classification of Books in University 1279
Notice that each line is the same title rotated to a different starting word, the slash marks where the title wraps, and the code 1279 anchors every entry. The keyword, the context, and the identification code are all present in each line.
Step three: filing the entries
Finally, all the generated entries are merged into a single alphabetical sequence with entries from other documents. Once filed, a reader looking up “Library” will land on the relevant entry, see the full title for context, and use the code 1279 to retrieve the document.
Selecting keywords: separating the significant from the trivial
The quality of a keyword index depends almost entirely on which words are chosen as keywords. The goal is to keep words that carry subject meaning and discard those that do not.
The stop list
Non-significant words are removed using a stop list – a stored list of articles, prepositions, conjunctions, and other common words that should never become keywords. When the index is prepared by computer, the machine compares each title word against the stop list and rejects any match. Words like “the,” “of,” “in,” “and,” and “a” are typical stop-list entries. This is the same principle behind the way a KWIC index makes every non-stop word in a title searchable.
Editorial intervention
Selection can also be done by a human editor who marks the significant words before the title is processed. This editorial step matters because automatic selection is not perfect. A purely mechanical system may keep a word that is grammatically significant but useless for retrieval, or it may struggle with titles that do not clearly express their subject. Where a title is vague, the words alone may not represent the true content of the document, which is a known limitation of title-based indexing.
Contextual integrity: keeping meaning intact
One reason the context component exists is to protect against ambiguity. A single keyword pulled out on its own can be badly misleading. The word “Library” could refer to a building, a software library, or a collection. By displaying the keyword alongside the rest of its title, the index preserves the sense in which the word was originally used.
This is precisely what distinguishes the in-context approach from its variant, KWOC (Keyword Out of Context). In a KWOC arrangement, the keyword is lifted out and placed separately, often as a heading, with the title shown beside or below it. The trade-off is real: removing the keyword from its natural place in the title can make entries quicker to scan but weakens the immediate sense of context. The in-context format keeps the keyword embedded in its title precisely so that contextual integrity is never lost.
Identification codes and document retrieval
The identification code deserves a closer look because it is the part that makes the whole system practically useful. Finding a relevant title is only half the job; the reader still needs to get to the document.
What the code points to
The serial number in each entry is not random. It links to a separate master list, sometimes called a bibliography or document register, where the full citation lives – author, full title, publisher, year, and shelf location. The index itself stays compact because it does not repeat all those details in every entry. Instead, the short code does the work of pointing to them. This separation keeps the index lean and easy to scan while still guaranteeing that every entry leads back to a complete record.
Why serial numbers work well
Serial numbers are ideal as location codes because they are short, unique, and easy to sort. A number takes up little space at the right edge of an entry, and it can be matched instantly against the master register. Because the same code is repeated across all the entries generated from one title, a reader who finds the document under any of its keywords arrives at the same identification number and therefore the same document. This consistency is what allows a user to approach a single document from several different access points and still retrieve it reliably.
Putting the structure and format together
The structure and format of a keyword index are not arbitrary conventions; each element solves a specific problem. The keyword answers “what did I search for,” the context answers “what does this actually mean,” and the identification code answers “where do I find it.” The format – keyword positioned left or centre, the slash marking the title’s end, and the code pinned to the right – exists to make those three answers instantly readable. When all three parts are present and correctly laid out, a simple list of title words becomes a powerful, self-generating retrieval tool that pointed the way toward the automated search systems we rely on today.
What do you think? If automatic keyword selection depends so heavily on the words an author happens to choose for a title, how reliable can a title-derived index really be for documents with vague or clever titles? And in an age of full-text search engines, does the disciplined three-part structure of keyword indexing still have lessons worth preserving?

Leave a Reply