Turning a printed page into a digital file sounds like a single button press, but it is actually a chain of decisions. Before you scan anything, you have to choose what kind of digital object you want to end up with: a flat picture of the page, a searchable document, or fully editable text. These choices are the input and output options of digitisation, and getting them right determines whether your digital collection becomes a useful resource or just a pile of pretty but unusable images. This post breaks down the scanning methods available, how Optical Character Recognition fits in, and why retaining the original layout matters so much.
Table of Contents
- What input and output mean in digitisation
- Scanning options: the input side
- Image-only scanning
- Scanning with OCR
- Colour modes: bitonal, grayscale, and colour
- Resolution and DPI
- How OCR works
- Searchable text versus fully editable text
- OCR and layout retention
- Zonal OCR
- Output options: choosing the right file format
- Bringing the choices together
What input and output mean in digitisation
Digitisation is the process of capturing analog materials, such as books, manuscripts, photographs, and old records, as digital images. The input is the physical item and the way you capture it, usually through a scanner or a camera. The output is the digital file you finally save and share. Between these two ends sit a set of options that shape quality, file size, searchability, and long-term usability.
The key point is that a scan and a usable document are not the same thing. As one digitisation guide from the University of California, Berkeley explains, scanning captures analog material as a digital image, and a separate process is needed to read that image and turn it into searchable text. Understanding this distinction is the first step in planning any digitisation project.
Scanning options: the input side
When you place a document on a scanner, the software almost always offers you a set of capture choices. These determine what the raw image looks like and how big the file becomes.
Image-only scanning
The simplest option is image-only scanning. Here the scanner produces a picture of the page and nothing more. You can view it, zoom into it, and print it, but you cannot select, search, or copy any text on it. The computer sees a single image, not letters and words.
This mode is perfect for material where the visual appearance is the whole point, such as photographs, maps, artwork, illustrations, and old manuscripts with decorative scripts. For these items, faithful reproduction matters more than searchable text. A scanned archive that stops at image-only, however, can behave like a locked cabinet: the content is preserved but you cannot search it, quote it, or feed it to assistive software, a limitation highlighted in this overview of OCR-based digitisation.
Scanning with OCR
The second major option is scanning with Optical Character Recognition (OCR). In this mode the scanner still captures the image, but the software then analyses that image, identifies the characters, and adds a layer of machine-readable text. The result is a document you can search and, depending on the output format, edit. Many modern scanners and multifunction devices can do this in one step, producing a searchable PDF directly from the page.
This is the right choice for printed books, newspapers, journals, government reports, and typed letters, where the words themselves are the value. We will look at OCR in more detail below, because it is the heart of most text digitisation work.
Colour modes: bitonal, grayscale, and colour
Beyond image-only versus OCR, scanners offer several capture modes that affect quality and file size.
Bitonal (black and white) mode records each pixel as either pure black or pure white. It produces the smallest files and works well for clean printed text, which is why bitonal images are often the default for business documents. Grayscale captures shades of gray and suits typed pages with faded ink, pencil notes, or subtle tonal detail. Colour mode captures the full range of hues and is essential for photographs, illustrations, coloured seals, and any item where colour carries meaning.
Resolution and DPI
Resolution, measured in dots per inch (DPI) or pixels per inch, decides how much fine detail your scan captures. A higher DPI records more detail but creates a larger file. For ordinary text documents, around 200 to 300 DPI is generally sufficient. For archival master copies meant to replace or outlive the original, institutions often scan higher. Preservation guidance from organisations such as the U.S. National Archives and Records Administration sets out detailed resolution standards for different material types, and many archives now treat around 400 DPI as a sensible minimum for true archival scanning. The general rule is to capture the highest resolution you can reasonably store, because you can always create smaller copies later but you cannot add detail that was never captured.
How OCR works
Optical Character Recognition is the technology that converts a scanned image of text into machine-readable characters. As Adobe describes it, OCR transforms static, picture-based content into text you can search, copy, edit, and highlight. A quick way to check whether a PDF has been OCR-processed is to try selecting the text: if you cannot highlight it, you are looking at a plain image.
The software examines the shapes in the image, matches them to known letters and numbers, and builds a text layer. Accuracy depends heavily on the input. Clean, high-resolution scans of standard printed fonts give excellent results, while poor image quality, unusual fonts, or handwriting reduce accuracy. This is why scanning carefully at the input stage pays off at the output stage. For handwritten material, a related technology called Handwritten Text Recognition (HTR) is used instead, as noted in this guide from Yale University Library.
Searchable text versus fully editable text
OCR can give you two different kinds of output, and the difference matters. A searchable PDF keeps the original page image on top and hides the recognised text in an invisible layer beneath it. The page looks exactly like the scan, but you can now search and copy the words. This is ideal for preserving the appearance of historical documents while making them findable.
Fully editable output, such as a Word or plain text file, rebuilds the document as actual text you can rewrite. This is useful when you need to reuse or update the content, but it discards the original page image. Choosing between the two depends on whether faithful appearance or easy editing is your priority.
OCR and layout retention
Recognising the words is only half the challenge. The other half is preserving how those words were arranged on the page. This is called layout retention, and it separates basic OCR from advanced OCR.
A document is rarely just one block of running text. It has columns, headings, tables, footnotes, page numbers, captions, and images, all positioned deliberately. Simple OCR may correctly read every word yet dump them into one jumbled column, collapsing tables and scrambling the reading order. Advanced OCR engines, by contrast, analyse the structure of the page and reproduce it. Professional tools are valued precisely because they preserve layouts, styles, and structure rather than just extracting raw characters. Adobe’s OCR, for example, aims to retain the original layout, fonts, and formatting so the output resembles the source document.
Zonal OCR
A specialised approach to layout is Zonal OCR, sometimes called template OCR. Instead of reading the whole page, the operator divides the document into predefined zones, such as an invoice number, a date, or a name field, and OCR is applied only to those areas. This technique, explained in this guide to zonal OCR for PDF parsing, is widely used for structured forms where the same fields appear in the same place on every page. It is a powerful way to pull specific data out of large batches of similar documents automatically.
Output options: choosing the right file format
Once recognition is done, you must decide what file format to save. Each format serves a different purpose.
TIFF is an uncompressed, stable format widely preferred for archival master images because it preserves maximum quality, as outlined in these digitisation format guidelines. Its drawback is large file size. JPEG uses lossy compression to create small files, which makes it convenient for sharing and web display but unsuitable for archival masters because some detail is lost each time it is saved.
PDF is the most common delivery format for text documents. Unlike pure image formats, it can hold searchable text, metadata, and multiple pages in one file. Its archival variant, PDF/A, is an ISO standard designed for long-term preservation; it disables features like embedded audio and font linking that could make a file unreadable in the future. For editable output, OCR tools can export to Word (DOCX) or plain text (TXT) files. A common professional workflow keeps a high-quality TIFF master for preservation while producing a compressed, OCR-enabled PDF for everyday access.
Bringing the choices together
Every digitisation project is a set of linked decisions. You choose a capture mode and resolution based on the material, decide whether OCR is needed, select between searchable and editable output, weigh how much layout retention you require, and finally pick a file format for both preservation and access. A photograph might be scanned in colour at high resolution and saved as TIFF with no OCR at all. A printed report might be scanned in bitonal mode, OCR-processed with layout retention, and delivered as a searchable PDF/A. Matching the options to the purpose is what turns digitisation from mere copying into genuine access.
What do you think? If you were digitising a rare regional-language newspaper from the 1950s, would you prioritise faithful image reproduction or searchable text, and how would you balance the two? When is layout retention worth the extra processing effort, and when is plain searchable text enough for your users?
References
- https://digitalhumanities.berkeley.edu/resources/digitization-workflows-scanning-ocr-and-audio-transcription
- https://www.digitaldividedata.com/blog/optical-character-recognition-ocr-digitization
- https://www.archives.gov/files/preservation/technical/guidelines-1998.pdf
- https://www.adobe.com/acrobat/online/ocr-pdf.html
- https://guides.library.yale.edu/dh/ocr
- https://parsio.io/blog/zonal-ocr/
- https://www.wisconsinhistory.org/pdfs/la/Digitization-State/9_Digitization-Format-Guidelines.pdf

Leave a Reply