Behind every successful digitisation project lies a carefully chosen set of tools. Converting a fragile manuscript, a stack of old journals, or a collection of photographs into reliable digital files is not a single action but a chain of steps, and each step depends on the right equipment and the right software. Choose poorly, and you end up with blurry scans, unsearchable text, and files that may not survive the next decade. Choose well, and you create a digital collection that serves readers for generations. This post walks through the essential tools that make digitisation work, from the scanners that capture the image to the software that cleans, converts, and preserves it.
Table of Contents
- Scanners: the capture devices
- Flatbed scanners
- Drum scanners
- Handheld scanners
- Choosing the right scanner
- Software and image editing: turning scans into usable files
- Capture and image editing software
- Optical Character Recognition software
- File formats and conversion
- Tools in practice: large-scale digitisation
Scanners: the capture devices
A scanner is the device that converts a physical document or image into a digital file. It works by passing a light source over the material and using a sensor to record the reflected light as data. Most general-purpose scanners rely on either a Charge-Coupled Device (CCD) or a Contact Image Sensor (CIS) to capture the image line by line. The choice of scanner matters enormously, because different materials demand different handling. A bound rare book cannot be treated like a loose sheet of paper, and a delicate film negative needs far more precision than a printed office memo. This is why digitisation professionals work with several scanner types rather than relying on one machine for everything.
Flatbed scanners
The flatbed scanner is the most familiar type and the workhorse of most digitisation projects. The document is placed face down on a flat glass surface, and a scanning head moves beneath it to capture the image. Its biggest strength is versatility. A flatbed can handle books, newspapers, photographs, and artwork, and it accommodates a range of sizes from a standard A4 page to larger legal documents. Flatbed scanners also offer adjustable resolution, typically between 600 and 4800 dpi, which lets the operator choose the quality appropriate for each item. For libraries handling mixed collections of moderate value, a good flatbed is often the most practical and cost-effective option.
Drum scanners
Drum scanners take a completely different approach. The material is mounted onto a rotating cylinder, which spins at high speed while a sensor reads the image point by point. Instead of the CCD sensors found in flatbeds, drum scanners use a photomultiplier tube (PMT), which captures far finer detail. The result is extraordinary resolution, often exceeding 10,000 dpi, with exceptional colour accuracy and a wide dynamic range. Because photomultipliers can extract detail from very dark shadow areas that CCD sensors miss, drum scanners remain the tool of choice for high-end work such as fine art reproduction, professional printing, and the digitisation of film negatives and transparencies.
These advantages come at a price. Drum scanners are expensive, bulky, and require a trained operator who knows how to carefully mount materials and control the scanning parameters. They are not everyday office machines. In a library or archive setting, they are reserved for delicate or irreplaceable materials where capturing every detail in a single pass is essential, since the item may never be safe to scan again.
Handheld scanners
A handheld scanner is a small, portable device that the operator moves manually across the surface of a document. Its great advantage is portability. It can reach materials that simply will not fit inside a flatbed, and it can be used on items that cannot be removed from their location, such as a fragile bound volume in a special collection. This makes handheld scanners useful in libraries, museums, and archives where the material must stay where it is.
The trade-off is quality. Because the device depends on a steady human hand, results can suffer from distortion if the scanner is not moved evenly, and handheld scanners generally offer lower resolution than flatbed or drum models. They suit short texts, newspaper clippings, and quick capture jobs rather than archival-grade reproduction. Operators need some experience to use them well, particularly when scanning important documents.
Choosing the right scanner
The correct scanner depends entirely on the project. A digitisation programme might use a flatbed for the bulk of its printed material, reserve a drum scanner for valuable photographs and film, and keep a handheld device for awkward items that cannot be moved. It is worth noting that records larger than A3 cannot fit easily onto a standard flatbed scanner, which is one reason large maps and oversized documents are often captured using overhead scanners or high-resolution digital cameras instead. Matching the tool to the material is the single most important decision in the capture stage.
Software and image editing: turning scans into usable files
Capturing an image is only half the job. Raw scans are rarely ready for use straight out of the scanner. They may be crooked, have black borders, show uneven contrast, or simply be too large to share. This is where software takes over, handling everything from initial capture settings to editing, text recognition, and final format conversion. A typical digitisation workflow moves through three software stages: image editing, optical character recognition, and saving in appropriate file formats.
Capture and image editing software
Once an item is scanned, image editing software is used to correct and improve the file. Common tasks include deskewing (straightening crooked pages), rotating, cropping out scanning borders, and adjusting brightness and contrast so that text stands out clearly against the background. For large projects, batch editing is invaluable because it applies the same corrections to hundreds of images at once. The Code4Lib Journal describes a typical workflow in which scans are saved as high-resolution TIFFs, batch-edited in software such as Adobe Lightroom, and then combined into PDFs before text recognition is run. Open-source tools like GIMP offer many of the same editing capabilities without licensing costs, making them attractive for libraries working within tight budgets.
An important principle here is the distinction between master files and derivative files. A scanned master image should not be edited for any specific output and is preserved as a large, lossless file. Derivative files are the edited copies created for everyday use, such as cropped, compressed versions for display on a website. Keeping an untouched master means you can always generate new derivatives later without going back to the original physical item.
Optical Character Recognition software
A scanned page is, by default, just a picture of text. You cannot search it, copy from it, or index its contents. Optical Character Recognition (OCR) software solves this by analysing the image and converting the letters and words into machine-readable text. This transformation is what makes a digital collection truly useful, because it enables full-text searching, indexing, and editing.
OCR works through a series of steps: it pre-processes the image to remove specks and straighten lines, segments the individual characters, examines their shapes, and matches them against a library of known fonts before exporting the result as a searchable file. Popular OCR tools include the commercial ABBYY FineReader, which handles complex formatting and unusual characters well, Adobe Acrobat Pro for simpler documents, and the free, open-source Tesseract engine. Accuracy depends heavily on the quality of the source scan. A clear, high-resolution image of typeset text can reach accuracy of 80 to 90 percent or more, while faded, handwritten, or cursive documents remain extremely difficult for OCR to read. For most libraries, even an imperfect OCR result is a major improvement over having no searchable text at all.
File formats and conversion
The final software decision concerns the format in which files are saved, and this directly affects how long a collection survives. The widely accepted practice is to save archival master files as TIFF, a lossless format that preserves full image quality without discarding data. For access copies that need to load quickly over a network, a compressed format such as JPEG is used instead. As the Digital Preservation Handbook explains, lossless formats are best for creating and storing archival masters, while lossy formats should be reserved for delivery and access rather than long-term preservation. PDF/A, a version of PDF designed specifically for archiving, is another common choice for text documents, especially when combined with an OCR layer for searchability.
For institutions building large collections, following established standards keeps files compatible and sustainable across systems. Many digitisation programmes reference the technical guidelines developed by national initiatives to ensure their masters meet recognised benchmarks for resolution, bit depth, and embedded metadata.
Tools in practice: large-scale digitisation
The scale at which these tools operate becomes clear in major national projects. The National Digital Library of India, sponsored by the Ministry of Education through the National Mission on Education through Information and Communication Technology, has brought together content from a wide range of Indian digital repositories into a single searchable platform. Such projects depend on exactly the combination of tools described above: scanners to capture rare books and manuscripts, image editing software to correct the captured pages, OCR to make millions of pages searchable, and carefully chosen file formats to ensure the material remains accessible for years to come. The same principles apply whether a project digitises a handful of local records or an entire national heritage collection. The tools simply scale up.
What do you think? If your library had a limited budget and a mixed collection of printed books, old photographs, and fragile manuscripts, how would you prioritise which scanners and software to invest in first? And do you think the rising accuracy of AI-enhanced OCR will eventually make manual transcription of handwritten documents a thing of the past?
References
- https://shop.czur.com/blogs/blog/types-of-scanner
- https://en.wikipedia.org/wiki/Drum_scanner
- https://www.techgeekbuzz.com/blog/types-of-scanners/
- https://www.naa.gov.au/sites/default/files/2022-01/Preservation-Digitisation-Standards-2021.pdf
- https://journal.code4lib.org/articles/16132
- https://libguides.mst.edu/c.php?g=335435&p=2256780
- https://www.dpconline.org/handbook/technical-solutions-and-tools/file-formats-and-standards
- https://www.loc.gov/ndnp/guidelines/archive/NDNP_201113TechNotes.pdf
- https://www.ndl.gov.in/

Leave a Reply