Every digitisation project begins with a scanner or a recorder, but its real success is decided by a quieter choice: the file format. Pick the wrong one, and a beautifully scanned manuscript could become unreadable in fifteen years. Pick the right one, and the same file may still open smoothly long after the software that created it has vanished. File formats are the containers that hold digital content, and they directly shape whether that content survives, stays searchable, and remains usable. This post breaks down the standard formats for text, images, audio, and video, and explains how to match each format to its purpose.
Table of Contents
- Why file formats matter in digitisation
- Open versus proprietary formats
- Standard file formats for text
- PDF and PDF/A
- DOC and DOCX
- XML and HTML
- Image file formats for scanned documents
- TIFF
- JPEG
- GIF
- Audio and video file formats
- AVI and MPEG-4
- QuickTime (MOV) and RealVideo (RM)
- Best practices for file format selection
Why file formats matter in digitisation
A file format is essentially a set of rules for how data is encoded and stored. When a library digitises a rare book or an oral history recording, that content has to be saved in some structured way so that a computer can interpret it later. The format chosen determines two things that matter most in any preservation effort: accessibility and longevity.
Accessibility means the file can be opened and used by a wide range of devices and software without special tools. Longevity means the file remains readable decades into the future, even as technology changes. These goals can conflict. A highly compressed format may be easy to share today but may lose quality permanently, while a robust archival format may be too large to send over a network. The core skill in digitisation is balancing these trade-offs.
Preservation specialists draw a useful distinction here. A widely respected rule of thumb is to use lossless formats for archival master copies and lossy formats only for access copies meant for everyday viewing. The master file is the high-quality original you protect; the access copy is the lighter version users actually interact with.
Open versus proprietary formats
Another major factor is whether a format is open or proprietary. Open formats have publicly available specifications maintained by standards bodies, so any developer can build software to read them. Proprietary formats are controlled by a single company. The risk with proprietary formats is obsolescence: if the owner stops supporting the format or goes out of business, files can become orphaned. This is why national archives in countries such as Germany, the United States, and the United Kingdom recommend or mandate open, well-documented formats for permanent records. Choosing formats that are non-proprietary and based on open standards gives an institution a stronger foundation for long-term preservation.
Standard file formats for text
Text is the backbone of most library collections, from books and journals to government records and theses. Several formats dominate this space, each suited to a different need.
PDF and PDF/A
PDF (Portable Document Format) is the most familiar format for digitised documents. It preserves layout and formatting consistently across devices, supports text search, and is almost universally compatible. This makes it ideal for reports, forms, manuals, and any document where visual consistency matters. PDF was originally proprietary but later became an open ISO standard, which boosted its credibility for preservation work.
For archiving, however, the preferred variant is PDF/A. PDF/A is a constrained version of PDF designed specifically for long-term preservation. Its defining feature is that it is self-contained: everything needed to display the document, including fonts, must be embedded in the file rather than pulled from external sources. It also mandates the use of standardised metadata. There are several conformance levels, with PDF/A-1 and PDF/A-2 both considered suitable for long-term preservation, while the newer PDF/A-3 allows other file types such as XML to be embedded inside the document.
DOC and DOCX
DOC and DOCX are Microsoft Word formats, useful as working or editable files. DOCX is based on the Office Open XML standard, which became an international ISO standard (ISO/IEC 29500), making it easier for libraries to preserve content created in published, maintained specifications. That said, archives generally treat Word formats as having only limited preservation support and often recommend converting them to PDF/A for permanent storage, since editable proprietary formats are less stable over decades.
XML and HTML
XML (eXtensible Markup Language) is highly valued in preservation because it separates content from presentation and is fully open. It stores structured, machine-readable data and is enormously useful for encoding metadata and complex documents. For full preservation support, the schema or DTD should be included along with the well-formed XML file. The National Digital Library of India project, for instance, has used XML as a standard to allow seamless interchange of content between contributors.
HTML (HyperText Markup Language) is the language of the web and the natural choice for content meant to be accessed through browsers. For preservation, archives advise including a DOCTYPE declaration and packaging any referenced files, such as CSS stylesheets, alongside the HTML file so the page renders correctly in the future.
Image file formats for scanned documents
Digitising photographs, manuscripts, maps, and illustrations means working with image formats. The choice here has a direct effect on quality and storage cost.
TIFF
TIFF (Tagged Image File Format) is the long-standing favourite for preservation master images. It supports lossless storage, meaning images can be saved without discarding any data or detail. This makes it ideal for scanning historical documents, photographs, and archival materials where every detail matters. TIFF files retain their quality even after edits, and they integrate well with advanced image processing and OCR systems. A consortium of major institutions including the British Library, the Library of Congress, and the National Archives concluded that TIFF is widely regarded as the correct format for archiving master image files. The main drawback is size: TIFF files are large, which makes them difficult to store and transmit over the internet without compression.
JPEG
JPEG is the most common image format for everyday use because of its excellent compression. It produces small file sizes that are easy to share and is supported across virtually all devices and software. This makes JPEG perfect for access copies. The catch is that JPEG uses lossy compression, sacrificing some image quality for smaller size. For this reason, it is unsuitable as an archival master. Preservation guidance is clear that you should not store a JPEG as both the access and archival copy because of the irretrievable data loss involved.
GIF
GIF (Graphics Interchange Format) suits simple graphics, logos, line drawings, and small web images. Its strengths are support for transparency and animation. Its major limitation is that GIF is restricted to a palette of 256 colours, which makes it a poor choice for photographs or any image needing a rich colour range. In a digitisation context, GIF plays a niche supporting role rather than serving as a primary capture format.
Audio and video file formats
Libraries and cultural heritage institutions increasingly digitise more than text. Oral histories, recorded lectures, folk music, and film all need to be captured and preserved. Audiovisual formats are more complex than text or image formats because they often involve a “container” or wrapper that holds the actual encoded media inside it.
AVI and MPEG-4
AVI (Audio Video Interleave) is an older container format developed by Microsoft. It can hold high-quality, lightly compressed video, which historically made it useful for capture, though its large file sizes and ageing design make it less popular for new projects. MPEG-4 is a far more widely used standard today. It offers efficient compression that keeps file sizes manageable while retaining good quality, and it is supported across almost every platform. This combination makes MPEG-4 an excellent choice for access and distribution copies of video.
QuickTime (MOV) and RealVideo (RM)
QuickTime (MOV) is Apple’s container format and remains relevant in professional preservation workflows. Some archivists prefer to wrap high-quality uncompressed video inside a MOV container for archival video masters, pairing it with derivative MPEG copies for end-user delivery. RealVideo (RM) was once popular for streaming over slow internet connections, but it is a proprietary format that has largely fallen out of use. It illustrates the obsolescence risk clearly: content locked into a declining proprietary format becomes harder to access as support fades, which is exactly why preservation planning avoids such formats for masters.
It is worth noting that audiovisual format standards are still evolving. While stable preferences exist for text and images, the field has not fully settled on a single standard for archiving complex video content, with some institutions favouring uncompressed video in a MOV wrapper and others adopting open-source alternatives.
Best practices for file format selection
With so many formats available, a clear strategy prevents costly mistakes. The single most important principle is to separate the archival master from the access copy.
The archival master is your high-quality, long-term original. It should be created in a stable, lossless, open or well-documented format: TIFF for images, PDF/A or XML for text, and an appropriate lossless or uncompressed format for audiovisual content. This master is preserved carefully and rarely touched. The access copy is a lighter, compressed version generated from the master for daily use: JPEG for images, MPEG-4 for video, and standard PDF for documents. Users interact with the access copy, while the master stays protected.
Beyond this, a few guidelines consistently appear in preservation policy:
Favour open standards. Choosing non-proprietary formats based on open standards reduces the risk of obsolescence and keeps content accessible without dependence on a single vendor.
Match the format to the content. There is no universal best format. A preservation policy should recognise the needs of the collection and select the format that best preserves the qualities that matter for that material, whether that is exact visual fidelity, searchable text, or audio clarity.
Consider wide adoption and support. Formats that are widely used and supported across the community offer greater confidence, because there are more tools to read them and a larger pool of expertise to draw on. Looking at what comparable institutions use is a practical way to make safer choices.
Plan for metadata and migration. Good preservation is not just about the file itself. Embedding standardised metadata, as PDF/A requires, keeps essential descriptive information attached to the content. National-scale initiatives such as the National Digital Library of India rely on harmonised metadata schemas, combining standards like Dublin Core with learning resource and thesis schemas to integrate millions of records. Institutions should also plan for migration, periodically moving content to newer formats as older ones approach obsolescence.
Following these practices turns file format selection from a technical afterthought into a deliberate preservation decision. A scanned document is only as durable as the format it lives in, and a thoughtful choice today is what keeps digitised heritage usable for generations.
What do you think? If you were digitising a fragile regional-language manuscript collection with a limited storage budget, how would you balance the need for high-quality lossless masters against the practical cost of storing them? And as audiovisual standards continue to shift, how should libraries decide when it is time to migrate older recordings to newer formats?
References
- https://www.dpconline.org/handbook/technical-solutions-and-tools/file-formats-and-standards
- https://mapsoft.com/posts/pdf-pdfa-standard.html
- https://www.digitalpreservation.gov/documents/NDSA_PDF_A3_report_final022014.pdf
- https://www.digitalpreservation.gov/series/challenge/formats_challenge.html
- https://www.researchgate.net/publication/384260684_National_DIGITAL_LIBRARY_IN_INDIAITS_USES_IMPACTS_AND_IMPORTANCE
- https://www.goodreads.com/author_blog_posts/18421077-digital-archives-choosing-sustainable-file-formats
- https://www.dpconline.org/handbook/institutional-strategies/standards-and-best-practice
- https://cacm.acm.org/research/national-digital-library-of-india/

Leave a Reply