Digitisation has transformed how libraries, archives, and museums preserve knowledge and connect users with information. Yet the most important truth about any digitisation programme is rarely stated: you cannot digitise everything. Scanning is expensive, storage demands ongoing investment, and staff time is limited. This makes selection the single most consequential decision in the entire process. Choosing the wrong materials wastes resources, while overlooking valuable items can mean losing them to decay. This post walks through how institutions decide what deserves to be converted into a digital surrogate, and why legal clearance must sit at the heart of every selection decision.

Table of Contents

Why selection matters before scanning begins

A digitisation project is not simply about feeding pages into a scanner. It is a planned activity with defined goals, costs, and outcomes. As preservation specialists note, there are no absolute rules for selection, only a set of questions each institution must answer in light of its own mission. A national library preserving a country’s printed heritage will weigh priorities differently from a college library trying to support its current syllabus.

Selection asks three connected questions about any item: should it be digitised, may it be digitised, and can it be digitised? The first concerns value and demand, the second concerns legal rights, and the third concerns physical and technical feasibility. A strong candidate satisfies all three. Skipping any one of them invites either wasted effort or legal trouble.

Criteria for selection

Most established frameworks group selection criteria around a few core factors: the value of the material, the demand for it, and its physical condition. These work together rather than in isolation, and a thoughtful selector balances them against the cost of the project.

Intellectual value

The first question is how significant the material is. Rare manuscripts, unique local records, original research, government reports, and culturally important works generally rank high. These items often carry enduring value because they are scarce or irreplaceable. For example, an institution holding palm-leaf manuscripts, regional language literature, or records documenting local history possesses content that may exist nowhere else. Digitising such material protects it against permanent loss and opens it to scholars who could never have travelled to consult the original.

Value is not only about rarity. A heavily used textbook collection has high practical value even if it is not rare. Selectors therefore distinguish between research value, historical value, and instructional value, then match these against the goals of their specific project.

User demand

The second factor is how much the material is actually wanted. Because resources are limited, it makes sense to prioritise items that users request frequently or that would attract new audiences once available online. A collection that gathers dust on a shelf delivers little return when digitised, while a well-used set of journals or a popular special collection multiplies its reach the moment it goes online.

Demand should be assessed using real evidence: circulation statistics, reference enquiries, interlibrary loan requests, and feedback from researchers. A useful guideline is to consider both the current and the potential audience. Some materials see little use simply because few people know they exist or because the originals are too fragile to handle. Digitisation can unlock latent demand for exactly these hidden treasures.

Physical condition and fragility

The third factor pulls in two directions, which is why it needs care. Materials that are deteriorating, brittle, or at risk of loss are strong candidates because digitisation captures their content before it disappears. At the same time, an item must be able to withstand the physical handling that scanning requires. A manuscript so fragile that it crumbles on contact may need conservation treatment first, or specialised equipment, which raises the cost considerably.

Condition therefore influences not only whether to digitise but also how. Robust material may be scanned in-house on standard equipment, while fragile or oversized items may require outsourcing to specialists with overhead cameras and controlled environments. The format and size of the original – bound volumes, large maps, photographs, audio tapes – all shape the technical requirements and the budget.

Technical feasibility and cost

Even valuable, in-demand, well-preserved material can fail the feasibility test. Technical feasibility is among the most important practical criteria, because the physical characteristics of the source and the goals for capturing and storing the digital copy determine the equipment, file formats, and storage needed. Colour fidelity for art reproductions, optical character recognition for searchable text, and high resolution for detailed images each add complexity and expense.

Cost runs through every decision. Digitisation involves staff time for preparation, scanning, quality control, metadata creation, and long-term storage. For this reason, many institutions give higher priority to collections that are already organised and described with proper metadata, since much of the descriptive work is already done and digitisation becomes one more step toward access rather than a project from scratch. Selection, in practice, works as a matrix in which value, demand, condition, and cost are weighed together.

A material can score perfectly on value, demand, and condition and still be the wrong choice if you do not hold the right to copy and distribute it. This is why Intellectual Property Rights must be checked early in the selection process, not after scanning is complete. Resolving rights at the end of a project risks producing digital files that can never legally be made public.

The first legal step is to establish the copyright status of each item. Most material falls into one of three broad categories: works in the public domain, which are no longer protected by copyright; works still under copyright, where permission is required; and works whose status is unclear, often called orphan works because the rights holder cannot be identified or traced. In India, copyright in literary, dramatic, musical, and artistic works generally lasts for the lifetime of the author plus sixty years, after which the work enters the public domain and can be freely digitised.

Public domain and locally owned material – such as an institution’s own publications, government documents, or works for which it holds the rights – are the safest and easiest candidates. Copyrighted works require written permission or a licence from the rights holder, which can be slow and sometimes costly to obtain.

Fair dealing and library exceptions in India

Indian law does provide some breathing room. Section 52 of the Copyright Act, 1957 sets out a list of acts that do not amount to infringement. Two provisions are especially relevant to libraries. Section 52(1)(n) permits a non-commercial public library to store a work electronically for preservation, provided the library already holds a non-digital copy. Section 52(1)(o) allows such a library to make up to three copies of a book not available for sale in India for the use of the library.

These exceptions are deliberately narrow. They support preservation and limited research access, but they do not authorise large-scale digitisation of in-copyright works or open distribution over the internet. Section 52(1)(i) further allows reproduction in the course of instruction, a provision the Delhi High Court interpreted broadly in the well-known university photocopying case, where copying for educational purposes was held to fall within fair dealing. Even so, the statute does not explicitly address whether uploading a scanned chapter to a learning platform qualifies, which leaves genuine legal uncertainty that selectors should treat with caution.

Documenting permissions

Good practice is to record the rights status and any permissions for every selected item before scanning. Where works are in copyright, institutions should seek written agreements that specify how the digital copy may be stored, displayed, and shared. Where access must be restricted, the selection plan should note this so the digital files are delivered only to authorised users. Clear documentation protects the institution and ensures that effort spent digitising actually results in lawful access.

Selection in the Indian context

India’s major digitisation programmes illustrate how these criteria play out in practice. National efforts such as the National Digital Library of India, developed at IIT Kharagpur under the Ministry of Education, aggregate learning resources through a single search platform, while initiatives like Shodhganga for theses and the Traditional Knowledge Digital Library for indigenous medical knowledge target specific high-value content.

These programmes also reveal recurring challenges. Indian institutions have long struggled with the absence of a clear national digitisation and intellectual property policy, the limited availability of optical character recognition for Indian languages, and a shortage of trained personnel. For a selector, these constraints feed directly back into the feasibility criterion: a regional-language manuscript may be highly valuable, yet poor OCR support means full-text searching may not be achievable, affecting how the project is designed and prioritised.

What do you think? If your institution could digitise only one collection this year, would you prioritise a fragile but rarely consulted historical manuscript, or a heavily used set of in-copyright textbooks tangled in permissions? And how should libraries balance the urgency of preservation against the legal limits set by copyright law?

How useful was this post?

Click on a star to rate it!

Average rating 0 / 5. Vote count: 0

No votes so far! Be the first to rate this post.

We are sorry that this post was not useful for you!

Let us improve this post!

Tell us how we can improve this post?

References
  1. https://www.nedcc.org/free-resources/preservation-leaflets/6.-reformatting/6.6-preservation-and-selection-for-digitization
  2. https://guides.library.manoa.hawaii.edu/c.php?g=105219&p=687076
  3. https://www.sciencedirect.com/topics/computer-science/selecting-material
  4. https://www.margotnote.com/blog/digitization-criteria
  5. https://indiankanoon.org/doc/1013176/
  6. https://www.managingip.com/article/2dwcc61v1ul1gqqg8hou8/sponsored-content/the-internet-archive-case-implications-for-indias-copyright-landscape
  7. https://arxiv.org/pdf/2303.13594
  8. https://ebooks.inflibnet.ac.in/lisp8/chapter/digital-library-initiatives-in-india-part-i/

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

ICT Applications

1 Database- Concept and Components

  1. Database Approach
  2. Database Definition
  3. Different Approaches to Database
  4. Database Features
  5. Databases in Library and Information Science
  6. Database Functional Considerations
  7. Types of Databases
  8. Database Architecture

2 Data Structures, File Organisation and Physical Database Design

  1. Why Data Structures
  2. Memory Hierarchy
  3. RAID Technology
  4. Indexes
  5. Binary Search
  6. Linked Lists
  7. Inverted Lists
  8. B-Trees
  9. File Storage Concepts
  10. Sequential Access Method (SAM)
  11. Indexed Sequential Access Method (ISAM)
  12. Direct Access Method (DAM)
  13. Physical Database Design

3 Database Management Systems

  1. Data and Information
  2. Database and Database Management System (DBMS)
  3. Data Hierarchy
  4. Data Integrity
  5. Data Independence
  6. Objectives of DBMS
  7. Evolution of DBMS
  8. Functions and Components of a DBMS
  9. Architecture of a DBMS
  10. Entity-Relationship Model
  11. Types of Relationships in Data Modeling
  12. Relational Database Management Systems (RDBMS)
  13. Normalization of Relations
  14. Designing Databases
  15. Distributed Database Systems
  16. Database Systems for Management Support
  17. Artificial Intelligence and Expert Systems

4 Database Searching

  1. Introduction
  2. Information Retrieval
  3. Information Retrieval Versus Data Retrieval
  4. Parameters for Evaluation of Search Output
  5. Search Strategy
  6. Compound Queries
  7. Advanced Features
  8. Trends in Information Retrieval

5 Housekeeping Operations

  1. Overview of Library Housekeeping Operations
  2. Acquisition
  3. Processing
  4. Circulation
  5. Serials Control
  6. Maintenance
  7. Procedural Model of Library Housekeeping Operations
  8. Computerized Subsystems

6 Software Packages- Features

  1. Evolution of Library Automation Software
  2. General Functions of Library Automation Software
  3. Requirements for Library Automation Software
  4. Implementation of Library Automation Software
  5. Library Automation Software Packages Available in India
  6. Evaluation of Library Automation Software
  7. Trends and Future Directions

7 Digitization- Concept, Need, Methods and Equipment

  1. Digitisation: Basics
  2. Need for Digitisation
  3. Selection of Materials for Digitisation
  4. Steps in the Process of Digitisation
  5. Digitisation: Input and Output Options
  6. Technology of Digitisation
  7. Tools of Digitisation
  8. Digitisation of Audio and Video
  9. Organising Digital Images
  10. Digital Library Softwares
  11. Planning and Implementation

8 Alerting Services

  1. Current Awareness Service (CAS)
  2. Selective Dissemination of Information (SDI)
  3. Electronic Clipping Services (ECS)
  4. News Filtering Services
  5. New Directions for Alerting Services

9 Bibliographic Fulltext Services

  1. What is Bibliographic Fulltext Service?
  2. The Need for Bibliographic Fulltext Service
  3. Players in Bibliographic Fulltext Service
  4. Fulltext Sources
  5. Examples of Fulltext Databases
  6. Information Technology and Fulltext Resources
  7. Copyright and Licensing Issues
  8. Likely Future Trends

10 Document Delivery Services

  1. Historical Perspective
  2. Document Delivery Service
  3. Modes of Document Delivery Service
  4. Electronic Document Delivery Service
  5. Steps in Document Delivery
  6. Some Document Supplying Agencies
  7. Copyright Facilitators

11 Reference Services

  1. Reference Service
  2. Need for Reference Service
  3. Reference Service Process
  4. Digital Reference Service
  5. Evaluation of Digital Reference Service
  6. Major Digital Reference Services Projects
  7. Expert Systems in Reference Service
  8. Future of Reference Service

12 Basics of Internet

  1. History of Internet
  2. Growth of Internet
  3. Internet Architecture
  4. Accessing the Internet
  5. Internet Service Providers (ISPs)
  6. Hardware and Software for Internet
  7. Internet Protocols

13 Search Engines

  1. Search Engines: Definitions
  2. Search Engines: Evolution
  3. How Do Search Engines Work?
  4. Search Engines: Categories
  5. Choosing a Search Engine
  6. Searching the Web: Search Techniques
  7. Search Results
  8. Meta Tags
  9. Search Engines: Evaluation
  10. Important Search Engines

14 Internet Services

  1. World Wide Web
  2. Importance of the Web
  3. How does the Web Work?
  4. Web Servers
  5. Web Browsers
  6. Plug-ins or Helper Programs
  7. Using Web Browser
  8. Mark-up Languages
  9. SGML
  10. XML
  11. HTML

15 Internet Information Resources

  1. Internet Information Resources
  2. Types of Internet Resources
  3. Searching the Internet: Where to Start
  4. How to Keep Up-to-Date with New Internet Resources

16 Evaluation of Internet Resources

  1. Need for Evaluation
  2. Quality Assessment
  3. Evaluation Tools on the Net
  4. Evaluating Information Resources
  5. Generic Criteria for Evaluation
  6. Specific Criteria for Evaluation
  7. Process Criteria
  8. Other Key Indicators