Four steps with four different results
A PDF file extension does not tell you which of these results it contains. A PDF may consist of page images, images with recognised text, or digitally created text. Likewise, a folder of scans is not yet a structured catalogue.
| Step | Result | Example use |
|---|---|---|
| Scanning | Digital image of the printed page | View a historical illustration |
| OCR | Automatically recognised text | Search for a term within a work |
| Metadata | Descriptive information about a document | Filter by author or publication year |
| Article segmentation | Separated contributions with associations | Use an individual paper from a collected volume |
Example: making a journal volume searchable
Suppose a volume contains several papers. A complete scan first makes all pages readable digitally. OCR adds the ability to search recognised full text. Captured article boundaries then allow individual contributions to be identified.
Each contribution can have a title, author and page range recorded in a table. A unique identifier links the record to the correct file. This is a simplified workflow example, not a description of a particular client project.
How can quality be checked meaningfully?
Checks should address the particular output. Scans need to be checked for completeness, order and legibility, among other criteria. OCR text additionally needs comparison with the original. Metadata checks include correct field mapping and consistent treatment of missing information.
A general label such as “digitised” does not answer these questions. Before a project starts, the required outputs, acceptance criteria and treatment of difficult material should therefore be described.
- Are all agreed pages present and correctly associated?
- Is the text adequately captured for its intended purpose?
- Are missing or uncertain details identifiable?
- Do field names, file names and identifiers correspond?
- Can a sample delivery be used in the destination system?
What role do standards and file formats play?
Dublin Core defines metadata terms including title, creator, language and identifier. ALTO is an XML format for describing recognised text and page layout. Such specifications help express requirements precisely.
A particular standard or specialised export format is not automatically included in a digitisation order. The agreed deliverables remain decisive. Donauconnect clarifies the required fields and output formats before processing begins.