Digitisation explained

Scanning, OCR and metadata: what is the difference?

A scan shows the page, OCR makes its characters machine-readable and metadata describes the document. Article segmentation divides a larger publication into individual contributions. These outputs complement one another but serve different purposes.

Four steps with four different results

A PDF file extension does not tell you which of these results it contains. A PDF may consist of page images, images with recognised text, or digitally created text. Likewise, a folder of scans is not yet a structured catalogue.

StepResultExample use
ScanningDigital image of the printed pageView a historical illustration
OCRAutomatically recognised textSearch for a term within a work
MetadataDescriptive information about a documentFilter by author or publication year
Article segmentationSeparated contributions with associationsUse an individual paper from a collected volume

Example: making a journal volume searchable

Suppose a volume contains several papers. A complete scan first makes all pages readable digitally. OCR adds the ability to search recognised full text. Captured article boundaries then allow individual contributions to be identified.

Each contribution can have a title, author and page range recorded in a table. A unique identifier links the record to the correct file. This is a simplified workflow example, not a description of a particular client project.

How can quality be checked meaningfully?

Checks should address the particular output. Scans need to be checked for completeness, order and legibility, among other criteria. OCR text additionally needs comparison with the original. Metadata checks include correct field mapping and consistent treatment of missing information.

A general label such as “digitised” does not answer these questions. Before a project starts, the required outputs, acceptance criteria and treatment of difficult material should therefore be described.

  • Are all agreed pages present and correctly associated?
  • Is the text adequately captured for its intended purpose?
  • Are missing or uncertain details identifiable?
  • Do field names, file names and identifiers correspond?
  • Can a sample delivery be used in the destination system?

What role do standards and file formats play?

Dublin Core defines metadata terms including title, creator, language and identifier. ALTO is an XML format for describing recognised text and page layout. Such specifications help express requirements precisely.

A particular standard or specialised export format is not automatically included in a digitisation order. The agreed deliverables remain decisive. Donauconnect clarifies the required fields and output formats before processing begins.

Further technical reading

Discuss your project

Send us details of your material, volume and intended result. We will agree the scope of work and prepare a proposal.

info@donauconnect.com