Digitisation explained

Metadata and article records: which fields are useful?

Metadata describes a publication through fields such as title, author, publication year or language. Article records also connect contributions to page ranges and files. The necessary fields depend on the collection and destination system.

Which information is prepared?

The field set depends on your collection and destination system. Not every original contains all bibliographic details. Missing or unclear information is marked accordingly; we do not invent publication details or identifiers.

Field groupPossible informationPurpose
Bibliographic descriptionTitle, author, publisher, publication yearDescribe and distinguish publications
Content descriptionLanguage, existing subject classification, document typeFilter and organise collections
AssociationsCollection identifier, file name, page rangeConnect records to source documents
Processing statusMissing values, unclear associationsMake review needs traceable

What does data cleaning involve?

We structure existing information and standardise spellings and formats according to agreed rules. Permitted changes are defined beforehand. An empty value, an illegible entry and a field that has not been checked should remain distinguishable.

Work can start from existing spreadsheets and digital files. A sample file and a desired target structure help establish which fields should be retained, renamed or checked further.

How are individual articles identified?

A journal volume can contain many independent articles. Article segmentation identifies the boundaries of each contribution and links titles, page ranges and associated files. Scan page numbers do not necessarily match printed page numbers.

Traceable links between text, articles and source publications are useful for libraries, publishers and AI companies. Export fields and any additional processing are agreed for each project.

How does the data reach your system?

Possible outputs include CSV or JSON alongside the appropriate document files. The file extension alone is not enough: field names, character encoding, required values and associations must fit the destination system. A sample delivery helps test an import before larger quantities are processed.

Frequently asked questions

Does a publication need an ISBN to be catalogued?

No. Works without ISBNs can be described using their available bibliographic details and an agreed collection identifier.

Does a CSV export work with every catalogue?

No. The export structure must match the receiving system. A file format alone does not guarantee a successful import.

Can you clean data without scanning books?

Yes. Data cleaning, metadata preparation and article associations can also be based on existing files.

Further technical reading

Discuss your project

Send us details of your material, volume and intended result. We will agree the scope of work and prepare a proposal.

info@donauconnect.com