Which information is prepared?
The field set depends on your collection and destination system. Not every original contains all bibliographic details. Missing or unclear information is marked accordingly; we do not invent publication details or identifiers.
| Field group | Possible information | Purpose |
|---|---|---|
| Bibliographic description | Title, author, publisher, publication year | Describe and distinguish publications |
| Content description | Language, existing subject classification, document type | Filter and organise collections |
| Associations | Collection identifier, file name, page range | Connect records to source documents |
| Processing status | Missing values, unclear associations | Make review needs traceable |
What does data cleaning involve?
We structure existing information and standardise spellings and formats according to agreed rules. Permitted changes are defined beforehand. An empty value, an illegible entry and a field that has not been checked should remain distinguishable.
Work can start from existing spreadsheets and digital files. A sample file and a desired target structure help establish which fields should be retained, renamed or checked further.
How are individual articles identified?
A journal volume can contain many independent articles. Article segmentation identifies the boundaries of each contribution and links titles, page ranges and associated files. Scan page numbers do not necessarily match printed page numbers.
Traceable links between text, articles and source publications are useful for libraries, publishers and AI companies. Export fields and any additional processing are agreed for each project.
How does the data reach your system?
Possible outputs include CSV or JSON alongside the appropriate document files. The file extension alone is not enough: field names, character encoding, required values and associations must fit the destination system. A sample delivery helps test an import before larger quantities are processed.
Frequently asked questions
Does a publication need an ISBN to be catalogued?
No. Works without ISBNs can be described using their available bibliographic details and an agreed collection identifier.
Does a CSV export work with every catalogue?
No. The export structure must match the receiving system. A file format alone does not guarantee a successful import.
Can you clean data without scanning books?
Yes. Data cleaning, metadata preparation and article associations can also be based on existing files.