Identify individual contributions within larger publications
One file per volume is not sufficient for every workflow. If individual papers need to be catalogued, researched or incorporated into a dataset, associations are needed at article level.
Title pages, contents lists, supplements and differing page numbers can affect boundaries. We agree what counts as a separate unit and how these cases should be recorded.
How article segmentation works
- Define the source structure and the required unit
- Identify the beginning and end of individual contributions
- Capture the agreed title, author and page information
- Connect article files and records through clear identifiers
- Check boundaries and outputs to the agreed extent
Files and records that correspond
Possible outputs include individual article PDFs and associated metadata in CSV or JSON. Fields, identifiers, file names and the treatment of printed versus digital page numbers are agreed.
OCR can additionally make the content searchable. These associations can support further processing for libraries, publishers and AI companies. The actual fields and formats delivered depend on the agreed scope.
What helps us prepare
Send a representative volume or sample pages with a contents list. Tell us the approximate quantity, required metadata and destination system. Details of continuations or unclear article boundaries help us estimate the work.
Frequently asked questions
Is the printed page number the same as the PDF page?
Not always. Endpapers, covers and different numbering systems can cause differences. The page information required and its associations are therefore defined separately.
Can existing OCR files be incorporated?
Yes. We review the available files and establish how text, pages and articles can be connected.