Turn page images into usable full text
A page image alone initially contains no machine-readable characters. OCR produces a recognised text version. In a searchable PDF, the page image remains visible while an additional text layer supports searching.
If you need text independently of the PDF, we agree a suitable export and its association with the source document. Translation and capturing bibliographic metadata are additional services.
Recognition and checking suited to the material
- Review sample files, languages, fonts and typical page layouts
- Define the required text layer or separate export
- Run OCR and account for agreed special cases
- Check the results to the agreed extent and apply correction where included
Output quality and limitations
OCR can contain recognition errors. Columns, footnotes, historical typefaces and poor scan quality may require extra work. We therefore agree the level of checking and how difficult passages will be handled.
Broad full-text search has different requirements from text intended for further processing or precise quotation. Criteria are set for your intended use; universally error-free recognition is not promised.
Preparing your enquiry
Send representative pages, including some difficult examples. Page count, language, existing file formats and intended use are helpful. Tell us whether you need a searchable PDF, separately usable text or both.
Frequently asked questions
Can you use my existing scans?
Yes. We first assess the file structure and quality of the source material to define appropriate processing.
Are recognised tables automatically structured correctly?
No. Reliable associations between rows, columns and values can require additional preparation and checking.