How does a scan differ from OCR text?
A scan records the appearance of a page. OCR attempts to recognise the characters shown in it. A searchable PDF can combine the page image with an additional text layer. A separate text export can support other processing steps.
An OCR text layer is a recognised version of the content and may differ from the original. Passages used for precise quotation, or text from difficult originals, should be checked against the page image.
What affects recognition quality?
Skewed pages, blur, bleed-through and uneven backgrounds can make recognition more difficult. Fonts, languages, columns, tables and footnotes may also create different requirements. Historical publications should therefore be assessed using representative pages.
We agree the level of checking for the intended use. Basic full-text search may require a different level of work from a text collection intended for further processing. We do not promise universally error-free recognition.
- Review originals and representative problem pages
- Define the required text and file outputs
- Agree how reading order and unusual layouts should be handled
- Specify checking and any correction work
Which output is useful for your project?
A searchable PDF supports research within individual works. If text is to be imported into other systems, we agree a suitable separate output and its association with the source document. Metadata or article boundaries complement the full text when individual publications need to be discoverable.
For an enquiry, please send a few typical pages, information about the language and a description of the intended use. This helps define the required work more precisely.
Frequently asked questions
Can you process existing PDF files?
Yes. We first check whether the files contain page images, existing text or a mixture, and assess the quality of the source material.
Does OCR reliably translate text?
OCR recognises characters. Translation is a separate step whose target language, subject area and level of review are agreed separately.
Does OCR automatically turn tables into structured data?
Not necessarily. Recognised text does not automatically preserve correct rows, columns and relationships. Reliable table structure can require additional processing and checks.