How to run OCR on primary documents
Turn scanned, image-based PDFs into text-searchable documents from the Document tab in the Uwazi viewer.
Prerequisites
Before you start, make sure you have:
- Admin access to your Uwazi instance
- An OCR service connected to your instance, which an operator sets up at install time
- A primary document saved as a PDF, with a supported language set
OCR has two parts. First, an admin turns on the OCR trigger in settings. Then an admin or editor runs OCR from the entity's Document tab. The sections below cover each part.
Steps
Turn on the OCR trigger
-
Go to Settings > Collection.
-
Find the Services card.
If you don't see this card, your instance isn't connected to an OCR service.
-
Turn on the Document OCR trigger toggle.
-
Select Save.
Run OCR on a document
-
Open the entity and select the Document tab.
The tab behaves the same way in the left pane and in the right pane.
-
Look at the footer below the document. The OCR button sits on the left, and the page navigator on the right.
The button shows only for admin and editor roles, and only when the trigger is on.
-
Read the button's label. If it reads Unsupported OCR language, change the document's language first.
-
Select OCR PDF.
-
The button changes to In OCR queue and stops responding. The OCR service now works in the background, so you can leave the page.
-
Wait for the button to show OCR with a check mark. Uwazi reloads the entity on its own.
You can't undo OCR from the interface, and you can't run it twice on the same file.
Change a document's language
- Select the Files tab.
- Find the file under PRIMARY DOCUMENTS.
- Open the menu on its row and select Change language.
- In the right pane, pick a language in the Language field.
- Select Done.
- Go back to the Document tab and read the button again.
The list holds every Uwazi language, not only the ones your OCR service reads. So a language you pick here can still come back as unsupported.
Reading the OCR button
The button shows the document's current state. Use this table to read it:
| What you see | What it means |
|---|---|
| OCR PDF | Ready to start. Select it to begin. |
| In OCR queue | Submitted. The service is processing the document. |
| OCR with a check mark | Done. The document is now text-searchable. |
| Unsupported OCR language | The service can't read the document's language. |
| OCR error | The service couldn't process the document. |
Hover over the button for a tooltip. It gives the date of the last change, or the reason the run failed.
The button responds only when it reads OCR PDF. After an error, you can't start another run from the interface. Ask your instance administrator for help.
Result
Your primary document is now a text-searchable PDF. Readers can search and select its text. The original scan moves to SUPPORTING FILES on the Files tab, and Uwazi points any references in the old file at the new one.
See also
- How to extract metadata from your documents — pull field values from documents once they hold text
- How to extract paragraphs from documents — split a text-searchable document into paragraph entities
- How machine-learning (ML) extraction works — why OCR comes before extraction in the pipeline