Skip to main content

How to run OCR on primary documents

Turn scanned, image-based PDFs into text-searchable documents from the Document tab in the Uwazi viewer.

Prerequisites​

Before you start, make sure you have:

  • Admin access to your Uwazi instance
  • An OCR service connected to your instance, which an operator sets up at install time
  • A primary document saved as a PDF, with a supported language set

OCR has two parts. First, an admin turns on the OCR trigger in settings. Then an admin or editor runs OCR from the entity's Document tab. The sections below cover each part.

Steps​

Turn on the OCR trigger​

  1. Go to Settings > Collection.

  2. Find the Services card.

    If you don't see this card, your instance isn't connected to an OCR service.

  3. Turn on the Document OCR trigger toggle.

  4. Select Save.

Run OCR on a document​

  1. Open the entity and select the Document tab.

    The tab behaves the same way in the left pane and in the right pane.

  2. Look at the footer below the document. The OCR button sits on the left, and the page navigator on the right.

    The button shows only for admin and editor roles, and only when the trigger is on.

  3. Read the button's label. If it reads Unsupported OCR language, change the document's language first.

  4. Select OCR PDF.

  5. The button changes to In OCR queue and stops responding. The OCR service now works in the background, so you can leave the page.

  6. Wait for the button to show OCR with a check mark. Uwazi reloads the entity on its own.

warning

You can't undo OCR from the interface, and you can't run it twice on the same file.

Change a document's language​

  1. Select the Files tab.
  2. Find the file under PRIMARY DOCUMENTS.
  3. Open the menu on its row and select Change language.
  4. In the right pane, pick a language in the Language field.
  5. Select Done.
  6. Go back to the Document tab and read the button again.

The list holds every Uwazi language, not only the ones your OCR service reads. So a language you pick here can still come back as unsupported.

Reading the OCR button​

The button shows the document's current state. Use this table to read it:

What you seeWhat it means
OCR PDFReady to start. Select it to begin.
In OCR queueSubmitted. The service is processing the document.
OCR with a check markDone. The document is now text-searchable.
Unsupported OCR languageThe service can't read the document's language.
OCR errorThe service couldn't process the document.

Hover over the button for a tooltip. It gives the date of the last change, or the reason the run failed.

note

The button responds only when it reads OCR PDF. After an error, you can't start another run from the interface. Ask your instance administrator for help.

Result​

Your primary document is now a text-searchable PDF. Readers can search and select its text. The original scan moves to SUPPORTING FILES on the Files tab, and Uwazi points any references in the old file at the new one.

See also​