Data & MorePipelineOCR

How a document becomes text

How the ocr service turns a REQUEST_OCR document into text and identity signals, and which data-index fields that writes. RapidOCR is the page reader. Tesseract is used only for the passport MRZ.

Pixels in, signals out, nothing kept on disk.

4stages, own process pools
10PDF pages read, max
200DPI raster default
20 / 5MB gate, PDF / other
120 sexpress call timeout

Image dgm/ocr:amd64-d3c980e · state as of 2026-10-02

Section 01Position

Between collect and profile

OCR does not read customer sources itself. It polls Elasticsearch, downloads bytes from universal-collector, and writes the result back onto the same data document.

scan collect OCR profile classify Elasticsearch · data universal-collector poll · write back GET bytes
this functionpipeline orderdependency
Section 02Intake

How a document moves

Work is not a RabbitMQ queue. The scheduler queries data for DS_Status = REQUEST_OCR, PDFs first, then other types, in workspace-priority order (10 down to 1; priority 0 is excluded). A matching checksum in the same source and company is copied from the twin instead of being read again (COPY_DUPLICATES).

The download is GET universal-collector:8000/sources/{source_id}/file/{document_id}. Bytes are staged on the node only: /ocr-stage (emptyDir) for ordinary sources, /ocr-tmp (RAM) for Slack. The file is deleted when the result is saved. There is no shared OCR volume.

Gates · config index, doc id ocr
  • PDFs above max_pdf_size (default 20 MB) are not OCR'd.
  • Everything else above max_size (default 5 MB) is not OCR'd.
  • A source is only OCR'd when container_configuration.ocr_type is LOCAL_OCR.

Download result

ResultWhat happens to the document
404 · 410Terminal. REQUEST_PROFILING with no text, and a DS_StatusMessage the Kout OCR board buckets on.
401 · 403Same terminal path. The message says the source needs reconnecting.
429 · 5xx · timeoutStays THIS_IS_LOCKED_OCR and is retried. DS_FetchRetries counts attempts; after the cap the doc is SKIPPED.
Missing sourceSKIPPED. OCR never deletes the document.
Stage crash, no textERROR, DS_StatusMessage = OCR failed: …. An empty result is not saved, because the profiler would hide it as LOW_RISK.
Section 03Stages

The four stages

Each stage is its own process pool. A document is passed as a BaseTask (text, faces, signatures, MRZ) from one queue to the next. The pools scale on queue depth. RapidOCR is the large pool; its ONNX sessions use OCR_RAPIDOCR_THREADS (default 1) because each page runs three sessions.

INTAKE Elasticsearch DS_Status=REQUEST_OCR Scheduler PDFs first · prio 10→1 universal-collector GET /sources/…/file/… Node stage /ocr-stage · /ocr-tmp PDF image · HEIC STAGES STAGE 1 pdf_reader PyMuPDF · ≤10 pages STAGE 2 RapidOCR ONNX · rapid-layout STAGE 3 Combined signature · face · ID STAGE 4 MRZ PassportEye · tesserocr pixels ID hit text layer suffices pixels skipped no ID MRZ labels WRITE · SAME DATA DOCUMENT DS_Value · DS_Status= REQUEST_PROFILING · DM_OCR · DS_Pic_PII · DS_Signature_Number owner, path, name, tags untouched · staged file deleted
main pathconditional branchsuccess statewrite payload

Fig. 2 · Images skip stage 1. Only documents the ID cascade marks reach stage 4.

Stage 1 · pdf_reader

Read the text layer first

Only PDFs stop here. Images go straight to RapidOCR. PyMuPDF reads up to 10 pages. extract_pdf_text repairs fonts whose /ToUnicode map decodes everything to spaces, then reads the text layer. Form widget values are appended. Embedded images that the text layer already covers are not counted.

OCR of the pixels is skipped, and the text layer is saved, when any of these is true:

  • no uncovered images remain
  • the text contains a skip_pdf_text keyword from the values index
  • the file name contains a skip_pdf_name keyword

Otherwise the text layer is kept on the task (marked (OCR)) and the file continues, so a scan-behind-text still gets RapidOCR.

Stage 2 · RapidOCR

The text engine

pdf2image (Poppler) rasterises PDF pages at OCR_RAPIDOCR_DPI (default 200). Photos, including HEIC, are opened with Pillow.

RapidOCR (ONNX Runtime) returns lines and word boxes. rapid-layout (yolov8n_layout_general6) puts those lines back into reading order, but only when its regions cover at least 60% of the lines. Below that, the model has missed too much and the raw line order is kept.

OpenCV deskews, thresholds, and rotates. A zone that looks like an MRZ is rotated upright and upside down and scored later by the MRZ reader.

Stage 3 · Combined

Signature, face, ID

One worker decodes the page once and runs three detectors.

  • Signature. YOLOv8s (tech4humans/yolov8s-signature-detector). Optional: if the image was built without HUGGINGFACE_TOKEN, this detector is off and the other two still run. Dense pages (more than 1500 characters) skip it. Near-empty pages are still scanned, because a signed page looks empty to OCR.
  • Face. OpenCV YuNet. A face count from 1 to 5 sets DS_Pic_PII.
  • ID card. An OpenCV Haar cascade. A hit sends the document to the MRZ stage. Anything else is saved here.

Face and ID are both skipped when the page has nothing to look at, or the text is under 55 characters or over 1500.

Stage 4 · MRZ

Two reads, best check digits win

Only documents the ID cascade marked. The higher check-digit score wins.

  • PassportEye finds the MRZ band and OCRs the crop by shelling out to the tesseract binary (language mrz).
  • tesserocr reads the whole page in-process, upright and flipped, with a whitelist of A–Z, digits, and <.
A valid MRZ yields document number · personal number · name
Section 04Output

Fields written

All of these are an update on the existing data document. Owner, path, name, and tags are left alone. Text is written to DS_Value in Elasticsearch. Nothing is written to a text file on disk.

FieldWhat OCR puts there
DS_ValueExtracted text: the PDF text layer, RapidOCR text, or both. MRZ values that the company marks for hashing are replaced in this text by their hash.
DS_StatusREQUEST_PROFILING on success. ERROR if a stage crashed and there is no text. SKIPPED when the file can never be fetched.
DS_StatusMessageSet on the failure paths above. Empty on a normal save.
DS_DataOCRedTimestamp of this OCR pass.
DS_DataLoadedSame timestamp.
ocr_fromPRIMARY.
DS_Pic_PIITrue when the face count is 1–5.
DS_Pic_PII_NumberFace count from YuNet.
DS_Signature_NumberSignatures found by YOLO. Zero when the model is absent or the page was skipped.
DS_EntryType_ListLabel names, parallel with the two arrays below.
DS_EntryType_ValuesThe values. Hashed when the company's algorithm list says that label is a security value.
DS_EntryType_IndexAlways 0 for OCR-produced entries.
DS_FetchRetriesDownload attempts. Cleared on a successful fetch. Not part of the success write.
DS_OcrOwnerWhich OCR pod claimed the document (eu-ocr01, eu01, eu02 on EU).

Entry-type labels OCR itself adds

DM_OCRAI_IDPASSPORT_OCRID_CARD_OCRVISA_OCRPERSONAL_NUMBER_OCRFULL_NAME_OCRNAME_OCR
LabelWhen
DM_OCREvery successful save. Value [DM_OCR]. Tells the profiler this document was OCR'd, so it must not be downgraded to LOW_RISK.
AI_IDThe ID cascade hit. Value 1.
PASSPORT_OCR
ID_CARD_OCR
VISA_OCR
A valid MRZ whose type starts with P, I, or V. Value is the document number.
PERSONAL_NUMBER_OCRThe MRZ personal number, when present.
FULL_NAME_OCRGiven names plus surname from the MRZ.
NAME_OCROne entry per word of that name.

A face crop plus the MRZ is also posted to known-persons-service at POST /ocr/insert. That registry is separate from the data document.

Section 05Scope
Express OCR

Same stages, no index

POST /ocr/extract (and the MRZ and word-box variants) runs the same stages on uploaded bytes and returns the text and labels in the HTTP response. It does not read or write Elasticsearch.

graph-ingestion uses it when a DLP path has an image and no extracted text. The call waits up to 120 seconds.

Out of scope

What this function does not do

  • Profiling, classification, and policy actions happen after REQUEST_PROFILING. OCR does not call the profilers.
  • It does not delete documents.
  • It does not keep the file after the result is saved.

A PDF with a real text layer and no uncovered images never reaches RapidOCR. The text layer is the result. RapidOCR is for pixels the text layer does not already explain.