How the ocr service turns a REQUEST_OCR document into text and identity signals, and which data-index fields that writes. RapidOCR is the page reader. Tesseract is used only for the passport MRZ.
Pixels in, signals out, nothing kept on disk.
Image dgm/ocr:amd64-d3c980e · state as of 2026-10-02
OCR does not read customer sources itself. It polls Elasticsearch, downloads bytes from universal-collector, and writes the result back onto the same data document.
Work is not a RabbitMQ queue. The scheduler queries data for DS_Status = REQUEST_OCR, PDFs first, then other types, in workspace-priority order (10 down to 1; priority 0 is excluded). A matching checksum in the same source and company is copied from the twin instead of being read again (COPY_DUPLICATES).
The download is GET universal-collector:8000/sources/{source_id}/file/{document_id}. Bytes are staged on the node only: /ocr-stage (emptyDir) for ordinary sources, /ocr-tmp (RAM) for Slack. The file is deleted when the result is saved. There is no shared OCR volume.
max_pdf_size (default 20 MB) are not OCR'd.max_size (default 5 MB) is not OCR'd.container_configuration.ocr_type is LOCAL_OCR.| Result | What happens to the document |
|---|---|
| 404 · 410 | Terminal. REQUEST_PROFILING with no text, and a DS_StatusMessage the Kout OCR board buckets on. |
| 401 · 403 | Same terminal path. The message says the source needs reconnecting. |
| 429 · 5xx · timeout | Stays THIS_IS_LOCKED_OCR and is retried. DS_FetchRetries counts attempts; after the cap the doc is SKIPPED. |
| Missing source | SKIPPED. OCR never deletes the document. |
| Stage crash, no text | ERROR, DS_StatusMessage = OCR failed: …. An empty result is not saved, because the profiler would hide it as LOW_RISK. |
Each stage is its own process pool. A document is passed as a BaseTask (text, faces, signatures, MRZ) from one queue to the next. The pools scale on queue depth. RapidOCR is the large pool; its ONNX sessions use OCR_RAPIDOCR_THREADS (default 1) because each page runs three sessions.
Fig. 2 · Images skip stage 1. Only documents the ID cascade marks reach stage 4.
Only PDFs stop here. Images go straight to RapidOCR. PyMuPDF reads up to 10 pages. extract_pdf_text repairs fonts whose /ToUnicode map decodes everything to spaces, then reads the text layer. Form widget values are appended. Embedded images that the text layer already covers are not counted.
OCR of the pixels is skipped, and the text layer is saved, when any of these is true:
skip_pdf_text keyword from the values indexskip_pdf_name keywordOtherwise the text layer is kept on the task (marked (OCR)) and the file continues, so a scan-behind-text still gets RapidOCR.
pdf2image (Poppler) rasterises PDF pages at OCR_RAPIDOCR_DPI (default 200). Photos, including HEIC, are opened with Pillow.
RapidOCR (ONNX Runtime) returns lines and word boxes. rapid-layout (yolov8n_layout_general6) puts those lines back into reading order, but only when its regions cover at least 60% of the lines. Below that, the model has missed too much and the raw line order is kept.
OpenCV deskews, thresholds, and rotates. A zone that looks like an MRZ is rotated upright and upside down and scored later by the MRZ reader.
One worker decodes the page once and runs three detectors.
tech4humans/yolov8s-signature-detector). Optional: if the image was built without HUGGINGFACE_TOKEN, this detector is off and the other two still run. Dense pages (more than 1500 characters) skip it. Near-empty pages are still scanned, because a signed page looks empty to OCR.DS_Pic_PII.Face and ID are both skipped when the page has nothing to look at, or the text is under 55 characters or over 1500.
Only documents the ID cascade marked. The higher check-digit score wins.
tesseract binary (language mrz).<.All of these are an update on the existing data document. Owner, path, name, and tags are left alone. Text is written to DS_Value in Elasticsearch. Nothing is written to a text file on disk.
| Field | What OCR puts there |
|---|---|
| DS_Value | Extracted text: the PDF text layer, RapidOCR text, or both. MRZ values that the company marks for hashing are replaced in this text by their hash. |
| DS_Status | REQUEST_PROFILING on success. ERROR if a stage crashed and there is no text. SKIPPED when the file can never be fetched. |
| DS_StatusMessage | Set on the failure paths above. Empty on a normal save. |
| DS_DataOCRed | Timestamp of this OCR pass. |
| DS_DataLoaded | Same timestamp. |
| ocr_from | PRIMARY. |
| DS_Pic_PII | True when the face count is 1–5. |
| DS_Pic_PII_Number | Face count from YuNet. |
| DS_Signature_Number | Signatures found by YOLO. Zero when the model is absent or the page was skipped. |
| DS_EntryType_List | Label names, parallel with the two arrays below. |
| DS_EntryType_Values | The values. Hashed when the company's algorithm list says that label is a security value. |
| DS_EntryType_Index | Always 0 for OCR-produced entries. |
| DS_FetchRetries | Download attempts. Cleared on a successful fetch. Not part of the success write. |
| DS_OcrOwner | Which OCR pod claimed the document (eu-ocr01, eu01, eu02 on EU). |
| Label | When |
|---|---|
| DM_OCR | Every successful save. Value [DM_OCR]. Tells the profiler this document was OCR'd, so it must not be downgraded to LOW_RISK. |
| AI_ID | The ID cascade hit. Value 1. |
| PASSPORT_OCR ID_CARD_OCR VISA_OCR | A valid MRZ whose type starts with P, I, or V. Value is the document number. |
| PERSONAL_NUMBER_OCR | The MRZ personal number, when present. |
| FULL_NAME_OCR | Given names plus surname from the MRZ. |
| NAME_OCR | One entry per word of that name. |
A face crop plus the MRZ is also posted to known-persons-service at POST /ocr/insert. That registry is separate from the data document.
POST /ocr/extract (and the MRZ and word-box variants) runs the same stages on uploaded bytes and returns the text and labels in the HTTP response. It does not read or write Elasticsearch.
graph-ingestion uses it when a DLP path has an image and no extracted text. The call waits up to 120 seconds.
REQUEST_PROFILING. OCR does not call the profilers.A PDF with a real text layer and no uncovered images never reaches RapidOCR. The text layer is the result. RapidOCR is for pixels the text layer does not already explain.