tofu Preprocess Docs #
Classification: Confidential (describes source code, model architecture and infrastructure).
Prepares every document of an organization with the preprocess_docs feature (canary)
between bonsai-doc-convert and OCR/extraction. It turns
sideways and upside-down pages upright (see Page Orientation),
then detects how many documents are on a page image and splits a page that holds
more than one. Detection replaced two earlier approaches: the vision-LLM
classifier’s own bounding boxes, and the OCR-anchored refinement explored in
ENG-6272. The gate moved here from tofu-snippet when orientation arrived.
It sits in front of OCR and extraction and answers a routing question in about
a second a page. The expensive precise pass stays in
tofu Snippet’s tofu-snippet-main, a separate image that
runs only when a user asks for it. Keeping them apart is the whole design: the
precise model needs ~5Gi and ~37s per page, which cannot sit in front of every
upload.
Where it sits #
For an organization with preprocess_docs, bonsai-doc-convert sends every upload
with an extraction queue here at the end of a split, and decides whether detection
runs. Without the flag, or when the flags cannot be read, doc-convert publishes OCR and
the extraction job itself, as it did before this service. The worker orients the pages,
publishes OCR, then routes:
upload → doc-convert → tofu-preprocess-docs-v1 → tofu-preprocess-docs
orient pages → OCR job
├── one document → extraction queue
├── two or more → held for review
└── two or more, auto-extract entity
→ bonsapi cuts the snippets and
extracts them with the rest
The worker forwards the payload it was given rather than answering a question
about it. bonsai-doc-convert builds the extraction job up front and hands it
over inside the preprocess message, so the gate either publishes that payload
onward or holds the document. Nothing polls, and nothing waits on a timeout.
An entity on AUTO_EXTRACT_SNIPPET skips the review: bonsapi cuts each flagged page into its suggested snippets
and starts one extraction per snippet plus one for the document’s other pages, as
a reviewer accepting the boxes would. A page it cannot cut stays in review, and a
failed call holds the whole document as usual.
Bank statements are oriented but skip detection — one statement spanning many pages is a single document, so there is nothing to split.
bonsapi sends a snippet’s first extraction here too, and its next one after a re-crop, orientation only, so each snippet gets a rotation of its own rather than its parent page’s (see Page Orientation).
Detection follows the entity’s snippet_type, whatever the document’s length:
NO_SNIPPET: off.SKIP_MULTI_SNIPPET: a multi-document page is flagged and held, and the user draws the snippets. No boxes are suggested.AUTO_SUGGEST_SNIPPET: the page gets suggested boxes and is held for review.AUTO_EXTRACT_SNIPPET: the page is cut and extracted without review.
With preprocess_docs on, bonsai-invoice never runs its LLM snippet check, so this
worker is the only detector. Without it, uploads skip this worker and bonsai-invoice’s
LLM flags multi-document pages for manual drawing, as before preprocessing.
The two models #
Gate (Sam3SnippetGate) |
Predictor (Sam3SnippetPredictor) |
|
|---|---|---|
| Deployment | tofu-preprocess-docs |
tofu-snippet-main |
| Queue | tofu-preprocess-docs-v1 |
tofu-snippet-improve-v1 |
| Per page | ~0.9s | ~37s |
| Memory (request/limit) | 1280Mi / 2Gi | 5Gi / 6Gi |
| Output | plain rectangles | mask-fitted polygons |
The gate is a distilled EfficientSAM3 (TinyViT-11M backbone) with no segmentation head, so it can only draw rectangles. That is enough to decide how many documents are on a page, which is the only question on the critical path. Redrawing them precisely is the predictor’s job, triggered by the user’s “Improve” action.
Both are OpenVINO exports of ours, baked into their images at build time: the gate
and the orientation classifier for this service, the predictor for tofu-snippet,
all from tofu-snippet-model-${env}, shared on purpose. Nothing is pulled from HuggingFace at
build or run time. The predictor’s upstream lineage is an ONNX export of
facebook/sam3 under the Meta SAM License, converted offline. The ultralytics SAM3 wrapper is
deliberately avoided: it is AGPL-3.0.
Flatbed scans: the second stage #
PCS works on photographs, where each document is an object against a contrasting background. The gate does not split a flatbed scan of several receipts on one white sheet: with no background to segment against, the whole sheet legitimately matches “document” and returns one page-sized box, which the area-ratio guard drops.
Those pages get a second stage, the snippet detector: D-FINE-seg nano, fine-tuned
on our pages (flatbed scans included), on OpenVINO, 17MB and 120-230ms a page. It runs
only when the gate kept no box, so the gate’s drawing always wins where both could
answer. Its suggestions are marked suggested_snippets_source: "detector", and
Improve is not offered for them: the predictor refines gate boxes, and there are none.
The webapp hides the action and bonsapi refuses it.
Failing open #
Every extraction of an organization with preprocess_docs passes through this
service, so a dead worker must not stop them. A model that will not load, a page that will not read, and a bonsapi that
will not answer all end the same way: log it, forward anyway. An unoriented
page keeps the rotation it arrived with, and an undetected page has no suggestions;
neither is a document that stops moving. Only a broker or bonsapi failure during routing itself raises, because
that is the case a retry can fix.
This is why a pod whose models fail to load still reports ready — refusing readiness would stall uploads to protect a feature that is allowed to be absent.
Operations #
- Scaling. KEDA on queue depth, floor of one replica, no scale-to-zero: at zero, nothing extracts. The queue’s depth counts jobs: one per document, except that a document over 100 pages is split into 100-page batches that separate pods run at once. The batch that completes the set routes the whole document, on every batch’s results.
- Drains. Every PDB uses
maxUnavailable: 1rather thanminAvailable: 1, so a node can be drained at a single replica (ENG-7305). - Rollouts. A
readinessProbechecks a marker file written once the models have loaded, so a rollout cannot retire the old pod before the new one can consume. - Duplicate work. Jobs are claimed in Redis with an ownership token refreshed every 30s, so a worker that runs past its claim TTL cannot extend or delete the claim its successor now holds. A finished job’s claim stays a done mark for 15 minutes, so a duplicate delivery skips it instead of forwarding the document again; a failed job’s claim is released for the retry. A delivery that finds a claim held waits out its TTL once, so a hard-killed worker’s job runs on the next worker instead of being dropped.
- Failures. A failed page does not abort its siblings. A detection rerun then
raises so
bonsai-mqdead-letters it for replay; an upload routes anyway, the failed page undetected. - Threads. The inference pool is sized from
CONTAINER_CPU_LIMITdivided byRABBITMQ_MAX_CONCURRENT_JOBS, so changing an overlay’s cpu limit is enough.
Re-measuring memory #
tests/measure_memory.py measures the model path’s peak RSS; the envelopes above
come from cgroup memory.peak on real pages at 2 vCPU. Re-run both after changing a
model, a threshold, or the thread count — the envelopes are deliberately just above measured peak, so an
unmeasured model change risks an OOM kill mid-inference.
See also #
apps/tofu-preprocess-docs/README.md— implementation detail and model provenance.- tofu Snippet — the Improve worker and its predictor.
- ENG-7130 (detection engine), ENG-7248 (gate-first), ENG-7277 (orientation).