tofu Snippet #
Classification: Confidential (describes source code, model architecture and infrastructure).
Runs the precise snippet pass on demand. When a user presses Improve on a page
that the automatic gate flagged, tofu-snippet-main redraws that page’s suggested
snippets with the full SAM3 predictor: polygons fitted to each document’s mask
instead of the gate’s rectangles.
The automatic gate runs in tofu Preprocess Docs, in front of OCR and extraction. The two are separate images on purpose: the predictor needs ~5Gi and ~37s per page, which cannot sit in front of every upload.
Where it sits #
Improve → bonsapi (page → "improving") → tofu-snippet-improve-v1 → tofu-snippet-main
├── done → suggestions, "predictor"
└── failed → back to "gate"
bonsapi accepts Improve only while the document is awaiting review. It marks the
page’s suggested_snippets_source as improving and moves the document to
PROCESSING. The worker overwrites the page’s suggested_snippets and marks them
predictor. A page it could not finish goes back to gate, so its card never waits
on a pass that is not running. The document returns to NEEDS_REVIEW once none of
its pages is still improving.
The model #
| Image | tofu-snippet |
| Deployment | tofu-snippet-main |
| Queue | tofu-snippet-improve-v1 |
| Per page | ~37s |
| Memory (request/limit) | 5Gi / 6Gi |
| Weights | production_openvino/, baked from tofu-snippet-model-${env} at build time |
An OpenVINO export of facebook/sam3 (Meta SAM License), converted offline; the
AGPL ultralytics wrapper is deliberately avoided. It runs Promptable Concept
Segmentation with the prompt fixed at "document", from a 33KB embedding committed
as an .npz asset, so the 1.6GB language encoder is not a dependency.
Operations #
- Scaling. KEDA on queue depth, one message per page, 1 to 6 replicas in prod.
- Rollouts. The
readinessProbechecks a marker written once the predictor has loaded, so a rollout cannot retire the old pod before the new one can consume. - Memory.
tests/measure_memory.pyproduces the numbers the envelope is sized from. Re-run it after changing the model or the thread count.
See also #
apps/tofu-snippet/README.md: implementation detail, model provenance and the prompt-embedding regeneration recipe.- ENG-7130 (detection engine), ENG-7250 (Improve worker).