tofu Snippet

tofu Snippet #

Classification: Confidential (describes source code, model architecture and infrastructure).

Runs the precise snippet pass on demand. When a user presses Improve on a page that the automatic gate flagged, tofu-snippet-main redraws that page’s suggested snippets with the full SAM3 predictor: polygons fitted to each document’s mask instead of the gate’s rectangles.

The automatic gate runs in tofu Preprocess Docs, in front of OCR and extraction. The two are separate images on purpose: the predictor needs ~5Gi and ~37s per page, which cannot sit in front of every upload.

Where it sits #

Improve → bonsapi (page → "improving") → tofu-snippet-improve-v1 → tofu-snippet-main
                                                                      ├── done   → suggestions, "predictor"
                                                                      └── failed → back to "gate"

bonsapi accepts Improve only while the document is awaiting review. It marks the page’s suggested_snippets_source as improving and moves the document to PROCESSING. The worker overwrites the page’s suggested_snippets and marks them predictor. A page it could not finish goes back to gate, so its card never waits on a pass that is not running. The document returns to NEEDS_REVIEW once none of its pages is still improving.

The model #

Image tofu-snippet
Deployment tofu-snippet-main
Queue tofu-snippet-improve-v1
Per page ~37s
Memory (request/limit) 5Gi / 6Gi
Weights production_openvino/, baked from tofu-snippet-model-${env} at build time

An OpenVINO export of facebook/sam3 (Meta SAM License), converted offline; the AGPL ultralytics wrapper is deliberately avoided. It runs Promptable Concept Segmentation with the prompt fixed at "document", from a 33KB embedding committed as an .npz asset, so the 1.6GB language encoder is not a dependency.

Operations #

  • Scaling. KEDA on queue depth, one message per page, 1 to 6 replicas in prod.
  • Rollouts. The readinessProbe checks a marker written once the predictor has loaded, so a rollout cannot retire the old pod before the new one can consume.
  • Memory. tests/measure_memory.py produces the numbers the envelope is sized from. Re-run it after changing the model or the thread count.

See also #

  • apps/tofu-snippet/README.md: implementation detail, model provenance and the prompt-embedding regeneration recipe.
  • ENG-7130 (detection engine), ENG-7250 (Improve worker).