Page Orientation #
Classification: Confidential (describes source code and model architecture).
Detects pages uploaded sideways or upside down and records the clockwise rotation that makes them upright, before OCR or extraction sees them.
Where it sits #
Inside tofu-preprocess-docs, first thing on every document bonsai-doc-convert
hands it, bank statements included. The worker downloads each page, classifies it,
PATCHes document_page.rotation, and copies the rotation into the pages of the OCR
job and the extraction job it publishes next, since both read rotation from their
own message.
upload → doc-convert (split → render → register pages)
└── tofu-preprocess-docs (classify orientation → PATCH rotation)
├── OCR job (tofu-ocr-gpu rotates the page upright first)
├── snippet gate (unchanged, polygons stay in page space)
└── extraction (hinoki already rotates by document_page.rotation)
It runs a second time for snippets. The whole-page pass says nothing about the documents on a page: a scan holding two receipts at different angles needs one rotation per receipt. So a snippet’s first extraction goes through the same worker, orientation only, and the worker classifies the snippet’s own crop before forwarding the job. On knowledge v2 the oriented snippets also go to OCR, since only the page they were cut from was OCRed at upload:
accept suggestions / snip by hand → snippet page (starts at its parent's rotation)
└── first extraction → bonsapi → tofu-preprocess-docs (crop → classify → PATCH rotation)
└── extraction (hinoki crops, then rotates)
It has to sit ahead of the extraction rather than beside it: hinoki reads rotation from its own message, which bonsapi builds the moment the webapp starts the extraction. Improve needs nothing of its own; it only replaces suggestions before they are accepted.
It runs in tofu-preprocess-docs beside the snippet gate, a 9.1MB ONNX model baked into that
image and run on OpenVINO, ~76ms per page on one CPU thread, ~60MB RSS. Nothing about the stored page files or
thumbnails changes; rotation is the same field the manual rotate button writes,
and every consumer applies it at read time.
Method #
One MobileNetV2 predicts all four angles directly, in orientation.py:
- Letterbox — the page is scaled onto a grey-128 square at 320², never squashed. Snippet crops reach 1:5 aspect ratios, and squashing one of those to a square destroys the line structure the decision rests on. The resize goes through PIL, not cv2: PIL’s BILINEAR antialiases when shrinking and cv2’s does not, which is worth half a point of accuracy.
- Classify — four logits for 0°/90°/180°/270°, softmaxed.
- Average the four views — a page at angle
a, seen from viewk, sits ata + k, so the four rotations are four looks at one question. This is worth 2.4 points of accuracy and makes the answer rotation-consistent by construction. Cost is four inferences, and the page is a 320px letterbox rather than a full render, so the whole thing is 76ms p50 against 56ms for the single-view two-stage model it replaces.
The label is the clockwise angle the page currently sits at, so
rotation = (360 - label) % 360. Model licence, checksum and training recipe:
apps/tofu-preprocess-docs/models/doc_orient/README.md.
Where the weights live #
In s3://tofu-snippet-model-${env}/doc_orient/orientation4.onnx, next to the
snippet models, not in the repo: the weights are derived from customer page imagery.
The Dockerfile’s model-downloader stage bakes it into the image at build time from a
presigned URL and fails the build on a missing or empty file, so a missing model is a
failed build rather than a pod that quietly leaves every page as it arrived.
model_artifacts.py resolves it under PREPROCESS_DOCS_MODEL_DIR. DocOrient’s MIT
LICENSE stays in apps/tofu-preprocess-docs/models/doc_orient/ and is copied into
the image beside the weights.
What this replaces, and why #
The first implementation was DocOrient’s
two-stage method (MIT): a projection-profile energy test decided the axis, then a binary
MobileNetV2 decided upright against upside-down. The weights here are a fine-tune of that
model’s, so the MIT licence still applies and the LICENSE stays in the asset directory.
The axis stage is what failed. It binarises at mean brightness, so in a photograph the
desk or table becomes ink and the profile measures the background rather than the text.
Measured on 685 production pages with every prediction reviewed by hand, it rotated
12.5% of pages that were already upright, against the 0.5% bar below. On snippet
crops, 58.7%. Long till receipts photographed on wood were wrong almost every time,
and sparse clean invoices drew a confident 180.
PaddleOCR’s PP-LCNet_x1_0_doc_ori was rejected earlier for related reasons: on four real sideways bank statements it scored 0.47–0.72 and mislabelled one. A 4-way classifier on a 224px centre crop sees almost nothing on a sparse landscape table.
Behaviour #
- On for every page of a document with an extraction queue, for an organization with
the
preprocess_docsfeature (canary); without it uploads and snippets are extracted as they arrive. A detection rerun never re-orients: by then a user may have turned a page by hand. - A snippet is oriented on its first extraction, while its invoice, direct expense or
bank statement holds no AI-extracted data, and again on the next extraction after a
re-crop, which drops its record. Otherwise it is not re-oriented once it holds its own
public_metadata.orientation, and a rotation a user set by hand (rotated_by_user) never is. A new snippet drops its parent’s record and mark, which would otherwise make it look oriented. - The score is the averaged probability of the winning angle (0.25–1.0).
- Applied when the score reaches the threshold in
orientation.pyand the answer differs from the page’s current rotation:MIN_CONFIDENCE(0.95) for whole pages,MIN_CONFIDENCE_CROP(0.75) for snippets, where rotation is seven times more common and raising the bar only discards correct answers. A snippet starts at its parent’s rotation, so a confident upright answer turns it back to 0. Below the threshold the page is left alone: a wrong rotation costs a user a click on a page that was already fine. - Below the threshold the text lines get a say first. The model’s errors are all upright
against upside down, and a long receipt can fool it confidently from two of its four
views.
text_direction.pyruns PaddleOCR’s line detector and text-line direction classifier on the page turned to the model’s answer, reads each line as it stands and turned half round, and when one reading wins byMIN_VOTE_MARGINthe page turns to it (text_check=confirmed|flipped). Too few lines, lines running the wrong way or a close vote leave the page alone as before. Only unsure pages pay for it, about 250ms on one thread for a receipt. - Every prediction lands in
public_metadata.orientation(label,score,applied,skew,text_check,blank) and in the metricspreprocess_docs.orientation.predictions.totalandpreprocess_docs.orientation.score, taggedblankandpage_type(page|snippet) since the two run at different thresholds. - A user overriding an applied rotation is the false-positive signal to watch.
- Blank pages (under
MIN_INK_FRACTIONof pixels standingINK_CONTRASTgrey levels off the page’s median background, either way) are reported upright with score 1.0, recorded withblank: true, and never turned; they have no orientation, and without the guard all four views answer at random. - Fail open: a model that will not load, a page that will not classify or a
rotation bonsapi will not take keeps the rotation it arrived with, logs, and increments
preprocess_docs.orientation.failures.total(stage=unavailable|classify|patch|changed), so an outage shows as a rate, not as silence. - If it misfires broadly, raising
MIN_CONFIDENCEis the lever; a wrong rotation is one click to undo in review. No backfill of existing documents.
Downstream consumers #
| consumer | how it honors rotation |
|---|---|
| webapp | CSS rotate on the stored image and on box overlays; a snippet’s by rotation − skew, its boxes in the straightened crop |
| hinoki (extraction) | rotates the image before the LLM, Transform maps coordinates back; a snippet’s highlights stay in its straightened crop |
| tofu-ocr-gpu | rotates before detection, maps polygons back to the stored frame; a snippet is cut from its parent first and keeps its straightened crop |
| tofu-bbox | rotates before the VLM, Transform.to_highlight_normalized_from_llm_unit maps boxes back, to the same frames as tofu-ocr-gpu |
| snippet gate | unchanged: polygons are in stored-page space and the webapp rotates them |
Eval #
apps/tofu-preprocess-docs/tests/eval_orientation.py rotates a directory of upright
pages four ways and prints a confusion matrix, accuracy per threshold and latency.
Change the threshold only once upright pages rotated by mistake are ≤ 0.5% and rotated
pages recovered ≥ 95% at the chosen threshold.
Run it on real pages. The synthetic set the first implementation was accepted on did not show any of the failures above, because our own corpus is where they live.
Measured on production pages, 2026-09-16 #
2,255 documents pulled from production for snippet-model training, opted-in entities only, rendered by this same pipeline, every orientation reviewed by hand: 685 whole pages, and 1,570 snippet crops of which 413 were cut by the detector where its box matched a human label at IoU ≥ 0.75. Splits are entity-disjoint. Rotation is far more common after the split than before it — 2.8% of whole pages arrive rotated against 18.5% of crops — because documents get placed sideways on the flatbed.
Held-out validation, 339 documents scored at all four rotations:
| two-stage DocOrient | this model | |
|---|---|---|
| all | 0.605 | 0.985 |
| whole pages | 0.875 | 0.981 |
| snippet crops | 0.485 | 0.987 |
| threshold | upright disturbed | rotated recovered | |
|---|---|---|---|
| whole pages | 0.95 | 1 of 104 | 89% |
| snippet crops | 0.75 | 2 of 235 | 97% |
Every remaining error is upright-versus-upside-down or its 90/270 equivalent; the sideways axis is never wrong on either domain.
These numbers chose the model on validation, so they are optimistic, and the test split has not been read yet. 104 upright validation pages also cannot resolve a 0.5% rate — it resolves about 1% — so the false-positive claim needs the test split folded in before it is settled.