tofu-web Assistant #
This document is the architecture of the in-app AI assistant: the
generic, roster-driven orchestrator in apps/tofu-web/src/server/ai/, the roles it dispatches
to (support, sales, bookkeeper), and the channels that run turns through it (web, headless
CLI, WhatsApp). The operational rules for changing it — the four recipes — live next to the code
in apps/tofu-web/AGENTS.md. The review ruleset is apps/tofu-web/REVIEW.md. This document
explains the shape those two enforce.
1. The one-turn flow #
Every turn — web, headless or WhatsApp — goes through the same seven steps. Nothing in
orchestrator.ts branches on which role or which channel is running; a role contributes its own
AgentDefinition, toolScope and optional buildBrief/judgment, and a channel contributes its
own authenticate, tools, capabilityTools and optional per-role roleCapabilityTools. The
orchestrator only ever calls what the RoleDefinition/ChannelDefinition interfaces expose.
- Grounding. On every leg, continuation included,
readGroundingFactsreadsforwardedProps.grounding— what the screen (or, for a channel with no screen, the channel’s owngroundinghook) reports about what is currently on view — and turns it into a handful of sentences (groundingSentence, capped at 4,000 characters across every source) appended to the system prompt. - Judge.
prepareTurn(prepare-turn.ts) callsjudgeTurn(judge.ts), which sends one merged System One request per turn covering the whole roster: which role should answer (route), how much reasoning effort the turn needs (tier), whether it needs the user’s own data or is purely about the app (bundle), whether it is in scope at all, and — namespaced underrole.<id>.— every routable role’s own fact-extraction questions in the same call. If that request fails (or its circuit breaker is open),judgeTurnasksoptions.routeFallback, a second, cheaper model call that decides only the route and the scope; a turn neither could clear fails closed,inScope: false, which adds the out-of-scope note but still lets a genuine Tofu question be answered. Either way the judgment isdegraded, and if the routed role registers an extractor inFACT_FALLBACK_BY_ROLE,prepareTurnmakes a third model call to read that role’s facts out of the conversation. A continuation leg (see step 6) judges nothing: it reuses the decision stored for the leg it continues, and is refused if that decision is gone. - Role brief. The routed role’s own
buildBrief(facts, ...)(if it declares one) renders a block appended to its system prompt for this turn: what is already established, what the current play needs, how the person is doing, and the one move to make now.bookkeeper’s andsales’s briefs both follow this same four-beat shape. Its{capability}placeholders resolve against the channel’scapabilityToolsoverlaid withroleCapabilityTools[role.id], so one channel can bind{handoff}to a different tool per role. - Tool offer.
offerTools(tools/offer.ts) takes the MCP reads the routed role’s source can deliver plus the channel’s own client-tool descriptors, narrows them to the role’stoolScope(scopeMcpToolSource/scopeClientTools/toolIsInScope), folds in the role’s always-offered set and the judgment’s ranked names, and caps the result atMAX_OFFERED_TOOLS(20). The judgment itself only ranks when there are more thanMAX_TURN_TOOLS(12) candidates to narrow. - Model.
chatWithFallback/trackedChatruns the flattened agent (system prompt, grounding, brief, the role’s standing notes and the channel’spromptFragment; server tools, betas,maxOutputTokens) against the offered tool set, with therunidentity (threadId,runId,parentRunId,resume,state) always spread through —run-identity.test.tsscans the source and fails on achat(call that forgets to. - Interrupt / continuation. Every leg, fresh or continuation, stores its decision (role, tier,
facts and the exact tools offered) under its own
runIdbefore the model is asked. A client-tool call the browser must execute (or atool-approvalgate) comes back as a TanStack AI interrupt bound to thatrunId. The browser resolves it and posts a continuation withparentRunIdset;createOrchestratorreads the decision that leg stored (an in-process copy first, then Redis viaturn-decision-store.ts, keyed by the conversation and the run), so any pod can take the continuation, and dispatches straight to the same role and tools without re-judging. A continuation whose decision has expired is refused, never judged afresh.src/ai/session.ts(the one module that constructs aChatClientor touches its interrupts/queue/updateOptions, seeREVIEW.md) owns the browser side of this: resolvingtool-approvalinterrupts throughchannel.confirm(), asking once for each (a submission that fails waits forretry(), the chat pane’s Retry, and is not asked again), declining the pending ones indiscardPendingTool()(the chat pane calls it beforestop()), answering a client-tool call a page reload cut off with an error result rather than running it again, and pushingclient.updateOptions({ tools: channel.tools() })before every send. - Narration / post-turn. The UI narrates a command’s
calling/donei18n keys, since the model itself cannot observe when a tool call lands. When a leg’s stream ends — continuation legs included, whose resumed tool calls are read back off the interrupted message —runPostTurnreads the tool calls and results, applies the role’sevidenceRulesandremember()s any fact they prove, runshooks.afterTurn(only for a fully delivered, non-aborted reply, handed the set of tools the legoffered) andremember()s what it derives, then logshooks.kpi. The turn’s main fact write is not here but inprepareTurn, whichremember()s what the judgment learned before the model runs; these are up to two more. A hand-off play on a channel that binds no hand-off tool for the role is markedhandoffUnavailablebyafterTurnafter one turn, so it is not picked again every turn. The web binds each role’s{handoff}to its command and offers both as commands (commands/handoff.ts):book_sales_callputs HubSpot’s booking scheduler on the canvas andopen_support_chatthe support contact page. Support is offered both, and a standing note keeps help and setup requests off the sales booking. A broken hook or a Redis hiccup here is logged and swallowed; it must never turn an already-streamed reply into a 500.
flowchart TD
Start([Leg arrives]) --> Ground["1 · Grounding<br/>readGroundingFacts → groundingSentence<br/>(every leg)"]
Ground --> IsCont{Continuation leg?<br/>parentRunId set}
IsCont -- yes --> Cached["Reuse the decision stored<br/>for parentRunId — no re-judging<br/>(refused if it has expired)"]
Cached --> Brief
IsCont -- "no, fresh turn" --> Judge["2 · Judge<br/>judgeTurn: one merged System One call —<br/>route / tier / bundle / inScope / role.*.facts"]
Judge -. "call fails" .-> Fallback["routeFallback<br/>(2nd, cheaper call; inScope false if it fails too)<br/>+ the role's FACT_FALLBACK_BY_ROLE extractor"]
Fallback --> Brief
Judge -- "call succeeds" --> Brief["3 · Role brief<br/>role.buildBrief(facts)"]
Brief --> Offer["4 · Tool offer<br/>offerTools: MCP reads + client tools,<br/>scoped by role, ranked, capped at 20"]
Offer --> Model["5 · Model<br/>decision stored, then chatWithFallback / trackedChat —<br/>system prompt + grounding + brief, offered tools"]
Model --> HasInterrupt{Client-tool call or<br/>tool-approval gate?}
HasInterrupt -- yes --> Interrupt["6 · Interrupt<br/>browser resolves the call"]
HasInterrupt -- no --> Narrate["7 · Narration / post-turn<br/>runPostTurn: evidenceRules → facts,<br/>hooks.afterTurn/kpi, ≤ 2 FactStore writes"]
Interrupt --> Narrate
Interrupt -. "continuation posted<br/>(parentRunId set)" .-> Start
2. Tool taxonomy #
Every tool a turn can offer is one of two origins, and every tool of either origin carries a category and an effect.
Origins
- MCP tools — bonsapi operations reached through
tofu-external-mcp, declared once as anMCPToolinlibs/typescript/mcp-tool/src/tool-groups/*.tsand grouped into anMCPToolGroupwith one category for the whole group.mcp-client.tsallowlistsREAD_TOOL_NAMESat the transport level — every stable tool whosereadOnlyHintis set, derived rather than hand-picked, so “the connection can only read” is a property of the annotations, not of a list. No MCP write ever reaches the connection, and the orchestrator offers only reads; a role’stoolScopenarrows what it may actually call on top of that floor. - Client tools — commands the browser (or another channel’s own runtime) executes, declared as
ChannelTools on aChannelDefinition.tools.routes/api/chat.tsbuilds the web channel per request:GLOBAL_COMMANDSplus the feature-surface commands (ALL_SURFACES) the browser declares that turn, which is when that surface is mounted. WhatsApp’s is a short, channel-specific set (whatsapp/channel-tools.ts) with no browser to resolve anything against: each role’s hand-off tool (book_sales_call,open_support_chat) andselect_entity.
Categories (ToolCategory, libs/typescript/mcp-tool/src/tool-category.ts) — one per
MCPToolGroup (entity, invoice, bank_statement, direct_expense, extraction, document,
integration_accounting_connection, integration_setup, integration_trigger, knowledge,
organization, reference_data, user_account, and the six entity_accounting_* groups), plus
five client-only categories with no MCPToolGroup of their own: accounting (a client-side
accounting action, e.g. reconnect), navigation, layout, session and chat (the chrome and
conversation commands a channel exposes alongside its MCP tools). A role’s toolScope.categories
is a set membership check against this list — “every accounting read” or “no navigation” is one
set, not nineteen tool names.
Effects — read, reversible, irreversible, gesture or human_only (ToolEffect,
libs/typescript/mcp-tool/src/tool-effect.ts). For an MCP tool the effect is deriveToolEffect’s
output from its readOnlyHint/destructiveHint annotations. For a client tool it is tagged by the
channel that declares the ChannelTool: the web channel copies each command’s own category and
reversibility (read/reversible/irreversible, see AGENTS.md §5) in routes/api/chat.ts,
and WhatsApp tags its hand-off tools irreversible and select_entity reversible. gesture (a
client-side action with no server effect) and human_only (offered to the model but only ever run
by a person) are values a channel may tag with; no server-side channel does today. A role’s
toolScope never assigns an effect, it only admits: toolScope.effects narrows on top of
category — sales and support scope to {read}, bookkeeper to {read, reversible}, so no
irreversible MCP write is ever in their offered sets, regardless of category. For a client
command declared with defineCommand, commands/runner.ts refuses an irreversible one outright
when the caller is the assistant, returning requires_human so the model can describe the action
while a person performs it — but that refusal is commands/runner.ts’s own behaviour, not a
property of the irreversible effect itself. A channel tool built outside defineCommand carries
no such gate: WhatsApp’s hand-off tools (whatsapp/channel-tools.ts) are effect: 'irreversible'
yet are admitted by name through each role’s toolScope.include and called directly like any
other tool — safe only because what they do (send a hand-off notice) is not itself destructive,
not because anything stops an irreversible channel tool from running.
flowchart LR
subgraph "Origin — how a tool exists"
MCP["MCP tool<br/>tool-groups/*.ts,<br/>reached via tofu-external-mcp (reads only)"]
Client["Client tool<br/>ChannelDefinition.tools —<br/>web GLOBAL_COMMANDS + declared surfaces,<br/>or WhatsApp's own short set"]
end
subgraph "Category — ToolCategory (24 values)"
MCPCat["19 MCP-backed categories<br/>one per MCPToolGroup —<br/>entity, invoice, document, knowledge, …"]
ClientCat["5 client-only categories<br/>accounting, navigation, layout,<br/>session, chat"]
end
subgraph "Effect — ToolEffect (5 values)"
Derived["read / reversible / irreversible<br/>deriveToolEffect, from an MCP<br/>tool's readOnlyHint / destructiveHint"]
Declared["read / reversible / irreversible<br/>tagged on the ChannelTool by its channel —<br/>web copies the command's reversibility"]
Assigned["gesture / human_only<br/>a channel may tag a ChannelTool with them;<br/>none does today"]
end
MCP --> MCPCat
MCP --> Derived
Client --> ClientCat
Client --> Declared
Client -.-> Assigned
MCPCat -.-> Scope["role.toolScope.categories<br/>set-membership narrowing"]
ClientCat -.-> Scope
Derived -.-> ScopeEff["role.toolScope.effects<br/>set-membership narrowing"]
Declared -.-> ScopeEff
Assigned -.-> ScopeEff
toolScope also carries an include/exclude name allowlist, typed as
ToolScopeSpec<AssistantToolName> (tools/tool-name.ts: every MCP tool, client command and
declared channel capability tool name) so a misspelt name fails to compile. include admits a tool
whatever its category and effect — how support reaches every command and how a role admits a
client tool whose category or effect its scope leaves out, such as render_on_canvas for sales —
and exclude wins over it. effects is the axis that cannot be left open: it is required, and
toolIsInScope admits a tool only when the role names its effect, so an empty effects (or one
an untyped caller left out) admits nothing but include, which is what turns include into an
actual allowlist. An absent categories still means “any category”.
3. The per-turn budget #
Three numbers bound the cost of a turn, independent of how large the roster or the tool catalogue grows:
| Budget | Value | Where enforced |
|---|---|---|
| Model calls before the reply | ≤ 3 | judge.ts’s judgeTurn: one merged System One request per turn (skipped while its circuit breaker is open); only when that fails or is skipped, one routeFallback call; and only on such a degraded turn, if the routed role registers one in FACT_FALLBACK_BY_ROLE, prepare-turn.ts’s fact-fallback extraction. Steady state is one. A continuation leg makes none; it reuses the decision stored for its parentRunId (turn-decision-store.ts), or is refused. |
| Offered tools | ≤ 20 | tools/offer.ts’s MAX_OFFERED_TOOLS. The judgment’s own ranking step only narrows when there are more than MAX_TURN_TOOLS (12) candidates to choose from; below that it ranks nothing and offerTools still enforces the 20 cap as the hard ceiling. |
| Redis commands | ≤ 11 | Before the model, ≤ 7. Six at most in prepare-turn.ts’s module comment: (1) one GET for the active-role pointer (active-role-store.ts); (2) if that pointer names a role with a fact store, one GET for that role’s facts; (3) once routed, if that role has a fact store and is not the pointer’s, one GET for its own facts; (4) if the routed role learned anything, one remember() (a GET then a compare-and-set EVAL); (5) if the route changed, or a fresh turn could not clear it, one pointer write() (a bare SET). Steady state costs two. A conversation’s first turn skips (1)–(2) and clears what /reset clears in one DEL, so a new chat never inherits the last one’s role or facts. The seventh is the orchestrator’s: every leg stores its decision (one SET, turn-decision-store.ts) before the model is asked, so a continuation can resume on any pod; a continuation leg judges nothing and costs that SET and at most one GET, of the leg it continues (none when this pod still holds it in memory). After the reply, ≤ 4: runPostTurn runs on every leg and adds up to two more remember()s, one for what evidenceRules proved and one for what afterTurn derived. Each remember() repeats its GET/EVAL only when another turn wrote those facts first, the one way a leg goes past eleven. A channel’s own commands (WhatsApp’s per-number dedupe, lock and link reads) and a tool’s own writes (select_entity’s session fact) are outside this count. |
A roster or tool catalogue growing in size never moves any of these three numbers — that flatness is the point of merging the judgment and deriving the tool offer from scopes rather than per-role lists.
4. Where to go next #
- Changing this architecture — read
apps/tofu-web/AGENTS.mdfirst; it is the enforced, operational form of this document; where the two disagree, the lint config and the tests are the tiebreaker, not either document. - Adding a role, a client tool, an MCP read or a channel —
AGENTS.md’s four recipes. - The chat runtime spike and Platform Decisions — Chat Runtime Spike and Platform Decisions for the history and the alternatives considered before this shape was settled on.