Explain: shorter answers, streamed, remembered, and prepared for settings #59

Merged
jcoffey-dev merged 2 commits from feature/explain-faster into main 2026-09-26 23:30:08 +00:00
Owner

Implements amendment 1 of the Explain spec (inbuxa-drafts/specs/ai-explain.md, EX-22 to EX-28): John asked for all four speed-ups.

What changes

  • Shorter answers (EX-22): three or four sentences; max_tokens 400 -> 160; cut at 700 chars (was 1,200).
  • Streaming (EX-23): POST /api/explain (authenticated as JMAP is) answers text/event-stream: delta pieces, then done with the whole explanation, or one error with the same types as the JMAP SetErrors. All checks run first. A client that leaves doesn't stop the call, so its answer is remembered. inbuxa:Explanation/set is unchanged for non-streaming clients.
  • Remembered answers (EX-24, EX-25): per node, in memory only, 1,000 answers / 24 h, keyed by kind + facts + grounding + prompt version + model (xxh3, stable across builds). Shared by server-level admins only (tenants are refused before any lookup). No model call, no hourly count, no slot.
  • Prepared answers (EX-26): resources/explain/settings.json.gz ships 717 answers for settings at their defaults (keyed without the model), generated by the ignored test prepare_setting_explanations against Qwen3 4B Instruct 2507 on two local GPUs in 129 s. A changed value or description goes to the live model.
  • Provenance (EX-27): source (model / remembered / prepared), answeredAt, preparedFor on inbuxa:Explanation.
  • Prompt reuse (EX-28): the per-request marker and the per-question reference notes moved from the system prompt to the user message, so the system prompt is byte-identical per kind and llama.cpp reuses it.

Measured (Qwen3 4B Q4_K_M, llama.cpp b11160, 4 cores, same machine)

before, total after, first text after, total asked again
setting 10.8 s 1.5 s 4.4 s instant
spam verdict 14.8 s 2.7 s 10.2 s instant
log line 10.2 s 2.3 s 5.1 s instant

A setting at its default is now answered from the prepared set instantly. Checked end to end: default -> prepared; changed value -> model (6.6 s); repeat -> remembered.

Checks

cargo test -p inbuxa-features (93), -p jmap --lib inbuxa (incl. new memory, prompt and generator-shape tests); clippy clean on changed files; name/notice/context fork checks clean.

Note on the prepared texts

They're model-written and labeled as such in the console ("Prepared for release 2026.9.27 with qwen3-4b-instruct-2507", plus the existing caution). Spot checks are mostly accurate but some are generic. preparedFor says 2026.9.27; if the release gets another number, rerun the generator (it keeps existing answers, so it takes seconds).

Console: inbuxa-admin feature/explain-streaming. Docs: inbuxa.org docs/explain-faster (merge after release).

Implements amendment 1 of the Explain spec (`inbuxa-drafts/specs/ai-explain.md`, EX-22 to EX-28): John asked for all four speed-ups. ## What changes - **Shorter answers (EX-22):** three or four sentences; `max_tokens` 400 -> 160; cut at 700 chars (was 1,200). - **Streaming (EX-23):** `POST /api/explain` (authenticated as JMAP is) answers `text/event-stream`: `delta` pieces, then `done` with the whole explanation, or one `error` with the same types as the JMAP `SetError`s. All checks run first. A client that leaves doesn't stop the call, so its answer is remembered. `inbuxa:Explanation/set` is unchanged for non-streaming clients. - **Remembered answers (EX-24, EX-25):** per node, in memory only, 1,000 answers / 24 h, keyed by kind + facts + grounding + prompt version + model (xxh3, stable across builds). Shared by server-level admins only (tenants are refused before any lookup). No model call, no hourly count, no slot. - **Prepared answers (EX-26):** `resources/explain/settings.json.gz` ships 717 answers for settings at their defaults (keyed without the model), generated by the ignored test `prepare_setting_explanations` against Qwen3 4B Instruct 2507 on two local GPUs in 129 s. A changed value or description goes to the live model. - **Provenance (EX-27):** `source` (`model` / `remembered` / `prepared`), `answeredAt`, `preparedFor` on `inbuxa:Explanation`. - **Prompt reuse (EX-28):** the per-request marker and the per-question reference notes moved from the system prompt to the user message, so the system prompt is byte-identical per kind and llama.cpp reuses it. ## Measured (Qwen3 4B Q4_K_M, llama.cpp b11160, 4 cores, same machine) | | before, total | after, first text | after, total | asked again | |---|---|---|---|---| | setting | 10.8 s | 1.5 s | 4.4 s | instant | | spam verdict | 14.8 s | 2.7 s | 10.2 s | instant | | log line | 10.2 s | 2.3 s | 5.1 s | instant | A setting at its default is now answered from the prepared set instantly. Checked end to end: default -> prepared; changed value -> model (6.6 s); repeat -> remembered. ## Checks `cargo test -p inbuxa-features` (93), `-p jmap --lib inbuxa` (incl. new memory, prompt and generator-shape tests); clippy clean on changed files; name/notice/context fork checks clean. ## Note on the prepared texts They're model-written and labeled as such in the console (\"Prepared for release 2026.9.27 with qwen3-4b-instruct-2507\", plus the existing caution). Spot checks are mostly accurate but some are generic. `preparedFor` says 2026.9.27; if the release gets another number, rerun the generator (it keeps existing answers, so it takes seconds). Console: inbuxa-admin `feature/explain-streaming`. Docs: inbuxa.org `docs/explain-faster` (merge after release).
jcoffey-dev added 1 commit 2026-09-26 23:12:20 +00:00
Explain: shorter answers, streamed, remembered, and prepared for settings
ci / fork-checks (pull_request) Successful in 1m31s
ci / build (pull_request) Failing after 5m18s
7e7eca0883
ai-explain spec, amendment 1 (EX-22 to EX-28):
- answers are three or four sentences, max_tokens 160, cut at 700 chars;
- POST /api/explain streams the answer as server-sent events;
- each node remembers answers in memory (1,000, 24 h), keyed by the facts,
  prompt version and model, shared by server-level administrators;
- resources/explain/settings.json.gz ships answers for settings at their
  defaults, generated with prepare_setting_explanations (717 for 2026.9.27);
- the system prompt no longer carries the per-request marker, so a model
  server can reuse it;
- inbuxa:Explanation gains source, answeredAt and preparedFor.
jcoffey-dev added 1 commit 2026-09-26 23:25:35 +00:00
Calibration test: pass the new stream argument to request::body
ci / fork-checks (pull_request) Successful in 50s
ci / build (pull_request) Successful in 4m29s
ad648d8d12
jcoffey-dev merged commit fdbc72e574 into main 2026-09-26 23:30:08 +00:00
jcoffey-dev deleted branch feature/explain-faster 2026-09-26 23:30:08 +00:00
Sign in to join this conversation.
No Reviewers
No labels
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: inbuxa/inbuxa-server#59