jcoffey-dev is traveling from Thursday 1 October through Sunday 4 October. Issues and pull requests are welcome, and will get an answer after that. Thanks for your patience.
Implements amendment 1 of the Explain spec (inbuxa-drafts/specs/ai-explain.md, EX-22 to EX-28): John asked for all four speed-ups.
What changes
Shorter answers (EX-22): three or four sentences; max_tokens 400 -> 160; cut at 700 chars (was 1,200).
Streaming (EX-23):POST /api/explain (authenticated as JMAP is) answers text/event-stream: delta pieces, then done with the whole explanation, or one error with the same types as the JMAP SetErrors. All checks run first. A client that leaves doesn't stop the call, so its answer is remembered. inbuxa:Explanation/set is unchanged for non-streaming clients.
Remembered answers (EX-24, EX-25): per node, in memory only, 1,000 answers / 24 h, keyed by kind + facts + grounding + prompt version + model (xxh3, stable across builds). Shared by server-level admins only (tenants are refused before any lookup). No model call, no hourly count, no slot.
Prepared answers (EX-26):resources/explain/settings.json.gz ships 717 answers for settings at their defaults (keyed without the model), generated by the ignored test prepare_setting_explanations against Qwen3 4B Instruct 2507 on two local GPUs in 129 s. A changed value or description goes to the live model.
Prompt reuse (EX-28): the per-request marker and the per-question reference notes moved from the system prompt to the user message, so the system prompt is byte-identical per kind and llama.cpp reuses it.
Measured (Qwen3 4B Q4_K_M, llama.cpp b11160, 4 cores, same machine)
before, total
after, first text
after, total
asked again
setting
10.8 s
1.5 s
4.4 s
instant
spam verdict
14.8 s
2.7 s
10.2 s
instant
log line
10.2 s
2.3 s
5.1 s
instant
A setting at its default is now answered from the prepared set instantly. Checked end to end: default -> prepared; changed value -> model (6.6 s); repeat -> remembered.
Checks
cargo test -p inbuxa-features (93), -p jmap --lib inbuxa (incl. new memory, prompt and generator-shape tests); clippy clean on changed files; name/notice/context fork checks clean.
Note on the prepared texts
They're model-written and labeled as such in the console ("Prepared for release 2026.9.27 with qwen3-4b-instruct-2507", plus the existing caution). Spot checks are mostly accurate but some are generic. preparedFor says 2026.9.27; if the release gets another number, rerun the generator (it keeps existing answers, so it takes seconds).
Console: inbuxa-admin feature/explain-streaming. Docs: inbuxa.org docs/explain-faster (merge after release).
Implements amendment 1 of the Explain spec (`inbuxa-drafts/specs/ai-explain.md`, EX-22 to EX-28): John asked for all four speed-ups.
## What changes
- **Shorter answers (EX-22):** three or four sentences; `max_tokens` 400 -> 160; cut at 700 chars (was 1,200).
- **Streaming (EX-23):** `POST /api/explain` (authenticated as JMAP is) answers `text/event-stream`: `delta` pieces, then `done` with the whole explanation, or one `error` with the same types as the JMAP `SetError`s. All checks run first. A client that leaves doesn't stop the call, so its answer is remembered. `inbuxa:Explanation/set` is unchanged for non-streaming clients.
- **Remembered answers (EX-24, EX-25):** per node, in memory only, 1,000 answers / 24 h, keyed by kind + facts + grounding + prompt version + model (xxh3, stable across builds). Shared by server-level admins only (tenants are refused before any lookup). No model call, no hourly count, no slot.
- **Prepared answers (EX-26):** `resources/explain/settings.json.gz` ships 717 answers for settings at their defaults (keyed without the model), generated by the ignored test `prepare_setting_explanations` against Qwen3 4B Instruct 2507 on two local GPUs in 129 s. A changed value or description goes to the live model.
- **Provenance (EX-27):** `source` (`model` / `remembered` / `prepared`), `answeredAt`, `preparedFor` on `inbuxa:Explanation`.
- **Prompt reuse (EX-28):** the per-request marker and the per-question reference notes moved from the system prompt to the user message, so the system prompt is byte-identical per kind and llama.cpp reuses it.
## Measured (Qwen3 4B Q4_K_M, llama.cpp b11160, 4 cores, same machine)
| | before, total | after, first text | after, total | asked again |
|---|---|---|---|---|
| setting | 10.8 s | 1.5 s | 4.4 s | instant |
| spam verdict | 14.8 s | 2.7 s | 10.2 s | instant |
| log line | 10.2 s | 2.3 s | 5.1 s | instant |
A setting at its default is now answered from the prepared set instantly. Checked end to end: default -> prepared; changed value -> model (6.6 s); repeat -> remembered.
## Checks
`cargo test -p inbuxa-features` (93), `-p jmap --lib inbuxa` (incl. new memory, prompt and generator-shape tests); clippy clean on changed files; name/notice/context fork checks clean.
## Note on the prepared texts
They're model-written and labeled as such in the console (\"Prepared for release 2026.9.27 with qwen3-4b-instruct-2507\", plus the existing caution). Spot checks are mostly accurate but some are generic. `preparedFor` says 2026.9.27; if the release gets another number, rerun the generator (it keeps existing answers, so it takes seconds).
Console: inbuxa-admin `feature/explain-streaming`. Docs: inbuxa.org `docs/explain-faster` (merge after release).
ai-explain spec, amendment 1 (EX-22 to EX-28):
- answers are three or four sentences, max_tokens 160, cut at 700 chars;
- POST /api/explain streams the answer as server-sent events;
- each node remembers answers in memory (1,000, 24 h), keyed by the facts,
prompt version and model, shared by server-level administrators;
- resources/explain/settings.json.gz ships answers for settings at their
defaults, generated with prepare_setting_explanations (717 for 2026.9.27);
- the system prompt no longer carries the per-request marker, so a model
server can reuse it;
- inbuxa:Explanation gains source, answeredAt and preparedFor.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Implements amendment 1 of the Explain spec (
inbuxa-drafts/specs/ai-explain.md, EX-22 to EX-28): John asked for all four speed-ups.What changes
max_tokens400 -> 160; cut at 700 chars (was 1,200).POST /api/explain(authenticated as JMAP is) answerstext/event-stream:deltapieces, thendonewith the whole explanation, or oneerrorwith the same types as the JMAPSetErrors. All checks run first. A client that leaves doesn't stop the call, so its answer is remembered.inbuxa:Explanation/setis unchanged for non-streaming clients.resources/explain/settings.json.gzships 717 answers for settings at their defaults (keyed without the model), generated by the ignored testprepare_setting_explanationsagainst Qwen3 4B Instruct 2507 on two local GPUs in 129 s. A changed value or description goes to the live model.source(model/remembered/prepared),answeredAt,preparedForoninbuxa:Explanation.Measured (Qwen3 4B Q4_K_M, llama.cpp b11160, 4 cores, same machine)
A setting at its default is now answered from the prepared set instantly. Checked end to end: default -> prepared; changed value -> model (6.6 s); repeat -> remembered.
Checks
cargo test -p inbuxa-features(93),-p jmap --lib inbuxa(incl. new memory, prompt and generator-shape tests); clippy clean on changed files; name/notice/context fork checks clean.Note on the prepared texts
They're model-written and labeled as such in the console ("Prepared for release 2026.9.27 with qwen3-4b-instruct-2507", plus the existing caution). Spot checks are mostly accurate but some are generic.
preparedForsays 2026.9.27; if the release gets another number, rerun the generator (it keeps existing answers, so it takes seconds).Console: inbuxa-admin
feature/explain-streaming. Docs: inbuxa.orgdocs/explain-faster(merge after release).