Phase 7: AI-assisted query authoring (autocomplete, explain, fix, optimize, NL translation)

Adds a self-hosted (Ollama, qwen2.5-coder) model provider abstraction
with a pluggable opt-in cloud adapter, schema grounding, and a shared
cost/safety guard every AI-suggested query is assessed against --
compiling to and executing through the same unchanged Phase 2 IR/
compiler and Phase 4 tenant scoping as a hand-written query, no
parallel execution path.

Track A (built into the query bar): inline ghost-text autocomplete,
"Explain this query", "Fix this query" with a diff view, and a
rule-based "Optimize" suggestion. Track B: natural-language-to-query
translation, always a separate review step from execution, with
`sentryctl query --nl` requiring explicit confirmation to run.
Every accepted/dismissed translate-fix-optimize interaction is logged
into the same append-only audit_log table Phase 4 built.

Two real product bugs were found and fixed via live browser
verification (a Svelte effect re-running on every keystroke that
silently cancelled the ghost-text debounce; a ghost-text widget
positioned at document offset 0 instead of the cursor), and a real
costguard logic bug (unbounded-aggregation vs. raw-row) was caught by
its own test suite. New integration tests wire a real Ollama client
through the real HTTP handler against a mock server matching Ollama's
wire contract (hack/mock-ollama), keeping model-quality verification
out of CI as a disclosed, periodic human-run check instead.

See /docs/phase-7-ai-design.md and /docs/phase-7-runbook.md.
This commit is contained in:
2026-08-16 18:06:27 -07:00
parent 661568085e
commit 7d316f92db
37 changed files with 5230 additions and 20 deletions
+25
View File
@@ -17,6 +17,20 @@ type Config struct {
QueryTimeout time.Duration
CORSAllowedOrigin string
EnterpriseAuthURL string
AI AIConfig
}
// AIConfig gates Phase 7's AI-assisted query features (Track A/B) --
// off unless OllamaBaseURL is set, same "off unless configured"
// convention as EnterpriseAuthURL and everything else optional in this
// codebase. OllamaFastModel is the per-operation override for
// Complete's tight latency budget (/docs/phase-7-ai-design.md's
// per-operation provider/model config) -- empty means Complete uses
// OllamaModel too, same as every other operation.
type AIConfig struct {
OllamaBaseURL string
OllamaModel string
OllamaFastModel string
}
type ClickHouseConfig struct {
@@ -65,6 +79,17 @@ func Load() (Config, error) {
// base URL (e.g. "http://enterprise-auth:8081") to turn on
// real session/service-token enforcement.
EnterpriseAuthURL: getenv("ENTERPRISE_AUTH_URL", ""),
// Empty OllamaBaseURL means AI features are entirely disabled --
// /ai/* routes aren't even registered (see main.go), matching
// "no cloud dependency required for the default deployment" and,
// by the same reasoning, no *local* model dependency forced on a
// deployment that doesn't want one either. Model names default to
// the recommendation confirmed in /docs/phase-7-ai-design.md.
AI: AIConfig{
OllamaBaseURL: getenv("OLLAMA_BASE_URL", ""),
OllamaModel: getenv("OLLAMA_MODEL", "qwen2.5-coder:7b"),
OllamaFastModel: getenv("OLLAMA_FAST_MODEL", "qwen2.5-coder:1.5b"),
},
}
timeoutSec, err := strconv.Atoi(getenv("QUERY_TIMEOUT_SECONDS", "30"))