Phase 1: Windows log collection + full-text search
Extends the agent, ingest, storage, api, and web with Windows Event Log/ETW sourcing and Tantivy-backed free-text search, per the approved Phase 1 plan. - CLAUDE.md: materialized on disk (never existed as a file before) with a new Phase 1 "done looks like" section. - agent: Windows Event Log (EvtSubscribe) and ETW sources, Windows service wrapper (install/uninstall/run-service), both feature- and target_os-gated so Linux builds/tests/clippy stay unaffected. Also fixed two pre-existing Phase 0 clippy gaps (dead-code on default-features-only builds, a type-inference edge case) found while testing every feature combination properly for the first time. UNVERIFIED on real Windows -- no Windows toolchain existed anywhere in the build environment; flagged prominently in three places. - proto/ingest: new record_id field, assigned once server-side in ingest's gRPC front end so ClickHouse and Tantivy agree on the same ID for the same record. - storage: record_id column + bloom filter index, verified against a live ClickHouse. - search: new service, Tantivy index, rskafka consumer as an independent second consumer group on the same Redpanda topic ingest already reads. - api/web: new /search endpoint and page, sharing the query page's result-table shape and component. - hack/windows-fixture: sends realistic Windows-shaped data straight to ingest, so the pipeline's handling of it is verifiable without a Windows host. Verified end-to-end on the live docker-compose stack: the same record_id comes back from both /query and /search for the same log line, including for windows-fixture's synthetic Windows Event Log data. Real bugs found and fixed along the way: api/Dockerfile missing proto/ in its build context, search's logs being completely silent (RUST_LOG gap), and search/target/ missing from .gitignore/.dockerignore.
This commit is contained in:
@@ -0,0 +1,48 @@
|
||||
use std::sync::Arc;
|
||||
use tonic::{Request, Response, Status};
|
||||
|
||||
use crate::index::SearchIndex;
|
||||
use crate::searchv1;
|
||||
|
||||
const DEFAULT_LIMIT: usize = 100;
|
||||
|
||||
pub struct SearchServer {
|
||||
index: Arc<SearchIndex>,
|
||||
}
|
||||
|
||||
impl SearchServer {
|
||||
pub fn new(index: Arc<SearchIndex>) -> Self {
|
||||
Self { index }
|
||||
}
|
||||
}
|
||||
|
||||
#[tonic::async_trait]
|
||||
impl searchv1::search_service_server::SearchService for SearchServer {
|
||||
async fn search(
|
||||
&self,
|
||||
request: Request<searchv1::SearchRequest>,
|
||||
) -> Result<Response<searchv1::SearchResponse>, Status> {
|
||||
let req = request.into_inner();
|
||||
if req.query.trim().is_empty() {
|
||||
return Err(Status::invalid_argument("query must not be empty"));
|
||||
}
|
||||
|
||||
let limit = if req.limit == 0 {
|
||||
DEFAULT_LIMIT
|
||||
} else {
|
||||
req.limit as usize
|
||||
};
|
||||
|
||||
let index = Arc::clone(&self.index);
|
||||
let query = req.query.clone();
|
||||
// Tantivy's searcher is synchronous; run it on a blocking thread
|
||||
// so it doesn't stall the async runtime alongside the consumer
|
||||
// tasks.
|
||||
let record_ids = tokio::task::spawn_blocking(move || index.search(&query, limit))
|
||||
.await
|
||||
.map_err(|e| Status::internal(format!("search task panicked: {e}")))?
|
||||
.map_err(|e| Status::invalid_argument(format!("search failed: {e}")))?;
|
||||
|
||||
Ok(Response::new(searchv1::SearchResponse { record_ids }))
|
||||
}
|
||||
}
|
||||
Reference in New Issue
Block a user