From 95eaf01755d18c501fa8d1c5dae2255cf460b9a2 Mon Sep 17 00:00:00 2001 From: Martin Vogel Date: Sat, 26 Sep 2026 01:32:06 +0200 Subject: [PATCH] docs(readme): semantic_query is a search_graph parameter; team artifact is opt-in (#1297) The Search section introduced `semantic_query` as if it were an MCP tool. It is a parameter of search_graph (the tool's input schema in src/mcp/mcp.c). The Team-Shared Graph Artifact section said the artifact is written whenever you index. `persistence` defaults to false. The Best tier is written only by index_repository with persistence: true. Without it, an index only refreshes an artifact that already exists (Fast tier; see export_after_publish in src/pipeline/pipeline.c), and that applies to any index, not just the watcher's. The issue's other points are already correct on main: - the 17 registered tools match the README; - trace_call_path is a real alias of trace_path; - check_index_coverage is documented; - NOT EXISTS { (f)<-[:CALLS]-() } is supported since 08b62f03. Signed-off-by: Martin Vogel --- README.md | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) diff --git a/README.md b/README.md index 9c82c18aed..98b1e293e5 100644 --- a/README.md +++ b/README.md @@ -199,7 +199,7 @@ The install script placed beside the binary is **reported, not deleted** — uni - **Cypher-like queries**: `MATCH (f:Function)-[:CALLS]->(g) WHERE f.name = 'main' RETURN g.name` ### Search -- **Semantic search** (`semantic_query`): vector search across the entire graph, powered by bundled Nomic `nomic-embed-code` embeddings (40K tokens, 768d int8) compiled into the binary — no API key, no Ollama, no Docker. 11-signal combined scoring (TF-IDF, RRI, API/Type/Decorator signatures, AST profiles, data flow, Halstead-lite, MinHash, module proximity, graph diffusion). +- **Semantic search** (the `semantic_query` parameter of `search_graph`, not a separate tool): vector search across the entire graph, powered by bundled Nomic `nomic-embed-code` embeddings (40K tokens, 768d int8) compiled into the binary — no API key, no Ollama, no Docker. 11-signal combined scoring (TF-IDF, RRI, API/Type/Decorator signatures, AST profiles, data flow, Halstead-lite, MinHash, module proximity, graph diffusion). - **BM25 full-text search** via SQLite FTS5 with `cbm_camel_split` tokenizer (camelCase / snake_case aware) - **Structural search** (`search_graph`): regex name patterns, label filters, min/max degree, file scoping - **Code search** (`search_code`): graph-augmented grep over indexed files only @@ -243,12 +243,12 @@ The install script placed beside the binary is **reported, not deleted** — uni Commit a single compressed file to your repo and your teammates skip the reindex. -`.codebase-memory/graph.db.zst` is a zstd-compressed snapshot of the knowledge graph that lives next to your source. When you index, the artifact is written or refreshed; when a teammate clones the repo and runs `codebase-memory-mcp` for the first time, the artifact is decompressed and incremental indexing fills in their local diff. +`.codebase-memory/graph.db.zst` is a zstd-compressed snapshot of the knowledge graph that lives next to your source. The artifact is opt-in: `index_repository` writes it only with `persistence: true` (the default is `false`), and once it exists later indexes refresh it; when a teammate clones the repo and runs `codebase-memory-mcp` for the first time, the artifact is decompressed and incremental indexing fills in their local diff. - **Format**: SQLite database, indexes stripped, `VACUUM INTO` compacted, then zstd 1.5.7 compressed (8–13:1 ratio typical) - **Two tiers**: - - **Best** (`zstd -9` + index strip + `VACUUM INTO`) — written on explicit `index_repository` - - **Fast** (`zstd -3`) — written by the watcher for low-latency incremental updates + - **Best** (`zstd -9` + index strip + `VACUUM INTO`) — written by `index_repository` with `persistence: true` + - **Fast** (`zstd -3`) — refreshes an existing artifact on every other index (including the watcher's low-latency incremental updates) - **Bootstrap**: when no local DB exists but the artifact is present, `index_repository` imports the artifact first, then runs incremental indexing — avoiding the full reindex cost - **No merge pain**: a `.codebase-memory/.gitattributes` line with `merge=ours` is auto-created on first export, so concurrent edits don't produce conflicts on the binary artifact - **Commit it deliberately**: the artifact is rewritten on every index, including the watcher's Fast tier, and git stores each rewrite as a full new blob. Committing every refresh is what turns a 20 MB file into gigabytes of history — one team reached ~6 GB across ~350 commits of this single path. Pick a cadence (a release, a milestone, a nightly job) rather than committing every save.