feat(server): split camelCase identifiers in embeddings and stamp stored vectors - #26
Merged
Merged
Conversation
…s an identity Natural-language search missed symbols whose names carry the words being searched for. `getUserById` reaches the embedder as one rare token, so a query of "get user by id" had little to match. Identifier splitting Measured on a doc->symbol retrieval eval, BGE-small, pure semantic R@1: camelCase (TypeScript, 377 symbols) 0.355 -> 0.637 +79% PascalCase (Rust, 200 symbols) 0.520 -> 0.640 +23% snake_case (Rust, 723 symbols) 0.683 -> 0.692 +1.3% snake_case (SystemVerilog, 696) 0.510 -> 0.497 -2.5% So it applies only where words run together. snake_case and kebab-case already tokenise into the same words; a leading or trailing delimiter separates nothing, so `_handleClick` splits like `handleClick`. The Rust snake-vs-Pascal pair is the controlled comparison: same repo, same docs, only the casing differs. Exposed as --split-identifiers, codegraph.splitIdentifiers in VS Code and a checkbox in JetBrains, wired through the CLI, MCP builder, engine config, daemon, LSP initializationOptions and the cross-client parity check. Flags that could not be turned off `--full-body-embedding` was `#[arg(long, default_value = "true")]` on a bool, which clap parses as a flag that rejects a value: always true, and `=false` errored. Both it and --split-identifiers now take an optional value. Vector identity Stored vectors had no record of what produced them, so nothing could tell whether they were comparable to the ones a process was about to make. Switching --embedding-model was the sharpest case: there is no dimension check anywhere, and cosine_similarity zips its inputs, so a 768d query against a stored 384d vector silently scored a dot product over the first 384 dimensions against norms of different lengths. Semantic ranking was garbage, with no error, until a manual reindex. Present before this change. Vectors are now stamped with the schema, model, full-body and split settings that built them, and a mismatched set is never loaded. - A rebuild takes ownership of the project once, after it knows it has something to write, in one atomic step that replaces the stored set and writes its stamp. Saves only add; writes are chunked so a save under memory pressure does not triple the footprint; checkpoints keep a crashed rebuild resumable. - A store that cannot be read is left alone. graph.db is one RocksDB shared by every project, and a lock held elsewhere reads exactly like an absent store; treating it as absent claimed the project and cleared a valid set. An unreadable store now embeds in memory for the session. - A live --watch daemon owns its project's vectors; sessions never claim over it, and the daemon carries the same embed-text settings as the sessions that read from it. - An auto-spawned engine is passed --full-body-embedding and --split-identifiers with their values. Both are in the stamp, so an engine started with defaults would have served its client no vectors at all. An index written by 0.20.1 or earlier carries no stamp and is re-embedded the first time this version opens it. A --watch daemon left running across the upgrade keeps writing unstamped vectors that this build will not load; restart it after upgrading. Known trade-off: an LSP client that omits fullBodyEmbedding still defaults it off, unlike every other client. VS Code and JetBrains always send it; a bare nvim/emacs/helix client and an IDE client on the same project would replace each other's vector set. Aligning the default would move every bare client to ~3x slower indexing, so it is left and documented in place. Also: the eval harness read CODEGRAPH_SPLIT_IDS with is_ok(), so `=0` enabled it and an A/B run had two identical arms; StorageBackend gains scan_prefix_keys with a default body so out-of-tree backends still compile. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017rVbt7rENTwXkdHt3Bpgb5
🔍 CodeGraph PR Review30 files changed (+1244/−198, 68 functions) · Risk: 🔴 high Blast radius117 direct callers affected (55 breaking) across
|
0.21.0 rather than a patch: it adds --split-identifiers, changes what every camelCase symbol is embedded from, and re-embeds existing indexes on first open. Moves every release pin together: Cargo.toml and the workspace members in Cargo.lock, both ENGINE_VERSION pins (fetch-engine.js, shared with the VS Code client, and the JetBrains resolver), the npm package, both version fields in server.json, the VSIX version, pluginVersion, and the VSIX install examples in both READMEs. Left at 0.20.1 on purpose: lsp4ijVersion, which is the LSP4IJ dependency and only coincidentally matches; the comments and README lines that name 0.20.1 as the last release without vector stamps; and the test fixtures that publish and download their own fake release. The README gains an "Upgrading to 0.21" section. The upgrade guidance existed but sat inside the flags table, where nobody upgrading would look. It also says something nothing else did: graphs persisted by 0.20.1 can already hold files its watcher let in from excluded directories, and stale symbols from workspaces reached through a symlink. The fixes stop new damage but do not remove what is already stored, so those users need one reindex. The clients now ask for v0.21.0 release assets, which do not exist yet. package-npm.sh refuses to package until they are published. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017rVbt7rENTwXkdHt3Bpgb5
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
The developer wanted to finish the split-identifiers work, which had already been rebased and fixed, and ship it as a release. Their explicit instructions were to update the README, bump the version, and merge the open PR #26. That PR makes camelCase identifier splitting the default for embeddings and gives stored vectors an identity stamp covering schema, full-body, split, and model. The bump is a minor release, 0.20.1 to 0.21.0, applied to every tracked release-version field. The README should document the upgrade, including that a long-running --watch daemon must be restarted after upgrading. The developer's standing rules also apply: no em dashes, no manual edits to auto-generated files such as CHANGELOG, and pushes go through the no-mistakes gate rather than directly to origin.
What Changed
getUserByIdis embedded as "get user by id" too. This is on by default. You can turn it off with the new--split-identifiers=falseCLI flag or thecodegraph.splitIdentifierssetting in VS Code and JetBrains.--full-body-embedding=falsenow actually works; before, the flag parsed as always-true. Both settings are passed on to the shared engine (--connect) and the--watchdaemon.scan_prefix_keysmethod so callers can list keys without loading the values.server.json. The README gains an "Upgrading to 0.21" section, including the need to restart any running--watchdaemon, plus updated flag docs.🤖 Generated with Claude Code
Risk Assessment
✅ Low: The only new commit since the last review deletes the reindex paragraph from README's 'Upgrading to 0.21' section, exactly as the user instructed, with no em dashes; all release version fields (Cargo, Cargo.lock, vscode, mcp-package, server.json, jetbrains pluginVersion and ENGINE_VERSION pins) consistently read 0.21.0.
Testing
I built the 0.21.0 target and the 0.20.1 base and drove both through MCP over stdio with an isolated HOME and a TypeScript fixture workspace. I covered six things: the upgrade re-embed, vector reuse on restart, the --split-identifiers=false re-stamp, CLI flag parsing (including a rejected bad value), and the stale-daemon status before and after restarting the daemon on 0.21.0. I also parsed every release version field. Every scenario passed and logs and JSON results are in the evidence directory. Separately I found one bug that is not caused by this change: 0.20.1 does the same thing (details in the findings). This change has no UI surface, so there are no screenshots.
Evidence: Live scenario transcript summary
Evidence: Upgrade: 0.21.0 session re-embeds 0.20.1 vectors
Evidence: Restart reuses stamped vectors
Evidence: Split identifiers off re-stamps
Evidence: CLI flag parsing 0.20.1 vs 0.21.0
Evidence: Session attached to stale 0.20.1 daemon gets restart guidance
embedding_status: The --watch daemon's stored vectors were built with different embedding settings, so semantic matching is unavailable this session - results are from name/text search only. Restart the daemon with the same --full-body-embedding / --split-identifiers / --embedding-model flags as this session.Evidence: Session after daemon restart on 0.21.0 loads daemon vectors
Evidence: Release version fields
Evidence: MCP stdio driver
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
🔧 **Review** - 1 issue found → auto-fixed ✅
README.md:103- The new 'Upgrading to 0.21' section tells affected users to runcodegraph_reindex_workspaceonce to remove stale 0.20.1 symbols. Withoutforce, that tool is incremental, so it does not do this.Indexer::index_file(crates/codegraph-server/src/indexer.rs:736-743) returns early for any file whose hash is unchanged. The walk never visits excluded directories, so it never deletes their nodes. Old nodes are removed only for a re-parsed file and only under its current path spelling (line 748). As a result, symbols leaked in from excluded directories, and stale symbols stored under a symlinked path spelling, survive the recommended reindex. Only theforce: truebranch (mcp/server.rs:3413-3420) clears the graph and the hash cache. Fix: saycodegraph_reindex_workspacewithforce: true(a full rebuild).🔧 Fix applied.
✅ Re-checked - no issues remain.
crates/codegraph-server/src/mcp/server.rs- Not caused by this change: the 0.20.1 base build does the same. If the same project is indexed once with--workspace .and once with an absolute path, every symbol is stored twice, once under './src/lib.ts' and once under '/tmp/cgnm4/ws/src/lib.ts'. Both builds return duplicate symbol_search results, and 0.21.0 embeds 8 vectors for 4 functions. The relative workspace path is not made absolute before it is used as a node path. This is outside the scope of this release, so it is recorded here and not fixed.cargo build -p codegraph-serverat target 0621096 (0.21.0) and at base 1e6084d (0.20.1, built from a git archive export), both into /tmp/cg-nm-targetmcp-driver.py(MCP stdio JSON-RPC driver) running a 0.20.1 session to seed vectors, then the first 0.21.0 session on the same isolated HOME (upgrade re-embed)mcp-driver.pyrunning a 0.21.0 restart with the same flags (stored vectors are reused)mcp-driver.py '["--split-identifiers=false"]'(re-embeds under the split=off stamp; keyword and semantic results come back)codegraph-server --workspace . --full-body-embedding=false|--split-identifiers=false|--split-identifiers=maybe --run-tool codegraph_symbol_searchon both binariescg-base --watch --workspace /tmp/cgnm4/ws --extension-path /tmpdaemon, then a 0.21.0 MCP session attaching to it (status for the stale daemon)SIGTERM the 0.20.1 daemon, start a 0.21.0--watchdaemon, then a 0.21.0 MCP session loading the vectors the daemon maintainsPython tomllib/json parse of Cargo.toml, Cargo.lock, mcp-package/package.json, server.json, vscode/package.json, jetbrains/gradle.properties, plusgit grepfor ENGINE_VERSION pins and leftover 0.20.1Extracted the README 'Upgrading to 0.21' section and checked it has no em dash and no reindex paragraph✅ **Document** - passed
✅ No issues found.
crates/codegraph-memory/examples/embed_eval.rs-cargo fmt --all --checkreports formatting drift in about 60 files across the workspace: language crates, codegraph-harness, codegraph-memory and others. None of it is on lines this change touched. The one changed file affected, embed_eval.rs, drifts only in lines this change did not edit. The repo has no rustfmt config and no CI fmt gate. Proposed follow-up: a separatecargo fmt --allcommit rather than mixing a workspace-wide reformat into this release. Clippy on codegraph-server and codegraph shows 24 existing warnings, none on changed lines. The VS Code ESLint and tsc checks were skipped because vscode/node_modules is not installed; the TS change is two one-line additions.✅ **Push** - passed
✅ No issues found.