Ground interview prompts in a deterministic evidence ledger - #71
Merged
Merged
Conversation
jserv
force-pushed
the
deterministic
branch
from
September 20, 2026 16:56
7b8d389 to
70ac979
Compare
jserv
force-pushed
the
deterministic
branch
3 times, most recently
from
September 20, 2026 18:16
f8275b4 to
c5bfb86
Compare
jserv
force-pushed
the
deterministic
branch
2 times, most recently
from
September 20, 2026 19:27
eb71a7c to
2d8987c
Compare
jserv
force-pushed
the
deterministic
branch
2 times, most recently
from
September 21, 2026 00:46
bda1b5b to
a65821e
Compare
jserv
force-pushed
the
deterministic
branch
2 times, most recently
from
September 21, 2026 03:23
897f632 to
fd43594
Compare
jserv
force-pushed
the
deterministic
branch
2 times, most recently
from
September 21, 2026 19:51
3afd755 to
e701a50
Compare
jserv
force-pushed
the
deterministic
branch
2 times, most recently
from
September 21, 2026 22:47
de1e94a to
bc1d436
Compare
jserv
force-pushed
the
deterministic
branch
from
September 22, 2026 00:07
bc1d436 to
26514b3
Compare
The room loop waited for the farewell, up to eight seconds, before it built the report prompt and started the call, so the candidate waited for both in turn. The prompt is now frozen when the interview ends, with the candidate's last turn closed and the platform's own skips recorded, and the call runs beside the farewell. All the farewell can add is the skips a started behavioral round leaves to it, which the report scores as unassessed either way.
The report schema asked for every phase row's weakness tags and the prompt for a sorted plan, and the server overwrote both: output tokens, each a repeated improvement, that nothing read. They are gone from the schema and the prompt. A plan weakness that differs from its improvement only in case, spacing or a closing full stop is now written back as that improvement rather than costing the whole report a repair call; a paraphrase is still refused.
The hint rule was written out in the REACTO policy twice, again in flow 2 and again under the tools; the disclosure rule twice; a block of "question-specific" directions was the same six REACTO steps for every problem; and a session with no profile or documents read two sections announcing that on every turn. Each rule is now stated once, the empty sections are gone, and a coding-only session no longer carries the document grounding that only picks the behavioral question. The instructions shrink by about 380 tokens.
Both HTTP prompts opened on the per-interview material, the report's first sentence with the elapsed minutes, so no two calls shared a prefix and a repair resent the rules behind a different brief. The constant rules now go as the system instruction and the brief as the request. The report brief also gets the contract and constraints the tests grade in place of the published statement, and the interim note names the scenario rather than the published problem, whose title in a note the report's name check refused.
Every snapshot right-aligned its line numbers and numbered the blank lines an editor leaves at its end, a few tokens a line on every prompt; a buffer sent whole cut its long lines while calling itself all N lines, and a region past the line cap was labelled with the lines the change spanned. The watch excerpt was also the only bounded carrier: read_editor, a requested hint and a cold briefing numbered the whole buffer, so one pasted file stayed in the Live session at full size for the rest of the interview. Numbers are now unpadded, trailing blank lines go, a whole buffer is whole and the label names the range shown. The whole-editor carriers stop at about eight thousand tokens and cut a line at a thousand characters, with read_editor paging past the cut, and the cold briefing's transcript tail halves to six thousand bytes, since the round, the evidenced steps and the editor travel with it.
Every idle-window review resent up to four kilobytes of editor, most often the same buffer the previous note had already read, and a review could start seconds before the end that would abort it. A review whose code has not changed since the last one is now told so instead, and none starts once the interviewer has asked to close or the planned time runs out before the call could return.
The Live model answers no countTokens call, so the session's real spend was invisible: every turn is billed on the whole context it runs in, which counting the text this server sends cannot show. Live reports each turn's usage on the frame that completes it; the session now sums those and logs the total at the end, and each report and interim call logs its own usage.
The Live setup left thinking at the server's default and the end-of-speech sensitivity unset while its comment said both halves of endpointing were named. Thinking is now pinned off as the HTTP calls pin it; measured, the pin changes nothing today, and no Live reply reported a thought token. The end sensitivity becomes an optional GEMINI_END_SENSITIVITY, unset by default because nothing measured picks a side, and the comment says so.
A test reaction listed every failing case the browser reported, up to four with their expected and actual values, although it asks the candidate to pick one failure and reason about it; the rest were only material for the interviewer to say too much with. The reaction now reads the first failure and a count of the others, while read_editor and the report keep the full account.
The browser check required livekit-server-sdk, which is in neither package.json nor node_modules, from a function only the credentialed lanes reach: every dispatch run died on it after spending a whole interview. Listing and removing a participant is one signed POST each and the server already speaks that protocol, so the harness now does too rather than taking the first npm dependency that is neither a linter nor a browser.
node --check parses a script without resolving what it imports, and every other JavaScript file here is loaded by some gate, so a missing package fails that gate. The browser check is the exception: the lane that runs it needs credentials and a browser, which is how a require of an uninstalled package survived for months.
The cursor records what the last review showed the model, but the empty-editor branch left it holding the code from before. A candidate who clears the editor and then restores exactly what was there is then reported as unchanged, to a model whose last look found nothing.
One frame carries both, and the room loop reads them one at a time through a channel. Queued behind the completion, the last turn's tokens are still there when the interview tears down or the socket is replaced, and a replacement empties the queue: the session then bills itself for less than it spent.
The line count drops the blank lines an editor leaves at the end, so a buffer too large to send whole and holding nothing else measures zero of them. The excerpt then announced "lines 1-0 of 0", a range running backwards over a body already saying the editor is empty.
Three budgets bound what is quoted and then append their own notice, so a truncated result is that fixed string longer than the number. Reserving the width instead would drop a line of the candidate's code to buy an exact ceiling nobody is measuring against, so the caps keep their meaning and the comments now state it.
Gemini interrupts on its own voice activity detection, and the opening seconds of a room hand it a microphone opening or a chair rather than speech. Two interviews in a row lost ten seconds of the opening line that way, with nothing transcribed either time, which is the silence this path exists to explain. Until the transcript records the candidate saying something there is nobody who could be barging in, so the queued turn plays and the one after it is cut as before.
A job the runner kills for exceeding timeout-minutes is reported as "Cancelled after 45m", which in the checks list is indistinguishable from somebody cancelling it or a newer push superseding it, and this branch spent three runs looking like all three. The shard now holds its own budget below the job's and fails with a line naming what overran, keeps whatever it managed to test, and gets a ceiling that lets a large diff shard out to the width the plan already asks for.
The wrap-up loop hands every event Interruptible::No because the goodbye is the one turn that plays to the end, and take_stale_playout honors it, but the arm dispatching Gemini's own interruption dropped the mode on the floor. A candidate clearing their throat over the farewell cut the farewell, through the door the promise did not cover.
validate_livekit_url accepts userinfo, so wss://user:pass@host is a configuration the server runs on, and fetch refuses such a URL by quoting it back. Redaction knew the key and the secret by name and a JWT by shape, and a password inside a URL is neither.
Stripping every // to end of line also erased the rest of a line holding a URL in a string, and a require after one went unscanned. Whole-line comments only: a false alarm is visible and one line to fix, while a miss is the silent failure the gate exists to catch.
The parameter defaults to the variable, so passing undefined tested whatever the shell happened to export: the case passed on CI for the right reason and failed on a developer's machine for the wrong one.
The shard list is the assignment made before the run, so reporting it as what is left to test overstates the remainder on a shard that got most of the way. It is the size of the shard, and the artifact beside it is what says how far the run actually got.
The ledger is bounded at 256 entries and drops its oldest, so calling the sequence append-only promised a replay it cannot serve. The review gate suppresses renames as well as layout and comments. And a call is not always spelled with the word this doc named, which is why the kinds are matched in full rather than as substrings.
The pattern stopped at the first @, so a query string carrying an address was read as a credential and https://host?to=user@example.com came out masked on a host it never named. Query and fragment now end the authority, and userinfo runs to its last @, which is where the room service already cuts it: a password holding one is masked whole by both readers or by neither.
The scan erased from any /* to the next */, and this repository keeps
"/* setup {" as a fixture string six lines above a "/* seed */" one,
so a require between them vanished before the scan saw it: the exact
miss the gate exists to prevent, in the file it was written for.
Anchoring the close to a line end brought the same miss back, because
the lazy body then ran past the first */ to find one that ended a
line, so the strip stops at the first close instead.
The system instruction is the largest text the session sends and the only one the cost line did not name, so a reconnect-heavy interview under-reported what it handed the model by tens of kilobytes. It is counted where a socket is adopted, which makes the count the number of Live sessions these metrics cover, and a source test pins both openers because a room that really connected is the only other witness to a call site that stops calling.
jserv
force-pushed
the
deterministic
branch
from
September 26, 2026 17:12
c60f83c to
58a8318
Compare
moon-jam
added a commit
to moon-jam/codetrial
that referenced
this pull request
Sep 28, 2026
PR sysprog21#71 parses edits on the Rust server, but Enter needs syntax ranges synchronously in the browser. Load pinned browser grammars and use their comment, string, and regexp regions for auto-indentation. Scan unfinished regions and fall back to the scanner when parsing is unavailable. Cover incomplete input and browser Enter behavior with regression tests.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The interviewer's prompts were written from whatever text was nearest to hand: the raw editor buffer pasted into a silence nudge, the raw runner output pasted into a review. This branch reduces what the browser reports into a bounded, replayable record of what the runtime actually observed. The ledger holds digests, counts and closed enumerations, never editor text, runner output or transcript, and the one timestamp it does not take from the browser is the receipt, read from the server's clock. What a prompt gets is not the ledger itself but a few lines of plain text rendered from its aggregates, with no digest, timestamp or entry in them; the watch prompts add a line of what arrived since the last one and the editor fenced as untrusted, whole while it is short. Later commits measure what each model input costs, delimit the candidate's words in the report prompt, reclaim expired sessions, correct the change classifier, parse each editor update once and incrementally, seed the two HTTP calls, label the browser's test claims as unverified, cut what each prompt carries and how long each reply waits, move the platform's own bookkeeping out of the model's turns, and log what each session was billed.
Why a parser
Syntax facts carry every leaf token in document order, which is what makes an edit visible at all: operators are anonymous nodes in each grammar here, so facts drawn from named nodes alone cannot tell
a + bfroma - b, and facts compared as sets cannot tellf(a, b)fromf(b, a). Those are the edits an interview turns on. Verified against the pinned tree-sitter grammars for C, C++, Java, JavaScript and Python, including that reindenting, retabbing, blank lines and a Python backslash continuation read as formatting, and that a comment edit reads as a comment in every grammar, which means recognizing Java'sline_commentandblock_commentand JavaScript'shtml_commentas well ascomment.What the parser earns is the proactive-review gate. On
mainit is a size heuristic (more than 80 non-whitespace characters or 3 lines added or removed); here it is the ledger'ssemantic_revision, which moves on any parsed candidate edit that is not layout or a comment. The same Python edits through both, with the fine class the ledger records and the coarse class the model is told:mainsize gatesemantic_revisionrange(len(nums))torange(len(nums) - 1)target - nums[i]totarget + nums[i]returndedented into the loopseentoindexenumerateThe heuristic misses every small edit that introduces or fixes a bug and fires on whitespace and comments; the parser inverts both. Two things here need a parser rather than a normalized text diff: syntax validity, which holds reviews back on a half-typed buffer and keeps the last parsed buffer as the recovery baseline, and whitespace- and comment-insensitivity that still sees Python indentation, which is the program. Stripping whitespace would have missed the dedented
return, and stripping comments correctly needs a tokenizer per language anyway.Tokens and latency
Token counts are exact, from Gemini
countTokensongemini-3.1-flash-lite(the Live model does not servecountTokens; the same tokenizer family is assumed). The session is a simulated forty-five line Python LRU cache typed at the browser's publish rate, with turns and test runs, reviewed under the same timing gates on both sides.mainread_editorcall each review was told to makeWhat each prompt carries is only what the rest of it does not already state. A watch prompt carries the code and goes to a session that holds the conversation, so it gets the tests, diagnostics, hints, phases and session state, and only the lines that changed since the last watch prompt the same Live session holds, with a line that has gone sent as
none; with none changed it has no evidence section at all. The interim review adds the edit history and the time since the last program change; the report keeps the edit history, the test history and the session state. Test counts and diagnostics are labeled as the browser's unverified claims wherever they appear.Every carrier of the code shares one record of what the Live model was last shown: a watch prompt, a test reaction,
read_editor, a requested hint and a cold briefing. Each sends the change since it, or says the editor is unchanged, so a silence nudge no longer sends the model to read code it holds and a review after a nudge no longer resends it, and a test reaction carries the code that changed where it used to need a read first. Line numbers are unpadded, blank lines at the end are not numbered, a buffer sent whole is never cut, andread_editor, the hint reply and the cold briefing stop at 32,000 bytes. A test reaction lists one failing case and a count of the rest.The model no longer does the platform's bookkeeping. The server records the unasked STAR steps as skipped at the five-minute warning and at the end, where the prompts had asked for up to four tool calls in front of each, and the evidence tool no longer offers the skip. Code evidence needs no
read_editorfirst, since the server already refuses it without written code; the evidence reply names the phase instead of echoing the row, and a batch of calls is answered in one message. The report is written while the farewell plays, from a prompt frozen when the interview ends, so the candidate waits for the longer of the two rather than both. The live instructions state the hint and disclosure rules once and drop the constant per-question directions and the empty profile and grounding sections. The report and interim calls send their rules as the system instruction; the report no longer asks for the phase rows' weakness tags or a sorted plan, which the server derives, and a plan weakness that differs from its improvement only in case, spacing or a full stop is accepted rather than repaired.A proactive review is armed only by a change that is not a rename alone, in code that parses, which a buffer past the parser's 64 KiB limit never does. A rename has to map names one for one: pointing
return xatyis a change. Interim reviews come at most six a session, 150 seconds apart, after six new candidate turns, never once the end is due, and without the code when it has not changed since the last one.The JSON view spent more than half its tokens on SHA-256 digests. Sending only the changed lines instead of the editor was tried first; against the Live model the interviewer read the editor on every sampled review anyway, reaching first audio in about 950 ms against 565 ms for
main. With the whole buffer sent while it is short (eighty lines, 4,000 bytes), the prompt reached first audio in a median 580 ms over four samples, against 1,080 ms for the JSON view and 531 ms formain.Measured against the Live model in this round: pinning thinking off changes nothing today (506 against 500 ms median over six replies each, no thought tokens reported either way), so it is pinned only against a moving default; a
NON_BLOCKINGevidence tool made replies slower, 1,305 against 547 ms, so evidence is still recorded before speaking. Live reports each turn's usage once, on the frame that completes it, billed on the whole context; the session now logs the sum, and each report and interim call logs its own. On the reply path, endpointing waits 1,000 ms of silence where it waited 1,500, audio goes to Gemini in 40 ms batches, a requested hint returns the editor with its clue, andGEMINI_END_SENSITIVITYcan name the end-of-speech sensitivity, left at the API's default because nothing measured picks a side.A watch prompt the socket refuses puts back the code and evidence it showed and the review baseline, so the prompt after a resumption carries what the model never saw. The test line states the change in passing and failing counts as signed numbers, and only against an earlier executed run.
Known limits
listanddict, with syntactic groupings and destructuring patterns excluded. Before the last two commits an off-by-one read asmixed(two uncategorized kinds) and anenumeraterewrite asdata_structure(Python'spattern_list); both are fixed and pinned,mixedis gone, and a loop whose binding changes shape is control flow. Review of this PR found three earlier defects in the same code (Python dedent read as formatting, a catch parameter read as the function's signature, then that shield hiding a nested function's signature) and a mutation-lane survivor in the rename branch. Validation after that found two more of the same kind, both fixed and pinned: Java and JavaScript comment kinds other thancommentread as code and moved the review gate, and a Python\continuation read as an expression edit.every_comment_kind_the_grammars_name_is_a_commentwalks the node kinds of all five pinned grammars, so a grammar bump that names a new comment kind fails a test rather than moving the gate.formatting,comment,identifierorcodeand carries no node facts; the ledger keeps the full analysis for replay. The review gate depends on the class only through the layout-or-comment split, so a wrong fine class can neither mislead the interviewer nor move the gate.identifiermeans names were substituted one for one, not that a binding was renamed consistently across scopes:merge(left, right)tomerge(first, right)is the same class and may point the call at a different value. It still movessemantic_revision, but not the review gate.expression; only a change to what the loop binds is control flow. Comparing loop headers field by field is left until an evaluation shows the fine class matters.docs/install.mdpromises, and in exchange a CA an organization installs locally, a TLS-inspecting proxy's for one, is not trusted.Cost
Six exactly pinned crates (
tree-sitterand five grammars), about 29 MB of generated C compiled at build, and about 5.4 MB of object code linked (C++ is 3.3 MB of it). No new toolchain: WebRTC already requires a C/C++ compiler. One parser per grammar is reused, the last buffer's tree is kept, and each update parses only the new buffer, incrementally; a test holds the incremental tree to a fresh parse node by node. An update costs about 1 ms at fifty lines and 12 ms at five hundred with the crate optimized, run inline on the agent's task, and a buffer past 64 KiB is left unparsed.Verification
Each of the thirty-four commits builds and passes
cargo test(835 Rust tests at the tip),cargo clippy --all-targets -D warningsand the formatter check on its own, so the series bisects. Every view is asserted whole per prompt kind on a full ledger, which pins each line it writes; a local mutation run over the view and diffing commit caught 66 of 67 mutants with the other unviable, and CI mutation-tests the rest of the diff.tests/golden/evidence-ledger.jsonfreezes the ledger a fixture replays to, so a change to any reducer shows up as a diff rather than as a test that only proves the reducers consult no clock. Locally,./scripts/test.shfails five gates (gen-problems,gen-problems-tests,recording-provision-check-tests,recording-integration-harness-tests,study-plan-guards), all because the machine's Python is 3.8 and those scripts usefunctools.cacheanddict | dict;browser-testspasses (536 tests) now that the systemg++accepts-std=c++20.Deliberately left out: a compile failure is still counted as
otherrather than guessed into a syntax, type or linker error, because nothing structured in the response says which it was, and the categories no runner ever names are now omitted from the projection instead of being reported to the model as zeros. One consequence worth knowing: a semicolon added in JavaScript now reads as a change rather than as layout, which is the side to err on given the alternative is a per-grammar list of tokens to ignore.Summary by cubic
Grounds interviewer prompts in a deterministic evidence ledger instead of raw editor text, runner output, or transcripts. The ledger stores digests, counts, and closed enumerations—never candidate text—with server-side receipt timestamps, so replays reproduce the same prompts. Proactive reviews now gate on a tree-sitter parse of each editor update rather than the size heuristic, which missed edits that fix or break code and fired on whitespace and comments.
Evidence and gate
code,identifier,formatting, orcomment; a review arms only on changes that aren't a rename alone, in code that parses.Prompts and cost
none; the code the Live model last saw is tracked and sent as a change since it.read_editortakes a start line.GEMINI_END_SENSITIVITY. The greeting is never cut off until the candidate has been heard once, and the farewell is no longer cut by Gemini's own interruption.livekit-server-sdk; redaction masks credentials a LiveKit URL carries; and the require scan strips whole-line comments only.Written for commit 58a8318. Summary will update on new commits.