Skip to content

Ground interview prompts in a deterministic evidence ledger - #71

Merged
jserv merged 44 commits into
mainfrom
deterministic
Sep 26, 2026
Merged

jserv merged 44 commits into
mainfrom
deterministic

Conversation

@jserv

@jserv jserv commented Sep 20, 2026 •

Copy link
Copy Markdown
Contributor

The interviewer's prompts were written from whatever text was nearest to hand: the raw editor buffer pasted into a silence nudge, the raw runner output pasted into a review. This branch reduces what the browser reports into a bounded, replayable record of what the runtime actually observed. The ledger holds digests, counts and closed enumerations, never editor text, runner output or transcript, and the one timestamp it does not take from the browser is the receipt, read from the server's clock. What a prompt gets is not the ledger itself but a few lines of plain text rendered from its aggregates, with no digest, timestamp or entry in them; the watch prompts add a line of what arrived since the last one and the editor fenced as untrusted, whole while it is short. Later commits measure what each model input costs, delimit the candidate's words in the report prompt, reclaim expired sessions, correct the change classifier, parse each editor update once and incrementally, seed the two HTTP calls, label the browser's test claims as unverified, cut what each prompt carries and how long each reply waits, move the platform's own bookkeeping out of the model's turns, and log what each session was billed.

Why a parser

Syntax facts carry every leaf token in document order, which is what makes an edit visible at all: operators are anonymous nodes in each grammar here, so facts drawn from named nodes alone cannot tell a + b from a - b, and facts compared as sets cannot tell f(a, b) from f(b, a). Those are the edits an interview turns on. Verified against the pinned tree-sitter grammars for C, C++, Java, JavaScript and Python, including that reindenting, retabbing, blank lines and a Python backslash continuation read as formatting, and that a comment edit reads as a comment in every grammar, which means recognizing Java's line_comment and block_comment and JavaScript's html_comment as well as comment.

What the parser earns is the proactive-review gate. On main it is a size heuristic (more than 80 non-whitespace characters or 3 lines added or removed); here it is the ledger's semantic_revision, which moves on any parsed candidate edit that is not layout or a comment. The same Python edits through both, with the fine class the ledger records and the coarse class the model is told:

Edit main size gate semantic_revision Ledger class Model sees
range(len(nums)) to range(len(nums) - 1) no yes expression code
target - nums[i] to target + nums[i] no yes expression code
final return dedented into the loop no yes expression code
returned pair swapped no yes expression code
rename seen to index no yes identifier_only identifier
reindent 4 to 2 spaces no no formatting_only formatting
3 blank lines added yes no formatting_only formatting
4-line comment block added yes no comment_only comment
loop rewritten with enumerate no yes control_flow code
half-typed line no no none (syntax_invalid) none

The heuristic misses every small edit that introduces or fixes a bug and fires on whitespace and comments; the parser inverts both. Two things here need a parser rather than a normalized text diff: syntax validity, which holds reviews back on a half-typed buffer and keeps the last parsed buffer as the recovery baseline, and whitespace- and comment-insensitivity that still sees Python indentation, which is the program. Stripping whitespace would have missed the dedented return, and stripping comments correctly needs a tokenizer per language anyway.

Tokens and latency

Token counts are exact, from Gemini countTokens on gemini-3.1-flash-lite (the Live model does not serve countTokens; the same tokenizer family is assumed). The session is a simulated forty-five line Python LRU cache typed at the browser's publish rate, with turns and test runs, reviewed under the same timing gates on both sides.

main ledger as JSON (earlier in this PR) this branch
review input over the session 6,776 (14 reviews) 42,212 (15), plus a read_editor call each review was told to make 6,128 (14), code included, before the last round below
review prompt, golden 181 1,049 204
report prompt, golden (rules and brief) 2,666 3,674 2,431
interim prompt, golden 365, up to 12 a session 1,255 435, up to 6 a session
system instructions, once a session and again on every resume 4,278 4,278 3,917
five-minute warning / wrap-up 119 / 126 62 / 78
one analysis at 50 / 500 lines, crate optimized none 3.6 / 39 ms 1.0 / 12 ms

What each prompt carries is only what the rest of it does not already state. A watch prompt carries the code and goes to a session that holds the conversation, so it gets the tests, diagnostics, hints, phases and session state, and only the lines that changed since the last watch prompt the same Live session holds, with a line that has gone sent as none; with none changed it has no evidence section at all. The interim review adds the edit history and the time since the last program change; the report keeps the edit history, the test history and the session state. Test counts and diagnostics are labeled as the browser's unverified claims wherever they appear.

Every carrier of the code shares one record of what the Live model was last shown: a watch prompt, a test reaction, read_editor, a requested hint and a cold briefing. Each sends the change since it, or says the editor is unchanged, so a silence nudge no longer sends the model to read code it holds and a review after a nudge no longer resends it, and a test reaction carries the code that changed where it used to need a read first. Line numbers are unpadded, blank lines at the end are not numbered, a buffer sent whole is never cut, and read_editor, the hint reply and the cold briefing stop at 32,000 bytes. A test reaction lists one failing case and a count of the rest.

The model no longer does the platform's bookkeeping. The server records the unasked STAR steps as skipped at the five-minute warning and at the end, where the prompts had asked for up to four tool calls in front of each, and the evidence tool no longer offers the skip. Code evidence needs no read_editor first, since the server already refuses it without written code; the evidence reply names the phase instead of echoing the row, and a batch of calls is answered in one message. The report is written while the farewell plays, from a prompt frozen when the interview ends, so the candidate waits for the longer of the two rather than both. The live instructions state the hint and disclosure rules once and drop the constant per-question directions and the empty profile and grounding sections. The report and interim calls send their rules as the system instruction; the report no longer asks for the phase rows' weakness tags or a sorted plan, which the server derives, and a plan weakness that differs from its improvement only in case, spacing or a full stop is accepted rather than repaired.

A proactive review is armed only by a change that is not a rename alone, in code that parses, which a buffer past the parser's 64 KiB limit never does. A rename has to map names one for one: pointing return x at y is a change. Interim reviews come at most six a session, 150 seconds apart, after six new candidate turns, never once the end is due, and without the code when it has not changed since the last one.

The JSON view spent more than half its tokens on SHA-256 digests. Sending only the changed lines instead of the editor was tried first; against the Live model the interviewer read the editor on every sampled review anyway, reaching first audio in about 950 ms against 565 ms for main. With the whole buffer sent while it is short (eighty lines, 4,000 bytes), the prompt reached first audio in a median 580 ms over four samples, against 1,080 ms for the JSON view and 531 ms for main.

Measured against the Live model in this round: pinning thinking off changes nothing today (506 against 500 ms median over six replies each, no thought tokens reported either way), so it is pinned only against a moving default; a NON_BLOCKING evidence tool made replies slower, 1,305 against 547 ms, so evidence is still recorded before speaking. Live reports each turn's usage once, on the frame that completes it, billed on the whole context; the session now logs the sum, and each report and interim call logs its own. On the reply path, endpointing waits 1,000 ms of silence where it waited 1,500, audio goes to Gemini in 40 ms batches, a requested hint returns the editor with its clue, and GEMINI_END_SENSITIVITY can name the end-of-speech sensitivity, left at the API's default because nothing measured picks a side.

A watch prompt the socket refuses puts back the code and evidence it showed and the review baseline, so the prompt after a resumption carries what the model never saw. The test line states the change in passing and failing counts as signed numbers, and only against an earlier executed run.

Known limits

  • The fine class is a heuristic. Control-flow and interface kinds are named in full per grammar; data structures are matched on words such as list and dict, with syntactic groupings and destructuring patterns excluded. Before the last two commits an off-by-one read as mixed (two uncategorized kinds) and an enumerate rewrite as data_structure (Python's pattern_list); both are fixed and pinned, mixed is gone, and a loop whose binding changes shape is control flow. Review of this PR found three earlier defects in the same code (Python dedent read as formatting, a catch parameter read as the function's signature, then that shield hiding a nested function's signature) and a mutation-lane survivor in the rename branch. Validation after that found two more of the same kind, both fixed and pinned: Java and JavaScript comment kinds other than comment read as code and moved the review gate, and a Python \ continuation read as an expression edit. every_comment_kind_the_grammars_name_is_a_comment walks the node kinds of all five pinned grammars, so a grammar bump that names a new comment kind fails a test rather than moving the gate.
  • The model is not told the fine class. The prompt view names the last edit formatting, comment, identifier or code and carries no node facts; the ledger keeps the full analysis for replay. The review gate depends on the class only through the layout-or-comment split, so a wrong fine class can neither mislead the interviewer nor move the gate.
  • identifier means names were substituted one for one, not that a binding was renamed consistently across scopes: merge(left, right) to merge(first, right) is the same class and may point the call at a different value. It still moves semantic_revision, but not the review gate.
  • A changed loop condition or bound stays expression; only a change to what the loop binds is control flow. Comparing loop headers field by field is left until an evaluation shows the fine class matters.
  • Python parameter names count as a rename rather than an interface change, although keyword callers can see them; the judges call candidate functions positionally.
  • The latency figures above rest on three to six Live samples per arm. The savings from the shorter silence window, the smaller audio batches, the one-call hint, the server-recorded skips and the report beside the farewell are estimates from the settings and the round trips removed, not measurements; a probe of the five-minute warning without a real conversation left the model silent too often to measure.
  • The Live session is left unseeded on purpose, so a replay reproduces the evidence and every prompt byte for byte (the ledger folds them into one digest) but not what the interviewer says. The report and interim calls are seeded, and four calls with one prompt returned one answer.
  • The silence window is a trade, not a free saving: 700 ms cut second-language candidates off mid-sentence, 1,500 ms made every reply wait, and 1,000 ms sits between them without having been measured against those candidates.
  • Gemini's implicit prefix cache still reports no cached tokens on the report call, with its 1,346-token rules now first as the system instruction and the same request sent twice, so no caching saving is claimed. The interim review is not compared yet with none at all on report quality.
  • Every TLS client verifies against the bundled Mozilla roots, reqwest and the Gemini WebSocket alike, and LiveKit's clients inherit the same features. The binary needs no CA store on the host, as docs/install.md promises, and in exchange a CA an organization installs locally, a TLS-inspecting proxy's for one, is not trusted.

Cost

Six exactly pinned crates (tree-sitter and five grammars), about 29 MB of generated C compiled at build, and about 5.4 MB of object code linked (C++ is 3.3 MB of it). No new toolchain: WebRTC already requires a C/C++ compiler. One parser per grammar is reused, the last buffer's tree is kept, and each update parses only the new buffer, incrementally; a test holds the incremental tree to a fresh parse node by node. An update costs about 1 ms at fifty lines and 12 ms at five hundred with the crate optimized, run inline on the agent's task, and a buffer past 64 KiB is left unparsed.

Verification

Each of the thirty-four commits builds and passes cargo test (835 Rust tests at the tip), cargo clippy --all-targets -D warnings and the formatter check on its own, so the series bisects. Every view is asserted whole per prompt kind on a full ledger, which pins each line it writes; a local mutation run over the view and diffing commit caught 66 of 67 mutants with the other unviable, and CI mutation-tests the rest of the diff. tests/golden/evidence-ledger.json freezes the ledger a fixture replays to, so a change to any reducer shows up as a diff rather than as a test that only proves the reducers consult no clock. Locally, ./scripts/test.sh fails five gates (gen-problems, gen-problems-tests, recording-provision-check-tests, recording-integration-harness-tests, study-plan-guards), all because the machine's Python is 3.8 and those scripts use functools.cache and dict | dict; browser-tests passes (536 tests) now that the system g++ accepts -std=c++20.

Deliberately left out: a compile failure is still counted as other rather than guessed into a syntax, type or linker error, because nothing structured in the response says which it was, and the categories no runner ever names are now omitted from the projection instead of being reported to the model as zeros. One consequence worth knowing: a semicolon added in JavaScript now reads as a change rather than as layout, which is the side to err on given the alternative is a per-grammar list of tokens to ignore.


Summary by cubic

Grounds interviewer prompts in a deterministic evidence ledger instead of raw editor text, runner output, or transcripts. The ledger stores digests, counts, and closed enumerations—never candidate text—with server-side receipt timestamps, so replays reproduce the same prompts. Proactive reviews now gate on a tree-sitter parse of each editor update rather than the size heuristic, which missed edits that fix or break code and fired on whitespace and comments.

Evidence and gate

  • The ledger is bounded at 256 entries (dropping its oldest), keeps one parser per grammar, parses each update incrementally, and leaves buffers past 64 KiB unparsed.
  • Edits classify as code, identifier, formatting, or comment; a review arms only on changes that aren't a rename alone, in code that parses.
  • A rename must map each old name to a unique new name, and each changed reference must resolve, old and new, to the same local binding; anything else reads as code.
  • The model sees only the coarse class and the editor (whole while short, else the changed region), never node facts.

Prompts and cost

  • Each prompt gets only the evidence lines the rest of it doesn't state; gone lines are sent as none; the code the Live model last saw is tracked and sent as a change since it.
  • Unasked STAR steps are recorded server-side at the warning and end; a batch of tool calls is answered in one message; read_editor takes a start line.
  • Report and interim rules moved to system instructions; the report wraps candidate code and transcript in BEGIN/END and drops derived fields like weakness tags; a plan weakness matching its improvement in case, spacing, or full stop is accepted.
  • The report runs beside the farewell; the two HTTP calls are seeded while Live stays unseeded; billed usage and input bytes are logged per session and call.
  • Interim reviews cap at six per session, 150 seconds apart; expired sessions are swept on a timer.
  • Endpointing waits 1,000 ms silence, audio goes in 40 ms batches, thinking is pinned off, and end-of-speech sensitivity is an optional GEMINI_END_SENSITIVITY. The greeting is never cut off until the candidate has been heard once, and the farewell is no longer cut by Gemini's own interruption.
  • TLS clients verify bundled Mozilla roots; the browser check's room calls are hand-signed, dropping the uninstalled livekit-server-sdk; redaction masks credentials a LiveKit URL carries; and the require scan strips whole-line comments only.
  • A mutation shard that overruns its time fails naming what overran and keeps what it tested; a timed-out shard reports its assigned size; large diffs shard to the width the plan asks for.

Written for commit 58a8318. Summary will update on new commits.

Review in cubic

cubic-dev-ai[bot]

This comment was marked as resolved.

cubic-dev-ai[bot]

This comment was marked as resolved.

@jserv
jserv force-pushed the deterministic branch 3 times, most recently from f8275b4 to c5bfb86 Compare September 20, 2026 18:16
@sysprog21 sysprog21 deleted a comment from cubic-dev-ai Bot Sep 20, 2026
cubic-dev-ai[bot]

This comment was marked as resolved.

@jserv
jserv force-pushed the deterministic branch 2 times, most recently from eb71a7c to 2d8987c Compare September 20, 2026 19:27
@sysprog21 sysprog21 deleted a comment from cubic-dev-ai Bot Sep 20, 2026
cubic-dev-ai[bot]

This comment was marked as resolved.

@jserv
jserv force-pushed the deterministic branch 2 times, most recently from bda1b5b to a65821e Compare September 21, 2026 00:46
cubic-dev-ai[bot]

This comment was marked as resolved.

@sysprog21 sysprog21 deleted a comment from cubic-dev-ai Bot Sep 21, 2026
@jserv
jserv force-pushed the deterministic branch 2 times, most recently from 897f632 to fd43594 Compare September 21, 2026 03:23
@sysprog21 sysprog21 deleted a comment from cubic-dev-ai Bot Sep 21, 2026
cubic-dev-ai[bot]

This comment was marked as resolved.

@jserv
jserv force-pushed the deterministic branch 2 times, most recently from 3afd755 to e701a50 Compare September 21, 2026 19:51
@sysprog21 sysprog21 deleted a comment from cubic-dev-ai Bot Sep 21, 2026
cubic-dev-ai[bot]

This comment was marked as resolved.

@jserv
jserv force-pushed the deterministic branch 2 times, most recently from de1e94a to bc1d436 Compare September 21, 2026 22:47
@sysprog21 sysprog21 deleted a comment from cubic-dev-ai Bot Sep 21, 2026
cubic-dev-ai[bot]

This comment was marked as resolved.

@sysprog21 sysprog21 deleted a comment from cubic-dev-ai Bot Sep 22, 2026
The room loop waited for the farewell, up to eight seconds, before it
built the report prompt and started the call, so the candidate waited
for both in turn. The prompt is now frozen when the interview ends, with
the candidate's last turn closed and the platform's own skips recorded,
and the call runs beside the farewell. All the farewell can add is the
skips a started behavioral round leaves to it, which the report scores
as unassessed either way.
The report schema asked for every phase row's weakness tags and the
prompt for a sorted plan, and the server overwrote both: output tokens,
each a repeated improvement, that nothing read. They are gone from the
schema and the prompt. A plan weakness that differs from its
improvement only in case, spacing or a closing full stop is now written
back as that improvement rather than costing the whole report a repair
call; a paraphrase is still refused.
The hint rule was written out in the REACTO policy twice, again in flow
2 and again under the tools; the disclosure rule twice; a block of
"question-specific" directions was the same six REACTO steps for every
problem; and a session with no profile or documents read two sections
announcing that on every turn. Each rule is now stated once, the empty
sections are gone, and a coding-only session no longer carries the
document grounding that only picks the behavioral question. The
instructions shrink by about 380 tokens.
Both HTTP prompts opened on the per-interview material, the report's
first sentence with the elapsed minutes, so no two calls shared a
prefix and a repair resent the rules behind a different brief. The
constant rules now go as the system instruction and the brief as the
request. The report brief also gets the contract and constraints the
tests grade in place of the published statement, and the interim note
names the scenario rather than the published problem, whose title in
a note the report's name check refused.
Every snapshot right-aligned its line numbers and numbered the blank
lines an editor leaves at its end, a few tokens a line on every prompt;
a buffer sent whole cut its long lines while calling itself all N
lines, and a region past the line cap was labelled with the lines the
change spanned. The watch excerpt was also the only bounded carrier:
read_editor, a requested hint and a cold briefing numbered the whole
buffer, so one pasted file stayed in the Live session at full size for
the rest of the interview.

Numbers are now unpadded, trailing blank lines go, a whole buffer is
whole and the label names the range shown. The whole-editor carriers
stop at about eight thousand tokens and cut a line at a thousand
characters, with read_editor paging past the cut, and the cold
briefing's transcript tail halves to six thousand bytes, since the
round, the evidenced steps and the editor travel with it.
Every idle-window review resent up to four kilobytes of editor, most
often the same buffer the previous note had already read, and a review
could start seconds before the end that would abort it. A review whose
code has not changed since the last one is now told so instead, and
none starts once the interviewer has asked to close or the planned
time runs out before the call could return.
The Live model answers no countTokens call, so the session's real
spend was invisible: every turn is billed on the whole context it runs
in, which counting the text this server sends cannot show. Live
reports each turn's usage on the frame that completes it; the session
now sums those and logs the total at the end, and each report and
interim call logs its own usage.
The Live setup left thinking at the server's default and the
end-of-speech sensitivity unset while its comment said both halves of
endpointing were named. Thinking is now pinned off as the HTTP calls
pin it; measured, the pin changes nothing today, and no Live reply
reported a thought token. The end sensitivity becomes an optional
GEMINI_END_SENSITIVITY, unset by default because nothing measured
picks a side, and the comment says so.
A test reaction listed every failing case the browser reported, up to
four with their expected and actual values, although it asks the
candidate to pick one failure and reason about it; the rest were only
material for the interviewer to say too much with. The reaction now
reads the first failure and a count of the others, while read_editor
and the report keep the full account.
The browser check required livekit-server-sdk, which is in neither
package.json nor node_modules, from a function only the credentialed
lanes reach: every dispatch run died on it after spending a whole
interview. Listing and removing a participant is one signed POST each
and the server already speaks that protocol, so the harness now does
too rather than taking the first npm dependency that is neither a
linter nor a browser.
node --check parses a script without resolving what it imports, and
every other JavaScript file here is loaded by some gate, so a missing
package fails that gate. The browser check is the exception: the lane
that runs it needs credentials and a browser, which is how a require
of an uninstalled package survived for months.
The cursor records what the last review showed the model, but the
empty-editor branch left it holding the code from before. A candidate
who clears the editor and then restores exactly what was there is then
reported as unchanged, to a model whose last look found nothing.
One frame carries both, and the room loop reads them one at a time
through a channel. Queued behind the completion, the last turn's
tokens are still there when the interview tears down or the socket is
replaced, and a replacement empties the queue: the session then bills
itself for less than it spent.
The line count drops the blank lines an editor leaves at the end, so a
buffer too large to send whole and holding nothing else measures zero
of them. The excerpt then announced "lines 1-0 of 0", a range running
backwards over a body already saying the editor is empty.
Three budgets bound what is quoted and then append their own notice,
so a truncated result is that fixed string longer than the number.
Reserving the width instead would drop a line of the candidate's code
to buy an exact ceiling nobody is measuring against, so the caps keep
their meaning and the comments now state it.
Gemini interrupts on its own voice activity detection, and the opening
seconds of a room hand it a microphone opening or a chair rather than
speech. Two interviews in a row lost ten seconds of the opening line
that way, with nothing transcribed either time, which is the silence
this path exists to explain. Until the transcript records the
candidate saying something there is nobody who could be barging in, so
the queued turn plays and the one after it is cut as before.
A job the runner kills for exceeding timeout-minutes is reported as
"Cancelled after 45m", which in the checks list is indistinguishable
from somebody cancelling it or a newer push superseding it, and this
branch spent three runs looking like all three. The shard now holds
its own budget below the job's and fails with a line naming what
overran, keeps whatever it managed to test, and gets a ceiling that
lets a large diff shard out to the width the plan already asks for.
The wrap-up loop hands every event Interruptible::No because the
goodbye is the one turn that plays to the end, and take_stale_playout
honors it, but the arm dispatching Gemini's own interruption dropped
the mode on the floor. A candidate clearing their throat over the
farewell cut the farewell, through the door the promise did not cover.
validate_livekit_url accepts userinfo, so wss://user:pass@host is a
configuration the server runs on, and fetch refuses such a URL by
quoting it back. Redaction knew the key and the secret by name and a
JWT by shape, and a password inside a URL is neither.
Stripping every // to end of line also erased the rest of a line
holding a URL in a string, and a require after one went unscanned.
Whole-line comments only: a false alarm is visible and one line to
fix, while a miss is the silent failure the gate exists to catch.
The parameter defaults to the variable, so passing undefined tested
whatever the shell happened to export: the case passed on CI for the
right reason and failed on a developer's machine for the wrong one.
The shard list is the assignment made before the run, so reporting it
as what is left to test overstates the remainder on a shard that got
most of the way. It is the size of the shard, and the artifact beside
it is what says how far the run actually got.
The ledger is bounded at 256 entries and drops its oldest, so calling
the sequence append-only promised a replay it cannot serve. The review
gate suppresses renames as well as layout and comments. And a call is
not always spelled with the word this doc named, which is why the
kinds are matched in full rather than as substrings.
The pattern stopped at the first @, so a query string carrying an
address was read as a credential and https://host?to=user@example.com
came out masked on a host it never named. Query and fragment now end
the authority, and userinfo runs to its last @, which is where the
room service already cuts it: a password holding one is masked whole
by both readers or by neither.
The scan erased from any /* to the next */, and this repository keeps
"/* setup {" as a fixture string six lines above a "/* seed */" one,
so a require between them vanished before the scan saw it: the exact
miss the gate exists to prevent, in the file it was written for.
Anchoring the close to a line end brought the same miss back, because
the lazy body then ran past the first */ to find one that ended a
line, so the strip stops at the first close instead.
The system instruction is the largest text the session sends and the
only one the cost line did not name, so a reconnect-heavy interview
under-reported what it handed the model by tens of kilobytes. It is
counted where a socket is adopted, which makes the count the number
of Live sessions these metrics cover, and a source test pins both
openers because a room that really connected is the only other
witness to a call site that stops calling.
@jserv
jserv merged commit da4f8c0 into main Sep 26, 2026
21 checks passed
@jserv
jserv deleted the deterministic branch September 26, 2026 18:25
moon-jam added a commit to moon-jam/codetrial that referenced this pull request Sep 28, 2026
PR sysprog21#71 parses edits on the Rust server, but Enter needs syntax ranges
synchronously in the browser. Load pinned browser grammars and use
their comment, string, and regexp regions for auto-indentation.

Scan unfinished regions and fall back to the scanner when parsing is
unavailable. Cover incomplete input and browser Enter behavior with
regression tests.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant