feat(check): build inline simulation-run payloads and add a dry-run CLI - #61
Merged
Merged
Conversation
This was referenced Oct 1, 2026
Contributor
Author
This was referenced Oct 1, 2026
scott-lowe-vapi
marked this pull request as ready for review
October 1, 2026 23:52
vtkovapi
approved these changes
Oct 3, 2026
Contributor
Author
Merge activity
|
scott-lowe-vapi
changed the base branch from
feat/check-config
to
graphite-base/61
October 3, 2026 06:04
`npm run check -- <check>|--all --dry-run` reads vapi-checks.yml and, for each check target, builds the inline POST /eval/simulation/run body from the files on disk: the assistant or squad with its tools, handoffs and structured outputs, plus the scenarios, judges and personalities the check's suites and simulations name. Nothing is looked up on the platform and no key is needed; --print-payload writes the JSON. The builder follows push and the runtime so the check tests what push deploys: - tools in runtime order (model.tools, then toolIds, then toolRefs); knowledgeBase tools stay in toolIds by run-org UUID; - `##` comments stripped; UUID references resolved through state; - squad members inlined, handoffs to members by name; - hook toolIds, artifactPlan structuredOutputIds and linked structured outputs inlined; judges' structuredOutputId inlined; - credentials bound by name to the run org (promotion's bindings). Anything it can't place fails the build with the field named: missing files, handoffs leaving the target, legacy assistantDestinations by ID, tools by ID inside overrides (strict mocks), duplicate tool names, audio judges/hooks/no required judge over chat, leftover non-UUID references, and payloads over 4.5 MB. The tool mock policy and live runs follow in later changes; without --dry-run the command exits 2. Refs TEST-141 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
scott-lowe-vapi
force-pushed
the
feat/check-inline-payload
branch
from
October 3, 2026 06:06
d3e6a1b to
680a03d
Compare
scott-lowe-vapi
added a commit
that referenced
this pull request
Oct 3, 2026
#62) ## Value **V.A.L.U.E. tier:** project — PR 6 of 10 for inline simulation PR checks ([TEST-141](https://linear.app/vapi/issue/TEST-141/gitops-run-simulation-suites-against-pr-changes-inline-as-ci-checks)); **this is the safety-critical PR**, so it's kept separate for review. Still offline: nothing is sent until PR 7. - **Problem:** a PR check runs the branch's agents in a real org, with real LLM conversations that call tools. Without a policy, a check could send a real SMS, book a real appointment, transfer to a real number, or POST transcripts to a customer's webhook, on every push. - **Who it affects:** gitops users, who need checks safe by default; and the TEST-141 failure condition "any tool call reaching a real server under default settings". - **What changes:** `src/check-mocks.ts` runs as the last pass of `checkPayloadBuild` under `toolMocks: strict` (the default). | Class | Types | Treatment | |---|---|---| | No external side effect | `endCall`, `dtmf`, `voicemail`, `output`; `query` (same org only) | Sent as written | | Knowledge base | in `model.toolIds` by UUID | Same org: kept, with a warning. Cross-org or inline: fails | | Handoff | `handoff` | Only to a squad member by name, or an inline assistant (walked). `dynamic`, `squad`, non-members: fail | | Mockable | `function` by `function.name`; `apiRequest` by top-level `name` (`url` set to the dead host) | Scenario mock if present, otherwise `{"error":"vapi-gitops-ci: <tool> is not mocked in this scenario"}` | | Transfer | `transferCall` | Rewritten to a mocked dead-server `function` under its own name, so it can never connect | | Everything else | `sms`, `sipRequest`, `code`, `mcp`, `bash`, `computer`, `textEditor`, `transferCancel`, `transferSuccessful`, `google.*`, `slack.*`, `gohighlevel.*`, `ghl`, `make`, unknown | Fails the build, naming the tool | - **Structural rule:** any key in the exported `TOOL_BEARING_KEYS` outside a handled position fails with "unsupported tool position". Examples: `model.functions`, `model.toolRefs`, reasoner `skills`, `declineTool`, `tools:append` outside overrides. A future API field that carries tools fails instead of slipping through. - **Servers are replaced, never deleted**, because a deleted server falls back to the phone number's or the org's URL: - every assistant (target, members, inline handoff assistants, personalities) gets `server: {url: "https://vapi-gitops-ci.invalid", timeoutSeconds: 1}` and `serverMessages: []`; - overrides that set a server get the dead one; - every function tool gets the dead server, and `serverUrl` / `serverUrlSecret` are dropped; - scenario `webhook` hooks get the dead server. - **Default mocks go in each scenario's `toolMocks`**, never in assistant metadata, because handoffs rebuild the assistant. A user mock with `enabled: false` is replaced, and a mock naming no tool in the target produces a warning. - **Also fails:** - `model.knowledgeBaseId`, and custom-provider knowledge bases; - personality tools beyond the side-effect-free ones; - hook `transfer` actions, and hook `toolId`s; - `scenarioId` entries, and non-stock `personalityId`s. - **Hook-fired tools** (assistant `hooks[].do[]`) probably bypass scenario `toolMocks`, which apply on the LLM tool-call path. They're classified the same way and get the dead server, which is the real safeguard there. This is documented in the module header and goes into `simulations.md` in PR 8. - **Opt-outs:** - `toolMocks: off` skips the tool rules, for a dedicated CI org; - `stripWebhooks: false` keeps assistant servers, while tool servers are still replaced under strict mocks. ## Evidence of value **Dry run of the TEST-141 parity squad.** In a copy of the fixture: - every tool and the receptionist got real-looking `example.com` servers; - one scenario dropped its `book_appointment` mock; - the `--print-payload` output was then inspected: | Check | Result | |---|---| | `example.com` URLs left in the payload | **0** (8 dead URLs: 2 member assistants, 3 personality copies, 3 function tools including the `tools:append` one) | | Receptionist `server` / `serverMessages` | `{"url":"https://vapi-gitops-ci.invalid","timeoutSeconds":1}` / `[]` | | S1 and S3 mocks | all 3 tools from the scenario | | S2 mocks | `lookup_patient`, `check_availability` from the scenario; **`book_appointment` = default error mock** | | Same squad plus an `sms` tool on the receptionist | **exit 2**: `target.squad.members[0].assistant.model.tools[1]: sms tools can't be mocked; remove it, or set toolMocks: off with a dedicated CI org` | **`TOOL_BEARING_KEYS` audit** against the API's OpenAPI schema (`apps/dashboard/src/api/schema.json` in the monorepo), listing every property whose schema references a tool DTO or is named like a tool reference: - **Listed:** - `tools` (all model DTOs and `TransferAssistantModel`); - `tools:append` (`AssistantOverrides`); - `toolIds` and `toolRefs` (all model DTOs); - `declineTool` / `declineToolId` (`RecordingConsentPlanVerbal`); - `skills` (`OpenAIReasoner`); - `assistantDestinations` (`SquadMemberDTO`). - **Not listed, handled at their only position:** - `tool` / `toolId` (`ToolCallHookAction`) and `function` (`FunctionCallHookAction`) are hook actions; - `function` also appears on handoff DTOs, as the tool's own definition. - `functions` and `forwardingPhoneNumber(s)` aren't in the current schema, and stay listed as fail-closed. **Tests:** `npm test` goes from 430 to 450 passing; `npm run build` is clean. ## Testing plan - **`tests/check-mocks.test.ts`** (20 tests): - each class in the table, with every listed failing type asserted by name; - `query` and knowledge bases, same org vs cross-org; - every `TOOL_BEARING_KEYS` entry at an unhandled position, plus `functions` / `toolRefs` / `assistantDestinations` placement; - `parameters` / judge `schema` properties named `tools` not tripping the rule; - each handoff destination kind, with inline handoff assistants walked; - assistant and override servers; - each hook action kind, and scenario webhooks; - personality tools; - knowledge-base fields; - entry shapes; - `toolMocks: off` and `stripWebhooks: false`. - `tests/check-payload.test.ts` expectations were updated for the dead servers the policy now adds, and the parity end-to-end test still passes. - **Not tested:** - **Live runs:** that scenario `toolMocks` intercept every function and `apiRequest` call is shown for the parity run only, and the full transcript scan is PR 7. - The unverified `transferCall` and integration mock names. They're rewritten or refused rather than relied on. - Hook-fired tools bypassing mocks (assumed, and safe either way). Stacked on #61. Refs TEST-141 🤖 Generated with [Claude Code](https://claude.com/claude-code)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Value
V.A.L.U.E. tier: project — PR 5 of 10 for inline simulation PR checks (TEST-141); this PR adds an offline dry run, and nothing is sent.
POST /eval/simulation/runbody: the target with every tool, handoff and structured output, plus every scenario, judge and personality. It has to assemble that body the way push and the runtime would, or the check tests something other than what ships.npm run check -- core --dry-run --print-payload) before any minutes are spent. Every later PR (mock policy, live runs, workflow, promotion gate) builds on this payload.src/check-payload.ts(withcheck-payload-assistant.tsandcheck-payload-refs.ts) builds the body fromorgResourcesReadand the org states. It's pure, makes no network calls, and collects every problem so one dry run reports them all.model.tools, thentoolIds, thentoolRefs. This is whatcallAssistantsGetdoes, and the parity run's likelyendCalldifference came from getting it wrong;##comments stripped, and UUID references resolved through state;toolRefspin wins over a duplicatetoolIdsentry, with a warning that the version pin is ignored;knowledgeBasetools are kept by run-org UUID, since the API refuses them inline.do[].toolIdbecomes an inline tool;artifactPlan.structuredOutputIds, plus structured outputs that link the assistant through their ownassistant_ids, are inlined;structuredOutputIds are inlined;bind/omit).assistantDestinationsby ID;handoff_to_…but a handoff is auto-named, because generated names differ inline and stored.src/check-cmd.ts:npm run check -- <check>|--all --dry-run [--print-payload [dir]].VAPI_GITOPS_ROOT.--dry-runit exits 2; live runs land in PR 7.package.jsongets acheckscript, and the README and AGENTS.md command tables getnpm run checkrows.tests/fixtures/check-parity/: the TEST-141 parity squad written as gitops files.Evidence of value
The builder reproduces the payload that scored 15/15 in the parity run.
.mdassistants withtoolIds, a handoff tool byassistantId, judges bystructuredOutputId).npm run check -- core --dry-run --print-payloadwas run on it, and the output was diffed against the inline body the experiment sent, rebuilt fromparity.mjs.members[*].assistant.model.toolsorder[lookup_patient, handoff, endCall],[check_availability, endCall][endCall, lookup_patient, handoff],[endCall, check_availability]model.toolsthentoolIds). Same tools, byte-for-byte, order asidesquad.name,personality.nameinline,callerBright Smile Dental,Dental calleriterations,transportvapi.webchatvapi-checks.ymldefaultsassistantId: schedulercame out asassistantName: "Scheduler", exactly what the experiment sent.Tests:
npm testgoes from 404 to 430 passing (26 new), andnpm run buildis clean.Testing plan
tests/check-payload.test.ts(21 tests, temp-dir fixtures):.mdprompt, inline judges, entries, transport;.mdbody as the only system message;toolIds, andtoolIdsby UUID with server fields stripped;toolRefspin and warning;knowledgeBasekept by UUID, and failing with no UUID;artifactPlanand fromassistant_ids(by slug and by UUID);parametersignored;tests/check-cmd.test.ts(5 tests):--print-payload;--allwith one broken check exiting 2;Stacked on #60.
Refs TEST-141
🤖 Generated with Claude Code