Skip to content

ci: type-check tests/ and send a User-Agent on sim API requests - #57

Merged
scott-lowe-vapi merged 1 commit into
mainfrom
ci/test-workflow
Oct 3, 2026
Merged

scott-lowe-vapi merged 1 commit into
mainfrom
ci/test-workflow

Conversation

@scott-lowe-vapi

Copy link
Copy Markdown
Contributor

Value

V.A.L.U.E. tier: project — PR 1 of 10 for inline simulation PR checks (TEST-141); this PR is a small, behavior-preserving slice.

The CI workflow itself landed in #56, so this PR is smaller than PR 1 in the plan.

Evidence of value

Check main (69c7e83) This branch
tsc --noEmit with tests/ included 37 errors (6 test files) 0 errors
npm test 355 pass 357 pass (2 new, 1 rewritten)
User-Agent on POST /eval/simulation/run none vapi-gitops-sim/1.0.0 (asserted against a local HTTP server)

What the 37 errors were:

  • Fixture drift after the hash-store migration: state entries still carrying lastPulledHash/lastPushedHash, bare-string state values, and an untyped emptyLoaded().
  • One real gap: the reconcile-state-key harness never passed the required formatError, so any test reaching that error path would have thrown a TypeError instead of testing it.

Tests that used the removed hash fields as markers (state-merge, recanonicalize) now mark "which copy won" with distinct UUIDs or object identity, so they still check the same behavior.

Testing plan

  • npm run build (now covers src/ and tests/) and npm test locally on Node 22: green. CI on this PR runs both on Node 20 and 22.
  • New tests/user-agent.test.ts covers the header format against package.json's version, and the header actually sent on run create.
  • sim.test.ts now covers the legacy bare-string state value directly, replacing the old cast-based "forward-compat" test.
  • Not tested: a live run against the API (the header is asserted locally only), and Node 20 locally (left to CI). src/ behavior is unchanged apart from the added header.

Refs TEST-141

🤖 Generated with Claude Code

- tsconfig.json now includes tests/, so `npm run build` (already run by
  CI on every PR) type-checks the tests. tsx runs them without
  type-checking, which is how 37 type errors accumulated unnoticed.
- Fix those 37 errors. All were fixture drift after the hash-store
  migration (state entries carrying lastPulledHash/lastPushedHash or
  bare-string values, an untyped emptyLoaded fixture), plus one real
  gap: the reconcile-state-key harness never passed the required
  formatError, so any test reaching that error path would have thrown.
- New src/user-agent.ts; `npm run sim` sends
  `User-Agent: vapi-gitops-sim/<version>`, so gitops-started simulation
  runs can be counted in the platform's run-started analytics.

Refs TEST-141

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

scott-lowe-vapi commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor Author

Merge activity

  • Oct 3, 5:59 AM UTC: A user started a stack merge that includes this pull request via Graphite.
  • Oct 3, 5:59 AM UTC: @scott-lowe-vapi merged this pull request with Graphite.

@scott-lowe-vapi
scott-lowe-vapi merged commit d7b9a34 into main Oct 3, 2026
3 checks passed
scott-lowe-vapi added a commit that referenced this pull request Oct 3, 2026
…lse pass (#58)

## Value

**V.A.L.U.E. tier:** project — PR 2 of 10 for inline simulation PR checks ([TEST-141](https://linear.app/vapi/issue/TEST-141/gitops-run-simulation-suites-against-pr-changes-inline-as-ci-checks)). Stacked on #57.

- **Problem:** `npm run sim` reported every run as passed. It read a `results` field the simulation-run API doesn't return and counted `status === "pass"` (items are `passed`/`failed`), so it always summarised 0/0 and exited 0, failing runs included.
- **Who it affects:** anyone gating on `npm run sim`, locally or in CI. The PR check (later in this stack) reuses this verdict, so it has to be right.
- **What changes:**
  - A strict verdict (`src/sim-result.ts`).
  - Item fetching that handles both response shapes and late results.
  - The run link printed, plus each failing judge with expected vs extracted values.
  - `--timeout` and Ctrl-C both cancel the run.
  - Exit codes: 0 passed, 1 failed, 2 usage, 3 incomplete.
  - A config-free client that never retries run creation on a 5xx, because the run may already be queued.

## Evidence of value

**Same stub API, one passed and one failed item:**

| | Result |
|---|---|
| `main`'s `sim.ts` | `{"pass":0,"fail":0}` → exits 0 (false green) |
| This branch | `failed — 1 of 1 simulations failed` → exits 1 |

**Live, against a test org (chat transport):**

| Suite | Exit | Output |
|---|---|---|
| Designed to fail (judge: "open 24 hours?") | **1** | `✗ … open-24h (expected = true, got false)` — [run](https://dashboard.vapi.ai/simulations/run/9b933e3c-4ea6-4829-bedc-06bc795e6df3) |
| Designed to pass (judge: "open 8–5 on Fridays?") | **0** | `passed — 1 of 1 simulations passed` — [run](https://dashboard.vapi.ai/simulations/run/91bdb9d7-0bf5-4d07-87e3-bda9f5d1391b) |

The temporary resources were deleted afterwards.

## Testing plan

- `tests/sim-result.test.ts` is a verdict table covering:
  - the old false-green shape (no `results`);
  - 0 items, a short item list, and a count mismatch;
  - a failed item, with the failing judge listed;
  - canceled items;
  - all required evaluations skipped, and an optional skip alongside a scored required judge;
  - missing `itemCounts`, and a run that hasn't ended.
- `tests/sim-run.test.ts` runs `runSimulation` against a local HTTP stub:
  - pass, and fail using the bare-array item shape;
  - late item results;
  - timeout cancels the run, and an interrupt cancels it with the 400 "already ended" swallowed;
  - no retry of a 502 on create;
  - `--no-watch`;
  - pagination with overlapping pages deduped.
- `npm run build` and `npm test` pass (377 tests).
- **Not tested:** a live run that's still `running` when it's canceled (cancel was only exercised against the stub), and voice transport (the live runs used chat). Default behaviour change: unknown CLI arguments are now an error (exit 2) instead of being silently ignored.

Refs TEST-141

🤖 Generated with [Claude Code](https://claude.com/claude-code)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants