Skip to content

Count a routine's failures from when it was last switched on - #687

Merged
davidmckayv merged 2 commits into
CopilotKit:mainfrom
Chebaleomkar:fix/routine-fatigue-reset
Oct 2, 2026
Merged

davidmckayv merged 2 commits into
CopilotKit:mainfrom
Chebaleomkar:fix/routine-fatigue-reset

Conversation

@Chebaleomkar

Copy link
Copy Markdown
Contributor

What this changes

The fatigue rule (runner.ts:152-163) switches a routine off after ten failures in a row, and docs/routines.md says "nothing further fires until a person turns it back on". But consecutiveFailures (store.ts) counted every trailing failed run with no boundary, and nothing reset it on re-enable. So after a person fixed the cause and switched the routine back on, its next failure counted as the eleventh:

  • the routine was switched straight off again with "has failed ten times in a row";

  • the failures === 1 first-failure message was never posted;

  • the owner got one retry, not ten.

  • New column routines.enabled_at (timestamptz not null default now()), set on create and whenever update switches a routine on (enabling). It is not updatedAt, which every sweep that claims the routine moves.

  • It is stamped with sql\now()`, the same clock as routine_runs.started_at`, so skew between hosts cannot move a run across the line.

  • consecutiveFailures counts only runs with started_at >= enabled_at. The existing bound, the skip handling and the ordering are unchanged.

  • Migration 0048_routine_enabled_at.sql, generated by db:generate, is one ALTER TABLE ... ADD COLUMN. Existing rows get the time of the upgrade, so a streak already under way starts again from zero once. The CHANGELOG entry says so.

Where it runs

  • New state that outlives a request? One column on routines, in Postgres.
  • What happens on the second replica? It reads the same column; the count is one query.
  • Anything serialised? The enable write is unchanged and stays inside the existing withEnabledCapLock transaction.
  • Anything fanned out to a browser? No.
  • New listener, port, or schedule? No.

Boundary and audit

  • Not touched.

Changelog

  • A line in CHANGELOG.md under Unreleased, including the one-time reset at upgrade.

Proof

Against pgvector/pgvector:pg17 with every migration applied, including 0048:

  • New test in routines-store.integration.test.ts: ten failures, then switched off and on, gives a count of 0; one more failure gives 1.
  • With upstream's store.ts the same test fails: Expected: 0, Received: 10.
  • routines-store.integration, routine-sweep.integration, routine-runner and routine-routes pass 106 of 106.
  • migration-journal.test.ts passes 8 of 8, and drizzle-kit check reports "Everything's fine".
  • tsc --noEmit on server/ exits 0, and biome is clean.

The fatigue rule switches a routine off after ten failures in a row and leaves it to a person to switch back on, but consecutiveFailures counted across that, so the first failure after re-enabling read as the eleventh and switched it straight off again. A new routines.enabled_at, set on create and whenever a routine is switched on (on the database clock), bounds the count.
…eset

# Conflicts:
#	server/drizzle/meta/0048_snapshot.json
#	server/drizzle/meta/_journal.json

@davidmckayv davidmckayv left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The current head 0e91aeb matches the accepted source disposition from the refreshed triage. Required CI must pass before landing.

@davidmckayv
davidmckayv enabled auto-merge (squash) October 2, 2026 16:52
@davidmckayv
davidmckayv merged commit d28aa31 into CopilotKit:main Oct 2, 2026
19 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants