fix(mothership): answer a retried task wake whose turn already ran - #8484
Conversation
The worker retries a task wake under the same run ID until its own run appears. When sim ended that turn without reaching the worker (a usage-limit refusal), every retry reopened the turn, hit the unique stream-id constraint, and failed behind a generic message, so the worker retried forever. Under the chat lock, a wake whose run ID already has a sim run now releases the lock and answers not-found, which the worker treats as a refusal and dismisses the notification. An in-flight turn still holds the lock and answers busy. The headless run-record catch-all now logs the underlying insert error.
The retried-wake check reads copilot_runs after taking the chat lock. If that read threw, the lock stayed held until its TTL because the wake turn that releases it never started. Release with the exact lease on any throw after the acquire.
|
@cubic-dev-ai review this PR |
|
The latest updates on your projects. Learn more about Vercel for GitHub. |
@waleedlatif1 I have started the AI code review. It will take a few minutes to complete. |
There was a problem hiding this comment.
All reported issues were addressed across 3 files
Reply with feedback, questions, or to request a fix.
Fix all with cubic | Re-trigger cubic
|
There was a problem hiding this comment.
All reported issues were addressed across 3 files
Requires human review: Auto-approval blocked because this review re-detected 1 unresolved issue already reported by Cubic.
Re-trigger cubic
|
@cubic-dev-ai review this PR |
@waleedlatif1 I have started the AI code review. It will take a few minutes to complete. |
|
@cubic-dev-ai review this PR |
@waleedlatif1 I have started the AI code review. It will take a few minutes to complete. |
Summary
Problem. The worker retries a task wake under the same run ID until it sees its own run for that ID. When sim accepted the wake but the headless turn ended inside sim without reaching the worker (for example, a usage-limit refusal), the worker never saw a run. It retried the wake, sim accepted it again, and the loop never ended.
Root cause.
prepareTaskWakeonly checked the chat stream lock. It had no idea that a turn under this run ID had already run, so every retry looked like a new wake.Fix.
prepareTaskWaketakes the chat lock, it looks up the latest run for the wake's run ID. If a turn already ran under that ID, it answersnot_found(404), and the worker drops the notification instead of retrying. The check runs while the lock is held, so a turn that is still running under that ID keeps answering busy (409).ensureHeadlessRunIdentitynow logs the underlying error and keeps it as thecauseof its generic "execution record is unavailable" error, so the failure can be diagnosed.Behaviour changes
Test plan
prepare-wake.integration.tsruns against real PostgreSQL and Redis through the wake route. It covers the not-found retry, a busy retry while the turn still holds the chat, and releasing the lock when the lookup fails.prepare-wake.tsreverted to staging, the not-found and lock-release tests fail. The busy test passes both before and after the fix; it is there to keep the new check from turning an in-flight turn's 409 into a 404.lib/mothership/tasks,lib/mothership/request/lifecycle,app/api/mothership/wake.bun run type-check(apps/sim), biome on the changed files, andbun run check:test-patternsall pass.