Skip to content

fix(search): bound Search retirement pages so the migration finishes under its statement timeout - #8460

Merged
waleedlatif1 merged 5 commits into
stagingfrom
fix/retire-embeddings-migration-timeout
Sep 30, 2026
Merged

waleedlatif1 merged 5 commits into
stagingfrom
fix/retire-embeddings-migration-timeout

Conversation

@waleedlatif1

@waleedlatif1 waleedlatif1 commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

The Search retirement script migration (0029_retire_all_search_embeddings, which runs the 0027 retirement code) failed during a deploy with canceling statement due to statement timeout.

Each retirement page read 25,000 IDs and updated or deleted every target row among them in a single statement. The slow part is the writes, not the scan:

  • Documents: retiring a document is a non-HOT update, because user_excluded appears in partial-index predicates, so every index on document gets a new entry.
  • Chunks: deleting a chunk cascades into its vector, keyword and provenance projections.

When one captured knowledge base holds most of a table, nearly every row in a page needs to be written. Those writes, plus checkpoint full-page writes and the synchronous-replication commit wait, pushed a page past the migration's own two-minute statement_timeout. The query plan is fine: the scan portion of a page is a primary-key range scan that takes a few seconds.

Changes

  • Row cap per page. Pages mutate at most a row limit of target rows and read at most four IDs per row of that limit (never more than 25,000), so re-reading after a capped page stays proportional to the work even when the limit shrinks. A capped page advances the cursor only to its last mutated row. Already-retired documents never count against the limit.
  • Adaptive limit. Each page is timed from the start of its transaction through commit.
    • The limit starts at 2,000 rows.
    • It halves after a page slower than 30 s.
    • It doubles, up to 8,000, after a page under 7.5 s.
    • The floor is 25 rows.
  • Pacing. After each page the migration pauses for as long as the page took, up to 5 s, to leave the primary headroom.
  • Timeout handling.
    • Only a timeout from the page's mutating statement triggers a retry with a smaller page. That page rolls back with its cursor.
    • Any other statement timeout, such as a completion recheck, fails the run at once instead of shrinking the page several times for nothing.
    • The completion rechecks, which walk every captured knowledge base once, get a 30-minute timeout, including the revalidation when a run resumes after retirement already completed.
  • Logging. The periodic progress log now includes phase, cursor, rows mutated and row limit. The migration also logs slow-page halvings, phase changes and the start of the completion recheck.
  • Docs. search-embedding-retirement.md is updated to match.

The retirement stays idempotent and resumes from its saved cursor. 0029 is edited in place, following the earlier 0027 fixes: the runner records script migrations by name only, and deployments that already completed 0029 do not need to run it again.

Tests

Real-PostgreSQL integration tests in 0027_retire_search_embeddings.integration.ts:

  • Split pages. Statement-level triggers raise the real statement-timeout error for any statement touching more than 300 rows. On the previous code the migration fails exactly as it did in the deploy. With this change it:
    • completes;
    • never commits a statement over the bound;
    • never mutates a row twice;
    • leaves ordinary rows untouched.
  • Timeout outside the mutation. The completion checkpoint times out. The run fails after exactly one attempt; the previous code retried eight times.
  • Already-retired documents. A run of retired documents longer than the row limit is crossed in a single page. Without the unretired filter the test fails, taking four pages instead of two.
  • Scan window. With the row limit pinned at its floor, the rows read across all pages stay a small constant per document; a fixed 25,000-ID window re-reads the rest of the table on every attempt.
  • Resumed completion. Revalidating a completed retirement on resume runs under the 30-minute recheck timeout rather than the two-minute page timeout.
  • Rollback test. It now asserts a geometry-independent invariant: every committed delete is at or behind the cursor, and a failed page leaves every row past it. Exact counts no longer work because page boundaries now adapt.

Unblocking the release

The remaining cleanup, plus the index rebuild and vacuum that follow it, will likely take longer than the migrations job's five-hour limit. The retirement is resumable and safe to kill at any point:

  • Every page commits together with its cursor.
  • An interrupted page rolls back when its connection closes.
  • An interrupted REINDEX CONCURRENTLY leaves invalid _ccnew/_ccold indexes, which the next run drops before retrying.
  • Completed rebuild and vacuum steps are checkpointed.
  • Nothing is recorded in script_migrations until retirement and maintenance both finish.

Option A: run it out of band first (recommended, so it does not hold up deploys)

  1. Merge this PR to staging.

  2. From a checkout that contains this change, run the standalone entry point. It uses the same resolution as the deploy runner (MIGRATION_DATABASE_URL, falling back to DATABASE_URL):

    bun install --frozen-lockfile
    MIGRATION_DATABASE_URL='<direct, non-pooled writer DSN>' \
      bun --no-env-file run packages/db/script-migrations/0027_retire_search_embeddings.ts
    • The DSN must be a direct or session-pooled connection. PgBouncer transaction pooling is not supported.
    • --no-env-file stops Bun from loading a local .env into the process.
    • The command resumes the saved cursor. It runs retirement and then maintenance through 0029_retire_all_search_embeddings, and on success records 0029 and the names it supersedes, exactly as the deploy runner would.
    • It can be stopped and re-run as often as needed.
  3. While it runs, do not start the migrations workflow. The standalone command does not take the deploy runner's advisory lock, so a concurrent deploy would retire pages in parallel. That is correct, because pages serialize on the progress row, but it doubles the load, and the second maintenance worker fails on the maintenance lock. Stop the standalone run before any deploy, then resume it afterwards.

  4. Once 0029_retire_all_search_embeddings is recorded, cut the release to main. The migrations job skips the completed script migration.

Option B: let the deploy run it

  1. Merge this PR to staging.
  2. Cut a new release to main that includes it. Re-running the failed migrations workflow run will not work, because it checks out the old commit.
  3. The migrations job resumes 0029 from its saved cursor. Pages completed before the failure stay committed and are not redone.
  4. If the job hits its five-hour limit, re-run the migrations workflow for the same release. Expect one or two re-runs.

In both cases, keep writers to the captured Search knowledge bases stopped until 0029_retire_all_search_embeddings is recorded. Progress can be inspected with SELECT * FROM search_embedding_cleanup_progress, and the migration logs its phase, cursor and rows mutated every ten pages.

Test plan

  • 0027_retire_search_embeddings.integration.ts against PostgreSQL 17 with pgvector
  • packages/db unit tests
  • bun run type-check (packages/db)
  • bun run lint
  • bun run check:audits

@vercel

vercel Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
docs Ready Ready Preview Sep 30, 2026 5:18pm UTC

Request Review

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 3 files

Reply with feedback, questions, or to request a fix.

Fix all with cubic | Re-trigger cubic

Comment thread packages/db/script-migrations/0027_retire_search_embeddings.ts
@greptile-apps

greptile-apps Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[Critical risk] Modifies database migration logic for search embeddings cleanup.

The PR appears safe to merge based on the reviewed migration changes.

Summary

The PR bounds Search-retirement writes per page, adapts the row limit after slow or timed-out pages, paces retries, and gives completion rechecks a longer timeout. It adds PostgreSQL integration coverage and updates the operational documentation.

Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[Read bounded ID window] --> B[Mutate up to row limit]
  B --> C{Mutation timed out?}
  C -- Yes --> D[Roll back page and cursor]
  D --> E[Halve limit and retry]
  C -- No --> F[Commit page and cursor]
  F --> G[Adjust limit and pace next page]
Loading

Reviews (3) · Last reviewed commit: "fix(search): never grow a Search retirem..."

Comment thread packages/db/script-migrations/0027_retire_search_embeddings.ts Outdated
@waleedlatif1

Copy link
Copy Markdown
Collaborator Author

@greptile

@waleedlatif1

Copy link
Copy Markdown
Collaborator Author

@cubic-dev-ai review this PR

@cubic-dev-ai

cubic-dev-ai Bot commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

@cubic-dev-ai review this PR

@waleedlatif1 I have started the AI code review. It will take a few minutes to complete.

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 3 files

Reply with feedback, questions, or to request a fix.

Fix all with cubic | Re-trigger cubic

…der the statement timeout

A retirement page read 25,000 IDs and updated or deleted every target row among them in one
statement. Retiring a document is a non-HOT update that writes every index on `document`, and a
deleted chunk cascades into its projections, so on a KB that dominates the table a page's write
cost, not its scan, outran the two-minute statement timeout and failed the deploy migration.

Each page now mutates at most a row limit of its target rows. A page that reaches the limit
advances the cursor only to its last mutated row, and already-retired documents never spend the
limit. The limit starts at 2,000, halves after a slow page or a statement timeout (the timed-out
page rolls back with its cursor and is retried), and doubles after a fast full page. The completion
rechecks, which walk every captured KB once, run with a 30-minute timeout. The retirement stays
idempotent and resumes from its saved cursor.
…irement pages

Only a statement timeout from a page's mutating statement now halves the row limit and retries
the rolled-back page. Any other timeout, such as a completion recheck, fails the run at once
instead of repeating the same statement at every smaller limit.

Pages are timed around the whole call, commit included, so the synchronous-replication wait
counts toward the slow-page threshold. Each page is followed by a pause as long as the page, up
to five seconds, and the row limit is capped at 8,000. Phase changes no longer adjust the limit.

The progress log now carries the phase, cursor and rows mutated, and the migration logs slow-page
halvings, phase changes and the start of the completion recheck.
…limit

A run of already-retired Search documents longer than the row limit must be crossed in one page. The test counts documents-phase statements and fails if the page filter on unretired rows is removed.
…time out resumed rechecks like completion

A capped page re-reads its scan from its last mutated row, so a fixed 25,000-ID window re-read
most of the same IDs on every page once the row limit shrank. Each page now reads at most four IDs
per row of its limit, capped at 25,000, and any fast page doubles the limit so sparse stretches
widen the window again.

A retry that finds retirement already complete revalidates every captured KB, as completion does,
so it now runs under the same 30-minute timeout instead of the two-minute page timeout.
@waleedlatif1
waleedlatif1 force-pushed the fix/retire-embeddings-migration-timeout branch from cf63591 to f202d9b Compare September 30, 2026 17:16
…timed out, and pause after a timeout

A timed-out page halved the row limit, but one fast page doubled it straight back, so the run
alternated between the size that timed out and half of it, rolling back a full statement-timeout
page each time. The limit now grows only up to half of the smallest size that timed out, and a
timed-out page is followed by the same pause as any other page.
@waleedlatif1
waleedlatif1 force-pushed the fix/retire-embeddings-migration-timeout branch from f202d9b to 1e7d6fd Compare September 30, 2026 17:16
@waleedlatif1

Copy link
Copy Markdown
Collaborator Author

@greptile

@waleedlatif1

Copy link
Copy Markdown
Collaborator Author

@cubic-dev-ai review this PR

@cubic-dev-ai

cubic-dev-ai Bot commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

@cubic-dev-ai review this PR

@waleedlatif1 I have started the AI code review. It will take a few minutes to complete.

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No issues found across 3 files

Confidence score: 5/5

  • Automated review surfaced no issues in the provided summaries.
  • No files require special attention.

Re-trigger cubic

@waleedlatif1
waleedlatif1 merged commit 294c505 into staging Sep 30, 2026
32 checks passed
@waleedlatif1
waleedlatif1 deleted the fix/retire-embeddings-migration-timeout branch September 30, 2026 18:09

This branch was successfully deployed

1 active deployment
Preview — 1e7d6fd5 Deployed Sep 30, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant