Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
9 changes: 6 additions & 3 deletions docs/capabilities.md
Original file line number Diff line number Diff line change
Expand Up @@ -75,7 +75,7 @@ visible in review rather than only in production.
| Chains per project | 1 | `schemas/project.schema.json` (`maxItems`) |
| Contract addresses | 20 | `schemas/project.schema.json` (`maxItems`) |
| Blocks per approved run | 100_000, `policy.block_budget` to change. A block count, not a duration: 100k blocks is about 14 days on Ethereum (12 s blocks), 2.3 days on Base (2 s), 7 hours on Arbitrum One (0.25 s) — set it for the chain you index | `src/plan/generate.ts` (`DEFAULT_BLOCK_BUDGET`) |
| Query deadline | 60 s | `src/query/runQuery.ts` (`DEADLINE_MS`, SIGKILL) |
| Query deadline | 60 s per query, and 60 s for loading snapshots and building models, which a build does once for all its queries | `src/query/runQuery.ts` (`DEADLINE_MS`, SIGKILL) |
| DuckDB memory | 1 GiB, spills to a temp dir | `src/query/workerMain.ts` (`MEMORY_LIMIT`) |
| Returned rows | 10_000, `policy.row_limit` to change | `src/project/limits.ts`, enforced in `src/query/workerMain.ts` (the reader stops at the limit) |
| RPC job wall clock | 30 min, resumable | `src/ingest/rindexer/runBounded.ts` |
Expand All @@ -86,8 +86,11 @@ visible in review rather than only in production.
| `fork` per-request timeout | 30 s | `src/fork/fetchGuard.ts` (`FORK_LIMITS`) |
| `fork` redirect hops | 0 | `src/fork/fetchGuard.ts` (refused outright) |

`CHAINPLOT_QUERY_MEMORY_LIMIT` overrides the memory figure; the query still
spills to disk rather than failing when it goes over.
`CHAINPLOT_QUERY_MEMORY_LIMIT` overrides the memory figure, as a size such as
`2GB` or `1536MB`; anything else is refused by name. The query still spills to
disk rather than failing when it goes over. The default stays small because a
forked recipe is untrusted; a project with millions of rows behind its models
builds far faster with 2–3 GB, where the memory is there to give.

The row limit bounds the *download*, not the rendering. The table is
virtualised — 5,000 rows put 26 in the DOM — so a wide result no longer
Expand Down
25 changes: 17 additions & 8 deletions docs/security.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,24 +16,32 @@ than the file format.
## Containment

Query execution happens in a forked child process (`src/query/workerMain.ts`),
never in the CLI process:
never in the CLI process. `build` runs all of a release's queries in one such
process: snapshots load and models build once, then each query runs in turn
against the same session.

| Control | Where |
|---|---|
| Separate process, env stripped to `PATH`/`HOME`/`LANG` | `src/query/runQuery.ts` |
| In-memory DuckDB; no database file on disk | `workerMain.ts` |
| File reads allowed for exactly the declared snapshots (`allowed_paths`), not their directories | `workerMain.ts` |
| Extension autoinstall and autoload disabled | `workerMain.ts` |
| `enable_external_access=false` **before any project SQL runs** | `workerMain.ts` |
| Single-SELECT admission control, via DuckDB's parser | `src/query/sqlGuard.ts` |
| 60 s deadline, SIGKILL on expiry | `runQuery.ts` |
| 60 s deadline for loading and models, then 60 s per query; SIGKILL on expiry, naming the query | `runQuery.ts` |
| Row limit enforced by stopping the reader, not by truncating after | `workerMain.ts` |

Ordering matters and is the part that was wrong before 2026-09-15. Snapshots
are read first, because `read_parquet` needs filesystem access. External
access is then disabled, and only after that are models materialized and the
query run. Models are project-supplied SQL like any other, so they must land
Ordering matters and is the part that was wrong before 2026-09-15. The
snapshot files are allowlisted first, by exact path, and external access is
then disabled; DuckDB refuses both to widen that list and to re-enable access
afterwards. Each snapshot is a view read in place, so a model scans only the
columns it uses rather than a copy of every column held in memory. Only after
that are models materialized and the queries run. Models are project-supplied SQL like any other, so they must land
on the closed side of that door; DuckDB does not allow external access to be
re-enabled within a session.
re-enabled within a session, so the door stays shut for every query in the
batch, not only the first. Sharing the session gives one query nothing over
another: each is still admitted only as a single SELECT, which cannot change
the session the next one runs in, and all of them come from the same recipe.

### Admission control

Expand Down Expand Up @@ -70,7 +78,8 @@ so a value cannot carry markup or a scheme into an `href`.
to the bucket can serve a consistent, hostile release. **Fork only from
buckets you would trust with the data itself.**
- **Denial of service by a hostile recipe.** Bounded, not eliminated: a forked
query gets 60 s, a row limit, and a 1 GiB memory cap that spills to a temp
query gets 60 s, as does loading the snapshots and building its models, plus a
row limit and a 1 GiB memory cap that spills to a temp
directory rather than failing. A release can still make your build slow, and
can still fill that temp directory.
- **Secrets you place inside the recipe directories.** The source bundle is an
Expand Down
39 changes: 23 additions & 16 deletions src/publish/writeRelease.ts
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ import type { CommandError } from "../cli/envelope.js";
import { loadProject } from "../project/load.js";
import { validateProject } from "../project/validate.js";
import { topoSortModels } from "../project/modelGraph.js";
import { runQuery } from "../query/runQuery.js";
import { runQueries } from "../query/runQuery.js";
import { isComplete, requiredEnd } from "../ingest/coverage.js";
import { readCoverageFile, segmentsFor } from "../ingest/coverageStore.js";
import { lastProvenCompleteBlock } from "../ingest/coverage.js";
Expand Down Expand Up @@ -178,7 +178,7 @@ export async function buildRelease(
);
}

// Models materialize in dependency order for every query.
// Models materialize in dependency order, once per build.
const modelOrder = topoSortModels(models);
const modelSql: { id: string; sql: string }[] = modelOrder.map((id) => {
const model = models.find((m) => m.id === id)!;
Expand All @@ -198,29 +198,36 @@ export async function buildRelease(
const files: string[] = [];

try {
// Queries (with models materialized first).
for (const query of queries) {
// Queries: one isolated session for the whole release. Snapshots load and
// models build once, then every query runs against them in turn.
//
// Every dataset is in scope, not just the declared one, so a model over
// dataset A can feed a query on dataset B and a query can join across
// datasets. `query.dataset` still names the provenance.
const planned = queries.map((query) => {
const dataset = datasetById.get(query.dataset);
if (!dataset) {
throw error("validation", `unknown dataset: ${query.dataset}`, {
resource_id: query.id,
pointer: "/queries",
});
}
const sqlPath = path.resolve(projectDir, query.file);
const data = await runQuery({
sql: fs.readFileSync(sqlPath, "utf8"),
// Every dataset is in scope, not just the declared one. Models are
// materialized into each query's session, so loading one table meant a
// model over dataset A failed every query on dataset B — which made
// models unusable in any multi-dataset project. It also lets a query
// join across datasets. `query.dataset` still names the provenance.
tables: allTables,
const sql = fs.readFileSync(path.resolve(projectDir, query.file), "utf8");
return { query, dataset, sql };
});
const results = await runQueries({
tables: allTables,
models: modelSql,
queries: planned.map(({ query, sql }) => ({
id: query.id,
sql,
rawAmountColumns: rawAmountNames(query.raw_amount_columns),
rowLimit: rowLimitFor(project),
models: modelSql,
});
})),
});

for (const [i, { query, dataset, sql }] of planned.entries()) {
const data = results[i]!;
const rel = path.join("results", `${query.id}.json`);
writeJson(path.join(staging, rel), {
schema_version: 1,
Expand All @@ -231,7 +238,7 @@ export async function buildRelease(
columns: decorateColumns(data.columns, query.raw_amount_columns),
rows: data.rows,
snapshot: dataset.snapshot,
query_digest: sha256(fs.readFileSync(sqlPath, "utf8")),
query_digest: sha256(sql),
raw_amount_columns: rawAmountNames(query.raw_amount_columns),
});
files.push(rel);
Expand Down
Loading
Loading