Preserve Python and Node symbols in memory flamegraphs - #556
not-matthias wants to merge 13 commits into
Conversation
Merging this PR will degrade performance by 10.47%
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | WallTime | memtrack track tar |
8.4 s | 9.4 s | -10.71% |
| ❌ | Memory | encode_events_realistic[8] |
169.6 MB | 189 MB | -10.24% |
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing cod-3654-support-memory-flamegraphs-for-pythonnode (a01c776) with main (214c040)
Footnotes
-
6 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
9108e6c to
1b0d6fd
Compare
|
366917f to
25472c3
Compare
a01c776 to
5e00d30
Compare
GuillaumeLagrange
left a comment
There was a problem hiding this comment.
olgtm, second round will be quick
Module artifact extraction checks that a mapped path still names the file that was mapped by comparing the perf MMAP2 (dev, ino) with stat(path). On nested overlayfs the two disagree for unchanged files: perf and /proc/PID/maps report the overlay superblock's device, while stat can return a per-layer pseudo device when the outer overlay cannot encode its layer in the inode's high bits (the inner overlay's xino already uses them). Every system library was rejected as changed, so libc and ld.so shipped no symbols or unwind data and their frames showed as unresolved addresses. Map each module once and read the device and inode of that mapping from /proc/self/maps, which the kernel derives the same way as MMAP2. Replaced files still get a different inode and are still rejected. Symbols, load bias and unwind data are parsed from the same mapping, so a path replaced during extraction cannot pair the verified identity with another file's contents.
Python and Node write runtime symbols to /tmp/perf-<pid>.map, and Python JIT dumps carry the unwind data needed to walk through interpreter trampolines. Collect both for the benchmark processes before saving the memtrack metadata, reusing the walltime artifact pipeline, so offline allocation stacks keep their runtime frames.
Memory stacks can only name Python and Node frames from the runtime perf maps. Set PYTHONPERFSUPPORT and the Node perf options per benchmark command, as simulation mode already does. Memory mode also passes --interpreted-frames-native-stack, since interpreted JS frames otherwise all resolve to the shared V8 interpreter trampoline.
Track a native allocation made from a Python function (perf trampoline) and a Node function (V8 perf-basic-prof). Assert the allocation carries a captured stack and that the runtime perf map names the allocating function, which offline attribution needs.
A process can map the same file at several addresses at once. V8 remaps its embedded builtins out of the node binary into its code range, so node runs from both its original text and that copy. Each placement has its own load bias, but only the last mapping per (path, pid) was kept, so frames in every earlier placement lost their symbols and unwind data. In a Node memory profile, all native node frames showed up as unresolved addresses. Record every distinct load bias and the unwind data of every executable mapping for each process, and emit them all in the artifact metadata. The walltime perf path shares this bookkeeping and gets the same fix. Add a sample that maps its own text a second time and allocates through both copies. The test feeds that process's real mappings into the artifact pipeline and checks that both copies resolve. Refs COD-1377
A failed perf-map append aborted the whole JIT harvest, discarding unwind data already collected for other pids while their symbols stayed written. Handle it per pid like the other per-dump failures. The function can no longer fail, so return the map directly; wall-time no longer aborts the benchmark save on this error.
5e00d30 to
1f73202
Compare
…rapper The runner (`helpers/env.rs`) and exec-harness each defined the environment a benchmark process needs: `CODSPEED_RUNNER_MODE`, the Python hash seed and perf-map switches, and the Java tool options. Move them into one `runner_shared::runtime_env` module, keyed on `MeasurementMode`, which moves to runner-shared as well. The runner converts its `RunnerMode` and extends its injected env from the shared list; the values are unchanged except that memory mode now also sets `PYTHONPERFSUPPORT=1`, so memory flamegraphs can name Python frames. Add a `node` wrapper script that the module installs on `PATH`. It runs the real node with the V8 flags codspeed-node's `getV8Flags()` would request for the current `CODSPEED_RUNNER_MODE` and node major version: the deterministic analysis set for simulation and memory, and the perf-prof/log-code set for walltime. Most of these flags are rejected in `NODE_OPTIONS`, so exec-harness targets without a codspeed-node integration had no way to get them before. The wrapper may run under valgrind with the simulation preload library in `LD_PRELOAD`, which reports a benchmark result from every process that loads it. The script therefore drops `LD_PRELOAD` for its only subprocess (`node --version`), restores it for the final `exec`, and avoids command substitutions whose forked subshells would report on exit. Under callgrind only the real node process emits the benchmark dump. The install writes a staging file and renames it into place, once per process: writing over an executing script, or a second in-process writer racing with a concurrent fork+exec, fails with `ETXTBSY`.
Set the benchmark environment once in `execute_benchmarks`, before the mode dispatch, through `runner_shared::runtime_env::apply_to_process`. Every spawned command inherits it, so the per-command `set_perf_map_env` and `set_node_options` calls in the memory, simulation and walltime loops go away, together with the local `node.rs` and its `NODE_OPTIONS` subset. Node targets now go through the shared `node` wrapper and receive the full codspeed-node V8 flag set for the mode, including indirect launches through npm, npx or `#!/usr/bin/env node` scripts. `MeasurementMode` is re-exported from runner-shared, so the CLI is unchanged.
codspeed-node's analysis flag set has no `--perf-basic-prof`, so the node wrapper in memory mode produced no `/tmp/perf-<pid>.map` and memory flamegraphs lost their JS frames. Add the flag for memory mode only. Switch the memtrack interpreter tests to the shared runtime env instead of hand-picked interpreter flags, so they exercise the same Python env and node wrapper the runner injects.
1f73202 to
2b88350
Compare
Summary
NODE_OPTIONS.Verification
cargo test --release --bin codspeed writes_keyed_artifacts_and_metadata_for_a_streamed_mappingcargo fmt --all --check