Skip to content

CI performance dashboard #1182

Description

Maintained by the CI profiler routine, which runs every 4h. Each run replaces this body and carries the history rows over.

Headline (run 2026-10-09 08:17 UTC)

Window: 2026-10-08 08:17 → 2026-10-09 08:17. Now = runs created since the previous profile (04:17).

Sample:

  • Job detail for 90 runs (12,459 jobs):
    • 52 CI merge_group runs: every completed one since 04:17, plus every 5th before that
    • 13 CI PR runs (8 non-draft) and 3 CI push runs
    • 1–2 runs for each compatibility workflow
  • Run counts per workflow × event come from the total_count of filtered run listings.
  • 14 PR timelines.
metric 24h now
PRs merged 69 13 between 06:56 and 08:01 (one burst)
Merge queue: last enqueue → merge, p50 / p90 (14 PRs, 03:14–08:01) 64 / 69 min the burst drained in waves of 5 (#1248, new)
Merge queue: first enqueue → merge, p50 / p90 65 / 82 min –
Evictions in timelines #1215 1×, #1205 1× –
merge_group CI success, created → done, p50 / p90 23.4 / 46.1 min (n=80) 21.5 / 23.0 min (n=16)
merge_group CI runs: success / failure / cancelled 80 / 16 / 54 (of 157) see below
PR push → ci-ok (CI success runs over 10 min), p50 / p90 31.7 / 34.2 min (n=58, 23:25–07:47) –
macOS queue wait, p50 / p90 (worst hour) 0.2 / 1.8 min (15h: 0.2 / 5.2) peak 50 concurrent macOS jobs at 06:42
Linux queue wait, p50 / p90 (worst hour) 0.1 / 1.4 min (20h: 0.9 / 5.8) –
Windows queue wait, p50 / p90 (worst hour) 0.0 / 0.7 min 08h: 3.1 / 6.3

Waste on merge_group runs (52-run sample):

run type runs job-min avg / run
successful 27 12,795 474
cancelled purely as rebuilds behind an evicted entry 15 4,311 287
cancelled or failed with a genuine failing job 10 4,110 –

Job-minutes/day. Estimated as per-run sample × run counts.

Critical paths now.

Merge-queue settings (ruleset 24668460): max_entries_to_build 5, max_entries_to_merge 5, ALLGREEN, 120-min timeout.

History

run (UTC) merges/24h MQ enq→merge p50/p90 (min) mg CI success p50 PR push→ci-ok p50/p90 mac queue p50/p90 linux queue p50/p90 job-min/day L/W/M
2026-10-09 08:17 69 first 65 / 82 · last 64 / 69 (burst) 23.4 (21.5 since 04:17) 31.7 / 34.2 (n=58) 0.2 / 1.8 0.1 / 1.4 ~180k / ~40k / 10.1k (Gradle compat not re-measured)
2026-10-09 04:16 ~66 first 100 / 380 · last 86 / 142 (47 since 21:35) 25 (22 since 21:35) 31 / 33 (n=7) 0.2 / 1.9 0.1 / 3.3 ~263k / ~89k / 13.4k (Gradle carried over)
2026-10-09 00:17 72 first 48 / 290 · last 26 / 78 39.1 (22.6 after #1143) 47.2 / 56.9 (31.7 / 33.8 now) 0.2 / 1.7 0.4 / 3.7 295k / 92k / 10k (all workflows, 24h)
2026-10-08 21:50 49 80 / 157 44.5 (23.9 post-#1133) 43.2 / 71.9 (29.0 / 30.3 post-#1133) 22.0 / 55.5 1.2 / 12.8 97.5k / 28.3k / 12.0k (7-day avg)
2026-10-08 19:36 43 83 / 177 42.6 (22–25 after #1133) 41.6 / 78.5 1.1 / 69.8 0.1 / 7.6 139k / 21k / 21k (CI only, 24h; L/M/W order)

Open ci-perf issues by ROI

ROI on one scale: (thousands of job-min/day weighted Linux ×1, Windows ×2, macOS ×3, plus 1 per critical-path or PR-latency minute) × confidence ÷ effort (S=1, M=2, L=3). Issues #1176–#1181 carry markers on a different scale; this table re-ranks them.

ROI issue saving status
~40 #1177 Gradle compat: PR runs repeat ci.yml's Gradle tier ~58k Linux + 38k Windows PR #1166 in the merge queue (08:17)
27 #1198 the other 10 compatibility workflows run on ~80% of PR pushes ~66k Linux + 22k Windows PR #1206 open
~20 #1170 CI on push to main re-tests the SHA the queue passed 17k Linux + 2.6k Windows + 2.4k macOS (re-measured)
9.0 #1199 coverage-docker compiles the dependency graph twice per leg ~8–10k Linux
7 #1248 (settings) merge queue builds 5 entries at a time (new) ~21 min per PR past the 5th in a burst, about +5k job-min/day human decision
~4.3 #1181 test-release is the whole PR critical path ~8–10 min off PR push→ci-ok
~4.2 #1178 ci.yml e2e: ~108 legs/run, Run e2e tests 1.3 min mean ~12k Linux/Windows PR #1234 in the merge queue
2.6 #1225 coverage-docker (sbt): image rebuild, merge-queue critical path ~1–2 min/merge_group run until #1173/#1171 land, ~0.7k Linux refreshed
2.1 #1176 e2e-macos: 23 one-minute macOS jobs ~1k macOS; gates #1248 refreshed
2.1 #1171 Gradle e2e shards unbalanced ~1–2 min/merge_group run
1.9 #1172 e2e fan-out (overlaps #1176 + #1178) see #1176/#1178
1.3 #1173 test (windows/macos): redundant Build step ~2 min/leg; on the merge_group critical path in 10 / 27 runs
0.42 #1174 coverage runs llvm-cov report twice ~530 Linux

Merged ci-perf / ci-janitor PRs (last 7 days) and measured effect

PR merged before → after
#1143 decouple CI platform builds 10-08 21:34 merge_group CI success p50 25 → 22 → 21.5 min (n=16 since 04:17); p90 46 → 23.0 min. Delivered, holding.
#1146 fail-fast eviction 10-08 21:05 Evicted runs average 287 job-min (rebuilds) to 348 (real failures), against 474 for a full run. Delivered.
#1208, #1189 Maven Central / reactor retries 10-09 07:19 / 00:13 No Maven eviction since 00:26 (2 genuine failures since 04:17, neither Maven). Looks delivered; re-check next run.
#1226, #1229 sbt warm-up retries; #1215 macOS Composer setup-php 10-09 07:41–08:01 Too early. One sbt eviction at 06:34 (coverage-docker (sbt), a Docker path that #1238 targets, not these).
#1201 macOS e2e unpack Broken pipe 10-09 02:12 No macOS unpack failure in the sampled runs since 04:17. Likely delivered.
#1159, #1162, #1165, #1169, #1154, #1148, #1192 (flake fixes) 10-08 21:04 → 10-09 02:12 Partly delivered. A Windows test flake (#1237 open) still evicted a 5-run wave at 06:19.
#1179 e2e artifact size −53% 10-08 23:51 "Download the e2e binaries" 0.08 min/job (576 min over 7,316 jobs). Holding.
#1133 shard Gradle e2e + test, skip test-release in queue 10-08 16:23 merge_group CI p50 48.0 → 21.5 min. Delivered.
#1093 skip drafts, macOS off the PR path 10-08 02:16 macOS queue p50 22.0 → 0.2 min, p90 1.8. PR CI runs for drafts finish with 1 job. Delivered.
#1139, #1102, #1087, #1080, #1059, #1018, #619, #897, #892, #874 10-02 → 10-08 See earlier rows. Delivered, or not a time change.

Method notes

  • GraphQL is blocked in this environment, so merge-queue state came from REST timelines of 14 merged PRs plus merge_group run head branches.
  • Paginated job listings follow Link headers to numeric-ID repositories/{id} URLs, which the proxy rejects. Pages 2+ have to be fetched explicitly with &page=N.
  • Queued-then-cancelled jobs with no runner count as queue time, not run time.
  • About 300 GitHub REST requests this run.

Generated by Claude Code

Activity

  1. mikolalysenko commented on Oct 8, 2026

    @mikolalysenko
    CollaboratorAuthor

    [agent] There are now two open dashboard issues: this one and #1175, both titled "CI performance dashboard" and labeled ci-perf-dashboard. The body says it is maintained in place, so the profiler routine should update one issue rather than open a new one each run. I'm not closing either one, because I can't tell which issue the routine will update next. Whoever owns the profiler should close the copy it no longer updates.


    Generated by Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions