Windows x64 server, client and processing-only control for the shared benchmark. The repository README covers campaigns and results; the contract defines the protocol, validation, load model and result schema.
From the repository root in PowerShell (Visual Studio 2026 C++ tools, Windows SDK, CMake with the Visual Studio 18 2026 generator):
cmake -S src/cpp -B src/cpp/build -G "Visual Studio 18 2026" -A x64 -DBENCH_ARCH=AVX512
cmake --build src/cpp/build --config Release
./src/cpp/build/Release/tcpbench.exe selftest --corpus tests/fixtures/golden.bin
./tests/run_all.ps1CMake verifies the vendored simdjson 3.12.3 files against pinned SHA-256 hashes and requires C++23 on x64. Release is /O2 without /GL//LTCG: whole-program optimization measured about 30% slower JSON processing with simdjson on MSVC 19.51. All validation and integrity checks stay enabled in Release.
BENCH_ARCH sets MSVC's /arch for the whole program and simdjson (AVX512, AVX2 or empty). bench/build.ps1 picks the best one the CPU supports, because the .NET JIT also compiles for the host CPU. It matters for simdjson: without /arch, MSVC defines none of the macros simdjson uses to choose its compile-time kernel, so the On-Demand front end becomes the generic fallback (measured 1-2% slower here). Every result records build.arch, build.parser_kernel and build.parser_builtin.
AddressSanitizer build for correctness only (never performance):
cmake -S src/cpp -B src/cpp/build-asan -G "Visual Studio 18 2026" -A x64 -DBENCH_ASAN=ON
cmake --build src/cpp/build-asan --config RelWithDebInfo
./tests/run_all.ps1 -AsanWindows ThreadSanitizer is not available; no race-freedom claim is made beyond review and the concurrent stress in the suites.
| File | Responsibility |
|---|---|
| bench.hpp | Shared types: options, JSON writer, histogram, epoch/result, owned rows, budget, processor, corpus, resources |
| main.cpp | Option parsing, clocks, JSON writer, histogram, build/resource metadata, socket helpers, global operator new counting |
| processing.cpp | Strict parsing with simdjson On-Demand, canonical digest, retention and eviction, corpus loading, processing-only control, selftest |
| fastjson.cpp | The alternative schema-specific parser (--parser fast) and its UTF-8 validator: a line-by-line port of the C# FastJson |
| selftest_fuzz.cpp | Differential fuzzing: fast parser vs the simdjson path, the validator vs a scalar reference, inputs placed against a no-access guard page |
| server.cpp | IOCP server: one operation per connection, inline synchronous completions, per-frame requeue, deadline scanner, measured interval |
| client.cpp | Load generator: per-connection sender and ACK reader around a lock-free slot ring, closed-loop and scheduled modes, pacing, accounting |
- Server. Each connection has exactly one outstanding overlapped operation, owned through
pending; a completion is dequeued before its buffers can be released, and the registry keeps the connection alive until then. Sockets useFILE_SKIP_COMPLETION_PORT_ON_SUCCESS, so synchronous completions continue inline within a frame (as .NET sockets do). After each response the connection re-posts itself to the IOCP, so connections take turns instead of one worker serving a connection whose next frame is already buffered. A 10 ms scanner closes expired connections; processing checks the same frame/idle budget every 1,024 records. - Client. Each connection has a sender and an ACK reader sharing a slot ring (
tailproduced,sent,headacknowledged) and a single wait/notify word. Closed loop: the sender refills its own connection when an acknowledgement frees a slot. Scheduled: a raised-priority generator paces arrivals (high-resolution waitable timer plus at most 1 ms of spin), admits them against global frame/byte caps, and never waits. Sends gather header and payload; the reader reads batches of acknowledgements but never past what is owed. Sockets are overlapped handles: a non-overlapped socket would serialize the sender's blocking send behind the reader's blocking receive. On failure or the drain deadline, sockets are shut down and pending I/O cancelled. - Processing. Keys are compared raw unless they contain an escape; integers are validated and parsed from the raw token in one pass (On-Demand does not validate unread scalars); stage 1 validates all UTF-8. Retained rows and text use the same chunked layout as C# (first chunk grows from 16 rows / 256 bytes; later chunks 1,024 rows / 64 KiB; messages never straddle), so large batches never need one large block or a large copy.
- Allocation counters count global
operator newcalls and bytes in per-thread cache-line shards; they do not see allocations made outsideoperator new(for example by the OS or_aligned_malloc). ASAN builds leave them at zero.
Diagnostic options (--pause-every/--pause-ms busy spin, --io-cap partial I/O) are documented in the contract and never used in performance cells except the labelled fragmentation stress.
- Fast parser.
--parser fastswaps simdjson for fastjson.cpp, a line-by-line port of the C# FastJson, so both languages can run the same algorithm: one UTF-8 validation per frame (Keiser-Lemire), 64-byte string scans, integer key matching, eight-digit integer parsing and speculative 32-byte copies when decoding escapes. The functions C# marksAggressiveInliningare__forceinlinehere, and the throw helper is out of line, like the managedFail. It accepts and rejects exactly what the simdjson path does. The selftest fuzzes one against the other with guard pages and scratch canaries, and the whole suite also runs under AddressSanitizer. It needs AVX2 (checked at startup) and uses AVX-512 for the wide scan when the CPU has it.BENCH_PARSER=fastmakes it the default for the shared test suites.