Skip to content
View JonSnow1807's full-sized avatar

Block or report JonSnow1807

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
JonSnow1807/README.md
Chinmay Shrivastava — Software Engineer. Backend, distributed systems and ML infrastructure.

LinkedIn Email Chinmay Hugging Face

I build high-performance backend and ML infrastructure, and I validate it against references instead of trusting it.

At Reliance Jio, I work on production 4G/5G analytics, deterministic query compilation and GPU-backed inference on Kubernetes/OpenShift, including a 13-node, air-gapped, IPv6-only deployment. Outside work, I contribute to PyTorch and bpfilter, and build tools for GPU performance and distributed coordination.

01 / Upstream contributions

PyTorch core — two accepted changes
Enabled existing dynamic-shape support for foreach operations by default under torch.compile, with a test update (#158985). Exposed rearrange through the torch.func public API, with tests and documentation (#173183).

Meta’s bpfilter — two merged pull requests
Added IPv4 Type of Service and IPv6 Traffic Class matchers, including command-line parsing, BPF code generation and unit tests (#364, #369). The IPv6 implementation uses endianness-safe byte loads.

PyTorch landing commits: foreach · rearrange.

02 / Selected builds

Mustard Watch Party — Multi-instance video synchronization. 48 ms; P95 player-reported drift. 3 Chrome clients / ~300 ms RTT. Steady state / 240-second test. Open the source repository.Mustard Watch Party — open Live: mustard.watch.

Inside Mustard · Clock synchronization, protocol correctness and runtime trade-offs

Mustard: clients estimate a clock offset over the socket ack and correct drift with predictive rate control; room state lives in Redis behind atomic Lua scripts. Chart: steady-state P95 drift 49, 48, 79, 29 and 86 ms across five network scenarios; relay memory 14.5 KB (Rust) versus 40 KB (Go) per connection.

Mustard is a multi-instance watch-party platform: clients connect over WebSockets and use NTP-style clock estimation with predictive drift correction, and updates go through atomic Redis Lua scripts. The 48 ms P95 is a specific claim: steady-state player-reported drift, measured across three Chrome clients at approximately 300 ms RTT during a 240-second test. It is not a physical audio-output measurement.

Protocol correctness was a separate question. TLA+/TLC model checking exposed a stale-epoch bug after store resets, leading to an ordered-epoch fix, and idempotency keys plus atomic updates keep control commands from being applied twice within the deduplication window.

I also built Go and Rust relays to test protocol conformance and runtime costs. They are study implementations, not the production backend. In a local 10,000-connection test, the Rust relay used 14.5 KB of memory per connection versus 40 KB for Go; the extreme-tail latency comparison was inconclusive.

Live site · Source · Sync design · TLA+ findings · Go/Rust study

Fused CUDA Operators — LayerNorm / RMSNorm for PyTorch. 1.04–1.76×; vs. compiled PyTorch. RMSNorm-to-FP8 / A100 / FP16. Kernel-time comparison. Open the source repository.Fused CUDA Operators — open Results: A100 · kernel time.

Inside Fused CUDA operators · Custom ops, deterministic gradients and FP8

Fused CUDA operators: the eager composite runs residual-add, normalization and an FP8 cast as three kernels with HBM round-trips; the fused version does all three in one kernel. Chart: dynamic RMSNorm to FP8 kernel time is 4.9 to 7.2 times faster than the eager composite and 1.04 to 1.76 times faster than the compiled composite across six shapes on an A100 in FP16.

These LayerNorm/RMSNorm kernels combine fused residual-add, FP8 outputs and deterministic backward reductions, and come as drop-in PyTorch modules that work under torch.compile without graph breaks. Parameter gradients use fixed-order reductions rather than atomic accumulation, which is where the determinism comes from.

For dynamic RMSNorm-to-FP8, I measured 4.9–7.2× over the eager PyTorch composite and 1.04–1.76× over the compiled composite, across tested FP16 shapes on an NVIDIA A100. Those are kernel-time comparisons, not end-to-end model speedups. For a fuller picture, the repository also documents H100 results and configurations where PyTorch wins.

Source · A100 results · Methodology

PyTorch AutoTune — Workload-specific training optimization. 2.7–6.7×; training-step speedup after tuning. vs. PyTorch-default FP32 eager. 3 workloads / NVIDIA A100. Open the source repository.PyTorch AutoTune — open PyPI: pytorch-autotune.

Inside PyTorch AutoTune · Measure the configuration instead of guessing it

PyTorch AutoTune: the baseline is measured first, candidates across precision, compile mode, memory format and fused optimizer run real training steps under a budget, the winner must beat the incumbent by more than 3 percent, and the report and cache record every trial. Chart: 2.71 times on ResNet-50, 3.84 times on ResNet-18 and 6.65 times on a six-layer Transformer versus the PyTorch-default fp32 eager baseline; the Transformer gain is 2.55 times against an eager baseline with TF32 matmuls enabled.

Rather than guess at a training configuration, the tuner benchmarks precision, compile mode, memory format and fused-optimizer configurations on the target model and batch. Search runs under a budget, configurations are cached for reuse, and reports show trial results, compilation cost and estimated break-even time.

I measured 2.7–6.7× training-step speedups after tuning versus the PyTorch-default FP32 eager baseline across ResNet-50, ResNet-18 and a six-layer Transformer on an NVIDIA A100, with the tuner using torch.compile alongside the other optimizations. The baseline matters, though: on the Transformer, the gain is 2.55× against an eager baseline with TF32 matmuls enabled, rather than 6.65× against the defaults. Search costs and numerical trade-offs are documented as well.

PyPI · Source · Benchmark artifacts

ChatGPT Spark — Conversation capture and semantic search. Published; Chrome Web Store. React dashboard / FastAPI backend. ChromaDB vector search. Open the source repository.ChatGPT Spark — open Web Store: Chrome extension.

Inside ChatGPT Spark · A shipped extension and its retrieval backend

ChatGPT Spark: a Chrome extension captures a conversation, a FastAPI backend stores it, ChromaDB holds the embeddings, and a React dashboard browses and semantically searches them. Published on the Chrome Web Store.

Spark is a Chrome extension I built and published for capturing and searching conversation history. A React/TypeScript dashboard lets you browse saved conversations, FastAPI handles backend requests, and ChromaDB stores the embeddings behind semantic search.

Source · Chrome Web Store

Also built: 3D Point Cloud Viewer — C++/OpenGL, octree-based culling and level-of-detail rendering.

03 / Toolbox

Area Tools I use
Languages C++, C, Python, Rust, Go, CUDA, TypeScript, SQL
Backend & data FastAPI, NestJS, WebSockets, SSE, ClickHouse, PostgreSQL/PostGIS, Redis, ChromaDB
Infrastructure Linux, Docker, Kubernetes/OpenShift, AWS, CI/CD
ML & performance PyTorch, torch.compile, vLLM, GPU profiling, kernel fusion, benchmark design
Correctness TLA+/TLC, reference-based tests, regression tests, protocol-conformance checks

Have a systems problem worth digging into?
I’m interested in backend, distributed-systems, ML-infrastructure and performance-engineering work.

Frisco, Texas · Open to relocation
Email · LinkedIn · Hugging Face

Pinned Loading

  1. Fused-LayerNorm-CUDA-Operator Fused-LayerNorm-CUDA-Operator Public

    Fused normalisation CUDA kernels for PyTorch — residual-add+LayerNorm/RMSNorm, fp8 outputs, fast deterministic backward, torch.compile-native, with reproducible A100/H100 benchmarks

    Python 1

  2. Mustard-Watch-Party Mustard-Watch-Party Public

    Real-time video synchronization for YouTube watch parties. React + NestJS + Socket.IO, PostgreSQL/Prisma, Redis. JWT auth with token revocation, sliding sessions and Google OAuth; exactly-once sync…

    TypeScript 1

  3. pytorch-autotune pytorch-autotune Public

    Measurement-driven PyTorch training autotuner — tries precision × torch.compile × memory-format configs on your model, keeps the fastest, never returns one slower than your baseline. Honest, commit…

    Python 1

  4. Medical-Prescription-OCR Medical-Prescription-OCR Public

    OCR system for handwritten medical prescriptions using Donut transformer and zero-shot classification

    Jupyter Notebook 25 6