AI Engineer — LLM agents, RAG and evaluation · Solo founder, Delhi
I build AI systems that run in production, and I measure them before I make claims about them. My support agent answers real customers for the D2C brand I run; my productivity app is my daily driver. Not tutorials, not demos — tools solving real problems, most of them mine first.
🎯 Open to remote AI engineering roles — LLM agents, RAG, evaluation, applied AI products.
| Project | What it does | Stack | |
|---|---|---|---|
| GlamShelf Twin | AI support agent for my own D2C brand — answers Instagram & WhatsApp DMs, routes each one AUTO / DRAFT / ESCALATE, and hands off to me on Telegram. RAG over store data, an output guard on prices and policy claims, and an LLM-as-judge eval harness |
Python · Flask · DeepSeek · fastembed + sqlite-vec · Telegram | |
| Prism | AI-native productivity PWA — tasks, notes, reminders, workout logging, SM-2 flashcards generated from notes, PDFs and YouTube. Offline-first | Next.js 14 · TypeScript · Supabase · Groq | Live demo ↗ |
| The Glam Shelf | D2C false-eyelash brand I run solo on Shopify — 1,000+ orders | Shopify · Razorpay · Shiprocket | Store ↗ |
| Project | What it does | Stack |
|---|---|---|
| taste-engine | Music recommender over a year of my YouTube history (40,619 plays). Beats a most-played baseline on nDCG@20 in all 3 held-out periods under nested tuning — and the README documents every time the harness caught its own measurement error, plus six ideas tested and rejected | Python · sentence-transformers · HDBSCAN · YouTube Data API |
| agentgrade | Rubric-driven LLM-as-judge CLI for support-agent transcripts — tone, accuracy, compliance, resolution. Works with any OpenAI-compatible provider | Python · LLM-as-judge |
| Project | What it does | Stack |
|---|---|---|
| AI Intelligence Daily | 90-second daily AI briefing delivered to Telegram every morning — curated for founders | Python · Groq · GitHub Actions |
| Daily Digest Newsbot | Twice-daily AI-summarised news briefing to Telegram — RSS-backed, multi-source | Python · Gemini 2.5 Flash · GitHub Actions |
- Evaluate before claiming — held-out splits, human-labelled eval sets, and judges checked against human verdicts. When a fancier option loses to a simpler one on the numbers, the simpler one ships.
- Agents that ship — intent routing (
AUTO/DRAFT/ESCALATE), human-in-the-loop approval, output guards, real webhooks and payment/fulfilment integrations. - Debug from the data — production incidents traced through database state, cron and HTTP logs, not guesses.
- End-to-end ownership — from Postgres schema and Row Level Security to PWA service workers, deploys and CI.
- Agentic-IDE fluent — Claude Code and Google Antigravity, driven by written briefs and verified live.
Languages — Python · TypeScript · SQL AI — DeepSeek · Groq (gpt-oss-120b) · Claude API · Gemini · Hugging Face Transformers · LoRA / PEFT · Ollama Retrieval & ML — fastembed (ONNX) · sqlite-vec · sentence-transformers · HDBSCAN · LLM-as-judge Web — Next.js 14 · React · Flask · FastAPI · Tailwind CSS Data — Supabase · PostgreSQL · SQLite Infra — Render · Vercel · GitHub Actions Commerce — Shopify · Razorpay · Shiprocket · WATI


