I build the browser side of web scraping: infrastructure that renders, survives and scales on pages designed to block automation.
Right now I'm a Lead Software Engineer at Crustdata, where I built and own the browser platform behind our SERP, page fetch and Google Maps APIs. It runs as a fleet of browser workers on Kubernetes and handles around a million requests a day. Before that I worked on fault-tolerant crawling pipelines for enterprise clients at Zyte, and before that I co-founded a company that ran scrapers against Instagram, TikTok and Facebook for about a thousand paying customers.
I created and maintain Pydoll, an async Python library that drives Chromium straight over the DevTools Protocol, with no WebDriver and nothing that gives the automation away. It started as a rewrite of a production crawler that Selenium could no longer keep alive, and reached #1 on GitHub trending and the front page of Hacker News at launch.
Pydoll has 7,105 stars and 408 forks, latest release 3.0.0 on Sep 29, 2026.
Recently merged
- docs: explain why the User-Agent major must equal the binary, and say… in
autoscrape-labs/pydoll, Sep 29, 2026 - fix(playwright): fractional clicks, geolocation accuracy, stale frames, wire headers and console handles in
autoscrape-labs/pydoll, Sep 29, 2026 - feat(sync): generated synchronous API for pydoll in
autoscrape-labs/pydoll, Sep 29, 2026 - Fix/fingerprint header order and coherence in
autoscrape-labs/pydoll, Sep 16, 2026 - Fix/fingerprint worker webgpu and sw refetch in
autoscrape-labs/pydoll, Sep 15, 2026
Updated 2026-09-30 12:09 UTC by a GitHub Action that reads this data with Pydoll itself.
Python (asyncio, FastAPI, Django), the Chrome DevTools Protocol, Kubernetes on EKS with KEDA, Redis, PostgreSQL, Docker, and OpenTelemetry with Grafana and Datadog for seeing what's actually happening in production.
Other things I've built
Browser worker platform. Seven worker types on a single Helm chart, queue-based scheduling with a circuit breaker in front of the workers, a four-level health recovery hierarchy, and autoscaling on queue depth. A Bing SERP service on top of it sustains around 1,000 requests per minute with stable tail latency while browsers restart and proxies rotate underneath.
Document processing pipeline. Takes multi-gigabyte archives of scanned traffic fines, runs CPU-optimized OCR, uses an LLM to pull out fields from documents with no standard layout, and reconciles everything against the database. Cut the manual work, previously done by freelancers, by more than 80%.
Signature fraud detection. A YOLO model trained on a hand-labeled dataset to find and crop signatures on driver's licenses in any format or orientation, followed by an LLM comparison guided by the rules of the manual process. Served through FastAPI and ran in production at close to 100% accuracy.
thalissfernandes99@gmail.com. Fala português? Pode mandar mensagem em português mesmo.





