Skip to content
View thalissonvs's full-sized avatar
🎯
Focusing
🎯
Focusing

Sponsors

@LambdaTest-Inc
@Alex-Byteful
@TheWebScrapingClub
@ManagerNodemaven

Organizations

@autoscrape-labs

Block or report thalissonvs

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
thalissonvs/README.md

Hi, I'm Thalison 👋

I build the browser side of web scraping: infrastructure that renders, survives and scales on pages designed to block automation.

Right now I'm a Lead Software Engineer at Crustdata, where I built and own the browser platform behind our SERP, page fetch and Google Maps APIs. It runs as a fleet of browser workers on Kubernetes and handles around a million requests a day. Before that I worked on fault-tolerant crawling pipelines for enterprise clients at Zyte, and before that I co-founded a company that ran scrapers against Instagram, TikTok and Facebook for about a thousand paying customers.

Pydoll

I created and maintain Pydoll, an async Python library that drives Chromium straight over the DevTools Protocol, with no WebDriver and nothing that gives the automation away. It started as a rewrite of a production crawler that Selenium could no longer keep alive, and reached #1 on GitHub trending and the front page of Hacker News at launch.

Stars PyPI downloads Docs

Live

Pydoll has 7,105 stars and 408 forks, latest release 3.0.0 on Sep 29, 2026.

Recently merged

Updated 2026-09-30 12:09 UTC by a GitHub Action that reads this data with Pydoll itself.

Snake eating my contribution graph

What I work with

Python (asyncio, FastAPI, Django), the Chrome DevTools Protocol, Kubernetes on EKS with KEDA, Redis, PostgreSQL, Docker, and OpenTelemetry with Grafana and Datadog for seeing what's actually happening in production.

Other things I've built

Browser worker platform. Seven worker types on a single Helm chart, queue-based scheduling with a circuit breaker in front of the workers, a four-level health recovery hierarchy, and autoscaling on queue depth. A Bing SERP service on top of it sustains around 1,000 requests per minute with stable tail latency while browsers restart and proxies rotate underneath.

Document processing pipeline. Takes multi-gigabyte archives of scanned traffic fines, runs CPU-optimized OCR, uses an LLM to pull out fields from documents with no standard layout, and reconciles everything against the database. Cut the manual work, previously done by freelancers, by more than 80%.

Signature fraud detection. A YOLO model trained on a hand-labeled dataset to find and crop signatures on driver's licenses in any format or orientation, followed by an LLM comparison guided by the rules of the manual process. Served through FastAPI and ran in production at close to 100% accuracy.

Get in touch

thalissfernandes99@gmail.com. Fala português? Pode mandar mensagem em português mesmo.

Pinned Loading

  1. autoscrape-labs/pydoll autoscrape-labs/pydoll Public

    Pydoll is a library for automating chromium-based browsers without a WebDriver, offering realistic interactions.

    HTML 7.1k 408

  2. autoscrape-labs/injekta autoscrape-labs/injekta Public

    Lightweight, type-safe dependency injection for Python. One decorator, zero dependencies, full type inference.

    Python 6