Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
-
Updated
Sep 25, 2026 - Python
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
To extract article from given URL
Readability / Html Content / Article Extractor & Web Scrapping library written in PHP
Use LLMs to robustly extract web data
MinerU-HTML: An SLM-powered HTML main content extractor that outputs clean HTML bodies. Perfect for Deep Research Agents, RAG applications, and training data generation.
SmartReader is a library to extract the main content of a web page, based on a port of the Readability library by Mozilla
Parse markdown article, download images and replace images URL's with local paths
Reddit bot to preview and post hyperlinks as comments
NLP Web Service
The best HTML to Markdown library, A esm-native & Useful Utilities with simple, lightweight and epic quality.
Laravel wrapper for common NLP tasks
This is a small and easy-to-use desktop application that allows exporting Web of Science API Expanded and InCites API data in Excel/CSV/JSON/XML with a configurable and flexible data export structure.
Extract article or news by url or html, parse the title and content, output in markdown format.
open-news is a Python library for news discovery and extraction. It fetches live news via DuckDuckGo, searches Google News, discovers RSS feeds, crawls websites, and extracts article text. Features filtering, deduplication, ranking, and optional JavaScript rendering. Includes Python API, CLI, and terminal UI.
This Python package can be used to systematically extract multiple data elements (e.g., title, keywords, text) from news sources around the world in over 50 languages.
Involution King Fun Book (IKFB, Chinese: 快卷, 卷王快乐本) is an integrated management system for papers and literature. Powered by Electron.
【 Spring Boot 实战开发】10 分钟快速构建一个自己的技术文章博客
A web page content extractor
📚 Сборник полезных штук из Natural Language Processing: Определение языка текста, Разделение текста на предложения, Получение основного содержимого из html документа
To associate your repository with the article-extractor topic, visit your repo's landing page and select "manage topics."