#data-extraction (13 Repositories)
Ranked open-source repositories tagged with #data-extraction, scored by pull request acceptance likelihood and maintainer engagement velocity.
46.8%
23.1h
13 repositories tagged #data-extraction
bzsanti/oxidizePdf
Pure Rust PDF library for AI/RAG: structure-aware chunking, no ML, no C deps.
us/crw
Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scrape, /crawl, /search). 2.3x faster than Tavily, 1.5x faster than Firecrawl in 1K-URL benchmarks. 6 MB RAM, single binary. Self-host or use managed cloud.
MontFerret/ferret
Declarative data automation language and Go runtime for structured extraction workflows.
apify/apify-sdk-python
Apify SDK for Python—The official library for building Apify Actors: serverless cloud programs for web scraping, browser automation, data processing, and AI agents. Manages the Actor lifecycle, storages (datasets, key-value stores, request queues), events, proxies, and pay-per-event monetization. Built on top of the the Apify API Client.
soxoj/socid-extractor
⛏️ The extraction engine behind Maigret: turn any profile URL into a structured OSINT record across 150+ sites
Xquik-dev/x-twitter-scraper
X (Twitter) Scraper API and X API Alternative. You do not need an official X developer account. You do not need to connect or use an X account for supported scraping. Not affiliated with X Corp.
yfedoseev/pdf_oxide
The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.
cporter202/coreclaw-api-directory
An unofficial, categorized directory of 118 CoreClaw Worker APIs for web scraping, automation, lead generation, e-commerce, social data, and more.
a-maliarov/amazoncaptcha
Pure Python, lightweight, Pillow-based solver for Amazon's text captcha.
py-pdf/benchmarks
Benchmarking PDF libraries
mrshu/github-statuses
The "Missing GitHub Status Page" -- a Flat Data attempt at historically documenting GitHub statuses
extractus/article-extractor
To extract article from given URL
shcherbak-ai/contextgem
ContextGem: Effortless LLM extraction from documents