Back to Topics Directory
Topic Hub

#data-extraction (13 Repositories)

Ranked open-source repositories tagged with #data-extraction, scored by pull request acceptance likelihood and maintainer engagement velocity.

Topic Avg Merge Rate

46.8%

Avg Review Latency

23.1h

Filter by language

13 repositories tagged #data-extraction

S TierRust 185

bzsanti/oxidizePdf

Pure Rust PDF library for AI/RAG: structure-aware chunking, no ML, no C deps.

96.7%
Merge Rate
3h
First Review
100%
1st-Timers
2
Maintainers
A TierRust 741

us/crw

Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scrape, /crawl, /search). 2.3x faster than Tavily, 1.5x faster than Firecrawl in 1K-URL benchmarks. 6 MB RAM, single binary. Self-host or use managed cloud.

94.6%
Merge Rate
3d
First Review
75%
1st-Timers
4
Maintainers
A TierGo 6.0k

MontFerret/ferret

Declarative data automation language and Go runtime for structured extraction workflows.

93.2%
Merge Rate
2d
First Review
100%
1st-Timers
2
Maintainers
A TierPython 176

apify/apify-sdk-python

Apify SDK for Python—The official library for building Apify Actors: serverless cloud programs for web scraping, browser automation, data processing, and AI agents. Manages the Actor lifecycle, storages (datasets, key-value stores, request queues), events, proxies, and pay-per-event monetization. Built on top of the the Apify API Client.

88.3%
Merge Rate
2d
First Review
100%
1st-Timers
6
Maintainers
B TierPython 1.1k

soxoj/socid-extractor

⛏️ The extraction engine behind Maigret: turn any profile URL into a structured OSINT record across 150+ sites

100.0%
Merge Rate
11h
First Review
100%
1st-Timers
1
Maintainers
B TierJavaScript 187

Xquik-dev/x-twitter-scraper

X (Twitter) Scraper API and X API Alternative. You do not need an official X developer account. You do not need to connect or use an X account for supported scraping. Not affiliated with X Corp.

67.4%
Merge Rate
1d
First Review
100%
1st-Timers
1
Maintainers
B TierRust 963

yfedoseev/pdf_oxide

The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.

67.5%
Merge Rate
3d
First Review
47%
1st-Timers
8
Maintainers
D TierPython 152

cporter202/coreclaw-api-directory

An unofficial, categorized directory of 118 CoreClaw Worker APIs for web scraping, automation, lead generation, e-commerce, social data, and more.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 493

a-maliarov/amazoncaptcha

Pure Python, lightweight, Pillow-based solver for Amazon's text captcha.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 338

py-pdf/benchmarks

Benchmarking PDF libraries

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierJavaScript 554

mrshu/github-statuses

The "Missing GitHub Status Page" -- a Flat Data attempt at historically documenting GitHub statuses

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierTypeScript 1.9k

extractus/article-extractor

To extract article from given URL

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 2.0k

shcherbak-ai/contextgem

ContextGem: Effortless LLM extraction from documents

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
Best Data-extraction Open Source Repositories & C-Rank™ | GetMerged