Back to Topics Directory
Topic Hub

#web-scraping (30 Repositories)

Ranked open-source repositories tagged with #web-scraping, scored by pull request acceptance likelihood and maintainer engagement velocity.

Topic Avg Merge Rate

39.0%

Avg Review Latency

33.6h

Filter by language

30 repositories tagged #web-scraping

S TierRust 137 4 GFIs

konippi/servo-fetch

A self-contained browser engine that fetches, renders, and extracts web content as Markdown, JSON, or screenshots — no Chromium, no API key, no setup.

97.0%
Merge Rate
1d
First Review
33%
1st-Timers
3
Maintainers
B TierTypeScript 434

figranium/figranium

Build complex browser workflows visually and execute them via API.

95.2%
Merge Rate
3h
First Review
100%
1st-Timers
0
Maintainers
B TierPython 211

ItamarZand88/CLI-Anything-WEB

Claude Code plugin that generates production-grade Python CLIs for any web app. 20 CLIs and counting.

100.0%
Merge Rate
24h
First Review
100%
1st-Timers
1
Maintainers
B TierJavaScript 1.8k

microlinkhq/browserless

The headless Chrome/Chromium driver on top of Puppeteer. Take screenshots, generate PDFs, extract text and HTML with a production-ready API.

85.3%
Merge Rate
1d
First Review
33%
1st-Timers
1
Maintainers
A TierPython 30.1k

ScrapeGraphAI/Scrapegraph-ai

Python scraper based on AI

68.8%
Merge Rate
19h
First Review
100%
1st-Timers
3
Maintainers
A TierGo 10.1k

pinchtab/pinchtab

High-performance browser automation bridge and multi-instance orchestrator with advanced stealth injection and real-time dashboard.

86.7%
Merge Rate
11d
First Review
80%
1st-Timers
9
Maintainers
B TierJavaScript 246

only-cli/oc

Turn any website into a compact CLI tailored for AI agents. Browse the web in hundreds of tokens, not tens of thousands.

50.0%
Merge Rate
23h
First Review
0%
1st-Timers
0
Maintainers
B TierPython 6.7k

adbar/trafilatura

Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML

100.0%
Merge Rate
3d
First Review
0%
1st-Timers
1
Maintainers
B TierTypeScript 1.2k

intoli/user-agents

A JavaScript library for generating random user agents with data that's updated daily.

91.7%
Merge Rate
-
First Review
100%
1st-Timers
0
Maintainers
B TierZig 355

justrach/kuri

Browser automation, web crawling, and iOS + Android device control for AI agents. Zig-native, token-efficient CDP snapshots, HAR recording, native adb wire-protocol client, and a standalone fetcher.

71.0%
Merge Rate
7d
First Review
25%
1st-Timers
0
Maintainers
B TierPython 6.4k 1 GFIs

lexiforest/curl_cffi

Python binding for curl-impersonate fork via cffi. A http client that can impersonate browser tls/ja3/http2 fingerprints.

46.6%
Merge Rate
22h
First Review
38%
1st-Timers
25
Maintainers
B TierMAMakefile 8.1k

lorien/awesome-web-scraping

List of libraries, tools and APIs for web scraping and data processing.

62.5%
Merge Rate
2d
First Review
100%
1st-Timers
3
Maintainers
B TierPython 64.1k

scrapy/scrapy

Scrapy, a fast high-level web crawling & scraping framework for Python.

48.1%
Merge Rate
2d
First Review
17%
1st-Timers
16
Maintainers
B TierGo 2.1k

Anakin-Inc/anakin

Open-source web scraping API. Turn any website into clean markdown or structured JSON. Anti-detect browser, proxy auto-selection, self-hosted. One command: make up

26.2%
Merge Rate
1d
First Review
25%
1st-Timers
4
Maintainers
C TierPython 77.3k

D4Vinci/Scrapling

🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!

41.5%
Merge Rate
3d
First Review
41%
1st-Timers
13
Maintainers
C TierPython 3.3k

oxylabs/oxylabs-ai-studio-py

Structured data gathering from any website using AI-powered scraper, crawler, and browser automation. Scraping and crawling with natural language prompts. Equip your LLM agents with fresh data. AI Studio python SDK for intelligent web data gathering.

100.0%
Merge Rate
6d
First Review
0%
1st-Timers
1
Maintainers
D TierPython 340

Bin-Huang/camoufox-cli

Anti-detect browser automation CLI & Skills for AI agents — Camoufox-powered fingerprint spoofing, no bot-detectable Playwright leaks

0.0%
Merge Rate
-
First Review
0%
1st-Timers
1
Maintainers
D TierGo 7.1k

go-rod/rod

A Chrome DevTools Protocol driver for web automation and scraping.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierHTML 474

davidteather/everything-web-scraping

Learn everything web scraping with David Teather Codes on YouTube

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 234

jordantete/OddsHarvester

A python app designed to scrape and process sports betting data directly from oddsportal.com 🎯

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierHTML 116

Omarshraf/MangaVolt-Archive-Engine

Best MangaFox Downloader Script 2026: Batch Manga Grabber Tool

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierMulti-language 794

oxylabs/agent-skills

Official Agent skills of Oxylabs products

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierJavaScript 1.6k

gildas-lormeau/single-file-cli

CLI tool for saving a faithful copy of a complete web page in a single HTML file (based on SingleFile)

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierGo 5.6k

gosom/google-maps-scraper

scrape data from Google Maps. Extracts data such as the name, address, phone number, website URL, rating, reviews number, latitude and longitude, reviews,email and more for each place

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPHP 371

crwlrsoft/crawler

Library for Rapid (Web) Crawler and Scraper Development

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierJava 486

VIDA-NYU/ache

ACHE is a web crawler for domain-specific search.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 792

z0m31en7/Uscrapper

Uscrapper Vanta: Dive deeper into the web with this powerful open-source tool. Extract valuable insights with ease and efficiency, from both surface and deep web sources. Empower your data mining and analysis with Vanta's advanced capabilities. Fast, reliable, and user-friendly, Uscrapper Vanta is the ultimate choice for researchers and analysts.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierTypeScript 326

oxylabs/oxylabs-ai-studio-js

Structured data gathering from any website using AI-powered scraper, crawler, and browser automation. Scraping and crawling with natural language prompts. Equip your LLM agents with fresh data. AI Studio JS SDK for intelligent web data gathering.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierC 414

serpapi/nokolexbor

High-performance HTML5 parser for Ruby based on Lexbor, with support for both CSS selectors and XPath.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierHTML 378

City-Bureau/city-scrapers

Scrape, standardize and share public meetings from local government websites

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
Best Web-scraping Open Source Repositories & C-Rank™ | GetMerged