Back to Topics Directory
Topic Hub

#scraping (30 Repositories)

Ranked open-source repositories tagged with #scraping, scored by pull request acceptance likelihood and maintainer engagement velocity.

Topic Avg Merge Rate

50.9%

Avg Review Latency

52.7h

Filter by language

30 repositories tagged #scraping

S TierRust 1.2k 2 GFIs

plabayo/rama

modular service framework to move and transform network packets

94.5%
Merge Rate
16h
First Review
57%
1st-Timers
5
Maintainers
B TierDODockerfile 619

umihico/docker-selenium-lambda

The simplest demo of chrome automation by python and selenium in AWS Lambda

100.0%
Merge Rate
-
First Review
100%
1st-Timers
0
Maintainers
B TierTypeScript 372

sqdshguy/wreq-js

HTTP client for Node.js with browser TLS fingerprint impersonation

87.1%
Merge Rate
5h
First Review
33%
1st-Timers
1
Maintainers
A TierPython 176

apify/apify-sdk-python

Apify SDK for Python—The official library for building Apify Actors: serverless cloud programs for web scraping, browser automation, data processing, and AI agents. Manages the Actor lifecycle, storages (datasets, key-value stores, request queues), events, proxies, and pay-per-event monetization. Built on top of the the Apify API Client.

88.3%
Merge Rate
2d
First Review
100%
1st-Timers
6
Maintainers
B TierPython 1.1k

soxoj/socid-extractor

⛏️ The extraction engine behind Maigret: turn any profile URL into a structured OSINT record across 150+ sites

100.0%
Merge Rate
11h
First Review
100%
1st-Timers
1
Maintainers
B TierPython 2.0k

feder-cr/invisible_playwright

Free antidetect browser stealth for Playwright: undetected headless Firefox fingerprint. Python scraping, recaptcha and bot detection bypass. Open source

88.7%
Merge Rate
2d
First Review
50%
1st-Timers
0
Maintainers
B TierTypeScript 2.6k

botswin/BotBrowser

Advanced Privacy Browser Core with Unified Fingerprint Defense: Cloudflare, Akamai, Kasada, Shape, DataDome, PerimeterX, hCaptcha, FunCaptcha, Imperva, reCAPTCHA, ThreatMetrix, Adscore

94.7%
Merge Rate
<1h
First Review
0%
1st-Timers
1
Maintainers
B TierPython 610

AnonCatalyst/Ominis-OSINT

This Python application is an OSINT (Open Source Intelligence) tool called "Ominis OSINT - Web Hunter." It performs online information gathering by querying Google for search results related to a user-inputted query. The tool extracts relevant information such as titles, URLs, and potential mentions of the query in the results.

100.0%
Merge Rate
<1h
First Review
100%
1st-Timers
1
Maintainers
B TierPython 139

masterFuf/taktik-bot

Instagram & TikTok automation via real Android devices. Likes, follows, DMs, scraping. No API abuse. Built with Python, uiautomator2 & ADB.

100.0%
Merge Rate
-
First Review
100%
1st-Timers
0
Maintainers
A TierPython 30.1k

ScrapeGraphAI/Scrapegraph-ai

Python scraper based on AI

68.8%
Merge Rate
19h
First Review
100%
1st-Timers
3
Maintainers
B TierC# 183

kameleo-io/kameleo

Anti-detect browser for web scraping and automation. Engine-level fingerprint masking for Chromium and Firefox. Self-hosted, Docker-ready. Integrates with Selenium, Playwright, and Puppeteer via SDKs in Python, JavaScript, and C#.

87.2%
Merge Rate
19d
First Review
33%
1st-Timers
0
Maintainers
B TierPHP 207

playwright-php/playwright

Playwright PHP library for browser automation: navigation, E2E tests, assertions, screenshots, and so much more!

89.6%
Merge Rate
6d
First Review
75%
1st-Timers
2
Maintainers
B TierJavaScript 363

microlinkhq/top-user-agents

Always up-to-date list of the top 100 most common browser user-agents for HTTP clients.

100.0%
Merge Rate
<1h
First Review
0%
1st-Timers
0
Maintainers
B TierC++ 11.3k 2 GFIs

daijro/camoufox

🦊 Anti-detect browser

44.2%
Merge Rate
7h
First Review
20%
1st-Timers
42
Maintainers
B TierTypeScript 174.2k

firecrawl/firecrawl

The context API to search, scrape, and interact with the web at scale. 🔥

44.3%
Merge Rate
3d
First Review
26%
1st-Timers
41
Maintainers
B TierMAMakefile 8.1k

lorien/awesome-web-scraping

List of libraries, tools and APIs for web scraping and data processing.

62.5%
Merge Rate
2d
First Review
100%
1st-Timers
3
Maintainers
B TierPython 64.1k

scrapy/scrapy

Scrapy, a fast high-level web crawling & scraping framework for Python.

48.1%
Merge Rate
2d
First Review
17%
1st-Timers
16
Maintainers
B TierRuby 4.4k

sparklemotion/mechanize

Mechanize is a ruby library that makes automated web interaction easy.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
1
Maintainers
C TierTypeScript 100

daniel-hauser/moneyman

Automatically save transactions from all major Israeli banks and credit card companies, using GitHub actions (or a self hosted docker image)

38.2%
Merge Rate
11d
First Review
40%
1st-Timers
1
Maintainers
C TierPython 77.3k

D4Vinci/Scrapling

🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!

41.5%
Merge Rate
3d
First Review
41%
1st-Timers
13
Maintainers
C TierPython 144

cumbucadev/cinemaempoa

Site que agrega filmes em cartaz em algumas das diversas salas de cinema de Porto Alegre.

50.0%
Merge Rate
13d
First Review
0%
1st-Timers
1
Maintainers
D TierPython 30.3k

feder-cr/Jobs_Applier_AI_Agent_AIHawk

Open source AI job application bot in Python: browser automation and web scraping to read job postings, then auto-apply with a tailored resume and cover letter for each posting.

0.0%
Merge Rate
<1h
First Review
0%
1st-Timers
0
Maintainers
D TierPython 448

danieldotnl/ha-multiscrape

Home Assistant custom component for scraping (html, xml or json) multiple values (from a single HTTP request) with a separate sensor/attribute for each value. Support for (login) form-submit functionality.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 607

l4rm4nd/LinkedInDumper

Python 3 script to dump/scrape/extract company employees from LinkedIn API

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPython 607

Yuvi9587/Kemono-Downloader

Kemono pawchive Downloader is a fast, PyQt5 app for archiving content from a wide array of sites, including Kemono, Coomer, Bunkr, Erome, Saint2.su, nhentai, and Discord. It features a Creator Browser, update checker, and supports multi-threaded, multi-part downloads. Sessions can pause, resume, or recover. filters

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierOCOCaml 409

aantron/lambdasoup

Functional HTML scraping and rewriting with CSS in OCaml

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierElixir 336

Anonyfox/elixir-scrape

Scrape any website, article or RSS/Atom Feed with ease!

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierTypeScript 2.6k

apify/fingerprint-suite

Browser fingerprinting tools for anonymizing your scrapers. Developed by Apify.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierGo 4.8k

bjesus/pipet

Swiss-army tool for scraping and extracting data from online assets, made for hackers

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierMulti-language 6.6k

proxifly/free-proxy-list

🚀 Free HTTP, SOCKS4, & SOCKS5 proxy list * Updated every 5 minutes * and rotating proxy API (100+ countries)

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
Best Scraping Open Source Repositories & C-Rank™ | GetMerged