Back to Topics Directory
Topic Hub

#crawler (30 Repositories)

Ranked open-source repositories tagged with #crawler, scored by pull request acceptance likelihood and maintainer engagement velocity.

Topic Avg Merge Rate

59.8%

Avg Review Latency

82.4h

Filter by language

30 repositories tagged #crawler

S TierPHP 77

eliashaeussler/cache-warmup

🔥 PHP library to warm up caches of URLs located in XML sitemaps

87.5%
Merge Rate
<1h
First Review
100%
1st-Timers
2
Maintainers
B TierC# 142 1 GFIs

KamiYomu/KamiYomu

A self-hosted, extensible manga reader and download tool with plug-in support.

97.1%
Merge Rate
2d
First Review
100%
1st-Timers
0
Maintainers
S TierPython 626 2 GFIs

BlessedRebuS/Krawl

Krawl is a customizable, lightweight, cloud-native web deception server and anti-crawler that creates fake web applications with low-hanging vulnerabilities using realistic, randomly generated decoy data and AI-generated HTML templates.

87.5%
Merge Rate
2h
First Review
50%
1st-Timers
2
Maintainers
B TierPHP 2.4k

JayBizzle/Crawler-Detect

🕷 CrawlerDetect is a PHP class for detecting bots/crawlers/spiders via the user agent

100.0%
Merge Rate
<1h
First Review
100%
1st-Timers
0
Maintainers
A TierRust 741

us/crw

Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scrape, /crawl, /search). 2.3x faster than Tavily, 1.5x faster than Firecrawl in 1K-URL benchmarks. 6 MB RAM, single binary. Self-host or use managed cloud.

94.6%
Merge Rate
3d
First Review
75%
1st-Timers
4
Maintainers
B TierJava 1.1k

codelibs/fess

Open-source, self-hosted enterprise & site search server built on OpenSearch. Crawls web / file / DB / cloud sources, 20+ languages, REST API, and AI/RAG & semantic search. Apache-2.0.

92.0%
Merge Rate
3d
First Review
100%
1st-Timers
0
Maintainers
A TierJava 1.4k

dadoonet/fscrawler

Elasticsearch File System Crawler (FS Crawler)

80.8%
Merge Rate
6h
First Review
50%
1st-Timers
3
Maintainers
A TierJava 2.0k 1 GFIs

TeamNewPipe/NewPipeExtractor

NewPipe's core library for extracting data from streaming sites

85.7%
Merge Rate
9h
First Review
100%
1st-Timers
4
Maintainers
A TierPython 148

z-mio/ParseHub

轻量、异步、开箱即用的社交媒体聚合解析库

100.0%
Merge Rate
<1h
First Review
100%
1st-Timers
2
Maintainers
A TierPython 30.1k

ScrapeGraphAI/Scrapegraph-ai

Python scraper based on AI

68.8%
Merge Rate
19h
First Review
100%
1st-Timers
3
Maintainers
A TierPython 9.5k

apify/crawlee-python

Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.

78.9%
Merge Rate
3d
First Review
63%
1st-Timers
19
Maintainers
B TierC# 183

kameleo-io/kameleo

Anti-detect browser for web scraping and automation. Engine-level fingerprint masking for Chromium and Firefox. Self-hosted, Docker-ready. Integrates with Selenium, Playwright, and Puppeteer via SDKs in Python, JavaScript, and C#.

87.2%
Merge Rate
19d
First Review
33%
1st-Timers
0
Maintainers
B TierC# 3.9k 1 GFIs

hardkoded/puppeteer-sharp

Headless Chrome .NET API

82.8%
Merge Rate
4d
First Review
100%
1st-Timers
2
Maintainers
B TierJavaScript 363

microlinkhq/top-user-agents

Always up-to-date list of the top 100 most common browser user-agents for HTTP clients.

100.0%
Merge Rate
<1h
First Review
0%
1st-Timers
0
Maintainers
B TierPython 6.7k

adbar/trafilatura

Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML

100.0%
Merge Rate
3d
First Review
0%
1st-Timers
1
Maintainers
B TierJavaScript 44.5k

NaiboWang/EasySpider

A visual no-code/code-free web crawler/spider易采集:一个可视化浏览器自动化测试/数据采集/网页爬虫软件,可以无代码图形化的设计和执行爬虫任务。别名:ServiceWrapper面向Web应用的智能化服务封装系统。

66.7%
Merge Rate
9h
First Review
0%
1st-Timers
1
Maintainers
B TierTypeScript 174.2k

firecrawl/firecrawl

The context API to search, scrape, and interact with the web at scale. 🔥

44.3%
Merge Rate
3d
First Review
26%
1st-Timers
41
Maintainers
B TierPython 64.1k

scrapy/scrapy

Scrapy, a fast high-level web crawling & scraping framework for Python.

48.1%
Merge Rate
2d
First Review
17%
1st-Timers
16
Maintainers
B TierPython 357

MarshalX/telegram-crawler

🕷 Automatically detects changes to the official Telegram sites, beta clients, MTProto servers and mini apps

100.0%
Merge Rate
44d
First Review
0%
1st-Timers
1
Maintainers
B TierPython 8.4k

TeamWiseFlow/xiaobei

为OPC/中小微企业量身打造的自媒体获客智能体

50.0%
Merge Rate
<1h
First Review
0%
1st-Timers
1
Maintainers
C TierPython 77.3k

D4Vinci/Scrapling

🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!

41.5%
Merge Rate
3d
First Review
41%
1st-Timers
13
Maintainers
C TierRust 173

solarhell/font_obfuscator

字体混淆服务

100.0%
Merge Rate
-
First Review
100%
1st-Timers
0
Maintainers
D TierElixir 956

fredwu/crawler

A high performance web crawler / scraper in Elixir.

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierGo 31.7k

iawia002/lux

👾 Fast and simple video download library and CLI tool written in Go

0.0%
Merge Rate
12d
First Review
0%
1st-Timers
2
Maintainers
D TierHTML 117

jowelrana120/hentai-daily-crawler

Best Multi-Source Hentai Daily Aggregator 2026

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierScala 1.2k

seveniruby/AppCrawler

基于appium的app自动遍历工具

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPHP 10.0k

guyueyingmu/avbook

AV 电影管理系统, avmoo , javbus , javlibrary 爬虫,线上 AV 影片图书馆,AV 磁力链接数据库,Japanese Adult Video Library,Adult Video Magnet Links - Japanese Adult Video Database

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierPHP 2.1k

zorlan/skycaiji

蓝天采集器是一款开源免费的爬虫系统,仅需点选编辑规则即可采集数据,可运行在本地、虚拟主机或云服务器中,几乎能采集所有类型的网页,无缝对接各类CMS建站程序,免登录实时发布数据,全自动无需人工干预!是网页大数据采集软件中完全跨平台的云端爬虫系统

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierTypeScript 338

xiyuan-fengyu/ppspider

web spider built by puppeteer, support task-queue and task-scheduling by decorators,support nedb / mongodb, support data visualization; 基于puppeteer的web爬虫框架,提供灵活的任务队列管理调度方案,提供便捷的数据保存方案(nedb/mongodb),提供数据可视化和用户交互的实现方案

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
D TierGo 1.6k

Autumn-27/ScopeSentry

ScopeSentry-Cyberspace mapping, subdomain enumeration, port scanning, sensitive information discovery, vulnerability scanning, distributed nodes

0.0%
Merge Rate
-
First Review
0%
1st-Timers
0
Maintainers
Best Crawler Open Source Repositories & C-Rank™ | GetMerged