#crawler (30 Repositories)
Ranked open-source repositories tagged with #crawler, scored by pull request acceptance likelihood and maintainer engagement velocity.
59.8%
82.4h
30 repositories tagged #crawler
eliashaeussler/cache-warmup
🔥 PHP library to warm up caches of URLs located in XML sitemaps
KamiYomu/KamiYomu
A self-hosted, extensible manga reader and download tool with plug-in support.
BlessedRebuS/Krawl
Krawl is a customizable, lightweight, cloud-native web deception server and anti-crawler that creates fake web applications with low-hanging vulnerabilities using realistic, randomly generated decoy data and AI-generated HTML templates.
JayBizzle/Crawler-Detect
🕷 CrawlerDetect is a PHP class for detecting bots/crawlers/spiders via the user agent
us/crw
Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scrape, /crawl, /search). 2.3x faster than Tavily, 1.5x faster than Firecrawl in 1K-URL benchmarks. 6 MB RAM, single binary. Self-host or use managed cloud.
codelibs/fess
Open-source, self-hosted enterprise & site search server built on OpenSearch. Crawls web / file / DB / cloud sources, 20+ languages, REST API, and AI/RAG & semantic search. Apache-2.0.
dadoonet/fscrawler
Elasticsearch File System Crawler (FS Crawler)
TeamNewPipe/NewPipeExtractor
NewPipe's core library for extracting data from streaming sites
z-mio/ParseHub
轻量、异步、开箱即用的社交媒体聚合解析库
ScrapeGraphAI/Scrapegraph-ai
Python scraper based on AI
apify/crawlee-python
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
kameleo-io/kameleo
Anti-detect browser for web scraping and automation. Engine-level fingerprint masking for Chromium and Firefox. Self-hosted, Docker-ready. Integrates with Selenium, Playwright, and Puppeteer via SDKs in Python, JavaScript, and C#.
hardkoded/puppeteer-sharp
Headless Chrome .NET API
microlinkhq/top-user-agents
Always up-to-date list of the top 100 most common browser user-agents for HTTP clients.
adbar/trafilatura
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
NaiboWang/EasySpider
A visual no-code/code-free web crawler/spider易采集:一个可视化浏览器自动化测试/数据采集/网页爬虫软件,可以无代码图形化的设计和执行爬虫任务。别名:ServiceWrapper面向Web应用的智能化服务封装系统。
firecrawl/firecrawl
The context API to search, scrape, and interact with the web at scale. 🔥
scrapy/scrapy
Scrapy, a fast high-level web crawling & scraping framework for Python.
MarshalX/telegram-crawler
🕷 Automatically detects changes to the official Telegram sites, beta clients, MTProto servers and mini apps
TeamWiseFlow/xiaobei
为OPC/中小微企业量身打造的自媒体获客智能体
D4Vinci/Scrapling
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
solarhell/font_obfuscator
字体混淆服务
fredwu/crawler
A high performance web crawler / scraper in Elixir.
iawia002/lux
👾 Fast and simple video download library and CLI tool written in Go
jowelrana120/hentai-daily-crawler
Best Multi-Source Hentai Daily Aggregator 2026
seveniruby/AppCrawler
基于appium的app自动遍历工具
guyueyingmu/avbook
AV 电影管理系统, avmoo , javbus , javlibrary 爬虫,线上 AV 影片图书馆,AV 磁力链接数据库,Japanese Adult Video Library,Adult Video Magnet Links - Japanese Adult Video Database
zorlan/skycaiji
蓝天采集器是一款开源免费的爬虫系统,仅需点选编辑规则即可采集数据,可运行在本地、虚拟主机或云服务器中,几乎能采集所有类型的网页,无缝对接各类CMS建站程序,免登录实时发布数据,全自动无需人工干预!是网页大数据采集软件中完全跨平台的云端爬虫系统
xiyuan-fengyu/ppspider
web spider built by puppeteer, support task-queue and task-scheduling by decorators,support nedb / mongodb, support data visualization; 基于puppeteer的web爬虫框架,提供灵活的任务队列管理调度方案,提供便捷的数据保存方案(nedb/mongodb),提供数据可视化和用户交互的实现方案
Autumn-27/ScopeSentry
ScopeSentry-Cyberspace mapping, subdomain enumeration, port scanning, sensitive information discovery, vulnerability scanning, distributed nodes