#web-crawler (15 Repositories)
Ranked open-source repositories tagged with #web-crawler, scored by pull request acceptance likelihood and maintainer engagement velocity.
26.0%
19.0h
15 repositories tagged #web-crawler
internetarchive/Zeno
State-of-the-art web crawler 🔱
ScrapeGraphAI/Scrapegraph-ai
Python scraper based on AI
lexmount/moli
Best headless browser for AI agents. Lite, Fast, High-Compatibility. Built in Rust
MarginaliaSearch/MarginaliaSearch
Internet search engine for text-oriented websites. Indexing the small, old and weird web.
firecrawl/firecrawl
The context API to search, scrape, and interact with the web at scale. 🔥
Anakin-Inc/anakin
Open-source web scraping API. Turn any website into clean markdown or structured JSON. Anti-detect browser, proxy auto-selection, self-hosted. One command: make up
webrecorder/browsertrix-crawler
Run a high-fidelity browser-based web archiving crawler in a single Docker container
infinilabs/crawler
🕷️ An easy-to-use spider written in Golang. (previous named GOPA.)
s0rg/crawley
The unix-way web crawler
VIDA-NYU/ache
ACHE is a web crawler for domain-specific search.
lewisdonovan/google-news-scraper
Lightweight scraper for Google News
crwlrsoft/crawler
Library for Rapid (Web) Crawler and Scraper Development
Norconex/crawler
Norconex Crawlers (or spiders) are flexible web and filesystem crawlers for collecting, parsing, and manipulating data from the web or filesystem to various data repositories such as search engines.
gildas-lormeau/single-file-cli
CLI tool for saving a faithful copy of a complete web page in a single HTML file (based on SingleFile)
crawlab-team/crawlab-lite
Lite version of Crawlab. 轻量版 Crawlab 爬虫管理平台