#scraping (30 Repositories)
Ranked open-source repositories tagged with #scraping, scored by pull request acceptance likelihood and maintainer engagement velocity.
50.9%
52.7h
30 repositories tagged #scraping
plabayo/rama
modular service framework to move and transform network packets
umihico/docker-selenium-lambda
The simplest demo of chrome automation by python and selenium in AWS Lambda
sqdshguy/wreq-js
HTTP client for Node.js with browser TLS fingerprint impersonation
apify/apify-sdk-python
Apify SDK for Python—The official library for building Apify Actors: serverless cloud programs for web scraping, browser automation, data processing, and AI agents. Manages the Actor lifecycle, storages (datasets, key-value stores, request queues), events, proxies, and pay-per-event monetization. Built on top of the the Apify API Client.
soxoj/socid-extractor
⛏️ The extraction engine behind Maigret: turn any profile URL into a structured OSINT record across 150+ sites
feder-cr/invisible_playwright
Free antidetect browser stealth for Playwright: undetected headless Firefox fingerprint. Python scraping, recaptcha and bot detection bypass. Open source
botswin/BotBrowser
Advanced Privacy Browser Core with Unified Fingerprint Defense: Cloudflare, Akamai, Kasada, Shape, DataDome, PerimeterX, hCaptcha, FunCaptcha, Imperva, reCAPTCHA, ThreatMetrix, Adscore
AnonCatalyst/Ominis-OSINT
This Python application is an OSINT (Open Source Intelligence) tool called "Ominis OSINT - Web Hunter." It performs online information gathering by querying Google for search results related to a user-inputted query. The tool extracts relevant information such as titles, URLs, and potential mentions of the query in the results.
masterFuf/taktik-bot
Instagram & TikTok automation via real Android devices. Likes, follows, DMs, scraping. No API abuse. Built with Python, uiautomator2 & ADB.
ScrapeGraphAI/Scrapegraph-ai
Python scraper based on AI
kameleo-io/kameleo
Anti-detect browser for web scraping and automation. Engine-level fingerprint masking for Chromium and Firefox. Self-hosted, Docker-ready. Integrates with Selenium, Playwright, and Puppeteer via SDKs in Python, JavaScript, and C#.
playwright-php/playwright
Playwright PHP library for browser automation: navigation, E2E tests, assertions, screenshots, and so much more!
microlinkhq/top-user-agents
Always up-to-date list of the top 100 most common browser user-agents for HTTP clients.
daijro/camoufox
🦊 Anti-detect browser
firecrawl/firecrawl
The context API to search, scrape, and interact with the web at scale. 🔥
lorien/awesome-web-scraping
List of libraries, tools and APIs for web scraping and data processing.
scrapy/scrapy
Scrapy, a fast high-level web crawling & scraping framework for Python.
sparklemotion/mechanize
Mechanize is a ruby library that makes automated web interaction easy.
daniel-hauser/moneyman
Automatically save transactions from all major Israeli banks and credit card companies, using GitHub actions (or a self hosted docker image)
D4Vinci/Scrapling
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!
cumbucadev/cinemaempoa
Site que agrega filmes em cartaz em algumas das diversas salas de cinema de Porto Alegre.
feder-cr/Jobs_Applier_AI_Agent_AIHawk
Open source AI job application bot in Python: browser automation and web scraping to read job postings, then auto-apply with a tailored resume and cover letter for each posting.
danieldotnl/ha-multiscrape
Home Assistant custom component for scraping (html, xml or json) multiple values (from a single HTTP request) with a separate sensor/attribute for each value. Support for (login) form-submit functionality.
l4rm4nd/LinkedInDumper
Python 3 script to dump/scrape/extract company employees from LinkedIn API
Yuvi9587/Kemono-Downloader
Kemono pawchive Downloader is a fast, PyQt5 app for archiving content from a wide array of sites, including Kemono, Coomer, Bunkr, Erome, Saint2.su, nhentai, and Discord. It features a Creator Browser, update checker, and supports multi-threaded, multi-part downloads. Sessions can pause, resume, or recover. filters
aantron/lambdasoup
Functional HTML scraping and rewriting with CSS in OCaml
Anonyfox/elixir-scrape
Scrape any website, article or RSS/Atom Feed with ease!
apify/fingerprint-suite
Browser fingerprinting tools for anonymizing your scrapers. Developed by Apify.
bjesus/pipet
Swiss-army tool for scraping and extracting data from online assets, made for hackers
proxifly/free-proxy-list
🚀 Free HTTP, SOCKS4, & SOCKS5 proxy list * Updated every 5 minutes * and rotating proxy API (100+ countries)