Distributed crawling infrastructure running on top of severless computation, cloud storage (such as S3) and sophisticated queues.
☆437Dec 30, 2022Updated 3 years ago
Alternatives and similar repositories for Crawling-Infrastructure
Users that are interested in Crawling-Infrastructure are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Cloud crawler functions for scrapeulous☆44Feb 24, 2021Updated 5 years ago
- Javascript scraping module based on puppeteer for many different search engines...☆570Dec 30, 2022Updated 3 years ago
- Minimal set of tools to conduct stealthy scraping.☆166Apr 21, 2023Updated 3 years ago
- Module that extracts structured information from a rendered html site and outputs JSON. HTML to JSON.☆70Jun 8, 2021Updated 5 years ago
- A Python module to scrape several search engines (like Google, Yandex, Bing, Duckduckgo, ...). Including asynchronous networking support.☆2,869Jul 3, 2021Updated 5 years ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- This repository contains instructions how to use the free IP Address API. The databases are: ASN database, Geolocation database, hosting …☆117Updated this week
- Scrapoxy has been discontinued.☆2,414Feb 7, 2026Updated 5 months ago
- 📡 Renew the IP address of a tethered Android device via Node asynchronously.☆76Aug 3, 2023Updated 2 years ago
- 💯 Teach puppeteer new tricks through plugins.☆7,381Jul 18, 2024Updated 2 years ago
- Analysis of Bot Protection systems with available countermeasures 🚿. How to defeat anti-bot system 👻 and get around browser fingerprint…☆5,109May 12, 2026Updated 2 months ago
- Solution to stop sites from fingerprinting your puppeteer☆129Jun 12, 2026Updated last month
- ☆174Dec 30, 2022Updated 3 years ago
- The web scraper that's nearly impossible to block - now called @ulixee/hero☆737Mar 7, 2023Updated 3 years ago
- ☆20Apr 21, 2020Updated 6 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- POC code to crash Windows Event Logger Service☆27Oct 16, 2020Updated 5 years ago
- Repo for hosting various scripts for creating users for password spraying and other password attacks.☆11Jul 9, 2020Updated 6 years ago
- The BlogDB Webservice☆13Feb 1, 2022Updated 4 years ago
- Distributed crawler powered by Headless Chrome☆5,642Apr 29, 2023Updated 3 years ago
- Microsoft Applocker evasion tool☆39Nov 26, 2019Updated 6 years ago
- Docker kinsing malware bitcoin/xmr miner☆21Feb 18, 2021Updated 5 years ago
- Event Data Collector☆40Mar 23, 2026Updated 4 months ago
- Puppeteer Pool, run a cluster of instances in parallel☆3,517Mar 1, 2026Updated 4 months ago
- A web crawler. Supercrawler automatically crawls websites. Define custom handlers to parse content. Obeys robots.txt, rate limits and con…☆381Dec 30, 2022Updated 3 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- 🛡🎭 A conceptual patch which modifies some vanilla puppeteer files to decrease detection rates.☆56Mar 6, 2021Updated 5 years ago
- ☆11Dec 18, 2018Updated 7 years ago
- C# application that allows you to quick run SSH commands against a host or list of hosts☆42Sep 21, 2020Updated 5 years ago
- BH Cypher Queries picked up from random places☆41Dec 12, 2018Updated 7 years ago
- List of free and checked http, https, socks4 and socks5 proxies☆22Updated this week
- ☆117Mar 16, 2024Updated 2 years ago
- Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data …☆24,966Updated this week
- A simple, quick, and dirty websocket shell for PowerShell.☆20Jun 5, 2017Updated 9 years ago
- Article extraction benchmark: dataset and evaluation scripts☆376May 29, 2026Updated last month
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A test suite of common scraper detection techniques. See how detectable your scraper stack is.☆139Oct 31, 2022Updated 3 years ago
- How to detect puppeteer with 100% accuracy☆108May 30, 2021Updated 5 years ago
- Web application to visualize GreyNoise API data☆21Dec 4, 2018Updated 7 years ago
- This Scrapy project uses Redis and Kafka to create a distributed on demand scraping cluster.☆1,226Nov 7, 2023Updated 2 years ago
- OSINT Resources for Politics☆14Aug 13, 2018Updated 7 years ago
- Scrapy rotation proxy package with advanced functions☆94Jul 4, 2022Updated 4 years ago
- Additional module to use with 'puppeteer' for setting proxies per page basis.☆449Jun 9, 2024Updated 2 years ago