Distributed crawling infrastructure running on top of severless computation, cloud storage (such as S3) and sophisticated queues.
☆438Dec 30, 2022Updated 3 years ago
Alternatives and similar repositories for Crawling-Infrastructure
Users that are interested in Crawling-Infrastructure are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Cloud crawler functions for scrapeulous☆44Feb 24, 2021Updated 5 years ago
- Javascript scraping module based on puppeteer for many different search engines...☆570Dec 30, 2022Updated 3 years ago
- Module that extracts structured information from a rendered html site and outputs JSON. HTML to JSON.☆70Jun 8, 2021Updated 5 years ago
- A Python module to scrape several search engines (like Google, Yandex, Bing, Duckduckgo, ...). Including asynchronous networking support.☆2,882Jul 3, 2021Updated 5 years ago
- ☆12May 7, 2023Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Scrapoxy has been discontinued.☆2,411Feb 7, 2026Updated 7 months ago
- In-Memory Key-Value Database with Persistent File Storage☆16Sep 24, 2022Updated 3 years ago
- 📡 Renew the IP address of a tethered Android device via Node asynchronously.☆77Aug 3, 2023Updated 3 years ago
- 💯 Teach puppeteer new tricks through plugins.☆7,399Jul 18, 2024Updated 2 years ago
- Search google, bing, yahoo, and other search engines with python☆671Apr 2, 2025Updated last year
- Analysis of Bot Protection systems with available countermeasures 🚿. How to defeat anti-bot system 👻 and get around browser fingerprint…☆5,138Jul 27, 2026Updated last month
- Solution to stop sites from fingerprinting your puppeteer☆132Jun 12, 2026Updated 3 months ago
- ☆174Dec 30, 2022Updated 3 years ago
- The web scraper that's nearly impossible to block - now called @ulixee/hero☆736Mar 7, 2023Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆20Apr 21, 2020Updated 6 years ago
- POC code to crash Windows Event Logger Service☆27Oct 16, 2020Updated 5 years ago
- Repo for hosting various scripts for creating users for password spraying and other password attacks.☆11Jul 9, 2020Updated 6 years ago
- #️⃣ 🕸️ 👤 HTTP Headers Hashing☆12Aug 27, 2023Updated 3 years ago
- Distributed crawler powered by Headless Chrome☆5,637Apr 29, 2023Updated 3 years ago
- Microsoft Applocker evasion tool☆39Nov 26, 2019Updated 6 years ago
- Event Data Collector☆40Mar 23, 2026Updated 5 months ago
- Puppeteer Pool, run a cluster of instances in parallel☆3,513Mar 1, 2026Updated 6 months ago
- A web crawler. Supercrawler automatically crawls websites. Define custom handlers to parse content. Obeys robots.txt, rate limits and con…☆382Dec 30, 2022Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Site Hound (previously THH) is a Domain Discovery Tool☆24Apr 8, 2026Updated 5 months ago
- 🛡🎭 A conceptual patch which modifies some vanilla puppeteer files to decrease detection rates.☆56Mar 6, 2021Updated 5 years ago
- ☆10Dec 18, 2018Updated 7 years ago
- C# application that allows you to quick run SSH commands against a host or list of hosts☆42Sep 21, 2020Updated 5 years ago
- BH Cypher Queries picked up from random places☆41Dec 12, 2018Updated 7 years ago
- Nodejs lib to parse Google SERP html pages☆48Jul 27, 2023Updated 3 years ago
- Chromium Binary for AWS Lambda and Google Cloud Functions☆3,285Sep 3, 2024Updated 2 years ago
- A JavaScript library for generating random user agents with data that's updated daily.☆1,190Updated this week
- ☆118Mar 16, 2024Updated 2 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data …☆25,797Updated this week
- Scrapy extension that gives you all the scraping monitoring, alerting, scheduling, and data validation you will need straight out of the…☆38Apr 23, 2026Updated 4 months ago
- A simple, quick, and dirty websocket shell for PowerShell.☆20Jun 5, 2017Updated 9 years ago
- Article extraction benchmark: dataset and evaluation scripts☆378May 29, 2026Updated 3 months ago
- Streaming web crawler with WebSocket API☆48May 15, 2026Updated 4 months ago
- A test suite of common scraper detection techniques. See how detectable your scraper stack is.☆139Oct 31, 2022Updated 3 years ago
- BloodHound Cypher Queries Ported to a Jupyter Notebook☆53Jun 20, 2020Updated 6 years ago