Wget-compatible web downloader and crawler.
☆614Apr 29, 2024Updated 2 years ago
Alternatives and similar repositories for wpull
Users that are interested in wpull are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The archivist's web crawler: WARC output, dashboard for all crawls, dynamic ignore patterns☆1,612May 23, 2025Updated last year
- ArchiveBot, an IRC bot for archiving websites☆422Sep 8, 2026Updated last week
- Grabbing all news.☆60Dec 23, 2019Updated 6 years ago
- WARC writing MITM HTTP/S proxy☆467Jun 17, 2026Updated 3 months ago
- Tool and library for handling Web ARChive (WARC) files.☆166Oct 11, 2024Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Core Python Web Archiving Toolkit for replay and recording of web archives☆1,700Sep 11, 2026Updated last week
- brozzler - distributed browser-based web crawler☆820Aug 20, 2026Updated last month
- Collect and revisit web pages.☆1,549Jul 22, 2026Updated last month
- Wget-AT is a modern Wget with Lua hooks, Zstandard (+dictionary) WARC compression and URL-agnostic deduplication.☆138Mar 19, 2026Updated 6 months ago
- Webrecorder Player for Desktop (OSX/Windows/Linux). (Built with Electron + Webrecorder)☆445Sep 17, 2020Updated 6 years ago
- Archiving public telegram messages.☆18Aug 12, 2026Updated last month
- An Awesome List for getting started with web archiving☆2,648Updated this week
- WarcMiddleware lets users seamlessly download a mirror copy of a website when running a web crawl with the Python web crawler Scrapy.☆48Mar 19, 2018Updated 8 years ago
- The OpenWayback Development☆529Jan 3, 2024Updated 2 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Web archiving using Google Chrome☆44Dec 30, 2019Updated 6 years ago
- Convert Directories, Files and ZIP Files to Web Archives (WARC)☆102Aug 28, 2026Updated 3 weeks ago
- Serving content from a WARC☆61Jan 5, 2013Updated 13 years ago
- Nondestructive warc-in-tar to warc conversion☆27Apr 21, 2013Updated 13 years ago
- NOTE: This project is no longer being actively developed.. Check out https://replayweb.page / https://github.com/webrecorder/replayweb.pa…☆203Jan 22, 2025Updated last year
- Streaming WARC/ARC library for fast web archive IO☆474Jun 10, 2026Updated 3 months ago
- Making a reusable toolkit for writing seesaw scripts☆78Jul 17, 2026Updated 2 months ago
- Parse WARC (Web Archive Files) as a node.js stream☆23Oct 20, 2014Updated 11 years ago
- A Python and Command-Line Interface to Archive.org☆1,907Updated this week
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Archiving GitHub☆11Aug 5, 2025Updated last year
- Use yt-dlp to download video/metadata and upload to the Internet Archive.☆518Aug 12, 2026Updated last month
- Convert HTTP Archive (HAR) -> Web Archive (WARC) format☆55Oct 21, 2018Updated 7 years ago
- Chrome extension to "Create WARC files from any webpage"☆232Aug 21, 2026Updated 3 weeks ago
- Web Archiving Integration Layer: One-Click User Instigated Preservation☆401Jun 19, 2026Updated 3 months ago
- Run a high-fidelity browser-based web archiving crawler in a single Docker container☆1,143Updated this week
- We back up a lot of stuff from around the web; now it's time to back up the Internet Archive, just in case.☆92Jul 13, 2020Updated 6 years ago
- Command line tools and libraries for handling and manipulating WARC files (and HTTP contents)☆178Aug 18, 2025Updated last year
- Archiving Google+.☆26Apr 4, 2019Updated 7 years ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- A Tool To Push Web Resources Into Web Archives☆436Jan 23, 2024Updated 2 years ago
- A dockerized, queued high fidelity web archiver based on Squidwarc☆62Jul 9, 2024Updated 2 years ago
- The little things give you away... A collection of various small helper stuff – Mirror repo only, no longer kept in sync, refer to gitea.…☆24Sep 11, 2020Updated 6 years ago
- Heritrix is the Internet Archive's open-source, extensible, web-scale, archival-quality web crawler project.☆3,321Updated this week
- 🗃 Open source self-hosted web archiving. Takes URLs/browser history/bookmarks/Pocket/Pinboard/etc., saves HTML, JS, PDFs, media, and mor…☆28,548Updated this week
- Wget with Lua extension☆24Dec 17, 2015Updated 10 years ago
- Serverless replay of web archives directly in the browser☆985Sep 3, 2026Updated 2 weeks ago