Modern robots.txt Parser for Python
☆196Jan 12, 2024Updated 2 years ago
Alternatives and similar repositories for reppy
Users that are interested in reppy are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- URL Transformation, Sanitization☆103Jan 16, 2024Updated 2 years ago
- Robot exclusion protocol in C++☆11Jul 26, 2024Updated 2 years ago
- mltk - Moz Language Tool Kit☆12Mar 6, 2015Updated 11 years ago
- Alternative robots parser module for Python☆22Jun 19, 2026Updated 2 months ago
- Python API for Various DB-Backed Simhash Clusters☆64Mar 16, 2017Updated 9 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Extract embedded metadata from HTML markup☆971Apr 1, 2026Updated 4 months ago
- A fast URI parser that wraps Google's chromium URL canonicalization library☆16Aug 11, 2026Updated 2 weeks ago
- C++ bindings for url parsing and sanitization☆19May 2, 2024Updated 2 years ago
- A pure-Python robots.txt parser with support for modern conventions.☆93Aug 25, 2026Updated last week
- Site Hound (previously THH) is a Domain Discovery Tool☆24Apr 8, 2026Updated 4 months ago
- Ultimate Website Sitemap Parser☆256Jun 16, 2026Updated 2 months ago
- Python participant support for MsgFlo☆14May 23, 2020Updated 6 years ago
- Scrapy extension which writes crawled items to Kafka☆31Apr 8, 2026Updated 4 months ago
- Random Bingo Sheet for DB delays☆17Oct 3, 2024Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Just the facts -- web page content extraction☆1,274Jul 8, 2025Updated last year
- Accurately separates a URL’s subdomain, domain, and public suffix, using the Public Suffix List (PSL).☆2,015Aug 8, 2026Updated 3 weeks ago
- Fast multi-keyword search engine for text strings☆257Sep 14, 2024Updated last year
- Modularly extensible semantic metadata validator☆85Dec 10, 2015Updated 10 years ago
- Nginx configuration☆20Sep 17, 2025Updated 11 months ago
- Simhash and near-duplicate detection☆422May 15, 2023Updated 3 years ago
- Extract countries, regions and cities from a URL or text☆216Sep 10, 2020Updated 5 years ago
- Simple heuristic for measuring web page similarity (& data set)☆91Apr 8, 2026Updated 4 months ago
- This project deals with hierarchical classification of web pages based on dmoz dataset.☆14Apr 10, 2014Updated 12 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Training/test data for Dragnet☆42Jan 29, 2015Updated 11 years ago
- A plugin that highlights operator characters for every language.☆21Feb 8, 2015Updated 11 years ago
- Force-Atlas 2 graph layout in networkx☆22Sep 30, 2014Updated 11 years ago
- Extract city and country mentions from Text like GeoText without regex, but FlashText, a Aho-Corasick implementation.☆63Aug 24, 2026Updated last week
- URL normalization for Python☆100Apr 25, 2026Updated 4 months ago
- Collects multimedia content shared through social networks.☆19Feb 18, 2015Updated 11 years ago
- Decentralized DNS fuzzer to mitigate ISP Snooping☆12May 3, 2017Updated 9 years ago
- A scalable frontier for web crawlers☆1,332Jun 6, 2025Updated last year
- Event signalling for python and asyncio☆20Nov 12, 2015Updated 10 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A simple package allowing to use WebGraph data in Python (via the Jython interpreter).☆20Oct 21, 2020Updated 5 years ago
- Extract data from websites using basic statistical magic☆503Oct 2, 2020Updated 5 years ago
- Python module for Named Entity Recognition (NER) using natural language processing.☆13May 30, 2021Updated 5 years ago
- Calibre plugin. Searches for books metadata on bookradar.org☆10Nov 10, 2020Updated 5 years ago
- Data science tools from Moz☆23Jan 11, 2017Updated 9 years ago
- Front-end for the MediaCloud database☆16Apr 3, 2018Updated 8 years ago
- Python 3 AsyncIO powered scraping framework with batteries included☆20Sep 8, 2016Updated 9 years ago