Modern robots.txt Parser for Python
☆195Jan 12, 2024Updated 2 years ago
Alternatives and similar repositories for reppy
Users that are interested in reppy are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- URL Transformation, Sanitization☆104Jan 16, 2024Updated 2 years ago
- python library for extracting html microdata☆168May 8, 2023Updated 3 years ago
- Python API for Various DB-Backed Simhash Clusters☆64Mar 16, 2017Updated 9 years ago
- Extract embedded metadata from HTML markup☆966Apr 1, 2026Updated 3 months ago
- A fast URI parser that wraps Google's chromium URL canonicalization library☆16Oct 25, 2023Updated 2 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- A pure-Python robots.txt parser with support for modern conventions.☆90Updated this week
- A framework for rapidly building polymer apps.☆12Nov 19, 2015Updated 10 years ago
- Site Hound (previously THH) is a Domain Discovery Tool☆24Apr 8, 2026Updated 3 months ago
- ☆16Sep 13, 2016Updated 9 years ago
- 🖨️ Printer: Productivity Focused Next.js CLI Tool☆11Nov 24, 2023Updated 2 years ago
- 🖥 LinkedData based Applications generator☆19Updated this week
- Ultimate Website Sitemap Parser☆255Jun 16, 2026Updated last month
- Python participant support for MsgFlo☆14May 23, 2020Updated 6 years ago
- Tagging and annotation framework for scan data☆100Oct 16, 2018Updated 7 years ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Scrapy extension which writes crawled items to Kafka☆31Apr 8, 2026Updated 3 months ago
- Just the facts -- web page content extraction☆1,274Jul 8, 2025Updated last year
- A Corpus Data Retrieval Index using Lucene for Look-Ups☆20Jul 8, 2026Updated last week
- Python Bindings for qless☆47Sep 23, 2019Updated 6 years ago
- Accurately separates a URL’s subdomain, domain, and public suffix, using the Public Suffix List (PSL).☆2,011Apr 21, 2026Updated 3 months ago
- Bash flauvored lodash port☆11May 1, 2016Updated 10 years ago
- Experimental Distributed Web Crawling with Python + Gearman☆22May 2, 2012Updated 14 years ago
- Fast multi-keyword search engine for text strings☆258Sep 14, 2024Updated last year
- Modularly extensible semantic metadata validator☆85Dec 10, 2015Updated 10 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Tool to create image datasets for machine learning problems by scraping search engines like Google, Bing and Baidu.☆17Apr 20, 2019Updated 7 years ago
- Simhash and near-duplicate detection☆422May 15, 2023Updated 3 years ago
- Extract countries, regions and cities from a URL or text☆216Sep 10, 2020Updated 5 years ago
- Simple heuristic for measuring web page similarity (& data set)☆91Apr 8, 2026Updated 3 months ago
- This project deals with hierarchical classification of web pages based on dmoz dataset.☆14Apr 10, 2014Updated 12 years ago
- Force-Atlas 2 graph layout in networkx☆22Sep 30, 2014Updated 11 years ago
- htcap is a web application scanner able to crawl single page application (SPA) in a recursive manner by intercepting ajax calls and DOM c…☆18Sep 23, 2025Updated 9 months ago
- Statefull widgets for django upload☆15Oct 3, 2016Updated 9 years ago
- [DEPRECATED] Unofficial Python Pandas DataReader objects with requests and requests_cache☆16Mar 5, 2018Updated 8 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- A scalable frontier for web crawlers☆1,332Jun 6, 2025Updated last year
- Event signalling for python and asyncio☆20Nov 12, 2015Updated 10 years ago
- A recommender system for GitHub repositories☆14Jun 21, 2014Updated 12 years ago
- A simple package allowing to use WebGraph data in Python (via the Jython interpreter).☆20Oct 21, 2020Updated 5 years ago
- Extract data from websites using basic statistical magic☆505Oct 2, 2020Updated 5 years ago
- Python module for Named Entity Recognition (NER) using natural language processing.☆13May 30, 2021Updated 5 years ago
- Calibre plugin. Searches for books metadata on bookradar.org☆10Nov 10, 2020Updated 5 years ago