Scrapy pipeline to store chunked items into Amazon S3 or Google Cloud Storage bucket.
☆76Mar 18, 2022Updated 4 years ago
Alternatives and similar repositories for scrapy-s3pipeline
Users that are interested in scrapy-s3pipeline are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Sample of server-less crawler using AWS Fargate and Lambda☆12Dec 5, 2017Updated 8 years ago
- A library to make it easier to load input URLs to start scrapy processes☆14Feb 21, 2021Updated 5 years ago
- Scrapy schema validation pipeline and Item builder using JSON Schema☆45Mar 26, 2021Updated 5 years ago
- Creates a pipeline Airflow and Scrapy to output an average image composition of everyone's face in a given website☆43Oct 13, 2017Updated 8 years ago
- Bootable USB disk that lets you choose an ISO image☆16Oct 19, 2020Updated 5 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Library to populate items using XPath and CSS with a convenient API☆49Updated this week
- Utility functions for dbt projects running on Athena☆12Mar 25, 2025Updated last year
- Docker container running scrapyd with HTTP authentication☆41May 14, 2024Updated 2 years ago
- More flexible and featured Frontera scheduler for Scrapy☆36Jun 6, 2025Updated last year
- Step-by-step introduction to the traditional data warehousing with examples.☆11Mar 14, 2018Updated 8 years ago
- A simple component for infinite scroll☆20Apr 13, 2016Updated 10 years ago
- ☆16Apr 10, 2026Updated 3 months ago
- Performance-focused replacement for Python urllib☆21Apr 13, 2026Updated 3 months ago
- Library for annotation-based dependency injection☆24Jul 21, 2026Updated last week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Analyze scraped data☆47Dec 9, 2019Updated 6 years ago
- Simple Web UI for Scrapy spider management via Scrapyd☆50Jun 25, 2018Updated 8 years ago
- Scrapy Extension for monitoring spiders execution.☆562May 28, 2026Updated 2 months ago
- PyData Boston 2013 talks: "Intro to scikit-learn" & "Realtime Predictive Analytics: Using scikit-learn and RabbitMQ"☆11Jan 5, 2014Updated 12 years ago
- Use pyppeteer from a Scrapy spider☆59Feb 5, 2020Updated 6 years ago
- CLI: Delete GitHub Branches by pattern matching.☆16Aug 23, 2022Updated 3 years ago
- ☆20Oct 6, 2025Updated 9 months ago
- gzipstream allows Python to process multi-part gzip files from a streaming source☆23Feb 24, 2017Updated 9 years ago
- Library designed to replace the SQLite backend by a MongoDB backend on Scrapy queue management☆17Sep 2, 2017Updated 8 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- A daemon for scheduling Scrapy spiders☆65May 28, 2021Updated 5 years ago
- Scrapinghub Command Line Client☆129Jul 22, 2026Updated last week
- Page Object pattern for Scrapy☆127Jun 8, 2026Updated last month
- dlt-dagster-demo☆14Nov 6, 2023Updated 2 years ago
- Youtube crawler & scraper based on scrapy. Written in Python3.☆16Mar 13, 2026Updated 4 months ago
- AWS DynamoDB pipeline for Scrapy☆21Mar 26, 2025Updated last year
- Python dict-like interface for merging dicts with add to set property☆14Apr 13, 2026Updated 3 months ago
- A component that tries to avoid downloading duplicate content☆28Apr 8, 2026Updated 3 months ago
- use multiple proxies with Scrapy☆775Apr 8, 2026Updated 3 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- A curated list of awesome packages, articles, and other cool resources from the Scrapy community.☆561Dec 28, 2022Updated 3 years ago
- Module for pipelines concept in PySpark☆16Mar 27, 2024Updated 2 years ago
- Random User-Agent middleware based on fake-useragent☆688Sep 18, 2023Updated 2 years ago
- List of all Memcached Servers that are vulnerable to DDoS attack vector☆10Dec 21, 2020Updated 5 years ago
- I implement Flutter's "GetStarted" with using BLoC pattern (with RxDart)☆14Nov 12, 2020Updated 5 years ago
- An extension for pymongo that adds json schema validation and index management☆13Oct 19, 2019Updated 6 years ago
- Ideas for serverless applications to be published to the AWS Serverless Application Repository☆13Jun 10, 2018Updated 8 years ago