This project provides a comprehensive data pipeline solution to extract, transform, and load (ETL) Reddit data into a Redshift data warehouse. The pipeline leverages a combination of tools and services including Apache Airflow, Celery, PostgreSQL, Amazon S3, AWS Glue, Amazon Athena, and Amazon Redshift.
☆227Oct 23, 2023Updated 2 years ago
Alternatives and similar repositories for RedditDataEngineering
Users that are interested in RedditDataEngineering are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Data Engineering YouTube Analysis Project by Darshil Parmar☆248Dec 8, 2023Updated 2 years ago
- ☆345Aug 13, 2024Updated last year
- This project provides an end-to-end data processing and visualization of visa numbers in Japan using PySpark and Plotly. The spark cluste…☆11Oct 11, 2023Updated 2 years ago
- Dockerized analytics pipeline where Airflow moves mock-API data into Postgres, runs dbt transformations, and serves clean tables for Powe…☆81Oct 18, 2025Updated 9 months ago
- A data engineering project with Kafka, Spark Streaming, dbt, Docker, Airflow, Terraform, GCP and much more!☆894Apr 16, 2022Updated 4 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- An end-to-end data engineering pipeline that fetches real-time YouTube analytics and streams them through Kafka for processing with ksqlD…☆16Sep 19, 2023Updated 2 years ago
- ☆26Jan 31, 2023Updated 3 years ago
- An end-to-end data engineering pipeline that fetches data from Wikipedia, cleans and transforms it with Apache Airflow and saves it on Az…☆31Oct 2, 2023Updated 2 years ago
- Example end to end data engineering project.☆1,423Dec 8, 2022Updated 3 years ago
- An end-to-end data engineering pipeline that orchestrates data ingestion, processing, and storage using Apache Airflow, Python, Apache Ka…☆338Feb 14, 2025Updated last year
- Học Docker trên Ubuntu☆17Jul 16, 2025Updated last year
- This repository contains the necessary configuration files and DAGs (Directed Acyclic Graphs) for setting up a robust data engineering en…☆25Jan 26, 2024Updated 2 years ago
- This repository contains the code for a realtime election voting system. The system is built using Python, Kafka, Spark Streaming, Postgr…☆48Dec 11, 2023Updated 2 years ago
- An end-to-end GoodReads Data Pipeline for Building Data Lake, Data Warehouse and Analytics Platform.☆1,537Mar 9, 2020Updated 6 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- This repository contains an Apache Flink application for real-time sales analytics built using Docker Compose to orchestrate the necessar…☆51Dec 4, 2023Updated 2 years ago
- This project shows how to capture changes from postgres database and stream them into kafka☆41May 17, 2024Updated 2 years ago
- This project showcases how to integrate the world of DevOps, focusing on Continuous Integration (CI) and Continuous Deployment (CD) with …☆14Dec 27, 2023Updated 2 years ago
- This repository contains an end-to-end data engineering project using Apache Flink, focused on performing sales analytics. The project de…☆12Nov 18, 2023Updated 2 years ago
- An end-to-end, containerized data pipeline for near-real-time user event analytics using Kafka, ClickHouse, Airflow, and PySpark. Made to…☆81Sep 12, 2025Updated 10 months ago
- On-premises ELT Pipeline☆32Jul 10, 2025Updated last year
- ☆21Jan 13, 2024Updated 2 years ago
- A lakehouse to store data of books on Tiki☆22Apr 26, 2026Updated 3 months ago
- End-to-end ELT pipeline for 160K+ Skytrax airline reviews: Airflow orchestration, BeautifulSoup scraping, S3 staging, Snowflake wareho…☆13Jul 30, 2026Updated last week
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆17Dec 12, 2024Updated last year
- ☆15Dec 11, 2023Updated 2 years ago
- ☆402Jan 26, 2025Updated last year
- ☆23Feb 8, 2023Updated 3 years ago
- ☆27Dec 1, 2023Updated 2 years ago
- In this project, we will build and ETL(Extract,Transform,Load) pipeline using the Spotify API on AWS. The pipeline will retrieve data fro…☆25May 6, 2023Updated 3 years ago
- ☆19Feb 11, 2026Updated 5 months ago
- This project demonstrates real-time data streaming and processing architecture using Kafka, Spark Streaming, and Debezium for capturing C…☆16Oct 24, 2024Updated last year
- 🟣 Data Engineer interview questions and answers to help you prepare for your next machine learning and data science interview in 2026.☆99Jan 4, 2026Updated 7 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Data Engineering Zoomcamp is a free 9-week course on building production-ready data pipelines. The next cohort starts in January 2026. Jo…☆44,439Jun 10, 2026Updated 2 months ago
- ☆52Feb 27, 2026Updated 5 months ago
- Glue ETL job or EMR Spark that gets from data catalog, modifies and uploads to S3 and Data Catalog☆13Aug 26, 2023Updated 2 years ago
- Loan Default Prediction using PySpark, with jobs scheduled by Apache Airflow and Integration with Spark using Apache Livy☆22Dec 26, 2020Updated 5 years ago
- ☆71Aug 15, 2024Updated last year
- Python wrapper for Goodreads API☆30Feb 20, 2020Updated 6 years ago
- ☆64Mar 7, 2026Updated 5 months ago