Generate synthetic Spotify music stream dataset to create dashboards. Spotify API generates fake event data emitted to Kafka. Spark consumes and processes Kafka data, saving it to the Datalake. Airflow orchestrates the pipeline. dbt moves data to Snowflake, transforms it, and creates dashboards.
☆72Dec 17, 2023Updated 2 years ago
Alternatives and similar repositories for spotify-stream-analytics
Users that are interested in spotify-stream-analytics are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- velib-v2: An ETL pipeline that employs batch and streaming jobs using Spark, Kafka, Airflow, and other tools, all orchestrated with Docke…☆21Aug 12, 2025Updated 11 months ago
- ☆12May 27, 2024Updated 2 years ago
- ⚙️ Airflow data pipeline with Terraform, GCP BigQuery, dbt, Soda and Looker Studio.☆26Oct 19, 2023Updated 2 years ago
- ☆12Oct 10, 2023Updated 2 years ago
- ☆15Apr 14, 2026Updated 3 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- This repository contains the capstone project carried out as part of Machine Learning Zoomcamp course☆10Dec 26, 2022Updated 3 years ago
- This repository contains notebooks, homework, projects and notes done during Machine Learning Zoomcamp course.☆13Nov 13, 2024Updated last year
- Implementation of a Q&A Chatbot with RAG using the Data Engineering Community WhatsApp group messages.☆10Jan 17, 2025Updated last year
- Local development environment for python data projects, with Docker☆22Dec 14, 2022Updated 3 years ago
- This project is for demonstrating knowledge of Data Engineering tools and concepts and also learning in the process☆44Dec 1, 2022Updated 3 years ago
- This repo consists of all important concepts for data engineers.☆11Jun 2, 2026Updated 2 months ago
- ☆14Jan 12, 2017Updated 9 years ago
- A custom end-to-end analytics platform for customer churn☆10May 15, 2025Updated last year
- This is a capstone project associated with MLOps Zoomcamp. The end goal of the project is to build an end-to-end machine learning projec…☆13Sep 8, 2022Updated 3 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Testing Spark Structured Streaming anf Kafka with real data from traffic sensors☆17Nov 11, 2022Updated 3 years ago
- Using Apache Spark SQL, Spark ML, Pandas to analyse and predict using the Chicago crime dataset☆10Apr 6, 2018Updated 8 years ago
- Here I will be exploring various tools and methods that are used in data engineering process with Python.☆21Jan 4, 2021Updated 5 years ago
- End-to-end data platform leveraging the Modern data stack☆52Apr 10, 2024Updated 2 years ago
- Series follows learning from Apache Spark (PySpark) with quick tips and workaround for daily problems in hand☆55Sep 30, 2023Updated 2 years ago
- Code test for data engineering candidates☆47Mar 27, 2024Updated 2 years ago
- Candace's Data Engineering Zoomcamp files and notes☆18Jul 4, 2023Updated 3 years ago
- ☆15Oct 19, 2023Updated 2 years ago
- 네이버 쇼핑 리뷰 데이터를 통해 감성 분석하기(GRU, LSTM)☆10Sep 27, 2021Updated 4 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆14Aug 28, 2024Updated last year
- Code for my "Efficient Data Processing in SQL" book.☆63Aug 6, 2024Updated 2 years ago
- 📦 Starting box for Vagrant. Inside box Ubuntu 20.04 LTS with Git, Docker and Docker compose.☆19May 5, 2022Updated 4 years ago
- An AWS Data Engineering End-to-End Project (Glue, Lambda, Kinesis, Redshift, QuickSight, Athena, EC2, S3)☆17Sep 20, 2023Updated 2 years ago
- The goal of this project is to analyse the impact of Covid-19 on the Aviation industry through data engineering processes using technolog…☆13Jun 26, 2022Updated 4 years ago
- Data Engineer Project: An end-to-end Airflow data pipeline with BigQuery, dbt Soda, and more!☆14Dec 14, 2023Updated 2 years ago
- An open and introductory book for the Python API of Apache Spark (pyspark) 📚📖☆12Sep 19, 2025Updated 10 months ago
- ☆21Nov 4, 2023Updated 2 years ago
- (Python, PySpark)☆10Nov 15, 2020Updated 5 years ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- A simple and easy to use Data Quality (DQ) tool built with Python.☆51Sep 7, 2023Updated 2 years ago
- Set of Jupyter notebooks demonstrating Learning to Rank integrated with Solr and Elasticsearch☆17Jun 19, 2022Updated 4 years ago
- Case Study's from Danny Ma's Serious SQL Course☆19Aug 4, 2022Updated 4 years ago
- ☆16May 29, 2023Updated 3 years ago
- ☆12Jul 8, 2024Updated 2 years ago
- The goal of this project is to build a docker cluster that gives access to Hadoop, HDFS, Hive, PySpark, Sqoop, Airflow, Kafka, Flume, Pos…☆80Feb 27, 2023Updated 3 years ago
- AlvinToh Learning Repository for The Ultimate Hands-On Hadoop - Tame your Big Data!☆10May 23, 2018Updated 8 years ago