An end-to-end data engineering pipeline that fetches data from Wikipedia, cleans and transforms it with Apache Airflow and saves it on Azure Data Lake. Other processing takes place on Azure Data Factory, Azure Synapse and Tableau.
☆31Oct 2, 2023Updated 2 years ago
Alternatives and similar repositories for FootballDataEngineering
Users that are interested in FootballDataEngineering are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- This project provides an end-to-end data processing and visualization of visa numbers in Japan using PySpark and Plotly. The spark cluste…☆11Oct 11, 2023Updated 2 years ago
- This project showcases how to integrate the world of DevOps, focusing on Continuous Integration (CI) and Continuous Deployment (CD) with …☆14Dec 27, 2023Updated 2 years ago
- This project provides a comprehensive data pipeline solution to extract, transform, and load (ETL) Reddit data into a Redshift data wareh…☆227Oct 23, 2023Updated 2 years ago
- This repository contains the necessary configuration files and DAGs (Directed Acyclic Graphs) for setting up a robust data engineering en…☆25Jan 26, 2024Updated 2 years ago
- An end-to-end data engineering pipeline that orchestrates data ingestion, processing, and storage using Apache Airflow, Python, Apache Ka…☆351Feb 14, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Transform data from on-premises SQL Server to Azure Delta Lake Storage for Analytics and Visualization☆29Jul 16, 2023Updated 3 years ago
- This project shows how to capture changes from postgres database and stream them into kafka☆41May 17, 2024Updated 2 years ago
- Early detecting of lung cancer using the Luna data set with LIDC IDRI annotations using two models nodule classification"Googlent model" …☆17Sep 30, 2022Updated 3 years ago
- Few projects related to Data Engineering including Data Modeling, Infrastructure setup on cloud, Data Warehousing and Data Lake developme…☆12Feb 26, 2020Updated 6 years ago
- This repository contains an Apache Flink application for real-time sales analytics built using Docker Compose to orchestrate the necessar…☆51Dec 4, 2023Updated 2 years ago
- ☆16Oct 17, 2024Updated last year
- This repository contains the code for a realtime election voting system. The system is built using Python, Kafka, Spark Streaming, Postgr…☆49Dec 11, 2023Updated 2 years ago
- This project demonstrates how to use Apache Airflow to submit jobs to Apache spark cluster in different programming laguages using Python…☆48Mar 14, 2024Updated 2 years ago
- ☆17Aug 2, 2025Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Create a streaming data, transfer it to Kafka, modify it with PySpark, take it to ElasticSearch and MinIO☆66Jul 21, 2023Updated 3 years ago
- End-to-end ETL pipeline in the Microsoft Azure cloud - (Jun '24 - Jul '24)☆40Aug 9, 2024Updated 2 years ago
- ☆17Mar 10, 2025Updated last year
- Resources for tracking DP-700 exam prep progress.☆56Jun 14, 2025Updated last year
- ☆18Feb 1, 2025Updated last year
- Data pipeline from device to cloud☆11May 14, 2022Updated 4 years ago
- ☆18Feb 14, 2025Updated last year
- Toolset for detecting reflected xss in websites☆16Oct 6, 2018Updated 7 years ago
- Local SQL Database ---> Azure ---> Power BI☆15Oct 13, 2023Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆12Jan 14, 2023Updated 3 years ago
- ☆27Apr 16, 2026Updated 4 months ago
- ☆27Dec 31, 2024Updated last year
- ☆19May 11, 2023Updated 3 years ago
- End-to-end Data Project (DA/DS/DE/MLOps) - retail/e-commerce - interpretable dynamic clustering☆22Jul 12, 2025Updated last year
- With everything I learned from DEZoomcamp from datatalks.club, this project performs a batch processing on AWS for the cycling dataset wh…☆15Jan 4, 2026Updated 7 months ago
- Python wrapper for Goodreads API☆30Feb 20, 2020Updated 6 years ago
- Dockerized analytics pipeline where Airflow moves mock-API data into Postgres, runs dbt transformations, and serves clean tables for Powe…☆82Oct 18, 2025Updated 10 months ago
- This project demonstrates real-time data streaming and processing architecture using Kafka, Spark Streaming, and Debezium for capturing C…☆19Oct 24, 2024Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆14Oct 1, 2022Updated 3 years ago
- Add .waitForUrl() to Nightmare☆10Aug 23, 2016Updated 10 years ago
- ☆15Aug 3, 2022Updated 4 years ago
- a very powerful ai friendly web scrapper, can escape any kind of bot detection to read the web content trouhgfully.☆21Jun 11, 2026Updated 2 months ago
- A public repository of documents about Engineering at Iterable, communicating our values, teams, and processes. We hope this will be valu…☆17Nov 15, 2022Updated 3 years ago
- ELT Data Pipeline implementation in Data Warehousing environment☆31May 2, 2025Updated last year
- This project is a linear regression modeling of kings county housing price prediction. The data set was provided by Flatiron School for D…☆10Oct 24, 2020Updated 5 years ago