🦆 Batch data pipeline with Airflow, DuckDB, Delta Lake, Trino, MinIO, and Metabase. Full observability and data quality.
☆90Nov 5, 2025Updated 9 months ago
Alternatives and similar repositories for batch-data-pipeline
Users that are interested in batch-data-pipeline are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Declarative pipelines with multi-engine execution, quality checks, and built-in testing.☆17Jun 18, 2026Updated 2 months ago
- Personal project for setting up an open source data warehouse.☆32Jul 11, 2025Updated last year
- Fully operational local setup to experiment with a Spark-based ecosystem☆23Feb 27, 2025Updated last year
- ☆17Jan 23, 2026Updated 6 months ago
- ☆16Jul 25, 2025Updated last year
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Deploy a complete data stack in just a couple of minutes.☆15Mar 6, 2024Updated 2 years ago
- Open Data Stack Platform: a collection of projects and pipelines built with open data stack tools for scalable, observable data platform…☆22Updated this week
- ☆10Mar 6, 2026Updated 5 months ago
- Repository containing scripts for gathering, parsing, filtering, and visualizing declared taxes data present in Country-by-country report…☆11May 24, 2024Updated 2 years ago
- ☆10May 15, 2025Updated last year
- Apache Airflow advanced functionalities examples☆21Mar 22, 2024Updated 2 years ago
- learning-by-doing data model built with dbt-core☆17Apr 10, 2026Updated 4 months ago
- End-to-end data platform leveraging the Modern data stack☆52Apr 10, 2024Updated 2 years ago
- ☆16Feb 23, 2026Updated 5 months ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Trino Iceberg Metadata Insights via Streamlit☆17Apr 9, 2025Updated last year
- 📡 Real-time data pipeline with Kafka, Flink, Iceberg, Trino, MinIO, and Superset. Ideal for learning data systems.☆79Jan 18, 2025Updated last year
- TUI for google cloud storage (GCS)☆20Feb 2, 2026Updated 6 months ago
- Contains all of the code used in the GitHub Actions Basics Course that is hosted on YouTube here: https://www.youtube.com/watch?v=6FZEfoR…☆11Aug 19, 2024Updated last year
- Docker Big Data Tools: This docker-compose file is configured to run multiple nodes. This is a Hadoop Cluster that contains the necessary…☆31Jul 6, 2021Updated 5 years ago
- Some ways in which you can use DuckDB for streaming analytics☆29Oct 5, 2025Updated 10 months ago
- Repo for everything open table formats (Iceberg, Hudi, Delta Lake) and the overall Lakehouse architecture☆180May 22, 2026Updated 2 months ago
- Run an open-source data LakeHouse locally using Docker Compose☆12May 31, 2024Updated 2 years ago
- This repo contains examples of high throughput ingestion using Apache Spark and Apache Iceberg. These examples cover IoT and CDC scenario…☆30Jul 16, 2026Updated last month
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Comprehensive metrics, insights, and visualization for Agno and Crew AI applications☆26May 21, 2025Updated last year
- Draw like an artist.☆11Dec 18, 2023Updated 2 years ago
- ☆32Oct 4, 2024Updated last year
- Repo containing all of my Data engineering projects☆14May 4, 2025Updated last year
- DuckDB Kernel - analytical execution runtime for Jupyter☆24Jun 29, 2026Updated last month
- Tutorials and examples of how to deploy Presto and connect it to different data sources☆26Jul 21, 2026Updated 3 weeks ago
- Files for the Docker and Kubernetes on Google Cloud Hands-On labs☆11Mar 14, 2023Updated 3 years ago
- Sunona: Next-generation voice AI infrastructure. Orchestrate intelligent, action-oriented voice agents that sound human and execute compl…☆15Updated this week
- Spark In MapReduce (SIMR) - launching Spark applications on existing Hadoop MapReduce infrastructure☆44Mar 9, 2022Updated 4 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- chat-with-docs☆20Nov 28, 2024Updated last year
- Hexagonal (ports and adapters) architecture applied to Spark and Python data engineering project☆33Jul 26, 2023Updated 3 years ago
- The Lakehouse Engine is a configuration driven Spark framework, written in Python, serving as a scalable and distributed engine for sever…☆294Jun 3, 2026Updated 2 months ago
- Local Environment to Practice Data Engineering☆142Dec 30, 2024Updated last year
- Building Data Lakehouse by open source technology. Support end to end data pipeline, from source data on AWS S3 to Lakehouse, visualize a…☆43Dec 15, 2025Updated 8 months ago
- Real-time OLTP system for credit card fraud detection using AWS API Gateway, Kinesis, and RDS PostgreSQL. Features a scalable, serverless…☆25Dec 16, 2024Updated last year
- This project demonstrates how to integrate DuckLake, SQLMesh, and Neon PostgreSQL to create a modern data lakehouse architecture with ver…☆27Jun 3, 2025Updated last year