Repo for everything open table formats (Iceberg, Hudi, Delta Lake) and the overall Lakehouse architecture
☆180May 22, 2026Updated 3 months ago
Alternatives and similar repositories for awesome-lakehouse-guide
Users that are interested in awesome-lakehouse-guide are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- This repo contains examples of high throughput ingestion using Apache Spark and Apache Iceberg. These examples cover IoT and CDC scenario…☆30Jul 16, 2026Updated last month
- Monitoring and insights on your data lakehouse tables☆32Jul 27, 2026Updated last month
- Floe: Policy-based table maintenance for Apache Iceberg☆44May 6, 2026Updated 3 months ago
- A leightweight UI for Lakekeeper☆20Updated this week
- Apache XTable (incubating) is a cross-table converter for lakehouse table formats that facilitates interoperability across data processin…☆1,208Updated this week
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Iceberg Playground in a Box☆69Apr 8, 2026Updated 4 months ago
- Advanced tools for multi-armed and contextual bandits☆24May 11, 2026Updated 3 months ago
- Tutorials and examples of how to deploy Presto and connect it to different data sources☆26Jul 21, 2026Updated last month
- "Nature's economy shall be the base for our own, for it is immutable, but ours is secondary. An economist without knowledge of nature is …☆20May 31, 2021Updated 5 years ago
- Quick Guides from Dremio on Several topics☆90May 11, 2026Updated 3 months ago
- Local Environment to Practice Data Engineering☆142Dec 30, 2024Updated last year
- A tool to automate analytic platform evaluations. Barometer helps customers to get data points needed for service selection/service confi…☆19Jun 3, 2024Updated 2 years ago
- Dockerized runner, utilities, and functions for FlinkSQL applications☆33Updated this week
- 🦆 Batch data pipeline with Airflow, DuckDB, Delta Lake, Trino, MinIO, and Metabase. Full observability and data quality.☆90Nov 5, 2025Updated 9 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- This repository helps teach people how to correctly define and create cumulative tables!☆774Oct 29, 2024Updated last year
- Apache DataFusion Comet Spark Accelerator☆1,260Updated this week
- Apache Iceberg REST Catalog in Rust — access control, credential vending and audit for every engine and AI agent. Apache 2.0.☆1,430Updated this week
- ☆19May 15, 2026Updated 3 months ago
- Playground for Lakehouse (Iceberg, Hudi, Spark, Flink, Trino, DBT, Airflow, Kafka, Debezium CDC)☆71Sep 23, 2023Updated 2 years ago
- Docker envinroment to stream data from Kafka to Iceberg tables☆30Feb 27, 2024Updated 2 years ago
- Assets used in Cloudera Tutorials☆19Nov 22, 2021Updated 4 years ago
- A portable Datamart and Business Intelligence suite built with Docker, sqlmesh + dbtcore, DuckDB and Superset☆61Apr 5, 2026Updated 4 months ago
- This sample demonstrates building a logistics tracking agent using Strands Agents SDK and an OpenAI model hosted on Amazon Bedrock AgentC…☆26Jul 28, 2026Updated last month
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Apache Polaris, the interoperable, open source catalog for Apache Iceberg☆2,043Updated this week
- ☆44Jul 3, 2022Updated 4 years ago
- Fork of Apache Kafka implementing KIP-1150 -- Diskless Topics☆95Updated this week
- AWS Quick Start Team☆15Oct 3, 2024Updated last year
- Apache Spark Kubernetes Operator☆319Updated this week
- FinOps repository: An open-source framework for multi-cloud cost visibility. Extendable with dlt.☆35Feb 19, 2026Updated 6 months ago
- Configura containers do Spark (Master, Workers e History Server) + Jupyter☆21Jun 17, 2024Updated 2 years ago
- The native Rust implementation for Apache Hudi, with C++ & Python API bindings.☆280Aug 15, 2026Updated 2 weeks ago
- ☆15Oct 10, 2025Updated 10 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Open Control Plane for Tables in Data Lakehouse☆397Updated this week
- An Ansible collection for Cloudera Platform for cloud and Data Services☆23Aug 17, 2026Updated last week
- Delta Lake examples☆242Oct 8, 2024Updated last year
- Glue VSCode devcontainer setup☆14Jan 31, 2023Updated 3 years ago
- Deploy a complete data stack in just a couple of minutes.☆15Mar 6, 2024Updated 2 years ago
- New and extensible file format for storage of large columnar datasets.☆730Aug 7, 2026Updated 3 weeks ago
- MPC Server for PySpark inpired by the LakeSail☆18Feb 26, 2026Updated 6 months ago