Creation of a data lakehouse and an ELT pipeline to enable the efficient analysis and use of data
☆51Dec 2, 2023Updated 2 years ago
Alternatives and similar repositories for Building-Data-LakeHouse
Users that are interested in Building-Data-LakeHouse are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Big Data infrastructure with Hadoop, Spark, Hive and NiFi deployed using Docker Compose. https://doi.org/10.5281/zenodo.18968438☆21Mar 11, 2026Updated 4 months ago
- Sample Data Lakehouse deployed in Docker containers using Apache Iceberg, Minio, Trino and a Hive Metastore. Can be used for local testin…☆82Sep 2, 2023Updated 2 years ago
- ☆23Feb 5, 2024Updated 2 years ago
- ☆13Oct 4, 2023Updated 2 years ago
- This project serves as a comprehensive guide to building an end-to-end data engineering pipeline using TCP/IP Socket, Apache Spark, OpenA…☆45Jan 4, 2024Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆25Mar 15, 2024Updated 2 years ago
- Basic framework utilities to quickly start writing production ready Apache Spark applications☆36Dec 15, 2024Updated last year
- trino + hive + minio with postgres in docker compose☆28Aug 18, 2023Updated 2 years ago
- Đồ án tốt nghiệp | Data Lakehouse☆47Feb 9, 2026Updated 5 months ago
- Comprehensive study materials covering core data engineering concepts, tools, and practices.☆17Jan 20, 2026Updated 6 months ago
- Solutions & Code Related to Blog Posts☆11Nov 6, 2024Updated last year
- used Airflow, Postgres, Kafka, Spark, and Cassandra, and GitHub Actions to establish an end-to-end data pipeline☆32Oct 25, 2023Updated 2 years ago
- ☆42Jul 4, 2022Updated 4 years ago
- POC for all the stack of big data (kafka, spark, cassandra, hdfs, docker, springboot)☆12Dec 16, 2022Updated 3 years ago
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Repo which holds the materials for the EMR Zero To Hero☆28May 7, 2022Updated 4 years ago
- ☆19May 11, 2023Updated 3 years ago
- In this project I have built etl pipline which scraps the trending repository based on month,week and day LIVE extract other related info…☆12Sep 9, 2023Updated 2 years ago
- Spark data pipeline that processes movie ratings data.☆31Updated this week
- 🌟 An end-to-end full-stack data science project, including modelling, MLOps, and data storytelling. ✨☆16Aug 30, 2025Updated 10 months ago
- https://aka.ms/lakehouselab☆23Feb 14, 2023Updated 3 years ago
- A ready to go Big Data cluster (Hadoop + Hadoop Streaming + Spark + PySpark) with Docker and Docker Swarm!☆22May 20, 2025Updated last year
- This project uses Google Analytics 4 BigQuery Exports as its source data, and offers useful base transformations to provide report-ready …☆20Sep 30, 2022Updated 3 years ago
- ☆80Apr 23, 2025Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Fully automated csv to dashboard pipeline using Terraform, Google Cloud Storage, BigQuery, dbt, Prefect and Looker Studio. Peer ranked …☆53Nov 15, 2024Updated last year
- Spark Structured Streaming data pipeline that processes movie ratings data in real-time.☆14Jul 12, 2026Updated 2 weeks ago
- Data Guy Story commandline☆11Dec 2, 2022Updated 3 years ago
- Machine Learning DevOps Engineer Nanodegree☆11Jan 27, 2022Updated 4 years ago
- plan, design and implement enterprise data infrastructure solutions and create the blueprints for an organization’s data management syste…☆13Jun 25, 2023Updated 3 years ago
- ☆13Feb 27, 2024Updated 2 years ago
- ☆13Mar 30, 2024Updated 2 years ago
- Basin is a visual programming editor for building Spark and PySpark pipelines. Easily build, debug, and deploy complex ETL pipelines from…☆35Jan 5, 2023Updated 3 years ago
- StarCraft 2 Data Pipeline with Airflow, DuckDB and Streamlit☆17Mar 14, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆23Feb 8, 2023Updated 3 years ago
- Code for youtube channel☆10Apr 15, 2022Updated 4 years ago
- API for toxic text classification, utilized pre-trained Distilbert and trained on Kaggle datasets. It helps identify and handle toxic con…☆14Apr 30, 2024Updated 2 years ago
- Apache Beam Python examples and templates.☆14Dec 8, 2022Updated 3 years ago
- ☆10Nov 25, 2021Updated 4 years ago
- A Data Engineering Project that implements an ETL data pipeline using Dagster, Apache Spark, Streamlit, MinIO, Metabase, Dbt, Polars, Doc…☆25Nov 19, 2024Updated last year
- Guess what! ;)☆17Dec 16, 2025Updated 7 months ago