A Docker Compose template that builds a interactive development environment for PySpark with Jupyter Lab, MinIO as object storage, Hive Metastore, Trino and Kafka
☆47Dec 19, 2024Updated last year
Alternatives and similar repositories for lasagna
Users that are interested in lasagna are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Sample Data Lakehouse deployed in Docker containers using Apache Iceberg, Minio, Trino and a Hive Metastore. Can be used for local testin…☆82Sep 2, 2023Updated 3 years ago
- Repositório no Bootcamp de Engenharia de Dados da Stack Academy.☆44Feb 10, 2023Updated 3 years ago
- Generate DBT tests based on sample data☆38Feb 28, 2024Updated 2 years ago
- Playground for Lakehouse (Iceberg, Hudi, Spark, Flink, Trino, DBT, Airflow, Kafka, Debezium CDC)☆71Sep 23, 2023Updated 3 years ago
- Codebase for the backend of VUTTR (Very Useful Tools to Remember)☆12Jan 24, 2023Updated 3 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- used Airflow, Postgres, Kafka, Spark, and Cassandra, and GitHub Actions to establish an end-to-end data pipeline☆32Oct 25, 2023Updated 2 years ago
- Simple demo using "behave" and "pyspark" libraries to test data transformations in a human-readable way☆10Apr 5, 2019Updated 7 years ago
- 数据治理整体架构☆10Nov 11, 2019Updated 6 years ago
- Este é um projeto de exemplo que demonstra um processo de ETL (Extração, Transformação e Carga) de dados usando Python, Polars e AWS Loca…☆15Sep 25, 2023Updated 2 years ago
- Simple MLP in Python using Numpy☆10Nov 17, 2018Updated 7 years ago
- Docker compose and Google Colab demo to build a CDC with Delta Lake☆15Sep 7, 2022Updated 4 years ago
- VSCode Dev Container template for AWS Glue jobs development☆20Jul 25, 2024Updated 2 years ago
- Code to demonstrate data engineering metadata & logging best practices☆21Mar 12, 2024Updated 2 years ago
- Instalador autonomo do Apache Spark para Sistemas linux: based(Debian,RHEL)☆13Dec 10, 2024Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ☆13Mar 20, 2021Updated 5 years ago
- Traditionally, engineers were needed to implement business logic via data pipelines before business users can start using it. Using this …☆12Updated this week
- LobotoMl is a set of scripts and tools to assess production deployments of ML services☆10May 16, 2022Updated 4 years ago
- A minimal Python wrapper around the App Center REST API☆25Sep 15, 2026Updated last week
- minio as local storage and DynamoDB as catalog☆15May 14, 2024Updated 2 years ago
- A utility to inspect, validate, sign and verify machine learning model files.☆66Feb 5, 2025Updated last year
- Test data management tool for any data source, batch or real-time. Generate, validate and clean up data all in one tool.☆81Feb 14, 2026Updated 7 months ago
- This repository contains the necessary configuration files and DAGs (Directed Acyclic Graphs) for setting up a robust data engineering en …☆25Jan 26, 2024Updated 2 years ago
- Instructions and code for the workshop "From Big Data to NLP Insights: Unlocking the Power of PySpark and Spark NLP"☆12May 9, 2023Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- This is micro size AI RC Car projects. Using Tensorflow lite for microcontroller.☆11Aug 20, 2020Updated 6 years ago
- Code for the paper: Kernel Distributionally Robust Optimization☆13Feb 21, 2021Updated 5 years ago
- Wining solution and its further development for MICCAI 2017 Endoscopic Vision Challenge Angiodysplasia Detection and Localization☆16Jul 3, 2019Updated 7 years ago
- Real-time Data Warehouse with Apache Flink & Apache Kafka & Apache Hudi☆121Dec 15, 2023Updated 2 years ago
- ☆16May 30, 2024Updated 2 years ago
- Rope collision in cpp☆12Jun 2, 2025Updated last year
- A project for exploring how Great Expectations can be used to ensure data quality and validate batches within a data pipeline defined in …☆29Aug 30, 2022Updated 4 years ago
- Pytorch directly integrated to the cloud all through Bench AI!☆10Dec 10, 2023Updated 2 years ago
- Accompanying code for our NeurIPS 2019 paper☆11Nov 7, 2019Updated 6 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- ChatTube: A Retrieval QA System to Youtube Videos☆10Jun 6, 2023Updated 3 years ago
- HackerNews reader☆10Nov 13, 2015Updated 10 years ago
- 大数据环境搭建,本项目为大数据基础镜像组件,其中包括Hadoop、Spark、Hive、Tez、Hue、Flink、Zookeeper、Kafka、MySQL等,用于开发学习使用。☆13Sep 3, 2022Updated 4 years ago
- python and C++ projects - opencv☆12Nov 23, 2019Updated 6 years ago
- Portfolio optimization☆18Sep 10, 2026Updated last week
- Pytorch Implementation of the Explainable Conditional Adversarial Autoencoder using Saliency Maps and SHAP (J. of Imaging - MDPI)☆12Mar 5, 2025Updated last year
- Classify images of different kitchenware items☆11Apr 17, 2023Updated 3 years ago