Big Data Ecosystem Docker
☆429Apr 29, 2023Updated 3 years ago
Alternatives and similar repositories for bigdata_docker
Users that are interested in bigdata_docker are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Big Data Ecosystem Docker☆81May 17, 2022Updated 4 years ago
- Modern Data Stack☆63Aug 8, 2025Updated last year
- Hadoop, Hive, Spark, Zeppelin and Livy: all in one Docker-compose file.☆168Feb 4, 2021Updated 5 years ago
- Hadoop-Hive-Spark cluster + Jupyter on Docker☆84Jan 2, 2025Updated last year
- Notas das aulas da Aceleração Dev #4 da DIO sobre Engenharia de Dados, ministrado pela Everis.☆13Feb 6, 2021Updated 5 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Docker-compose contains the most common big data systems like: Apache Hadoop, Apache Hive, Apache Spark, Jupyter, Flink☆28Oct 9, 2023Updated 2 years ago
- Docker images for building hadoop3.2, hive 3.1, hbase2.3, presto 0.247, flink1.11.3 on yarn, etc.☆32Apr 25, 2023Updated 3 years ago
- Data Engineering made simple - An opinionated Data Engineering framework☆66Mar 20, 2024Updated 2 years ago
- Base Docker image with just essentials: Hadoop, Hive and Spark.☆67Feb 3, 2021Updated 5 years ago
- ETL e visualização do Censo escolar☆10May 3, 2023Updated 3 years ago
- Build Elastic Stack using Docker☆22Oct 31, 2025Updated 9 months ago
- ☆24Aug 9, 2023Updated 3 years ago
- ☆22May 16, 2023Updated 3 years ago
- ☆41Jul 23, 2024Updated 2 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆18Sep 17, 2021Updated 4 years ago
- Super Mario Demo Project☆10Dec 14, 2023Updated 2 years ago
- The demo of using Kafka, Spark, Hive, Cassandra, etc by using Docker. It produces the production ready environment for any kinds of big d…☆37Sep 27, 2019Updated 6 years ago
- Learn Apache Spark in Scala, Python (PySpark) and R (SparkR) by building your own cluster with a JupyterLab interface on Docker.☆512Nov 7, 2025Updated 9 months ago
- Sample Docker Compose files for running Apache Ambari☆11Oct 29, 2018Updated 7 years ago
- The goal of this project is to build a docker cluster that gives access to Hadoop, HDFS, Hive, PySpark, Sqoop, Airflow, Kafka, Flume, Pos…☆80Feb 27, 2023Updated 3 years ago
- ☆11Mar 15, 2025Updated last year
- Apache Spark docker image☆2,050Apr 20, 2026Updated 3 months ago
- Code for the "Cherry: A Distributed Task-Aware Shuffle Service for Serverless Analytics" paper for 2021 IEEE International Conference on …☆15Feb 3, 2022Updated 4 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Docker multi-nodes Hadoop cluster with Spark 2.4.1 on Yarn☆50Dec 7, 2020Updated 5 years ago
- Estudos e projetos.☆61Jan 14, 2022Updated 4 years ago
- Portfolio of projects and studies conducted in data engineering.☆34Feb 22, 2025Updated last year
- ☆12Feb 11, 2022Updated 4 years ago
- Related resources to the paper RoBERTaLexPT: A Legal RoBERTa Model pretrained with deduplication for Portuguese.☆22Mar 14, 2024Updated 2 years ago
- Building Real Time Data Pipeline using Apache Kafka, Apache Spark, Hadoop, PostgreSQL, Django and Flexomonster on Docker to track status …☆24Dec 29, 2020Updated 5 years ago
- ☆13Dec 28, 2023Updated 2 years ago
- ☆12Aug 11, 2021Updated 5 years ago
- PySpark course.☆10Feb 21, 2022Updated 4 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Deploy of Airflow 2.0 using ECS Fargate and AWS CDK.☆13Nov 5, 2021Updated 4 years ago
- Airflow plugins for implementing data pipelines. | Plugins do Airflow para implementação de pipelines de dados.☆52Apr 29, 2026Updated 3 months ago
- Docker Big Data Tools: This docker-compose file is configured to run multiple nodes. This is a Hadoop Cluster that contains the necessary…☆31Jul 6, 2021Updated 5 years ago
- ☆23Jun 30, 2024Updated 2 years ago
- Dockerizing an Apache Spark Standalone Cluster☆42Jun 29, 2022Updated 4 years ago
- Multi docker container images for main Big Data Tools. (Hadoop, Spark, Kafka, HBase, Cassandra, Zookeeper, Zeppelin, Drill, Flink, Hive, …☆35Dec 9, 2024Updated last year
- Zeppelin docker☆16Nov 16, 2020Updated 5 years ago