The goal of this project is to build a docker cluster that gives access to Hadoop, HDFS, Hive, PySpark, Sqoop, Airflow, Kafka, Flume, Postgres, Cassandra, Hue, Zeppelin, Kadmin, Kafka Control Center and pgAdmin. This cluster is solely intended for usage in a development environment. Do not use it to run any production workloads.
☆80Feb 27, 2023Updated 3 years ago
Alternatives and similar repositories for Big-Data-Cluster
Users that are interested in Big-Data-Cluster are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Spark all the ETL Pipelines☆37Aug 2, 2023Updated 3 years ago
- This project demonstrates real-time data streaming and processing architecture using Kafka, Spark Streaming, and Debezium for capturing C…☆16Oct 24, 2024Updated last year
- Demonstrating practical SQL skills through a curated portfolio of solved problems from top coding platforms.☆51Mar 18, 2026Updated 5 months ago
- This project aims to move the data from a Relational database system (RDBMS) to a Hadoop file system (HDFS)☆11Apr 29, 2022Updated 4 years ago
- ☆23Feb 5, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Here I will be exploring various tools and methods that are used in data engineering process with Python.☆21Jan 4, 2021Updated 5 years ago
- Docker Big Data Tools: This docker-compose file is configured to run multiple nodes. This is a Hadoop Cluster that contains the necessary…☆31Jul 6, 2021Updated 5 years ago
- ☆15May 1, 2024Updated 2 years ago
- Airflow Examples: code samples for Medium articles☆14Jan 10, 2021Updated 5 years ago
- This Repo contains Jupyter Notebooks to recap on RDD, DataFrame, Spark Streaming and ML operations using Pyspark☆11Nov 3, 2024Updated last year
- A shell script to automate the operations of sqoop☆11Mar 29, 2021Updated 5 years ago
- Scalable OLAP system for credit card transaction analysis, leveraging AWS S3, Databricks, and dbt. Features end-to-end batch processing p…☆26Oct 24, 2024Updated last year
- an example of how to handle the gowalla check in dataset☆14May 25, 2017Updated 9 years ago
- ☆20Mar 11, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Extract, transform, and load data for analytic processing using AWS Glue☆17May 2, 2021Updated 5 years ago
- Series follows learning from Apache Spark (PySpark) with quick tips and workaround for daily problems in hand☆55Sep 30, 2023Updated 2 years ago
- R package for Markov regime-switching models☆12Jan 23, 2018Updated 8 years ago
- Cloud Functions streaming insert to BigQuery (with Cloud Pub/Sub trigger). In this example, the function will make a REST API call to get…☆29Aug 28, 2023Updated 2 years ago
- A Python PySpark Projet with Poetry☆31May 2, 2026Updated 3 months ago
- Spark implementation of Slowly Changing Dimension type 2☆10Jan 8, 2019Updated 7 years ago
- Data pipeline for extracting, transforming, and visualising Covid-19 data☆14Apr 23, 2023Updated 3 years ago
- Public Docker Images for popular services☆57Sep 7, 2025Updated 11 months ago
- A minimal docker compose setup for experimenting with cloud agnostic Lakehouse Architectures Apache Spark with Hive Metastore + Delta Lak…☆34Apr 17, 2024Updated 2 years ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Project - Data Processing and Analysis in Python Course☆39Oct 10, 2018Updated 7 years ago
- This project aims to use the Hadoop framework to analyze unstructured data that we obtain from Twitter and perform sentiment and trend an…☆17May 8, 2020Updated 6 years ago
- TTS utility☆12Aug 2, 2020Updated 6 years ago
- ☆10Jul 24, 2024Updated 2 years ago
- Parameter Importance according to OpenML☆14Feb 23, 2022Updated 4 years ago
- audio, NLP, ML with huggingface, nvidia/nemo, speechbrain☆11Sep 4, 2023Updated 2 years ago
- Small data engineering tutorial☆10Oct 24, 2018Updated 7 years ago
- Dockerizing an Apache Spark Standalone Cluster☆42Jun 29, 2022Updated 4 years ago
- Solar Resource Assessment in Python☆12Jul 27, 2026Updated 3 weeks ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆16Apr 1, 2024Updated 2 years ago
- Big Data Inventory Management on AWS (Demand Forecasting, Machine Learning, Dashboarding) : Presented at Carlson School of Management dur…☆11Apr 15, 2020Updated 6 years ago
- Use `outlines` generators with Haystack.☆14Updated this week
- Big Data Ecosystem Docker☆429Apr 29, 2023Updated 3 years ago
- ☆14Oct 25, 2020Updated 5 years ago
- This repository contains an end-to-end data engineering project using Apache Flink, focused on performing sales analytics. The project de…☆12Nov 18, 2023Updated 2 years ago
- A component that will display some developer quotes in your Backstage app!☆10Jul 14, 2026Updated last month