Big Data ETL and Utilities for Hadoop Map Reduce, Spark and Storm
☆106Jan 22, 2024Updated 2 years ago
Alternatives and similar repositories for chombo
Users that are interested in chombo are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Apache Spark based ETL Engine☆71Oct 18, 2016Updated 9 years ago
- ☆10Aug 17, 2022Updated 4 years ago
- A dynamic data completeness and accuracy library at enterprise scale for Apache Spark☆30May 13, 2026Updated 3 months ago
- Scala API for Apache Spark SQL high-order functions☆15Aug 4, 2023Updated 3 years ago
- A pyspark lib to validate data quality☆19Nov 11, 2022Updated 3 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆21Oct 1, 2015Updated 10 years ago
- Library and a Framework for building fast, scalable, fault-tolerant Data APIs based on Akka, Avro, ZooKeeper and Kafka☆25Oct 16, 2020Updated 5 years ago
- Apache Amaterasu☆56Oct 18, 2019Updated 6 years ago
- ☆25Oct 12, 2016Updated 9 years ago
- flinksql-platform☆19Mar 22, 2021Updated 5 years ago
- My branch of Apache Flume with a generic JDBC sink (not yet licensed to Apache)☆11Feb 12, 2022Updated 4 years ago
- Extensible streaming ingestion pipeline on top of Apache Spark☆47Jul 17, 2025Updated last year
- Open source task scheduler with dependency management☆15Jul 1, 2018Updated 8 years ago
- Spark package to "plug" holes in data using SQL based rules ⚡️ 🔌☆28May 15, 2020Updated 6 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Real time and offline time series analysis with Spark, Spark Streaming and Storm☆21Oct 20, 2020Updated 5 years ago
- POC of PAAS on top of Hadoop YARN☆24Jun 19, 2012Updated 14 years ago
- ☆32Mar 21, 2018Updated 8 years ago
- Apache Spark ETL Utilities☆40Oct 23, 2024Updated last year
- Terraform provider for interacting with NiFi cluster☆51May 29, 2019Updated 7 years ago
- Build configuration-driven ETL pipelines on Apache Spark☆162Oct 4, 2022Updated 3 years ago
- An example of how to use the JDBC to issue Hive queries from a Java client application.☆11Apr 5, 2018Updated 8 years ago
- Lighthouse is a library for data lakes built on top of Apache Spark. It provides high-level APIs in Scala to streamline data pipelines an…☆64Sep 6, 2024Updated last year
- Distributed SQL query engine for running interactive analytic queries against big data sources.☆10Jul 1, 2016Updated 10 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Bulletproof Apache Spark jobs with fast root cause analysis of failures.☆73Mar 14, 2021Updated 5 years ago
- ☆16Jun 27, 2020Updated 6 years ago
- A containerized development environment for Go, includes automatic code reloads, test running, and vendored dependencies. Powered by Dock…☆12Feb 21, 2017Updated 9 years ago
- Using JRecord to build a mapred and mapreduce inputformat for HDFS, MAPREDUCE, PIG, HIVE, Spark, ...