Big Data ETL and Utilities for Hadoop Map Reduce, Spark and Storm
☆106Jan 22, 2024Updated 2 years ago
Alternatives and similar repositories for chombo
Users that are interested in chombo are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Apache Spark based ETL Engine☆71Oct 18, 2016Updated 9 years ago
- Scala API for Apache Spark SQL high-order functions☆15Aug 4, 2023Updated 2 years ago
- ☆21Oct 1, 2015Updated 10 years ago
- ☆10Apr 10, 2014Updated 12 years ago
- Library and a Framework for building fast, scalable, fault-tolerant Data APIs based on Akka, Avro, ZooKeeper and Kafka☆25Oct 16, 2020Updated 5 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- A simple Bloom Filter implementation in Java☆16Oct 21, 2012Updated 13 years ago
- flinksql-platform☆19Mar 22, 2021Updated 5 years ago
- Utilities for working with Hadoop and Cascading☆19Feb 8, 2011Updated 15 years ago
- Spark pipelines that correspond to a series of Dataflow examples.☆27May 5, 2019Updated 7 years ago
- Extensible streaming ingestion pipeline on top of Apache Spark☆47Jul 17, 2025Updated last year
- Open source task scheduler with dependency management☆15Jul 1, 2018Updated 8 years ago
- Plot live-stats as graph from ApacheSpark application using Lightning-viz☆18Jul 3, 2017Updated 9 years ago
- Spark Structured Streaming JDBC Sink☆16Apr 26, 2021Updated 5 years ago
- Spark package to "plug" holes in data using SQL based rules ⚡️ 🔌☆28May 15, 2020Updated 6 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Repository of Notebooks taken from https://neo4j.com/graph-algorithms-book/☆26Feb 21, 2020Updated 6 years ago
- Waimak is an open-source framework that makes it easier to create complex data flows in Apache Spark.☆76Apr 24, 2024Updated 2 years ago
- Real time and offline time series analysis with Spark, Spark Streaming and Storm☆21Oct 20, 2020Updated 5 years ago
- ☆32Mar 21, 2018Updated 8 years ago
- docs, codes and resources to prepare for the CRT020: Databricks Certified Associate Developer for Apache Spark 2.4 with Python 3 certific…☆10Sep 25, 2019Updated 6 years ago
- Smart Automation Tool for building modern Data Lakes and Data Pipelines☆129Updated this week
- Terraform provider for interacting with NiFi cluster☆51May 29, 2019Updated 7 years ago
- An example of how to use the JDBC to issue Hive queries from a Java client application.☆11Apr 5, 2018Updated 8 years ago
- Multi-stage, config driven, SQL based ETL framework using PySpark☆26Sep 16, 2019Updated 6 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Distributed SQL query engine for running interactive analytic queries against big data sources.☆10Jul 1, 2016Updated 10 years ago
- Bulletproof Apache Spark jobs with fast root cause analysis of failures.☆73Mar 14, 2021Updated 5 years ago
- 优化flink的多流操作(例如join),优化点不限于数据丢失问题,以及性能问题☆11Apr 8, 2019Updated 7 years ago
- ☆16Jun 27, 2020Updated 6 years ago
- Using JRecord to build a mapred and mapreduce inputformat for HDFS, MAPREDUCE, PIG, HIVE, Spark, ...☆19Dec 7, 2017Updated 8 years ago
- Yet Another Spark SQL JDBC/ODBC server based on the PostgreSQL V3 protocol☆34Sep 8, 2022Updated 3 years ago
- Spark package for checking data quality☆220Feb 28, 2020Updated 6 years ago
- Java event logs collector for hadoop and frameworks☆42Mar 25, 2025Updated last year
- Serde for Cobol Layout to Hive table☆24Feb 23, 2019Updated 7 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Machine Learning Stack for Big Data, Big Cluster and Big Challenges☆22Sep 6, 2018Updated 7 years ago
- Spatial search using Elastic Search☆12Dec 27, 2014Updated 11 years ago
- 使用spark + kudu的案例☆15Sep 13, 2017Updated 8 years ago
- Scala API for distributed closures on Apache Ignite☆11Jun 6, 2015Updated 11 years ago
- A light Kafka to HDFS/S3 ETL library based on Apache Spark☆40Jun 29, 2017Updated 9 years ago
- Indexing framework designed for the automated creation of structured knowledge bases in Azure AI Search☆15Jul 17, 2026Updated last week
- Cloud based Data Platform based on Apache Spark☆28Jun 30, 2026Updated 3 weeks ago