Useful reusable pipeline components for Crunch jobs
☆27Feb 10, 2015Updated 11 years ago
Alternatives and similar repositories for crunch-lib
Users that are interested in crunch-lib are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Mirror of Apache Crunch (Incubating)☆110Feb 2, 2021Updated 5 years ago
- Provides a simple archetype to create MapReduce jobs with Maven.☆24Dec 3, 2010Updated 15 years ago
- OLD - impyla now developed at `cloudera/impyla`☆23Apr 16, 2014Updated 12 years ago
- Kubernetes operator for Apache Druid. Deploy and run Druid clusters with the Stackable Data Platform (SDP).☆12Updated this week
- On demand presto cluster with mesos, marathon and docker.☆29Mar 7, 2018Updated 8 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Apache Pig plugin for Eclipse☆12Feb 28, 2017Updated 9 years ago
- Utilities to use Avro files from Hadoop Map/Reduce jobs and Streaming☆26Sep 10, 2013Updated 12 years ago
- Tools for working with parquet, impala, and hive☆135Jan 4, 2021Updated 5 years ago
- ☆29Nov 17, 2014Updated 11 years ago
- Code repository for Java Data Science Cookbook, published by Packt☆25Jan 30, 2023Updated 3 years ago
- HBase as a JSON Document Database☆27Jun 14, 2023Updated 3 years ago
- Recipes and examples for Apache Spark☆13Jan 21, 2015Updated 11 years ago
- Combination of Dockerized Hortonworks projects and other Hadoop ecosystem components☆10Oct 11, 2019Updated 6 years ago
- Build time tool for detecting link problems in java projects☆156Dec 17, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Periscope brings SLA policy based autoscaling to Hadoop☆35Jan 25, 2016Updated 10 years ago
- Timberlake is a Job Tracker for Hadoop.☆177Jan 24, 2020Updated 6 years ago
- DuckDB extension for MySQL☆15Mar 17, 2024Updated 2 years ago
- Cloudera Maven Archetypes☆18Sep 7, 2011Updated 14 years ago
- An Apache Mesos Framework that allows for replaying load over and over and over (and over) again☆10Aug 10, 2015Updated 11 years ago
- Hive + Avro. Serde for working with Avro in Hive☆60Dec 16, 2023Updated 2 years ago
- Scripts for running Apache Kafka on Mesosphere's Marathon☆14Dec 6, 2015Updated 10 years ago
- Integration for Cascading and Apache Hive☆25Oct 31, 2017Updated 8 years ago
- The released version of Astro(Spark SQL on HBase) has been moved to:☆16Jul 23, 2015Updated 11 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- A dbt adapter for Decodable☆11Sep 4, 2025Updated 11 months ago
- Muppet☆128May 7, 2021Updated 5 years ago
- Easy way to send Finagle metrics to Codahale Metrics library☆43Apr 2, 2020Updated 6 years ago
- Example code showing how to use a Kafka Spout in Storm 0.9.3☆12Dec 26, 2014Updated 11 years ago
- Schedoscope is a scheduling framework for painfree agile development, testing, (re)loading, and monitoring of your datahub, lake, or what…☆98Nov 14, 2019Updated 6 years ago
- Deploying apache-hadoop in a virtualized cluster as easy as 1-2-3.☆15Jul 17, 2019Updated 7 years ago
- Code for Springer Book: High Performance Distributed Computing: Case Studies with Hadoop, Scalding and Spark☆15Oct 6, 2017Updated 8 years ago
- local development sandbox containers☆19Jul 31, 2026Updated last week
- Hadoop YARN monitoring with R☆19Sep 16, 2014Updated 11 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Kubernetes operator for the Open Policy Agent (OPA). Deploy and run OPA for authorization with the Stackable Data Platform (SDP).☆21Updated this week
- Packages part of the M3DB project but outside the main tree with loose compatibility requirements☆17Apr 8, 2019Updated 7 years ago
- Probabilistic data structures server. The data model is key-value, where values are: Bloomfilters, LinearCounters, HyperLogLogs, CountMin…☆24Jan 25, 2016Updated 10 years ago
- Utilities for writing dropwizard integration tests.☆15Jul 14, 2015Updated 11 years ago
- Kubernetes operator for Apache HBase. Deploy and run HBase masters and region servers with the Stackable Data Platform (SDP).☆21Updated this week
- Luigi Workflow Engine integration for Treasure Data☆16May 14, 2018Updated 8 years ago
- Embedded Kafka for testing and quick prototyping.☆14Apr 19, 2016Updated 10 years ago