Twitter's collection of LZO and Protocol Buffer-related Hadoop, Pig, Hive, and HBase code.
☆1,133Apr 10, 2023Updated 3 years ago
Alternatives and similar repositories for elephant-bird
Users that are interested in elephant-bird are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Hadoop library for large-scale data processing, now an Apache Incubator project☆581Jul 8, 2014Updated 12 years ago
- Refactored version of code.google.com/hadoop-gpl-compression for hadoop 0.20☆548Apr 24, 2024Updated 2 years ago
- A platform for visualization and real-time monitoring of data workflows☆1,170Jan 22, 2020Updated 6 years ago
- Streaming MapReduce with Scalding and Storm☆2,123Jan 19, 2022Updated 4 years ago
- Mahout vector encoding for pig☆53Nov 20, 2022Updated 3 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Common metadata layer for Hadoop's Map Reduce, Pig, and Hive☆77Feb 17, 2011Updated 15 years ago
- Distributed and fault-tolerant realtime computation: stream processing, continuous computation, distributed RPC, and more☆8,766Aug 16, 2017Updated 8 years ago
- A Scala API for Cascading☆3,522May 28, 2023Updated 3 years ago
- Oozie - workflow engine for Hadoop☆373Jun 8, 2017Updated 9 years ago
- A grouping of Apache Pig examples.☆65Oct 13, 2020Updated 5 years ago
- Development in Shark has been ended.☆991Aug 11, 2015Updated 11 years ago
- WE HAVE MOVED to Apache Incubator. https://cwiki.apache.org/FLUME/ . Flume is a distributed, reliable, and available service for effici…☆942May 26, 2021Updated 5 years ago
- A bunch of utility classes for Java, Hadoop, HBase, Pig, etc.☆77Mar 31, 2014Updated 12 years ago
- Distributed database specialized in exporting key/value data from Hadoop☆558Jun 27, 2014Updated 12 years ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Piglet is a DSL for writing Pig scripts in Ruby☆83Jul 21, 2010Updated 16 years ago
- Machine Learning for Cascading☆85Jun 12, 2015Updated 11 years ago
- LinkedIn's previous generation Kafka to HDFS pipeline.☆878Aug 27, 2020Updated 5 years ago
- Lightning-fast cluster computing in Java, Scala and Python.☆1,418Apr 8, 2014Updated 12 years ago
- Bulk loading for elastic search☆186Dec 16, 2023Updated 2 years ago
- Scribe is a server for aggregating log data streamed in real time from a large number of servers. It is designed to be scalable, extensib…☆112May 17, 2011Updated 15 years ago
- Data processing on Hadoop without the hassle.☆1,373May 18, 2023Updated 3 years ago
- Hadoop log aggregator and dashboard☆190Oct 29, 2013Updated 12 years ago
- Data and example code for Programming Pig, by Alan F. Gates☆186Oct 15, 2016Updated 9 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Use Avro to store all your values in HBase instead of regular columns☆76Dec 1, 2017Updated 8 years ago
- Hive + Avro. Serde for working with Avro in Hive☆60Dec 16, 2023Updated 2 years ago
- A fully asynchronous, non-blocking, thread-safe, high-performance HBase client.☆609May 19, 2023Updated 3 years ago
- Remedy small files by combining them into larger ones.☆196Jul 1, 2022Updated 4 years ago
- A reporistory of User-defined functions for Apache Pig☆16Sep 20, 2010Updated 15 years ago
- Zohmg is a data store for aggregation of multi-dimensional time series data, built on top of Hadoop, Dumbo and HBase.☆173Oct 16, 2012Updated 13 years ago
- HBase as the backing store for the TF-IDF representations for Lucene☆110May 14, 2010Updated 16 years ago
- Examples of use of pig scripting languages capabilities☆39Aug 1, 2016Updated 10 years ago
- A set of examples and utilities for using Pig with Cassandra. For the latest jar release, check the Downloads link.☆84Aug 21, 2014Updated 11 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Cloud9 is a Hadoop toolkit for working with big data☆237Dec 15, 2015Updated 10 years ago
- Abstract Algebra for Scala☆2,298Nov 21, 2025Updated 8 months ago
- Stream summarizer and cardinality estimator.☆2,264Nov 28, 2019Updated 6 years ago
- A Python wrapper for Cascading☆220Dec 30, 2019Updated 6 years ago
- GoldenOrb is an open-source implementation of Pregel, Google's graph processing framework☆293Jun 29, 2022Updated 4 years ago
- Python module that allows one to easily write and run Hadoop programs.☆1,030Jan 9, 2018Updated 8 years ago
- ☆558Feb 12, 2022Updated 4 years ago