Example MapReduce jobs in Java, Hive, Pig, and Hadoop Streaming that work on Avro data.
☆114Nov 12, 2015Updated 10 years ago
Alternatives and similar repositories for avro-hadoop-starter
Users that are interested in avro-hadoop-starter are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A very simple example of using Hadoop's MapReduce functionality in Java.☆72Jun 18, 2013Updated 13 years ago
- Utilities to use Avro files from Hadoop Map/Reduce jobs and Streaming☆26Sep 10, 2013Updated 12 years ago
- A collection of tools that help me work with Avro☆23Jan 7, 2010Updated 16 years ago
- Hadoop log aggregator and dashboard☆190Oct 29, 2013Updated 12 years ago
- Pig on Apache Spark☆82Mar 23, 2015Updated 11 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- What happens on the wire when Hadoop RPC call is issued?☆13Jul 1, 2022Updated 4 years ago
- A grouping of Apache Pig examples.☆65Oct 13, 2020Updated 5 years ago
- A simple test of Avro 1.5 capabilities including dynamic typing, untagged (compact) data storage and schema evolution.☆36May 5, 2011Updated 15 years ago
- ☆15Feb 26, 2013Updated 13 years ago
- Example project to show how to use Spark to read and write Avro/Parquet files☆50Aug 21, 2013Updated 13 years ago
- Utilities for converting to and from JSON from Avro records via Hadoop streaming or Hive.☆30Oct 13, 2020Updated 5 years ago
- Pig Visualization framework☆465Mar 24, 2023Updated 3 years ago
- Tool which generates Avro schemas and Java bindings from XML schemas.☆40Aug 3, 2020Updated 6 years ago
- Examples and Slides for "Introduction to Spring for Apache Hadoop" at SpringOne2GX 2014☆16Jan 7, 2019Updated 7 years ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Source code to accompany the book "Hadoop in Practice", published by Manning.☆202Feb 11, 2020Updated 6 years ago
- [PROJECT IS NO LONGER MAINTAINED] Wirbelsturm is a Vagrant and Puppet based tool to perform 1-click local and remote deployments, with a …☆329Feb 21, 2022Updated 4 years ago
- DocId set compression and set operation library☆22Mar 7, 2014Updated 12 years ago
- An example of how to bulk import data from CSV files into a HBase table.☆65Apr 12, 2022Updated 4 years ago
- ☆18Mar 14, 2016Updated 10 years ago
- The GraphBuilder library provides functions to construct large scale graphs. It is implemented on Apache Hadoop.☆101Oct 9, 2014Updated 11 years ago
- Load testing tools for Flume☆18Jun 22, 2012Updated 14 years ago
- Kite SDK Examples☆99May 8, 2021Updated 5 years ago
- spark + drools☆100May 20, 2022Updated 4 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Combination of Dockerized Hortonworks projects and other Hadoop ecosystem components☆10Oct 11, 2019Updated 6 years ago
- Patched, refactored version of code.google.com/hadoop-gpl-compression for hadoop 0.20☆100Jan 10, 2012Updated 14 years ago
- recordbus: mysql binlog to apache kafka☆78Jul 31, 2015Updated 11 years ago
- Remedy small files by combining them into larger ones.☆196Jul 1, 2022Updated 4 years ago
- Collaborative filtering with MLLib on Spark based on data in Cassandra☆21Mar 11, 2022Updated 4 years ago
- Using Node.js to ingest into Accumulo via RabbitMQ and Java☆19Jun 3, 2012Updated 14 years ago
- Mesos on Mesos☆15Mar 11, 2015Updated 11 years ago
- DataNode Volumes Rebalancing tool for Apache Hadoop HDFS (HDFS-1312)☆23Dec 12, 2017Updated 8 years ago
- Apache Avro RPC Quick Start.☆413Feb 17, 2026Updated 6 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Dockerfiles and scripts for Spark and Shark Docker images☆259Jun 19, 2014Updated 12 years ago
- A Storm Based DRPC Search Engine☆29Aug 26, 2015Updated 11 years ago
- hRaven collects run time data and statistics from MapReduce jobs in an easily queryable format☆129Jan 14, 2022Updated 4 years ago
- ☆18Feb 22, 2026Updated 6 months ago
- Next-generation web analytics processing with Scala, Spark, and Parquet.☆330Mar 28, 2015Updated 11 years ago
- LinkedIn's previous generation Kafka to HDFS pipeline.☆878Aug 27, 2020Updated 6 years ago
- Utilities for working with Hadoop and Cascading☆19Feb 8, 2011Updated 15 years ago