Hadoop In Action Examples
☆40Apr 26, 2021Updated 5 years ago
Alternatives and similar repositories for hia-examples
Users that are interested in hia-examples are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Mahout Examples☆26Aug 2, 2016Updated 9 years ago
- Slinky, a high-performance web crawler / text analytics in Python, Redis, Hadoop, R, Gephi☆40Aug 30, 2010Updated 15 years ago
- Movie recommendations and more in MapReduce and Scalding☆117Feb 11, 2013Updated 13 years ago
- Spark Elastic MapReduce bootstrap and runnable examples.☆17Jun 26, 2013Updated 13 years ago
- The Scalding WordCountJob example as a standalone SBT project with Specs2 tests, runnable on Amazon EMR☆82Aug 28, 2014Updated 11 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A simple easy to use Hadoop map reduce workflow engine☆18Mar 30, 2012Updated 14 years ago
- Find implementation for Hadoop☆17Sep 9, 2015Updated 10 years ago
- Summingbird Workshop at Lambda Jam 2013.☆24Aug 21, 2018Updated 7 years ago
- Gephi plugin that allows to import graph data from any Tinkerpop Blueprints-compatible graph implementation☆25Nov 19, 2012Updated 13 years ago
- The Scalding tutorial as a standalone SBT project☆51Oct 16, 2017Updated 8 years ago
- Large scale k-nn experiments☆69Jul 31, 2024Updated last year
- A half-day workshop on Scalding, the Scala API for Cascading☆48Mar 21, 2016Updated 10 years ago
- Files to help make new spark EMR Bootstraps☆15Aug 4, 2013Updated 12 years ago
- Driving Spark stream with Scalaz-Stream☆26Mar 18, 2014Updated 12 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Hadoop Inroduction Presentation Demos☆22Jul 27, 2013Updated 12 years ago
- Utilities for converting to and from JSON from Avro records via Hadoop streaming or Hive.☆30Oct 13, 2020Updated 5 years ago
- Muppet☆128May 7, 2021Updated 5 years ago
- Google Analytics plugin for sending events to Snowplow☆17Sep 30, 2020Updated 5 years ago
- MapReduce examples☆20Nov 18, 2011Updated 14 years ago
- Notebooks for playing with NLTK parsing capabilities☆16Dec 1, 2014Updated 11 years ago
- cascading_ext is a collection of tools built on top of the Cascading platform which make it easy to build, debug, and run simple and high…☆58Feb 25, 2026Updated 4 months ago
- ☆28Feb 2, 2023Updated 3 years ago
- HBase DSL for Scala with MapReduce support☆128Jan 4, 2018Updated 8 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- iSAX Indexing persisted in HBase☆39Jul 26, 2011Updated 14 years ago
- Set of Hadoop, Spark and Storm based tools for web and customer analytic☆34Jun 7, 2021Updated 5 years ago
- Source code to accompany the book "Hadoop in Practice", published by Manning.☆202Feb 11, 2020Updated 6 years ago
- scala driver for launching Amazon EMR jobs☆40Feb 10, 2016Updated 10 years ago
- Sample Applications for getting started with Spring for Apache Hadoop☆42Apr 4, 2022Updated 4 years ago
- Zohmg is a data store for aggregation of multi-dimensional time series data, built on top of Hadoop, Dumbo and HBase.☆173Oct 16, 2012Updated 13 years ago
- Distributed Indices for Disco☆15Jan 10, 2012Updated 14 years ago
- Machine learning and natural language processing with Apache Pig☆53Dec 17, 2013Updated 12 years ago
- Dice.com tutorial on using black box optimization algorithms to do relevancy tuning on your Solr Search Engine Configuration from Simon H…☆29Mar 20, 2019Updated 7 years ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Cascading.Multitool is a sed and grep command line tool for Apache Hadoop.☆21May 1, 2012Updated 14 years ago
- Stand-alone ANSI SQL for Cascading on Apache Hadoop☆48Jan 25, 2018Updated 8 years ago
- Command-line utilities for data analysis.☆18Mar 4, 2011Updated 15 years ago
- ☆10Feb 10, 2017Updated 9 years ago
- EMI Music Hackathon entry☆23Jul 30, 2012Updated 13 years ago
- Lemur is a tool to launch hadoop jobs locally or on EMR, based on a configuration file, referred to as a jobdef. The jobdef file describe…☆84Oct 16, 2017Updated 8 years ago
- Uses TF-IDF and inverted search to cluster search results☆22Mar 10, 2011Updated 15 years ago