Python MapReduce library written in Cython. Visit us in #hadoopy on freenode. See the link below for documentation and tutorials.
☆243Jan 8, 2016Updated 10 years ago
Alternatives and similar repositories for hadoopy
Users that are interested in hadoopy are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Example code for "Web-Scale Computer Vision using MapReduce for Multimedia Data Mining"☆48Aug 2, 2010Updated 15 years ago
- Python module that allows one to easily write and run Hadoop programs.☆1,030Jan 9, 2018Updated 8 years ago
- A collection of classifiers with a standardized interface. Has a HTTP server interface that allows any language to access.☆18Jan 25, 2012Updated 14 years ago
- Computer vision in the cloud: CV + ML + Hadoop + HBase + REST.☆42Feb 21, 2014Updated 12 years ago
- A Python client for the HBase Avro interface☆50Feb 1, 2016Updated 10 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Run MapReduce jobs on Hadoop or Amazon Web Services☆2,610Apr 2, 2026Updated 3 months ago
- Image Feature Descriptors☆26Feb 15, 2013Updated 13 years ago
- A fast Python module for dealing with so called "typed bytes".☆15Mar 24, 2015Updated 11 years ago
- Parallel Algorithms in Python for Hadoop/Mapreduce☆55Aug 10, 2012Updated 13 years ago
- Python library for similarity search on text data (such as web pages). Currently intended primarily for pedagogical purposes.☆14Oct 8, 2011Updated 14 years ago
- ☆17Mar 6, 2012Updated 14 years ago
- Utilities to use Avro files from Hadoop Map/Reduce jobs and Streaming☆26Sep 10, 2013Updated 12 years ago
- A Python wrapper for Cascading☆220Dec 30, 2019Updated 6 years ago
- A reporistory of User-defined functions for Apache Pig☆16Sep 20, 2010Updated 15 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- a Map/Reduce framework for distributed computing☆1,631Jan 30, 2018Updated 8 years ago
- Lightning-fast cluster computing in Java, Scala and Python.☆1,419Apr 8, 2014Updated 12 years ago
- Python Client for WebHDFS REST API☆43May 8, 2015Updated 11 years ago
- Gisting is an open source Ruby implementation of Google's MapReduce programming paradigm☆18Oct 28, 2011Updated 14 years ago
- The Colossal Pipe framework for map/reduce processing.☆29Aug 19, 2014Updated 11 years ago
- Social Graph Analysis using Elastic MapReduce and PyPy☆55May 4, 2011Updated 15 years ago
- Development in Shark has been ended.☆992Aug 11, 2015Updated 10 years ago
- RHadoop☆760Nov 24, 2015Updated 10 years ago
- Refactored version of code.google.com/hadoop-gpl-compression for hadoop 0.20☆548Apr 24, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- MessagePack MessageQueue without AMQP☆15Mar 31, 2013Updated 13 years ago
- Python connector for ElasticSearch - the pythonic way to use ElasticSearch☆605Oct 13, 2021Updated 4 years ago
- Apache Pig utilities to build training corpora for machine learning / NLP out of public Wikipedia and DBpedia dumps.☆163Nov 8, 2022Updated 3 years ago
- Applied Parallel Computing tutorial material for PyCon 2013 (Minesh Amin, Ian Ozsvald)☆17Apr 2, 2013Updated 13 years ago
- Drivers and libraries for the Xbox Kinect device on WIndows, Linux, and OS X☆36Apr 14, 2011Updated 15 years ago
- Hadoop (Utilities, Patches and Examples)☆241Jun 21, 2016Updated 10 years ago
- r³ is a map-reduce engine written in python using redis as a backend☆346Sep 7, 2012Updated 13 years ago
- an implemetation of LDA in Python, from Heinrich's paper : http://www.arbylon.net/publications/text-est.pdf☆43Feb 13, 2010Updated 16 years ago
- ☆74Jun 18, 2013Updated 13 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Using Hadoop with Scala☆70Oct 5, 2013Updated 12 years ago
- A Python MapReduce and HDFS API for Hadoop☆241Jan 19, 2026Updated 6 months ago
- A Chinese Words Segmentation Tool Based on Bayes Model☆79Jun 21, 2013Updated 13 years ago
- Example code for running R on Hadoop☆132Oct 17, 2012Updated 13 years ago
- This software demonstrates one way to create and manage a cluster of Hadoop nodes running on Google Compute Engine.☆29Sep 23, 2015Updated 10 years ago
- Redis bulk-loader for Apache Pig☆40Apr 21, 2012Updated 14 years ago
- IMPORTANT: Data Brewery is now Bubbles: https://github.com/stiivi/bubbles This brewery repository is NOT MAINTAINED any more.☆133Jul 17, 2013Updated 13 years ago