Suite of parallel iterative algorithms built on top of Iterative Reduce
☆111Jun 24, 2014Updated 12 years ago
Alternatives and similar repositories for Metronome
Users that are interested in Metronome are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Data science repo to help others☆12Feb 10, 2016Updated 10 years ago
- Parallel Iterative Algorithm (SGD) on Hadoop's YARN framework☆42Jan 30, 2013Updated 13 years ago
- The Deep Learning training framework on Spark☆219May 3, 2025Updated last year
- Gust is a set of GPU extensions for Breeze.☆32Apr 10, 2015Updated 11 years ago
- A compiler and runtime for Google's Sawzall language, optimized for Hadoop☆43Apr 26, 2013Updated 13 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- scalding powered machine learning☆109Nov 18, 2014Updated 11 years ago
- DEPRECATED! Use https://github.com/h2oai/sparkling-water repository! H2O and Spark interoperability based on Tachyon.☆44Nov 25, 2014Updated 11 years ago
- CDAP Cube Dataset Guide☆11Aug 26, 2017Updated 9 years ago
- Scalable Machine Learning in Scalding☆358Feb 16, 2018Updated 8 years ago
- Mavuno: A Hadoop-Based Text Mining Toolkit☆48Feb 7, 2012Updated 14 years ago
- The fast and fun way to write YARN applications.☆135Nov 14, 2018Updated 7 years ago
- Load testing tools for Flume☆18Jun 22, 2012Updated 14 years ago
- Mahout vector encoding for pig☆53Nov 20, 2022Updated 3 years ago
- A Python wrapper for MADlib(http://madlib.net) - an open source library for scalable in-database machine learning algorithms☆63Nov 18, 2020Updated 5 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- Experimental logistic regression code supporting multiple result categories, many levels of categorical modeling variables, good optimiza…☆38Oct 14, 2020Updated 5 years ago
- Machine Learning for Cascading☆85Jun 12, 2015Updated 11 years ago
- Behemoth is an open source platform for large scale document analysis based on Apache Hadoop.☆282Apr 25, 2018Updated 8 years ago
- The Nak Machine Learning Library☆342Jul 18, 2017Updated 9 years ago
- Machine learning and natural language processing with Apache Pig☆53Dec 17, 2013Updated 12 years ago
- REST based interface for PIG execution☆25Dec 13, 2021Updated 4 years ago
- Package for Apache Pig support in Sublime Text 2☆16Apr 16, 2012Updated 14 years ago
- Factorization Machines for Julia☆11Aug 26, 2016Updated 10 years ago
- ☆20Jun 26, 2017Updated 9 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆23Feb 22, 2024Updated 2 years ago
- Reactive Factorization Engine☆105Feb 18, 2015Updated 11 years ago
- Easily calculates and explores common Apache Cassandra heap pressure issues.☆18Aug 21, 2014Updated 12 years ago
- Distributed Deep Learning on Spark☆403Oct 8, 2016Updated 9 years ago
- A Clojure DSL for Storm/Trident☆177Apr 3, 2017Updated 9 years ago
- A Hadoop toolkit for web-scale information retrieval research☆86Dec 12, 2014Updated 11 years ago
- A PL/Java Wrapper on Ark-Tweet-NLP (http://www.ark.cs.cmu.edu/TweetNLP/) - Twitter Parts-of-speech tagger in Postgres/Greenplum☆17Jul 25, 2014Updated 12 years ago
- distributed latent dirichlet allocation☆29Dec 15, 2011Updated 14 years ago
- Cassandra state implementation for Twitter Storm Trident API☆17Jan 21, 2013Updated 13 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Large scale k-nn experiments☆69Jul 31, 2024Updated 2 years ago
- Code and configuration files from my tutorial on Big Data infrastructure (VM, EC2, etc.) and RHadoop's rmr package☆39Sep 24, 2012Updated 14 years ago
- A Hivemall wrapper for Spark☆31Apr 21, 2016Updated 10 years ago
- SAMOA (Scalable Advanced Massive Online Analysis) is an open-source platform for mining big data streams.☆427Mar 28, 2016Updated 10 years ago
- An API for Distributed Machine Learning☆156Sep 22, 2016Updated 10 years ago
- Simplifying robust end-to-end machine learning on Apache Spark.☆472Apr 18, 2017Updated 9 years ago
- ☆29Nov 17, 2014Updated 11 years ago