Mavuno: A Hadoop-Based Text Mining Toolkit
☆48Feb 7, 2012Updated 14 years ago
Alternatives and similar repositories for mavuno
Users that are interested in mavuno are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- A Hadoop input format for sending lists of files as keys to a mapper. Set the list of files, and an input split will be created per file…☆16Apr 7, 2010Updated 16 years ago
- Behemoth is an open source platform for large scale document analysis based on Apache Hadoop.☆282Apr 25, 2018Updated 8 years ago
- A Hadoop toolkit for web-scale information retrieval research☆86Dec 12, 2014Updated 11 years ago
- Mahout vector encoding for pig☆53Nov 20, 2022Updated 3 years ago
- What happens on the wire when Hadoop RPC call is issued?☆13Jul 1, 2022Updated 4 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Cloud9 is a Hadoop toolkit for working with big data☆237Dec 15, 2015Updated 10 years ago
- Set of Hadoop, Spark and Storm based tools for web and customer analytic☆34Jun 7, 2021Updated 5 years ago
- Twitter Tools☆222Feb 18, 2018Updated 8 years ago
- Hadoop for archiving email☆22Sep 29, 2011Updated 14 years ago
- Utilities to use Avro files from Hadoop Map/Reduce jobs and Streaming☆26Sep 10, 2013Updated 13 years ago
- Machine learning and natural language processing with Apache Pig☆53Dec 17, 2013Updated 12 years ago
- Parallelizing Machine Learning-- Functionally.☆56Jun 14, 2012Updated 14 years ago
- Hadoop library for large-scale data processing, now an Apache Incubator project☆581Jul 8, 2014Updated 12 years ago
- Zookeeper management helpers for importing, exporting and purging of zk data☆17Nov 4, 2015Updated 10 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- A Hadoop job that runs GATE applications☆15Oct 16, 2013Updated 12 years ago
- Hadoop utility to quickly find large directories to clean up or small files to combine.☆15Jan 12, 2012Updated 14 years ago
- ☆19Mar 24, 2022Updated 4 years ago
- Open source framework for predictive modeling on Apache Hadoop☆35Aug 23, 2014Updated 12 years ago
- ☆50Sep 3, 2019Updated 7 years ago
- Search and Recommender Systems papers with Code☆21Nov 17, 2018Updated 7 years ago
- Apache Pig utilities to build training corpora for machine learning / NLP out of public Wikipedia and DBpedia dumps.☆163Nov 8, 2022Updated 3 years ago
- Large scale k-nn experiments☆69Jul 31, 2024Updated 2 years ago
- Twitter's fork of Apache Mahout (we intend to push changes upstream)☆62Jul 11, 2013Updated 13 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Suite of parallel iterative algorithms built on top of Iterative Reduce☆111Jun 24, 2014Updated 12 years ago
- Speaker Identity for Topic Segmentation (SITS)☆13Dec 14, 2014Updated 11 years ago
- Hadoop Inroduction Presentation Demos☆22Jul 27, 2013Updated 13 years ago
- NERD Machine Learner☆15Apr 11, 2013Updated 13 years ago
- EMI Music Hackathon entry☆23Jul 30, 2012Updated 14 years ago
- iSAX Indexing persisted in HBase☆39Jul 26, 2011Updated 15 years ago
- A system to generate SPARQL queries from natural language queries.☆30Feb 15, 2025Updated last year
- A number of algorithms for calculating string similarity in Java☆15Jan 23, 2011Updated 15 years ago
- Core libraries by the PRImA Research Lab☆16Jul 30, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Parallel Algorithms in Python for Hadoop/Mapreduce☆55Aug 10, 2012Updated 14 years ago
- Repository for the Apache Drill Workshop☆18Oct 31, 2016Updated 9 years ago
- Offline Recommender System Evaluation for Spark☆28Jul 2, 2017Updated 9 years ago
- Explorations relative to cloning FlumeJava☆94Oct 13, 2020Updated 5 years ago
- Text Analysis: Implementation of ULMFiT by Howard & Ruder on Twitter dataset☆10Feb 7, 2019Updated 7 years ago
- A bunch of utility classes for Java, Hadoop, HBase, Pig, etc.☆77Mar 31, 2014Updated 12 years ago
- A repository for Neural Document Ranking Models.☆82Sep 15, 2018Updated 8 years ago