example code for "Large-scale social media analysis with Hadoop" tutorial presented at ICWSM 2010
☆42Jul 16, 2010Updated 16 years ago
Alternatives and similar repositories for icwsm2010_tutorial
Users that are interested in icwsm2010_tutorial are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Optimized joins using bloom filters on Hadoop via Cascading.☆22Sep 25, 2009Updated 16 years ago
- Where 2.0 Workshop Code: Spatial Analysis of Tweets using Hadoop, Pig, Python & Mechanical Turk. Slides here: http://www.slideshare.net/…☆134Mar 31, 2010Updated 16 years ago
- Tuple MapReduce for Hadoop: Hadoop API made easy☆57Jun 27, 2022Updated 4 years ago
- Distributed database specialized in exporting key/value data from Hadoop☆559Jun 27, 2014Updated 12 years ago
- A program to use MapReduce and Graph Theory to efficiently and scalably find all words in a Boggle roll.☆16Sep 5, 2013Updated 13 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Crawling and analyzing data on Wikipedia☆17Mar 8, 2024Updated 2 years ago
- Code and slides in support of Data Bootcamp tutorial at Strata Conference 2011☆94Feb 1, 2011Updated 15 years ago
- ☆15Dec 14, 2010Updated 15 years ago
- A simple easy to use Hadoop map reduce workflow engine☆18Mar 30, 2012Updated 14 years ago
- John Langford's original release of Vowpal Wabbit -- a fast online learning algorithm☆57Aug 1, 2024Updated 2 years ago
- ☆10Jan 3, 2016Updated 10 years ago
- Zohmg is a data store for aggregation of multi-dimensional time series data, built on top of Hadoop, Dumbo and HBase.☆173Oct 16, 2012Updated 13 years ago
- A simple implementation of k-means clustering on the Spark cluster computing framework. See http://cs.berkeley.edu/~matei/spark.☆26Apr 9, 2011Updated 15 years ago
- Simulated large clusters for Kubernetes scheduler validation.☆15Jan 3, 2023Updated 3 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Example code for running R on Hadoop☆132Oct 17, 2012Updated 13 years ago
- Multichain MCMC framework and algorithms based on PyMC.☆17Feb 14, 2011Updated 15 years ago
- Fitting stochastic blockmodels to graphs☆17Jul 8, 2016Updated 10 years ago
- ChimpMARK-2010 is a collection of massive real-world datasets, interesting real-world problems, and simple example code to solve them. Le…☆17May 31, 2012Updated 14 years ago
- Hadoop Inroduction Presentation Demos☆22Jul 27, 2013Updated 13 years ago
- Probabilistic Data Structures in Python (originally presented at PyData 2013)☆55Jan 6, 2022Updated 4 years ago
- hive storage handler for connecting with MongoDB☆33Apr 14, 2023Updated 3 years ago
- GoldenOrb is an open-source implementation of Pregel, Google's graph processing framework☆293Jun 29, 2022Updated 4 years ago
- Rails app for tracking trends in server logs - powered by the Cloudera Hadoop Distribution on EC2☆358Aug 1, 2011Updated 15 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- This Project aims to implement an unofficial Android client for the service at https://getamen.com☆16Aug 11, 2012Updated 14 years ago
- Guide to Recommender Systems☆14Feb 24, 2012Updated 14 years ago
- Codec for Hadoop adding OpenPGP encryption using Bouncy Castle☆17Aug 18, 2011Updated 15 years ago
- S4 is a general-purpose, distributed, scalable, partially fault-tolerant, pluggable platform that allows programmers to easily develop ap…☆233Mar 4, 2011Updated 15 years ago
- Cascading is a feature rich API for defining and executing complex and fault tolerant data processing flows locally or on a cluster.☆355Apr 8, 2025Updated last year
- Neo4j POC to Integrate VisualSearch.js and Cypher☆18May 31, 2016Updated 10 years ago
- Emacs mode for Pig☆25Mar 8, 2023Updated 3 years ago
- Document clustering based on Latent Semantic Analysis☆96Apr 29, 2010Updated 16 years ago
- Marmalade is a fruit preserve made from the juice and peel of citrus fruits, boiled with sugar and water.☆16Aug 10, 2012Updated 14 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Text clustering service for the web☆25Mar 30, 2019Updated 7 years ago
- (DEPRECATED, migrated to main repo - hasktorch/hasktorch) Research code generation / FFI binding using libtorch 1.x for the next Hasktor…☆11Sep 13, 2019Updated 6 years ago
- Set of example algorithm implementations focused on statistics and machine learning☆30Apr 11, 2011Updated 15 years ago
- An R wrapper to the infochimps.com APIs☆38Mar 22, 2011Updated 15 years ago
- Rails REST web service and dashboard UI for launching MPI clusters on Amazon EC2 and running user submitted jobs☆47Jun 18, 2009Updated 17 years ago
- Elastic MapReduce instance optimizer☆30Mar 24, 2023Updated 3 years ago
- General information and docs about Crosscloud☆18Oct 30, 2014Updated 11 years ago