example code for "Large-scale social media analysis with Hadoop" tutorial presented at ICWSM 2010
☆42Jul 16, 2010Updated 16 years ago
Alternatives and similar repositories for icwsm2010_tutorial
Users that are interested in icwsm2010_tutorial are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Optimized joins using bloom filters on Hadoop via Cascading.☆22Sep 25, 2009Updated 16 years ago
- Where 2.0 Workshop Code: Spatial Analysis of Tweets using Hadoop, Pig, Python & Mechanical Turk. Slides here: http://www.slideshare.net/…☆134Mar 31, 2010Updated 16 years ago
- Tuple MapReduce for Hadoop: Hadoop API made easy☆57Jun 27, 2022Updated 4 years ago
- Re-Designing the classic email client.☆15Jul 25, 2012Updated 14 years ago
- A simple Python class for running multiple URL fetches in parallel☆40Jun 25, 2011Updated 15 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Distributed database specialized in exporting key/value data from Hadoop☆559Jun 27, 2014Updated 12 years ago
- Code and slides in support of Data Bootcamp tutorial at Strata Conference 2011☆94Feb 1, 2011Updated 15 years ago
- ☆15Dec 14, 2010Updated 15 years ago
- A simple easy to use Hadoop map reduce workflow engine☆18Mar 30, 2012Updated 14 years ago
- John Langford's original release of Vowpal Wabbit -- a fast online learning algorithm☆57Aug 1, 2024Updated 2 years ago
- An example project that demontrates real time big data stream processing using GigaSpaces☆19Feb 26, 2022Updated 4 years ago
- Zohmg is a data store for aggregation of multi-dimensional time series data, built on top of Hadoop, Dumbo and HBase.☆173Oct 16, 2012Updated 13 years ago
- A simple implementation of k-means clustering on the Spark cluster computing framework. See http://cs.berkeley.edu/~matei/spark.☆26Apr 9, 2011Updated 15 years ago
- Multichain MCMC framework and algorithms based on PyMC.☆17Feb 14, 2011Updated 15 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Hadoop Inroduction Presentation Demos☆22Jul 27, 2013Updated 13 years ago
- Probabilistic Data Structures in Python (originally presented at PyData 2013)☆55Jan 6, 2022Updated 4 years ago
- hive storage handler for connecting with MongoDB☆33Apr 14, 2023Updated 3 years ago
- GoldenOrb is an open-source implementation of Pregel, Google's graph processing framework☆293Jun 29, 2022Updated 4 years ago
- Rails app for tracking trends in server logs - powered by the Cloudera Hadoop Distribution on EC2☆359Aug 1, 2011Updated 15 years ago
- Implementation of Tyler Neylon's Locality-Specific Hash based on simplex tesselations☆28Oct 15, 2011Updated 14 years ago
- This Project aims to implement an unofficial Android client for the service at https://getamen.com☆16Aug 11, 2012Updated 14 years ago
- Data-driven modeling course, Applied Mathematics, Columbia University☆20Apr 23, 2012Updated 14 years ago
- Guide to Recommender Systems☆14Feb 24, 2012Updated 14 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Codec for Hadoop adding OpenPGP encryption using Bouncy Castle☆17Aug 18, 2011Updated 15 years ago
- S4 is a general-purpose, distributed, scalable, partially fault-tolerant, pluggable platform that allows programmers to easily develop ap…☆233Mar 4, 2011Updated 15 years ago
- Cascading is a feature rich API for defining and executing complex and fault tolerant data processing flows locally or on a cluster.☆355Apr 8, 2025Updated last year
- Neo4j POC to Integrate VisualSearch.js and Cypher☆18May 31, 2016Updated 10 years ago
- Document clustering based on Latent Semantic Analysis☆96Apr 29, 2010Updated 16 years ago
- Marmalade is a fruit preserve made from the juice and peel of citrus fruits, boiled with sugar and water.☆16Aug 10, 2012Updated 14 years ago
- Python module that allows one to easily write and run Hadoop programs.☆1,030Jan 9, 2018Updated 8 years ago
- Text clustering service for the web☆25Mar 30, 2019Updated 7 years ago
- An R wrapper to the infochimps.com APIs☆38Mar 22, 2011Updated 15 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Set of example algorithm implementations focused on statistics and machine learning☆30Apr 11, 2011Updated 15 years ago
- Rails REST web service and dashboard UI for launching MPI clusters on Amazon EC2 and running user submitted jobs☆47Jun 18, 2009Updated 17 years ago
- Extracts A Social Network From Cassandra NoSQL Data-store To The InfiniteGraph Graph Database For Analysis☆16Aug 26, 2010Updated 15 years ago
- General information and docs about Crosscloud☆18Oct 30, 2014Updated 11 years ago
- Open source framework for predictive modeling on Apache Hadoop☆35Aug 23, 2014Updated 11 years ago
- convert weibo(sina/tencent/netease) data source into an intermediate format supported by citespace☆10Sep 27, 2011Updated 14 years ago
- Python MCMC☆13Apr 18, 2011Updated 15 years ago