Optimized joins using bloom filters on Hadoop via Cascading.
☆22Sep 25, 2009Updated 16 years ago
Alternatives and similar repositories for cascading-batch-query
Users that are interested in cascading-batch-query are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆15Dec 14, 2010Updated 15 years ago
- Tool to help users migrate large relational databases into Hadoop clusters.☆67Mar 23, 2012Updated 14 years ago
- Annotations and Classes for managing and executing dependent processes☆39Apr 15, 2021Updated 5 years ago
- Cascalog functions for Midje.☆20Feb 16, 2013Updated 13 years ago
- A simple dependency management system for your projects.☆46Jan 24, 2011Updated 15 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Cascading is a feature rich API for defining and executing complex and fault tolerant data processing flows locally or on a cluster.☆355Apr 8, 2025Updated last year
- Open source framework for predictive modeling on Apache Hadoop☆35Aug 23, 2014Updated 12 years ago
- Simple bash functions for manipulating Amazon Elastic MapReduce clusters☆45Jan 5, 2016Updated 10 years ago
- Serializer and comparator for using Thrift objects in Cascading or Cascalog☆17Dec 31, 2014Updated 11 years ago
- Seamless integration of ElephantDB with Cascalog☆18Jan 3, 2012Updated 14 years ago
- example code for "Large-scale social media analysis with Hadoop" tutorial presented at ICWSM 2010☆42Jul 16, 2010Updated 16 years ago
- Learn Cascalog with Koans!☆27Mar 15, 2012Updated 14 years ago
- Zohmg is a data store for aggregation of multi-dimensional time series data, built on top of Hadoop, Dumbo and HBase.☆173Oct 16, 2012Updated 13 years ago
- Clojure library for serializing Clojure data using Kryo☆30Feb 12, 2016Updated 10 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- An example project that demontrates real time big data stream processing using GigaSpaces☆19Feb 26, 2022Updated 4 years ago
- Code for Strange Loop talk on Specter☆13Sep 26, 2015Updated 10 years ago
- A small C++ wrapper for Storm☆31May 20, 2013Updated 13 years ago
- ☆11Apr 19, 2018Updated 8 years ago
- Apache Camel component for Beanstalk☆16Sep 21, 2014Updated 11 years ago
- Demonstrate the some of features of gRPC☆14Dec 15, 2019Updated 6 years ago
- Talks at the <Programming> 2022 Conference in Porto, Portugal☆12Mar 30, 2022Updated 4 years ago
- Retrieves Ubuntu AMI information from Canonicals release list☆16May 6, 2022Updated 4 years ago
- A branch of the boilerpipe project☆15Mar 18, 2011Updated 15 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Write tweets from the Twitter streaming API to Hadoop☆15Feb 20, 2010Updated 16 years ago
- Clojure library for serializing Clojure data using Kryo☆59Jan 21, 2012Updated 14 years ago
- YourKit from the REPL☆13Sep 21, 2016Updated 9 years ago
- HBase secondary index using coprocessors☆21Jun 6, 2013Updated 13 years ago
- Neo4j POC to Integrate VisualSearch.js and Cypher☆18May 31, 2016Updated 10 years ago
- A simple Bloom Filter implementation in Java☆16Oct 21, 2012Updated 13 years ago
- My Emacs config☆27Aug 12, 2024Updated 2 years ago
- Text clustering service for the web☆25Mar 30, 2019Updated 7 years ago
- Composable metric reporters in Python.☆14Jun 6, 2024Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Extracts A Social Network From Cassandra NoSQL Data-store To The InfiniteGraph Graph Database For Analysis☆16Aug 26, 2010Updated 16 years ago
- A Redis PubSub Spout for Storm☆37Feb 29, 2012Updated 14 years ago
- Elastic MapReduce instance optimizer☆30Mar 24, 2023Updated 3 years ago
- HBase adapters for Cascading☆47Aug 9, 2009Updated 17 years ago
- Clusterless is a tool for scheduling decentralized, scalable, and secure data pipelines for continuously arriving data, across clouds.☆15Dec 22, 2025Updated 8 months ago
- Neo4j and the Crunchbase API mashup☆19Aug 15, 2013Updated 13 years ago
- Content-addressable storage, implemented over pyfilesystem2.☆17Jun 17, 2020Updated 6 years ago