Interactive Audience Analytics with Spark and HyperLogLog
☆55Oct 14, 2015Updated 10 years ago
Alternatives and similar repositories for spark-hyperloglog
Users that are interested in spark-hyperloglog are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Cantor provides utilities for estimating the cardinality of large sets.☆85Apr 12, 2022Updated 4 years ago
- Spark Extension : ML transformers, SQL aggregations, etc that are missing in Apache Spark☆144Jan 26, 2016Updated 10 years ago
- Starter project for building MemSQL Streamliner Pipelines☆32Apr 18, 2017Updated 9 years ago
- Experiments with the GDELT dataset and Cassandra schemas.☆25Feb 9, 2016Updated 10 years ago
- just some scripts that I use☆27Dec 19, 2012Updated 13 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Embedded Kafka for testing and quick prototyping.☆14Apr 19, 2016Updated 10 years ago
- Examples for Fast Data Processing with Spark☆59Sep 10, 2013Updated 13 years ago
- HDFS compatible Distributed Filesystem backed Cassandra☆25Sep 17, 2015Updated 11 years ago
- Coursera Machine Learning class examples in Spark☆42Feb 14, 2014Updated 12 years ago
- Social Media Data Mining and Analytics - HyperLogLog, BloomFilter and CountMinSketch with Scalding & Algebird☆27Oct 6, 2018Updated 7 years ago
- Locality Sensitive Hashing for Apache Spark☆198Nov 1, 2016Updated 9 years ago
- Sample of resteasy-netty project☆17Jun 25, 2015Updated 11 years ago
- Application that visualizes your google location history in form of a heatmap using Spark to aggregate the data.☆12Feb 19, 2015Updated 11 years ago
- Coding exercises for Apache Spark☆103Jun 4, 2015Updated 11 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- A tool for running Spark on Google Compute Engine☆16Jan 20, 2017Updated 9 years ago
- sparkhello: Scala to Spark - Hello World☆19Jul 12, 2017Updated 9 years ago
- Spark example of collecting tweets and loading into HDFS/S3☆42Oct 2, 2013Updated 12 years ago
- Ansible Role to install a Hadoop Cluster☆10Sep 21, 2020Updated 6 years ago
- Joins for skewed datasets in Spark☆58Aug 18, 2017Updated 9 years ago
- Integration testing with TorqueBox☆17Apr 8, 2015Updated 11 years ago
- Scala stuff☆18Jun 13, 2019Updated 7 years ago
- ☆11Oct 8, 2015Updated 10 years ago
- Philly ETE Reactive APIs talk☆17Aug 26, 2015Updated 11 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Scriptable scheduler for periodical Hadoop workflows☆22Feb 1, 2018Updated 8 years ago
- Collection of Interesting Algorithms☆16Oct 13, 2020Updated 5 years ago
- A Ruby DSL for creating Azkaban jobs using Rake☆18Jul 26, 2013Updated 13 years ago
- Example demonstrating a Scala project that builds using Gradle, produces a shadow jar suitable for spark-submit, and has tests using Scal…☆18Jun 18, 2015Updated 11 years ago
- On demand presto cluster with mesos, marathon and docker.☆29Mar 7, 2018Updated 8 years ago
- Based off the design of SparkOnHBase. This Repo will support Spark, Spark Streaming, and Spark SQL integration with Kudu.☆50May 19, 2016Updated 10 years ago
- Abstract Algebra for Scala☆2,296Nov 21, 2025Updated 10 months ago
- These are some code examples☆56Jan 12, 2020Updated 6 years ago
- Example project to show how to use Spark to read and write Avro/Parquet files☆50Aug 21, 2013Updated 13 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆19Dec 3, 2014Updated 11 years ago
- Sparse feature extraction with Spark☆29Jul 25, 2018Updated 8 years ago
- Automates Spark standalone cluster tasks with Puppet and Fabric.☆42Aug 14, 2014Updated 12 years ago
- Secondary sort and streaming reduce for Apache Spark☆77Jul 3, 2023Updated 3 years ago
- Dockerfile for Apache Zeppelin☆17Dec 9, 2015Updated 10 years ago
- Scala and SQL happy together.☆29Dec 13, 2016Updated 9 years ago
- Notes from 100 days with Kubernetes☆31Jan 25, 2019Updated 7 years ago