A Ruby toolkit for cloud-friendly ETL
☆37Jul 29, 2016Updated 10 years ago
Alternatives and similar repositories for sluice
Users that are interested in sluice are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Snowplow Request Validator☆10Apr 26, 2017Updated 9 years ago
- cascading.jruby build and execution tool☆16Sep 23, 2015Updated 10 years ago
- Example Scala/SBT event producer for Amazon Kinesis☆21Mar 29, 2015Updated 11 years ago
- UNMAINTAINED. 2013☆22Oct 17, 2013Updated 12 years ago
- A JRuby DSL for Cascading☆16Jan 11, 2015Updated 11 years ago
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- The missing dashboard builder for Snowplow Analytics.☆29Jun 30, 2014Updated 12 years ago
- Generate Redshift schemas from sample objects☆21Apr 25, 2014Updated 12 years ago
- An ETL (Extract-Transform-Load) library that uses a forking process model for concurrency.☆27May 14, 2016Updated 10 years ago
- Ruby-based programmatic access to Amazon's Elastic MapReduce service.☆105Feb 28, 2026Updated 5 months ago
- Example Scala/SBT event consumer for Amazon Kinesis☆22May 20, 2015Updated 11 years ago
- Cascading.Multitool is a sed and grep command line tool for Apache Hadoop.☆21May 1, 2012Updated 14 years ago
- Unix tee, but for Kinesis streams☆12Oct 19, 2021Updated 4 years ago
- NodeJS web analytics data collector for SnowPlow☆38Sep 25, 2014Updated 11 years ago
- JDBC adapter for Cascading☆23Jun 27, 2009Updated 17 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- A set of data structures I find useful built using redis as a backend☆27Mar 19, 2009Updated 17 years ago
- OLD - impyla now developed at `cloudera/impyla`☆23Apr 16, 2014Updated 12 years ago
- Files to help make new spark EMR Bootstraps☆15Aug 4, 2013Updated 13 years ago
- Services for Librato Metrics☆14Apr 22, 2020Updated 6 years ago
- Idiomatic, opinionated Scala library for AWS☆30May 13, 2017Updated 9 years ago
- In-memory grouping, summarization, pivoting, cross-tabulation etc. of datasets☆15Apr 30, 2012Updated 14 years ago
- Spark batch converter to convert AWS S3 server side logs to Parquet file format☆11Mar 24, 2023Updated 3 years ago
- A Serverless project to validate GitHub PRs against some specifications☆14Updated this week
- A utility for coordinating service rollouts☆15Feb 28, 2019Updated 7 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- Google Analytics plugin for sending events to Snowplow☆17Sep 30, 2020Updated 5 years ago
- Framework that makes processing arbitrary binary data in Hadoop easier☆22Apr 8, 2013Updated 13 years ago
- A we analytics and event tracking sleuth JavaScript library☆39Mar 29, 2017Updated 9 years ago
- A library for watching service configuration from various sources on the JVM☆14Jul 15, 2022Updated 4 years ago
- Store batched Kafka messages in S3.☆39Apr 13, 2022Updated 4 years ago
- Python SDK for working with Snowplow enriched events in Spark, AWS Lambda et al.☆22Nov 25, 2024Updated last year
- Declarative transformation of MongoDB documents into SQL tables (trees-to-rectangles)☆56Aug 25, 2015Updated 10 years ago
- A set of tools for working with Omniture daily data files (hit_data.tsv) in big or small tools like Spark, Hadoop or just Python.☆37May 14, 2019Updated 7 years ago
- scala driver for launching Amazon EMR jobs☆40Feb 10, 2016Updated 10 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- A Hive Deserializer for CloudFront access logs (supports download distribution files only)☆17Aug 7, 2012Updated 14 years ago
- Cookie getter/setter for browser in 340 bytes☆18Apr 9, 2018Updated 8 years ago
- Run templatable playbooks of SQL scripts in series and parallel on Redshift, PostgreSQL, BigQuery and Snowflake☆81May 16, 2025Updated last year
- This directory should help everyone who is looking for a tracking and analytics solution☆17Mar 7, 2022Updated 4 years ago
- A dynamic Rack server and helper methods to help testing Rack apps.☆17Jul 27, 2016Updated 10 years ago
- Validate an object against a redshift schema.☆22Mar 28, 2014Updated 12 years ago
- A rough prototype of a tool for discovering Apache Hive schemas from JSON documents.☆43Dec 16, 2023Updated 2 years ago