Dione - a Spark and HDFS indexing library
☆53Mar 26, 2026Updated 5 months ago
Alternatives and similar repositories for dione
Users that are interested in dione are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Spark-Radiant is Apache Spark Performance and Cost Optimizer☆25Dec 31, 2024Updated last year
- Spark-Dashboard is an open-source monitoring solution for Apache Spark that provides real-time performance dashboards using containers an…☆140Sep 18, 2026Updated last week
- Clink is a library that provides APIs and infrastructure to facilitate the development of parallelizable feature engineering operators th…☆30Feb 21, 2022Updated 4 years ago
- A Spark UI and Spark History Server alternative with CPU and Memory metrics! Delight is free, cross-platform, and open-source.☆344May 31, 2024Updated 2 years ago
- Port of TPC-DS dsdgen to Java☆23Jun 23, 2026Updated 3 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Spark to Tableau Extractor library☆19Oct 23, 2017Updated 8 years ago
- An open source indexing subsystem that brings index-based query acceleration to Apache Spark™ and big data workloads.☆430Jan 14, 2022Updated 4 years ago
- Tasks API for Stateful Functions on Flink☆13Aug 22, 2026Updated last month
- HDFS based on Java implementation as a remote ObjectStore for DataFusion☆10Feb 13, 2024Updated 2 years ago
- A Java connector for delta.io/sharing/ that allows you to easily ingest data on any JVM.☆15Updated this week
- Spark metrics related custom classes and sinks (e.g. Prometheus)☆186Aug 2, 2022Updated 4 years ago
- Queryable Window Example☆10Mar 9, 2016Updated 10 years ago
- Data Sketches for Apache Spark☆22Dec 22, 2022Updated 3 years ago
- ☆10Mar 25, 2021Updated 5 years ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Lighthouse is a library for data lakes built on top of Apache Spark. It provides high-level APIs in Scala to streamline data pipelines an…☆64Sep 6, 2024Updated 2 years ago
- Apache Spark Sentry Integration☆16Aug 13, 2021Updated 5 years ago
- Amundsen Gremlin☆22Sep 4, 2026Updated 3 weeks ago
- Activator showing integration of AspectJ to monitor the ActorSystem☆14May 31, 2017Updated 9 years ago
- Secure HDFS Access from Kubernetes☆61Jun 11, 2020Updated 6 years ago
- kafka-manager in Docker container☆19Dec 23, 2020Updated 5 years ago
- This component acts as a bridge between Spark and Vertica, allowing the user to either retrieve data from Vertica for processing in Spark…☆23Aug 2, 2026Updated last month
- Spark SQL index for Parquet tables☆134May 6, 2021Updated 5 years ago
- ☆18Feb 19, 2026Updated 7 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Native SQL Engine plugin for Spark SQL with vectorized SIMD optimizations.☆255Feb 21, 2023Updated 3 years ago
- Efficient, distributed downloads of large files from S3 to HDFS using Spark.☆17Apr 26, 2017Updated 9 years ago
- Shunting Yard is a real-time data replication tool that copies data between Hive Metastores.☆20Oct 11, 2021Updated 4 years ago
- Ansible playbook for automated HDP 2.x deployment install with Kerberos☆19Sep 8, 2016Updated 10 years ago
- Kafka support for Azure Schema Registry.☆17Jun 6, 2025Updated last year
- An Erlang ingester for GreptimeDB, which is compatible with GreptimeDB protocol and lightweight.☆16Updated this week
- [ARCHIVED] Moved to github.com/NVIDIA/spark-xgboost-examples☆72Jul 15, 2020Updated 6 years ago
- ☆38Sep 18, 2026Updated last week
- ☆11Aug 9, 2021Updated 5 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Example for simple Apache Arrow Flight service with Apache Spark and TensorFlow clients☆37Mar 9, 2021Updated 5 years ago
- Blazing-fast query execution engine speaks Apache Spark language and has Arrow-DataFusion at its core.☆11Apr 23, 2022Updated 4 years ago
- ☆23May 2, 2024Updated 2 years ago
- A heterogeneous Apache Spark framework.☆19Mar 2, 2017Updated 9 years ago
- Management and automation platform for Stateful Distributed Systems☆113Jun 17, 2026Updated 3 months ago
- An Extensible Data Skipping Framework☆50Jul 15, 2025Updated last year
- Load watcher is a cluster-wide aggregator of metrics, developed for Trimaran: Real Load Aware Scheduler in Kubernetes.☆78Jul 5, 2026Updated 2 months ago