Spark Terasort
☆121Apr 21, 2023Updated 3 years ago
Alternatives and similar repositories for spark-terasort
Users that are interested in spark-terasort are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Splash, a flexible Spark shuffle manager that supports user-defined storage backends for shuffle data storage and exchange☆131Dec 19, 2024Updated last year
- Use the TPC-DS benchmark to test Spark SQL performance☆185Apr 27, 2020Updated 6 years ago
- Various tools to help plan HDP and CDH upgrades to CDP☆14Dec 7, 2021Updated 4 years ago
- This is archive of SparkRDMA project. The new repository with RDMA shuffle acceleration for Apache Spark is here: https://github.com/Nvid…☆258May 13, 2019Updated 7 years ago
- TeraSort for Spark and Flink which uses a range partitioner based on sampling☆22Feb 5, 2016Updated 10 years ago
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- HiBench is a big data benchmark suite.☆1,485Dec 15, 2025Updated 9 months ago
- Artifact package for CBMM paper (ATC'22)☆11Jun 5, 2022Updated 4 years ago
- Benchmark Suite for Apache Spark☆242Apr 12, 2023Updated 3 years ago
- All the things about TPC-DS in Apache Spark☆110Jun 15, 2023Updated 3 years ago
- ☆14Mar 29, 2019Updated 7 years ago
- Additional useful algorithms that can be used with spark.☆24Dec 24, 2014Updated 11 years ago
- Parquet file generator☆22Apr 17, 2018Updated 8 years ago
- Spark* Shuffle plugin for support shuffling through remote persistent memory over fabrics, which leverages the RDMA network and remote pe…☆14Sep 18, 2023Updated 3 years ago
- MapReduce performance testing using teragen and terasort☆19Aug 26, 2021Updated 5 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- MLeap demo repository for use with MLeap blog posts☆11Jul 13, 2016Updated 10 years ago
- Spark cloud integration: tests, cloud committers and more☆20Jan 30, 2025Updated last year
- Alerting and monitoring tool for Apache Spark☆23May 20, 2022Updated 4 years ago
- Benchmarking suite for Apache Spark☆16Nov 24, 2017Updated 8 years ago
- Elastic ephemeral storage☆125Apr 12, 2022Updated 4 years ago
- A Study of Database Performance Sensitivity to Experiment Settings☆11May 31, 2022Updated 4 years ago
- 剥离的模块,用于查看Spark SQL生成的语法树☆90May 26, 2019Updated 7 years ago
- Mirror of Apache livy (Incubating)☆13Feb 8, 2024Updated 2 years ago
- Dr. Elephant is a job and flow-level performance monitoring and tuning tool for Apache Hadoop and Apache Spark☆1,370Aug 22, 2023Updated 3 years ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Grizzly: Efficient Stream Processing Through Adaptive Query Compilation☆17Jun 13, 2020Updated 6 years ago
- Fast I/O plugins for Spark☆42Dec 14, 2020Updated 5 years ago
- CoRM: Compactable Remote Memory over RDMA☆20Jun 18, 2021Updated 5 years ago
- NetEase Spark Courses☆15Sep 4, 2018Updated 8 years ago
- Bridging Immutable and Mutable Abstractions for Distributed Data Analytics☆12May 15, 2019Updated 7 years ago
- An Ansible collection for Cloudera Platform for on-premise and cloud Datahubs☆39Aug 26, 2025Updated last year
- Application to securely map users on a multi tenant Amazon EMR cluster to different IAM Roles and then assume the mapped Role.☆24Oct 24, 2023Updated 2 years ago
- Scripts to analyze Spark's performance☆136May 20, 2018Updated 8 years ago
- Apache Celeborn is an elastic and high-performance service for shuffle and spilled data.☆1,065Updated this week
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆39Oct 13, 2020Updated 5 years ago
- ESPBench - The Enterprise Stream Processing Benchmark☆15Dec 27, 2023Updated 2 years ago
- Neutron plugins for Ironic/Neutron integration. Mirror of code maintained at opendev.org.☆11Updated this week
- Albis: High-Performance File Format for Big Data Systems☆21Jul 12, 2018Updated 8 years ago
- Uniffle is a high performance, general purpose Remote Shuffle Service.☆454Sep 10, 2026Updated last week
- A Multiplatform benchmark designed to provide holistic, detailed and close-to-hardware view of memory system performance with family of b…☆46Oct 15, 2025Updated 11 months ago
- ☆21Sep 20, 2016Updated 9 years ago