Optimized communication collectives for the Cerebras waferscale engine
☆18Jun 5, 2024Updated 2 years ago
Alternatives and similar repositories for spatial-collectives
Users that are interested in spatial-collectives are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- WaferLLM: Large Language Model Inference at Wafer Scale☆121Jun 12, 2026Updated 3 months ago
- Official implementation of the ICLR'25 paper "QERA: an Analytical Framework for Quantization Error Reconstruction".☆14Feb 4, 2025Updated last year
- ☆11Nov 14, 2022Updated 3 years ago
- ☆17Apr 8, 2021Updated 5 years ago
- some tlb experimentation code: calculate L1, L2 miss penalties and show cross-HT interference.☆15Aug 30, 2019Updated 7 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- a robust metric (robust fidelity) for XGNN (ICLR24)☆12Jun 3, 2025Updated last year
- Running ahead of memory latency - Part II project☆10Jan 7, 2023Updated 3 years ago
- Artifacts for our ShowTime paper (AsiaCCS '23), including distinguishing cache hits and misses with the human eye.☆14Jul 21, 2023Updated 3 years ago
- Official implementation of ICML'24 paper "LQER: Low-Rank Quantization Error Reconstruction for LLMs"☆19Jul 11, 2024Updated 2 years ago
- ☆11Mar 17, 2021Updated 5 years ago
- ☆14May 26, 2022Updated 4 years ago
- ☆13Feb 16, 2023Updated 3 years ago
- C implementation of the PageRank algorithm, with and without parallelization.☆16Jun 21, 2016Updated 10 years ago
- A repo to store what M1 explainer does and ~Armv9~ Armv8.6 ISA on A16☆19Sep 28, 2024Updated last year
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Benchmark for memory store throughput☆23Jun 22, 2021Updated 5 years ago
- Tool for inferring cache replacement policies with automata learning. Uses LearnLib and Sketch.☆16Apr 21, 2020Updated 6 years ago
- ☆18Jun 26, 2025Updated last year
- This upload contains the artifacts for the paper "SLAP: Data Speculation Attacks via Load Address Prediction on Apple Silicon", to appear…☆26Jan 26, 2025Updated last year
- SLiM: One-shot Quantized Sparse Plus Low-rank Approximation of LLMs (ICML 2025)☆37Nov 28, 2025Updated 9 months ago
- Heron: Automatically Constrained High-Performance Library Generation for Deep Learning Accelerators☆24Jan 30, 2024Updated 2 years ago
- Distributed SDDMM Kernel☆13Jul 8, 2022Updated 4 years ago
- AutoTCL and Parametric Augmentation for Time Series Contrastive Learning(ICLR2024)☆25Mar 24, 2024Updated 2 years ago
- ☆31Feb 20, 2024Updated 2 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Scripts for fine-tuning an HPC Code LLM☆17Jul 19, 2024Updated 2 years ago
- Fast GPU error-bounded lossy compressor for floating-point data.☆74Jun 10, 2026Updated 3 months ago
- Rage Against The Machine Clear: A Systematic Analysis of Machine Clears and Their Implications for Transient Execution Attacks☆25Jun 11, 2021Updated 5 years ago
- Benchmark tests supporting the TiledCUDA library.☆19Nov 19, 2024Updated last year
- ☆20Sep 25, 2023Updated 2 years ago
- ☆12Sep 18, 2024Updated 2 years ago
- ☆26Oct 6, 2023Updated 2 years ago
- Streamline Covert Channel Attack (presented in ASPLOS'21)☆22Feb 18, 2021Updated 5 years ago
- FuseFL: One-Shot Federated Learning through the Lens of Causality with Progressive Model Fusion (NeurIPS 2024 Spotlight)☆15Mar 31, 2025Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Tutorials of Extending and importing TVM with CMAKE Include dependency.☆16Oct 11, 2024Updated last year
- HW interface for memory caches☆28Apr 21, 2020Updated 6 years ago
- TensorRT-in-Action 是一个 GitHub 代码库,提供了使用 TensorRT 的代码示例,并有对应 Jupyter Notebook。☆15Jun 1, 2023Updated 3 years ago
- Networking Template Library for Vivado HLS☆30Jul 12, 2020Updated 6 years ago
- ☆15Jun 19, 2025Updated last year
- ☆20Sep 28, 2024Updated last year
- Code repo for "CritiPrefill: A Segment-wise Criticality-based Approach for Prefilling Acceleration in LLMs".☆17Sep 15, 2024Updated 2 years ago