ArcLight: A Lightweight LLM Inference Framework
☆49May 30, 2026Updated 3 months ago
Alternatives and similar repositories for ArcLight
Users that are interested in ArcLight are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Supplemental materials for The ASPLOS 2025 / EuroSys 2025 Contest on Intra-Operator Parallelism for Distributed Deep Learning☆25May 12, 2025Updated last year
- ThinK: Thinner Key Cache by Query-Driven Pruning☆30Jun 2, 2026Updated 2 months ago
- ☆14Oct 3, 2024Updated last year
- 🛰️ A CLI tool for tracking token usage from OpenCode/Claude Code/Codex/Gemini CLI/Cursor IDE • Generate Your Wrapped 2025 🎉☆20Mar 14, 2026Updated 5 months ago
- ☆13Sep 19, 2024Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- The homepage of OneBit model quantization framework.☆208Feb 5, 2025Updated last year
- TPAMI 2025 Survey Paper☆34Mar 31, 2025Updated last year
- Fast and Flexible FPGA development using Hierarchical Partial Reconfiguration (FPT 2022)☆16Mar 21, 2024Updated 2 years ago
- We release Open Meditron, a fully open, clinician-audited medical training corpus and evaluation protocol that closes the open-vs-closed …☆17Aug 3, 2026Updated 3 weeks ago
- ☆34Jul 13, 2026Updated last month
- Tempo is a system for declarative, efficient, end-to-end compiled dynamic deep learning☆31Oct 21, 2025Updated 10 months ago
- CODO: An Automated Compiler for Comprehensive Dataflow Optimization☆36Jun 3, 2026Updated 2 months ago
- An experimental communicating attention kernel based on DeepEP.☆34Jul 29, 2025Updated last year
- 🫧 Code for Holistic Reasoning with Long-Context LMs: A Benchmark for Database Operations on Massive Textual Data (Maekawa*, Iso* et al.…☆12Feb 25, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A tool allowing students of Coursera's Heterogeneous Parallel Programming to work on homework using a machine without a CUDA GPU.☆11Mar 11, 2015Updated 11 years ago
- hostCC is a congestion control architecture which handles host congestion, along with in-network congestion☆57Updated this week
- Intent as Code | the workflow language for AI. One file, 4 verbs, one Rust binary. Local-first, any model, AGPL-3.0. 🦋☆60Updated this week
- C++ pipeline with OpenVINO native API for Stable Diffusion v1.5☆13Feb 23, 2024Updated 2 years ago
- ASPLOS'24: Optimal Kernel Orchestration for Tensor Programs with Korch☆41Mar 27, 2025Updated last year
- ☆34Mar 12, 2026Updated 5 months ago
- ☆54May 19, 2025Updated last year
- Research prototype of PRISM — a cost-efficient multi-LLM serving system with flexible time- and space-based GPU sharing.☆77Mar 17, 2026Updated 5 months ago
- 小彭老师推出 SyCL 2020 课程(施工中,日后会在直播中放出)☆15Sep 3, 2023Updated 2 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Some CS notes during Jiawei's undergrad.☆33Jan 6, 2022Updated 4 years ago
- 清华大学计算机系零字班计算机组成原理大实验作业。☆35Jan 27, 2023Updated 3 years ago
- Rust library implementing the Toorani-Beheshti signcryption scheme☆13Aug 15, 2023Updated 3 years ago
- A wiki designed for both humans and agents to read and write.☆68Aug 22, 2026Updated last week
- Free, Open-Source Playground for AI-to-AI Conversations☆42Aug 8, 2025Updated last year
- SC 2021, "LogECMem: Coupling Erasure-Coded In-Memory Key-Value Stores with Parity Logging"☆12Jul 12, 2021Updated 5 years ago
- DeeperGEMM: crazy optimized version☆85May 5, 2025Updated last year
- This repository is a read-only mirror of https://gitlab.arm.com/kleidi/kleidiai☆182Updated this week
- [ICPP'25] TD-Pipe: Temporally-Disaggregated Pipeline Parallelism Architecture for High-Throughput LLM Inference☆52Dec 24, 2025Updated 8 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- This is an implementation of paper "End-to-end Speech Translation via Cross-modal Progressive Training" (Interspeech2021)☆18May 1, 2022Updated 4 years ago
- ☆13Dec 11, 2020Updated 5 years ago
- study of Ampere' Sparse Matmul☆18Jan 10, 2021Updated 5 years ago
- ☆18Apr 21, 2024Updated 2 years ago
- TritonBench: Benchmarking Large Language Model Capabilities for Generating Triton Operators☆138Jun 14, 2025Updated last year
- ☆124May 19, 2025Updated last year
- [ACL 2025 main] FR-Spec: Frequency-Ranked Speculative Sampling☆55Jul 15, 2025Updated last year