☆37Sep 26, 2026Updated this week
Alternatives and similar repositories for awesome-papers
Users that are interested in awesome-papers are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Prefix-Aware Attention for LLM Decoding☆48May 26, 2026Updated 4 months ago
- An Open-Source RAG Workload Trace to Optimize RAG Serving Systems☆40Sep 1, 2026Updated 3 weeks ago
- ☆36Jul 27, 2026Updated 2 months ago
- Accommodating Large Language Model Training over Heterogeneous Environment.☆36Mar 13, 2025Updated last year
- 🎓Automatically Update LLM inference systems Papers Daily using Github Actions (Update Every 12th hours)☆12Updated this week
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆18Sep 21, 2025Updated last year
- The official implementation for the intra-stage fusion technique introduced in https://arxiv.org/abs/2409.13221☆32Apr 22, 2025Updated last year
- Paper reading and discussion notes, covering AI frameworks, distributed systems, cluster management, etc.☆75Mar 4, 2026Updated 6 months ago
- Horizontal Fusion☆24Jan 7, 2022Updated 4 years ago
- C++(From TJU)☆13Mar 4, 2025Updated last year
- Curated collection of papers in machine learning systems☆656Sep 15, 2026Updated last week
- My Paper Reading Lists and Notes.☆27Sep 20, 2026Updated last week
- Source code for Trinity(ASPLOS 2026)☆26Apr 24, 2026Updated 5 months ago
- ☆21Mar 17, 2026Updated 6 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆12May 19, 2025Updated last year
- ☆14May 28, 2019Updated 7 years ago
- Daily Arxiv Papers on LLM Systems☆82Updated this week
- ☆30Feb 27, 2025Updated last year
- CuTeDSL tutorials, tools, autotuner, profiler, etc.☆45Jun 27, 2026Updated 3 months ago
- >>> 异常中断 + 虚存页表 + 分支预测 + TLB + Cache + Flash + VGA + uCore☆20Nov 17, 2023Updated 2 years ago
- TPAMI 2025 Survey Paper☆34Mar 31, 2025Updated last year
- Empowering LLM Agents for Real-World Computer System Optimization☆17Sep 10, 2025Updated last year
- Large Language Model (LLM) Systems Paper List☆2,276Jul 25, 2026Updated 2 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆27Oct 1, 2025Updated 11 months ago
- ☆15Jun 26, 2024Updated 2 years ago
- 七夕孤寡助手☆13Aug 7, 2021Updated 5 years ago
- Injecting Adrenaline into LLM Serving: Boosting Resource Utilization and Throughput via Attention Disaggregation☆42Sep 21, 2026Updated last week
- This repository contains multiple implementations of Flash Attention optimized with Triton kernels, showcasing progressive performance im…☆13Mar 26, 2026Updated 6 months ago
- ☆38Nov 28, 2024Updated last year
- Reproducible Language Agent Research☆36Jun 25, 2025Updated last year
- ☆16Jan 7, 2018Updated 8 years ago
- 天津大学unipus帮助☆10Apr 20, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- ☆13Sep 19, 2024Updated 2 years ago
- Artifact Evaluation for SOSP 2025☆22Aug 16, 2025Updated last year
- Boosting GPU utilization for LLM serving via dynamic spatial-temporal prefill & decode orchestration☆53Updated this week
- A Library for intra-GPU/Inter-SM parallelsim☆13Aug 7, 2026Updated last month
- ☆12Aug 18, 2023Updated 3 years ago
- Fast and Flexible FPGA development using Hierarchical Partial Reconfiguration (FPT 2022)☆16Mar 21, 2024Updated 2 years ago
- ☆19Jan 10, 2023Updated 3 years ago