An End-to-End Benchmarking Framework for Retrieval-Augmented Generation Systems
☆29Mar 13, 2026Updated 5 months ago
Alternatives and similar repositories for RAGPerf
Users that are interested in RAGPerf are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆19Feb 9, 2026Updated 6 months ago
- Prefix-Aware Attention for LLM Decoding☆45May 26, 2026Updated 3 months ago
- An open-source simulator framework for neural processing units☆50Jul 25, 2026Updated last month
- CatRAG is a RAG framework builds on the HippoRAG 2 architecture and transforms the static KG into query-adaptive navigation structure. RA…☆29Aug 20, 2026Updated last week
- TPC-E Benchmark☆14Feb 16, 2016Updated 10 years ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- An FPGA-based full-stack in-storage computing system.☆37Nov 6, 2020Updated 5 years ago
- Clue-RAG: Towards Accurate and Cost-Efficient Graph-based RAG via Multi-Partite Graph and Query-Driven Iterative Retrieval☆26Mar 3, 2026Updated 5 months ago
- The code implementation of HyGRAG, accepted by WWW'26.☆15May 31, 2026Updated 2 months ago
- ☆17Jun 15, 2026Updated 2 months ago
- ☆37Apr 10, 2024Updated 2 years ago
- A guide on how to emulate an NVMe SPDM responder device with QEMU and Linux. Additionally, instructions on setting up and testing the (in…☆11Sep 3, 2024Updated last year
- Open-source MCP server — progressive tool discovery, code execution, intelligent routing & token optimization across 50+ tools☆15Jun 25, 2026Updated 2 months ago
- ☆20Mar 11, 2025Updated last year
- ☆26Jan 10, 2023Updated 3 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- Benchmarks, testbenches, and transformed codes for high-level synthesis research☆14Aug 18, 2017Updated 9 years ago
- [ASPLOS'26] HILOS: A Cost-Effective Near-Storage Processing Solution for Offline Inference of Long-Context LLMs☆22Jan 18, 2026Updated 7 months ago
- A low-cost, high-performance deep learning training framework that enables efficient 100B-scale model fine-tuning on a commodity server w…☆23Mar 21, 2025Updated last year
- DEDISbench: A disk I/O block-based benchmark for deduplication systems. Unlike other existing benchmarks, written content is generated i…☆14Jul 22, 2021Updated 5 years ago
- Agentic RAG Harness for long documents, Tree and Graph based reasoning. Cited answers down to the pixel☆68Apr 15, 2026Updated 4 months ago
- ☆14Jul 13, 2025Updated last year
- A modular, high-performance Retrieval-Augmented Generation framework with multi-path retrieval, graph extraction, and fusion ranking☆48Mar 4, 2026Updated 5 months ago
- PyCon mini 東海 2024 のトーク「Google Colaboratoryで試すVLM」で紹介したサンプル集☆12Nov 15, 2024Updated last year
- Doc Thinker: All-in-One RAG - document parsing, graph RAG, and evaluation☆32Jul 13, 2026Updated last month
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- multicast learning in network programming course☆10Oct 30, 2020Updated 5 years ago
- Convert regular expressions to minimized DFAs in AT&T FST format.☆15Updated this week
- Shielded Enclaves for Cloud FPGAs☆15Nov 24, 2021Updated 4 years ago
- ☆12Oct 25, 2022Updated 3 years ago
- AT2PO: Agentic Turn-based Policy Optimization via Tree Search☆22May 21, 2026Updated 3 months ago
- [ACL 2026] WildGraphBench: Benchmarking GraphRAG with Wild-Source Corpora☆17May 11, 2026Updated 3 months ago
- The Gtgraph library from Georgia Tech.☆18Mar 30, 2023Updated 3 years ago
- Code repository for the ICML 2026 Oral paper "Characterizing, Evaluating, and Optimizing Complex Reasoning".☆19Jun 21, 2026Updated 2 months ago
- vLLM fork with Marlin W4A8 SM121 patches + TMA module☆15Mar 23, 2026Updated 5 months ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- vLLM Qwen3.5-122B NVFP4 on DGX Spark (SM121) — full Docker build with 15 patches☆17Mar 17, 2026Updated 5 months ago
- Recursive unified ORAM☆15Sep 23, 2015Updated 10 years ago
- [KDD 2026] "Breaking Information Cocoons: A Hyperbolic Graph-LLM Framework for Exploration and Exploitation in Recommender Systems"☆16Jan 29, 2025Updated last year
- Example of applying CUDA graphs to LLaMA-v2☆11Aug 25, 2023Updated 3 years ago
- MLX-powered Qwen3 embedding server for Apple Silicon Macs. Features 0.6B/4B/8B models, 44K tokens/sec throughput, REST API, batch proce…☆17Aug 9, 2025Updated last year
- NVFP4 inference on Blackwell GeForce (RTX 5090/5080/5070 Ti/RTX PRO 6000) — SM120 patches for vLLM + FlashInfer + CUTLASS. 175 tok/s on Q…☆23Apr 27, 2026Updated 4 months ago
- Minimalist repo to do mlx vlm constrained decoding in batch mode☆18Apr 11, 2026Updated 4 months ago