This is the source code of our ICML25 paper, titled "Accelerating Large Language Model Reasoning via Speculative Search".
☆24Jun 1, 2025Updated last year
Alternatives and similar repositories for LLMReasoning-SpecSearch
Users that are interested in LLMReasoning-SpecSearch are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The code of paper *Learning Robust Policy against Disturbance in Transition Dynamics via State-Conservative Policy Optimization*.☆18Mar 26, 2022Updated 4 years ago
- The code for "AttentionPredictor: Temporal Pattern Matters for Efficient LLM Inference", Qingyue Yang, Jie Wang, Xing Li, Zhihai Wang, Ch…☆29Jul 15, 2025Updated last year
- This is the code of paper "De Novo Molecular Generation via Connection-aware Motif Mining". Zijie Geng, Shufang Xie, Yingce Xia, Lijun Wu…☆62Jan 7, 2025Updated last year
- This is the code of DiffILO, an unsupervised learning approach for predicting solutions to Integer Linear Programs (ILPs).☆39Oct 24, 2025Updated 9 months ago
- The code of paper "Learning Rule-Induced Subgraph Representations for Inductive Relation Prediction" in NeurIPS 2023.☆14Nov 25, 2023Updated 2 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆10Oct 5, 2023Updated 2 years ago
- [ICML 2025 Spotlight] RAPID: Long-Context Inference with Retrieval-Augmented Speculative Decoding☆23Mar 2, 2025Updated last year
- Official implementation of NeurIPS'24 paper "Reinforcement Learning Policy as Macro Regulator Rather than Macro Placer".☆20Aug 13, 2025Updated 11 months ago
- ☆18Apr 25, 2025Updated last year
- Pytorch implementation of our paper accepted by ICML 2023 -- "Bi-directional Masks for Efficient N:M Sparse Training"☆13Jun 7, 2023Updated 3 years ago
- Coco is a proactive co-assistant that connects user workspace with a broader ecosystem of AI agents.☆23Updated this week
- The official repo for the paper "Optimizing LLM Inference Throughput via Memory-aware and SLA-constrained Dynamic Batching"☆18Mar 17, 2025Updated last year
- Cross-Self KV Cache Pruning for Efficient Vision-Language Inference☆10Dec 15, 2024Updated last year
- Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding☆115Dec 2, 2025Updated 7 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- BESA is a differentiable weight pruning technique for large language models.☆17Mar 4, 2024Updated 2 years ago
- ☆16Feb 4, 2026Updated 5 months ago
- ☆16Jan 14, 2025Updated last year
- ☆16Jan 12, 2026Updated 6 months ago
- Official Repo for DAC-RL: Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability☆16Feb 26, 2026Updated 5 months ago
- channel pruning for accelerating very deep neural networks☆13Mar 8, 2021Updated 5 years ago
- When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning☆18Jun 2, 2026Updated last month
- This is the official Python version of CoreInfer: Accelerating Large Language Model Inference with Semantics-Inspired Adaptive Sparse Act…☆18Oct 25, 2024Updated last year
- Interpretable Contrastive Monte Carlo Tree Search Reasoning☆52Nov 9, 2024Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- 🧠Plan-and-Budget: Training-free test-time reasoning framework for adaptive token allocation in large language models (ICLR 2026).☆15Mar 2, 2026Updated 4 months ago
- This is the code for our paper "Reinforcement Learning within Tree Search for Fast Macro Placement".☆39Nov 13, 2024Updated last year
- Our Clone of Orca used for experimentation☆20Oct 15, 2024Updated last year
- This codebase is to reproduce the results of the paper "Grounded Test-Time Adaptation for LLM Agents".☆17Mar 4, 2026Updated 4 months ago
- GlimpRouter: Efficient Collaborative Inference by Glimpsing One Token of Thoughts☆16Apr 24, 2026Updated 3 months ago
- Mapper for digital PIM architectures☆29Sep 1, 2025Updated 10 months ago
- Official implementation of "SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching" (COLM 2025). A novel KV cache com…☆15Sep 29, 2025Updated 9 months ago
- Classifier Clustering and Feature Alignment for Federated Learning under Distributed Concept Drift [NeurIPS 2024]☆19Oct 25, 2024Updated last year
- ☆39Mar 17, 2025Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- [ICML 2025] CoreMatching: Co-adaptive Sparse Inference Framework for Comprehensive Acceleration of Vision Language Model☆16May 27, 2025Updated last year
- Official Code for paper "Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding""☆18Jun 2, 2026Updated last month
- The official implement of "Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings"☆18Dec 5, 2024Updated last year
- Source code for SWIFT, an efficient reward model.☆21Jan 13, 2026Updated 6 months ago
- Unite the knowledge of the world's top experts across every domain — to accelerate AI-driven scientific discovery.☆39May 4, 2026Updated 2 months ago
- This is the repository that introduces research topics related to protecting intellectual property (IP) of AI from a data-centric perspec…☆23Oct 30, 2023Updated 2 years ago
- PrefixKV: Adaptive Prefix KV Cache is What Vision Instruction-Following Models Need for Efficient Generation [NeurIPS 2025]☆19Oct 11, 2025Updated 9 months ago