[ICLR'26] The official code implementation for "Cache-to-Cache: Direct Semantic Communication Between Large Language Models"
☆426Mar 13, 2026Updated 4 months ago
Alternatives and similar repositories for C2C
Users that are interested in C2C are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICML 2026 Spotlight] Latent Collaboration in Multi-Agent Systems☆1,079Jun 18, 2026Updated last month
- [NeurIPS'25] The official code implementation for paper "R2R: Efficiently Navigating Divergent Reasoning Paths with Small-Large Model Tok…☆95Apr 7, 2026Updated 4 months ago
- [NeurIPS'25] KVCOMM: Online Cross-context KV-cache Communication for Efficient LLM-based Multi-agent Systems☆182Nov 3, 2025Updated 9 months ago
- [ICML'26] Official implementation of paper "Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models"☆79Jul 17, 2026Updated 3 weeks ago
- LatentMAS with kNN kv cache pruning | up to 40% more memory efficient and 30% faster☆19Dec 10, 2025Updated 8 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ICCV'25] The official code of paper "Combining Similarity and Importance for Video Token Reduction on Large Visual Language Models"☆75Jan 13, 2026Updated 6 months ago
- [NeurIPS'25] KVCOMM: Online Cross-context KV-cache Communication for Efficient LLM-based Multi-agent Systems☆17Nov 1, 2025Updated 9 months ago
- [NeurIPS'25 Oral] Query-agnostic KV cache eviction: 3–4× reduction in memory and 2× decrease in latency (Qwen3/2.5, Gemma3, LLaMA3)☆225Feb 11, 2026Updated 5 months ago
- Official implementation of "SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching" (COLM 2025). A novel KV cache com…☆15Sep 29, 2025Updated 10 months ago
- The official code implementation for paper "PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs"☆29May 24, 2025Updated last year
- Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs☆25Nov 11, 2025Updated 8 months ago
- ☆48Oct 16, 2025Updated 9 months ago
- [CoLM'25] The official implementation of the paper <MoA: Mixture of Sparse Attention for Automatic Large Language Model Compression>☆159Jan 14, 2026Updated 6 months ago
- [ICML 2026] Esoteric Language Models☆122Jul 13, 2026Updated 3 weeks ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- ACL'2025: SoftCoT: Soft Chain-of-Thought for Efficient Reasoning with LLMs. and preprint: SoftCoT++: Test-Time Scaling with Soft Chain-of…☆94May 30, 2025Updated last year
- 清华大学“荷塘雨课堂”助手,包含自动签到、答题等功能。☆24Dec 4, 2025Updated 8 months ago
- [NeurIPS 2025] Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning Chains☆98Jun 29, 2026Updated last month
- Inspect LLM's logprobs and perplexity over a piece of text, or compare two LLMs (like a git diff)☆19Aug 3, 2026Updated last week
- 🌿 DeepPrune: Parallel Scaling without Inter-trace Redundancy☆21Apr 20, 2026Updated 3 months ago
- A paper list of Awesome Latent Space.☆956Jul 13, 2026Updated 3 weeks ago
- [ECCV24] MixDQ: Memory-Efficient Few-Step Text-to-Image Diffusion Models with Metric-Decoupled Mixed Precision Quantization☆50Nov 27, 2024Updated last year
- KernelBench v2: Can LLMs Write GPU Kernels? - Benchmark with Torch -> Triton (and more!) problems☆24Jul 4, 2025Updated last year
- Code Repository of Evaluating Quantized Large Language Models☆135Sep 8, 2024Updated last year
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Implementation for the paper "CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards".☆55Jan 26, 2026Updated 6 months ago
- Benchmark Test-Time Scaling of General LLM Agents☆21Apr 14, 2026Updated 3 months ago
- [NeurIPS 25] InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding☆23Jan 25, 2026Updated 6 months ago
- Memory-bounded compressed sparse attention via streaming top-k. Triton kernels for the DeepSeek-V4 lightning indexer. 32x regime extensio…☆23May 5, 2026Updated 3 months ago
- Official implementation of the NeurIPS 2025 paper "Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space"☆347Jun 12, 2026Updated last month
- Official Repository for paper "HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding" [ACL 2026]☆94May 8, 2026Updated 3 months ago
- 2022 秋季学期清华大学电子系数据与算法课程 OJ 参考解答☆10Jun 18, 2023Updated 3 years ago
- LLM KV cache compression made easy☆1,163Updated this week
- [ICML 2026] Heima☆75May 20, 2026Updated 2 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- LatentMem: Customizing Latent Memory for Multi-Agent Systems☆49Feb 9, 2026Updated 6 months ago
- The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems☆28Mar 3, 2026Updated 5 months ago
- Official Repo for paper "VLCache: Computing 2% Vision Tokens and Reusing 98% for Vision-Language Inference"☆15Mar 28, 2026Updated 4 months ago
- Official implementation of "Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding"☆1,070May 30, 2026Updated 2 months ago
- ☆30Oct 2, 2025Updated 10 months ago
- [ICLR 2025] Distilled Decoding 1: One-step Sampling of Image Auto-regressive Models with Flow Matching☆55Apr 21, 2025Updated last year
- ☆51May 20, 2025Updated last year