[ICLR'26] The official code implementation for "Cache-to-Cache: Direct Semantic Communication Between Large Language Models"
☆435Mar 13, 2026Updated 5 months ago
Alternatives and similar repositories for C2C
Users that are interested in C2C are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICML 2026 Spotlight] Latent Collaboration in Multi-Agent Systems☆1,100Jun 18, 2026Updated 2 months ago
- [NeurIPS'25] The official code implementation for paper "R2R: Efficiently Navigating Divergent Reasoning Paths with Small-Large Model Tok…☆96Apr 7, 2026Updated 4 months ago
- [NeurIPS'25] KVCOMM: Online Cross-context KV-cache Communication for Efficient LLM-based Multi-agent Systems☆185Nov 3, 2025Updated 9 months ago
- [ICML'26] Official implementation of paper "Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models"☆81Jul 17, 2026Updated last month
- LatentMAS with kNN kv cache pruning | up to 40% more memory efficient and 30% faster☆19Dec 10, 2025Updated 8 months ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- [ICCV'25] The official code of paper "Combining Similarity and Importance for Video Token Reduction on Large Visual Language Models"☆77Jan 13, 2026Updated 7 months ago
- [NeurIPS'25] KVCOMM: Online Cross-context KV-cache Communication for Efficient LLM-based Multi-agent Systems☆18Nov 1, 2025Updated 9 months ago
- [NeurIPS'25 Oral] Query-agnostic KV cache eviction: 3–4× reduction in memory and 2× decrease in latency (Qwen3/2.5, Gemma3, LLaMA3)☆224Feb 11, 2026Updated 6 months ago
- Official implementation of "SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching" (COLM 2025). A novel KV cache com…☆15Sep 29, 2025Updated 11 months ago
- The official code implementation for paper "PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs"☆29May 24, 2025Updated last year
- Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs☆25Nov 11, 2025Updated 9 months ago
- ☆48Oct 16, 2025Updated 10 months ago
- [ICLR 2025] Linear Combination of Saved Checkpoints Makes Consistency and Diffusion Models Better☆16Feb 15, 2025Updated last year
- [CoLM'25] The official implementation of the paper <MoA: Mixture of Sparse Attention for Automatic Large Language Model Compression>☆160Jan 14, 2026Updated 7 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- [ICML 2026] Esoteric Language Models☆124Jul 13, 2026Updated last month
- ACL'2025: SoftCoT: Soft Chain-of-Thought for Efficient Reasoning with LLMs. and preprint: SoftCoT++: Test-Time Scaling with Soft Chain-of…☆95May 30, 2025Updated last year
- 清华大学“荷塘雨课堂”助手,包含自动签到、答题等功能。☆24Dec 4, 2025Updated 8 months ago
- [NeurIPS 2025] Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning Chains☆99Jun 29, 2026Updated 2 months ago
- 🌿 DeepPrune: Parallel Scaling without Inter-trace Redundancy☆21Apr 20, 2026Updated 4 months ago
- A paper list of Awesome Latent Space.☆961Jul 13, 2026Updated last month
- [ECCV24] MixDQ: Memory-Efficient Few-Step Text-to-Image Diffusion Models with Metric-Decoupled Mixed Precision Quantization☆50Nov 27, 2024Updated last year
- KernelBench v2: Can LLMs Write GPU Kernels? - Benchmark with Torch -> Triton (and more!) problems☆25Jul 4, 2025Updated last year
- Code Repository of Evaluating Quantized Large Language Models☆135Sep 8, 2024Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Implementation for the paper "CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards".☆59Jan 26, 2026Updated 7 months ago
- Benchmark Test-Time Scaling of General LLM Agents☆23Apr 14, 2026Updated 4 months ago
- [NeurIPS 25] InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding☆23Jan 25, 2026Updated 7 months ago
- Memory-bounded compressed sparse attention via streaming top-k. Triton kernels for the DeepSeek-V4 lightning indexer. 32x regime extensio…☆25May 5, 2026Updated 3 months ago
- Official implementation of the NeurIPS 2025 paper "Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space"☆352Jun 12, 2026Updated 2 months ago
- Official Repository for paper "HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding" [ACL 2026]☆101May 8, 2026Updated 3 months ago
- 2022 秋季学期清华大学电子系数据与算法课程 OJ 参考解答☆10Jun 18, 2023Updated 3 years ago
- LLM KV cache compression made easy☆1,196Aug 18, 2026Updated last week
- [ICML 2026] Heima☆76May 20, 2026Updated 3 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- LatentMem: Customizing Latent Memory for Multi-Agent Systems☆53Feb 9, 2026Updated 6 months ago
- The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems☆29Mar 3, 2026Updated 5 months ago
- Official Repo for paper "VLCache: Computing 2% Vision Tokens and Reusing 98% for Vision-Language Inference"☆15Mar 28, 2026Updated 5 months ago
- Chain of Agents implementation in Python and Swift☆32Jul 3, 2026Updated last month
- Official implementation of "Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding"☆1,085May 30, 2026Updated 3 months ago
- ☆30Oct 2, 2025Updated 10 months ago
- 清华大学计算机系面向对象程序设计基础 (OOP) 课程 2024 春编程作业☆19Jan 20, 2025Updated last year