[ICLR'26] The official code implementation for "Cache-to-Cache: Direct Semantic Communication Between Large Language Models"
☆612Mar 13, 2026Updated 6 months ago
Alternatives and similar repositories for C2C
Users that are interested in C2C are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [ICML 2026 Spotlight] Latent Collaboration in Multi-Agent Systems☆1,128Jun 18, 2026Updated 3 months ago
- [NeurIPS'25] The official code implementation for paper "R2R: Efficiently Navigating Divergent Reasoning Paths with Small-Large Model Tok…☆97Apr 7, 2026Updated 5 months ago
- [NeurIPS'25] KVCOMM: Online Cross-context KV-cache Communication for Efficient LLM-based Multi-agent Systems☆191Nov 3, 2025Updated 10 months ago
- [ICML'26] Official implementation of paper "Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models"☆84Jul 17, 2026Updated 2 months ago
- ☆21Apr 7, 2025Updated last year
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- LatentMAS with kNN kv cache pruning | up to 40% more memory efficient and 30% faster☆19Dec 10, 2025Updated 9 months ago
- [ICCV'25] The official code of paper "Combining Similarity and Importance for Video Token Reduction on Large Visual Language Models"☆78Jan 13, 2026Updated 8 months ago
- [NeurIPS'25] KVCOMM: Online Cross-context KV-cache Communication for Efficient LLM-based Multi-agent Systems☆18Nov 1, 2025Updated 10 months ago
- [NeurIPS'25 Oral] Query-agnostic KV cache eviction: 3–4× reduction in memory and 2× decrease in latency (Qwen3/2.5, Gemma3, LLaMA3)☆225Feb 11, 2026Updated 7 months ago
- Official implementation of "SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching" (COLM 2025). A novel KV cache com…☆15Sep 29, 2025Updated 11 months ago
- The official code implementation for paper "PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs"☆30May 24, 2025Updated last year
- Open source code in the field of semantic communication.☆360Mar 7, 2025Updated last year
- Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs☆25Nov 11, 2025Updated 10 months ago
- ☆48Oct 16, 2025Updated 11 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- [ICLR 2025] Linear Combination of Saved Checkpoints Makes Consistency and Diffusion Models Better☆16Feb 15, 2025Updated last year
- [CoLM'25] The official implementation of the paper <MoA: Mixture of Sparse Attention for Automatic Large Language Model Compression>☆160Jan 14, 2026Updated 8 months ago
- [ICML 2026] Esoteric Language Models☆125Jul 13, 2026Updated 2 months ago
- ACL'2025: SoftCoT: Soft Chain-of-Thought for Efficient Reasoning with LLMs. and preprint: SoftCoT++: Test-Time Scaling with Soft Chain-of…☆95May 30, 2025Updated last year
- 清华大学“荷塘雨课堂”助手,包含自动签到、答题等功能。☆24Dec 4, 2025Updated 9 months ago
- [NeurIPS 2025] Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning Chains☆97Jun 29, 2026Updated 2 months ago
- Inspect LLM's logprobs and perplexity over a piece of text, or compare two LLMs (like a git diff)☆18Aug 3, 2026Updated last month
- 🌿 DeepPrune: Parallel Scaling without Inter-trace Redundancy☆21Apr 20, 2026Updated 4 months ago
- Code implementation for 《Large AI Model Empowered Multimodal Semantic communication》☆27Jul 4, 2024Updated 2 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [ECCV24] MixDQ: Memory-Efficient Few-Step Text-to-Image Diffusion Models with Metric-Decoupled Mixed Precision Quantization☆50Nov 27, 2024Updated last year
- KernelBench v2: Can LLMs Write GPU Kernels? - Benchmark with Torch -> Triton (and more!) problems☆26Jul 4, 2025Updated last year
- Code Repository of Evaluating Quantized Large Language Models☆136Sep 8, 2024Updated 2 years ago
- Implementation for the paper "CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards".☆59Jan 26, 2026Updated 7 months ago
- Benchmark Test-Time Scaling of General LLM Agents☆25Apr 14, 2026Updated 5 months ago
- [NeurIPS 25] InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding☆24Jan 25, 2026Updated 7 months ago
- Memory-bounded compressed sparse attention via streaming top-k. Triton kernels for the DeepSeek-V4 lightning indexer. 32x regime extensio…☆26May 5, 2026Updated 4 months ago
- Official implementation of the NeurIPS 2025 paper "Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space"☆352Jun 12, 2026Updated 3 months ago
- Official Repository for paper "HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding" [ACL 2026]☆105May 8, 2026Updated 4 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- 2022 秋季学期清华大学电子系数据与算法课程 OJ 参考解答☆10Jun 18, 2023Updated 3 years ago
- LLM KV cache compression made easy☆1,208Updated this week
- [ICML 2026] Heima☆76May 20, 2026Updated 4 months ago
- LatentMem: Customizing Latent Memory for Multi-Agent Systems☆54Sep 7, 2026Updated last week
- The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems☆29Mar 3, 2026Updated 6 months ago
- Official Repo for paper "VLCache: Computing 2% Vision Tokens and Reusing 98% for Vision-Language Inference"☆15Mar 28, 2026Updated 5 months ago
- Chain of Agents implementation in Python and Swift☆32Aug 31, 2026Updated 2 weeks ago