Official implementation for "K-Forcing: Joint Next-K-Token Decoding via Push-Forward Language Modeling"
☆19Aug 16, 2026Updated last month
Alternatives and similar repositories for K-Forcing
Users that are interested in K-Forcing are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [NeurIPS 2024] The official implementation of ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification☆33Mar 30, 2025Updated last year
- [ICML 2025] This is the official PyTorch implementation of "ZipAR: Accelerating Auto-regressive Image Generation through Spatial Locality…☆52Mar 25, 2025Updated last year
- [CVPR2026] Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching☆19Jun 4, 2026Updated 3 months ago
- ☆16Sep 12, 2023Updated 3 years ago
- WorldOlympiad: Can Your World Model Survive a Triathlon?☆57Updated this week
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- [CVPRW 2026 Oral] Less Detail, Better Answers: Degradation-Driven Prompting for VQA☆20Apr 25, 2026Updated 5 months ago
- Official PyTorch implementation of [PSA: Pyramid Sparse Attention for Efficient Video Understanding and Generation](https://arxiv.org/abs…☆25Jan 25, 2026Updated 8 months ago
- [ICLR 2026] This is the official PyTorch implementation of "BLADE: Block-Sparse Attention Meets Step Distillation for Efficient Video Gen…☆53Oct 9, 2025Updated 11 months ago
- [NAACL'25 🏆 SAC Award] Official code for "Advancing MoE Efficiency: A Collaboration-Constrained Routing (C2R) Strategy for Better Expert…☆17Feb 4, 2025Updated last year
- Code of EMNLP 2025 paper 'UltraIF: Advancing Instruction Following from the Wild'.☆21Apr 3, 2025Updated last year
- PARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation (ICLR 26)☆35Jun 10, 2026Updated 3 months ago
- Latent Spatial Memory for Video World Models☆327Aug 27, 2026Updated last month
- Official implementation of "Streaming Communication in Multi-Agent Reasoning"☆36Jun 6, 2026Updated 3 months ago
- Repo of HawkLlama.☆16Jan 2, 2025Updated last year
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Post-Trained MoE Can Skip Half Experts via Self-Distillation☆42Sep 6, 2026Updated 3 weeks ago
- A Blender extension for automated, editable, and inter-shot consistent 3D storyboard production.☆32Sep 7, 2026Updated 3 weeks ago
- torch_quantizer is a out-of-box quantization tool for PyTorch models on CUDA backend, specially optimized for Diffusion Models.☆25Mar 29, 2024Updated 2 years ago
- [EMNLP2026 Main] Official implementation of “Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding”.☆138Jul 25, 2026Updated 2 months ago
- [ACL 2026 Findings] CoV: Chain-of-View Prompting for Spatial Reasoning☆63Apr 7, 2026Updated 5 months ago
- [ICLR2025] Are Large Vision Language Models Good Game Players?☆13Mar 3, 2025Updated last year
- [ICLR 2025] Official PyTorch implmentation of paper "T-Stitch: Accelerating Sampling in Pre-trained Diffusion Models with Trajectory Stit…☆107Feb 26, 2024Updated 2 years ago
- ☆17Apr 20, 2025Updated last year
- [ICML 2026] World-R1: Reinforcing 3D Constraints for Text-to-Video Generation☆427Jun 3, 2026Updated 4 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Official implementation for the paper "EBMDOCK: NEURAL PROBABILISTIC PROTEIN-PROTEIN DOCKING VIA A DIFFERENTIABLE ENERGY-BASED MODEL" (IC…☆14Apr 24, 2024Updated 2 years ago
- ☆17May 27, 2026Updated 4 months ago
- ☆21Jun 13, 2026Updated 3 months ago
- A comprehensive list of papers about Large-Language-Diffusion-Models.☆104Sep 1, 2026Updated last month
- ☆14Sep 7, 2024Updated 2 years ago
- a website for accessing many models through api(deepseek、Qwen、Hunyuan etc.)☆16Jul 12, 2025Updated last year
- Introduction To Computing Systems☆13May 5, 2015Updated 11 years ago
- high-performance linear attention kernel library built on TileLang☆718Updated this week
- [COLM 2026] Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes☆139May 19, 2026Updated 4 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [ACL'26] EvoToken-DLM (Beyond Hard Masks: Progressive Token Evolution for Diffusion Language)☆49Apr 7, 2026Updated 5 months ago
- 清华大学电子系科协学培部Sast Tutor共享仓库☆16Apr 27, 2022Updated 4 years ago
- A collection of specialized agent skills for AI infrastructure development, enabling Claude Code to write, optimize, and debug high-perfo…☆148Jul 9, 2026Updated 2 months ago
- Official repo for "TC-AE: Unlocking Token Capacity for Deep Compression Autoencoders"☆28Apr 9, 2026Updated 5 months ago
- ☆19Apr 16, 2025Updated last year
- TriSplat: Simulation-Ready Feed-Forward 3D Scene Reconstruction☆360Jun 12, 2026Updated 3 months ago
- Benchmarking Attention Mechanism in Vision Transformers.☆20Oct 10, 2022Updated 3 years ago