Kinetics: Rethinking Test-Time Scaling Laws
β87Jul 11, 2025Updated last year
Alternatives and similar repositories for Kinetics
Users that are interested in Kinetics are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- β65Jun 12, 2025Updated last year
- [NAACL'25 π SAC Award] Official code for "Advancing MoE Efficiency: A Collaboration-Constrained Routing (C2R) Strategy for Better Expertβ¦β16Feb 4, 2025Updated last year
- Vortex: Programmable Sparse Attention for Agents as Algorithm Designersβ68Aug 22, 2026Updated last week
- [ICML 2026] Less Is More: Training-Free Sparse Attention with Global Locality for Efficient Reasoningβ35Sep 12, 2025Updated 11 months ago
- β17Jun 15, 2026Updated 2 months ago
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- β83Jun 8, 2026Updated 2 months ago
- Official repository for Parallax (Parameterized Local Linear Attention)β68Jul 30, 2026Updated last month
- β56Jul 7, 2025Updated last year
- β86Aug 19, 2026Updated last week
- Official github repo for "Fractional Reasoning via Latent Steering Vectors Improves Inference Time Compute"β17Jun 30, 2025Updated last year
- Compact and Agent-Native MoE Training Systemβ344Aug 22, 2026Updated last week
- Code for the paper "Stack Attention: Improving the Ability of Transformers to Model Hierarchical Patterns"β18Mar 15, 2024Updated 2 years ago
- β98Updated this week
- An efficient implementation of the NSA (Native Sparse Attention) kernelβ134Jun 24, 2025Updated last year
- End-to-end encrypted cloud storage - Proton Drive β’ AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- [ICML 2024] Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inferenceβ404Jul 10, 2025Updated last year
- Official code for the paper "HEXA-MoE: Efficient and Heterogeneous-Aware MoE Acceleration with Zero Computation Redundancy"β15Mar 6, 2025Updated last year
- β34Oct 13, 2025Updated 10 months ago
- β88Jun 16, 2025Updated last year
- β16May 27, 2026Updated 3 months ago
- The official implementation of HybridNorm: Towards Stable and Efficient Transformer Training via Hybrid Normalizationβ19Mar 7, 2025Updated last year
- Simple & Scalable Pretraining for Neural Architecture Researchβ346Mar 31, 2026Updated 4 months ago
- [ICML 2025 Spotlight] ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inferenceβ313May 1, 2025Updated last year
- The official repository for SkyLadder: Better and Faster Pretraining via Context Window Schedulingβ43Dec 29, 2025Updated 8 months ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Ring is a reasoning MoE LLM provided and open-sourced by InclusionAI, derived from Ling.β109Aug 5, 2025Updated last year
- β36Mar 12, 2025Updated last year
- β138May 29, 2025Updated last year
- β14Oct 3, 2024Updated last year
- ArcherCodeR is an open-source initiative enhancing code reasoning in large language models through scalable, rule-governed reinforcement β¦β44Aug 6, 2025Updated last year
- π₯ A minimal training framework for scaling FLA modelsβ413Apr 22, 2026Updated 4 months ago
- [ICMLβ25] Official code for paper "Occult: Optimizing Collaborative Communication across Experts for Accelerated Parallel MoE Training anβ¦β14Apr 17, 2025Updated last year
- VibeRL is a Reinforcement Learning framework built essentially through vibe coding with Kimi K2.β18Updated this week
- [EMNLP 25] An effective and interpretable weight-editing method for mitigating overly short reasoning in LLMs, and a mechanistic study unβ¦β19Updated this week
- End-to-end encrypted cloud storage - Proton Drive β’ AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Code for "Reasoning to Learn from Latent Thoughts"β134Mar 28, 2025Updated last year
- Chain of Experts (CoE) enables communication between experts within Mixture-of-Experts (MoE) modelsβ231Nov 4, 2025Updated 9 months ago
- β124May 19, 2025Updated last year
- β113Aug 26, 2024Updated 2 years ago
- Physics of Language Models: Part 4.2, Canon Layers at Scale where Synthetic Pretraining Resonates in Realityβ360Aug 21, 2026Updated last week
- [ICML2024 Spotlight] Fine-Tuning Pre-trained Large Language Models Sparselyβ24Jun 26, 2024Updated 2 years ago
- Dataflow-Oriented Reinforcement Learning for (Multi-)Agentic LLMs