Kinetics: Rethinking Test-Time Scaling Laws
β87Jul 11, 2025Updated last year
Alternatives and similar repositories for Kinetics
Users that are interested in Kinetics are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- β65Jun 12, 2025Updated last year
- [NAACL'25 π SAC Award] Official code for "Advancing MoE Efficiency: A Collaboration-Constrained Routing (C2R) Strategy for Better Expertβ¦β16Feb 4, 2025Updated last year
- Vortex: Programmable Sparse Attention for Agents as Algorithm Designersβ68Updated this week
- [ICML 2026] Less Is More: Training-Free Sparse Attention with Global Locality for Efficient Reasoningβ35Sep 12, 2025Updated 10 months ago
- β16Jun 15, 2026Updated last month
- Bare Metal GPUs on DigitalOcean Gradient AI β’ AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- β82Jun 8, 2026Updated 2 months ago
- Official repository for Parallax (Parameterized Local Linear Attention)β68Jul 30, 2026Updated last week
- β56Jul 7, 2025Updated last year
- β79May 29, 2026Updated 2 months ago
- Official github repo for "Fractional Reasoning via Latent Steering Vectors Improves Inference Time Compute"β17Jun 30, 2025Updated last year
- Compact and Agent-Native MoE Training Systemβ327Jul 31, 2026Updated last week
- Code for the paper "Stack Attention: Improving the Ability of Transformers to Model Hierarchical Patterns"β18Mar 15, 2024Updated 2 years ago
- β93Updated this week
- An efficient implementation of the NSA (Native Sparse Attention) kernelβ134Jun 24, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- [ICML 2024] Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inferenceβ400Jul 10, 2025Updated last year
- Official code for the paper "HEXA-MoE: Efficient and Heterogeneous-Aware MoE Acceleration with Zero Computation Redundancy"β15Mar 6, 2025Updated last year
- β34Oct 13, 2025Updated 9 months ago
- β88Jun 16, 2025Updated last year
- β16May 27, 2026Updated 2 months ago
- The official implementation of HybridNorm: Towards Stable and Efficient Transformer Training via Hybrid Normalizationβ19Mar 7, 2025Updated last year
- Simple & Scalable Pretraining for Neural Architecture Researchβ341Mar 31, 2026Updated 4 months ago
- [ICML 2025 Spotlight] ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inferenceβ311May 1, 2025Updated last year
- The official repository for SkyLadder: Better and Faster Pretraining via Context Window Schedulingβ43Dec 29, 2025Updated 7 months ago
- GPU virtual machines on DigitalOcean Gradient AI β’ AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- Ring is a reasoning MoE LLM provided and open-sourced by InclusionAI, derived from Ling.β109Aug 5, 2025Updated last year
- β36Mar 12, 2025Updated last year
- β136May 29, 2025Updated last year
- β14Oct 3, 2024Updated last year
- ArcherCodeR is an open-source initiative enhancing code reasoning in large language models through scalable, rule-governed reinforcement β¦β44Aug 6, 2025Updated last year
- π₯ A minimal training framework for scaling FLA modelsβ409Apr 22, 2026Updated 3 months ago
- [ICMLβ25] Official code for paper "Occult: Optimizing Collaborative Communication across Experts for Accelerated Parallel MoE Training anβ¦β13Apr 17, 2025Updated last year
- VibeRL is a Reinforcement Learning framework built essentially through vibe coding with Kimi K2.β17Updated this week
- [EMNLP 25] An effective and interpretable weight-editing method for mitigating overly short reasoning in LLMs, and a mechanistic study unβ¦β19Dec 17, 2025Updated 7 months ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Code for "Reasoning to Learn from Latent Thoughts"β134Mar 28, 2025Updated last year
- Chain of Experts (CoE) enables communication between experts within Mixture-of-Experts (MoE) modelsβ231Nov 4, 2025Updated 9 months ago
- β125May 19, 2025Updated last year
- β113Aug 26, 2024Updated last year
- Physics of Language Models: Part 4.2, Canon Layers at Scale where Synthetic Pretraining Resonates in Realityβ357May 20, 2026Updated 2 months ago
- [ICML2024 Spotlight] Fine-Tuning Pre-trained Large Language Models Sparselyβ24Jun 26, 2024Updated 2 years ago
- Dataflow-Oriented Reinforcement Learning for (Multi-)Agentic LLMsβ99Jul 26, 2026Updated 2 weeks ago