Kinetics: Rethinking Test-Time Scaling Laws
β87Jul 11, 2025Updated last year
Alternatives and similar repositories for Kinetics
Users that are interested in Kinetics are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- β65Jun 12, 2025Updated last year
- [NAACL'25 π SAC Award] Official code for "Advancing MoE Efficiency: A Collaboration-Constrained Routing (C2R) Strategy for Better Expertβ¦β17Feb 4, 2025Updated last year
- Vortex: Programmable Sparse Attention for Agents as Algorithm Designersβ69Aug 22, 2026Updated last month
- [ICML 2026] Less Is More: Training-Free Sparse Attention with Global Locality for Efficient Reasoningβ36Sep 12, 2025Updated last year
- β17Jun 15, 2026Updated 3 months ago
- Virtual machines for every use case on DigitalOcean β’ AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- β84Jun 8, 2026Updated 4 months ago
- Official repository for Parallax (Parameterized Local Linear Attention)β69Jul 30, 2026Updated 2 months ago
- β57Jul 7, 2025Updated last year
- β88Aug 19, 2026Updated last month
- Official github repo for "Fractional Reasoning via Latent Steering Vectors Improves Inference Time Compute"β17Jun 30, 2025Updated last year
- Compact and Agent-Native MoE Training Systemβ355Updated this week
- β100Sep 14, 2026Updated 3 weeks ago
- An efficient implementation of the NSA (Native Sparse Attention) kernelβ135Jun 24, 2025Updated last year
- [ICML 2024] Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inferenceβ410Jul 10, 2025Updated last year
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Official code for the paper "HEXA-MoE: Efficient and Heterogeneous-Aware MoE Acceleration with Zero Computation Redundancy"β15Mar 6, 2025Updated last year
- Code for the paper "Stack Attention: Improving the Ability of Transformers to Model Hierarchical Patterns"β19Mar 15, 2024Updated 2 years ago
- β34Oct 13, 2025Updated 11 months ago
- β88Jun 16, 2025Updated last year
- β17May 27, 2026Updated 4 months ago
- The official implementation of HybridNorm: Towards Stable and Efficient Transformer Training via Hybrid Normalizationβ19Mar 7, 2025Updated last year
- Simple & Scalable Pretraining for Neural Architecture Researchβ347Mar 31, 2026Updated 6 months ago
- [ICML 2025 Spotlight] ShadowKV: KV Cache in Shadows for High-Throughput Long-Context LLM Inferenceβ314May 1, 2025Updated last year
- The official repository for SkyLadder: Better and Faster Pretraining via Context Window Schedulingβ44Dec 29, 2025Updated 9 months ago
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Ring is a reasoning MoE LLM provided and open-sourced by InclusionAI, derived from Ling.β108Aug 5, 2025Updated last year
- β36Mar 12, 2025Updated last year
- β138May 29, 2025Updated last year
- β14Oct 3, 2024Updated 2 years ago
- ArcherCodeR is an open-source initiative enhancing code reasoning in large language models through scalable, rule-governed reinforcement β¦β44Aug 6, 2025Updated last year
- π₯ A minimal training framework for scaling FLA modelsβ420Apr 22, 2026Updated 5 months ago
- [ICMLβ25] Official code for paper "Occult: Optimizing Collaborative Communication across Experts for Accelerated Parallel MoE Training anβ¦β14Apr 17, 2025Updated last year
- VibeRL is a Reinforcement Learning framework built essentially through vibe coding with Kimi K2.β18Sep 28, 2026Updated last week
- [EMNLP 25] An effective and interpretable weight-editing method for mitigating overly short reasoning in LLMs, and a mechanistic study unβ¦β20Aug 24, 2026Updated last month
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Code for "Reasoning to Learn from Latent Thoughts"β135Mar 28, 2025Updated last year
- Chain of Experts (CoE) enables communication between experts within Mixture-of-Experts (MoE) modelsβ233Nov 4, 2025Updated 11 months ago
- β125May 19, 2025Updated last year
- β112Aug 26, 2024Updated 2 years ago
- Physics of Language Models: Part 4.2, Canon Layers at Scale where Synthetic Pretraining Resonates in Realityβ361Aug 21, 2026Updated last month
- [ICML2024 Spotlight] Fine-Tuning Pre-trained Large Language Models Sparselyβ24Jun 26, 2024Updated 2 years ago
- Dataflow-Oriented Reinforcement Learning for (Multi-)Agentic LLMsβ105Jul 26, 2026Updated 2 months ago