☆119Jul 23, 2025Updated last year
Alternatives and similar repositories for grokking-at-the-edge-of-numerical-stability
Users that are interested in grokking-at-the-edge-of-numerical-stability are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official repository for the paper "Grokfast: Accelerated Grokking by Amplifying Slow Gradients"☆582Jun 28, 2024Updated 2 years ago
- ☆34Updated this week
- Reference implementation of "Softmax Attention with Constant Cost per Token" (Heinsen, 2024)☆25Jun 6, 2024Updated 2 years ago
- Exercises of the reinforcement learning course from Hugging Face☆13Mar 1, 2023Updated 3 years ago
- Code to generate figures of paper "When do spectral gradient updates help in deep learning?"☆17Dec 3, 2025Updated 9 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Code for NeurIPS'24 paper 'Grokked Transformers are Implicit Reasoners: A Mechanistic Journey to the Edge of Generalization'☆242Jul 19, 2025Updated last year
- Code for☆28Dec 16, 2024Updated last year
- ☆25Dec 9, 2024Updated last year
- ROSA+: RWKV's ROSA implementation with fallback statistical predictor☆36Oct 13, 2025Updated 11 months ago
- Kolmogorov–Arnold Networks with modified activation (using MLP to represent the activation)☆108Oct 4, 2025Updated 11 months ago
- ☆25Dec 13, 2024Updated last year
- Enable Next-sentence Prediction for Large Language Models with Faster Speed, Higher Accuracy and Longer Context☆42Aug 16, 2024Updated 2 years ago
- Github Repository for the HOI4 ULTRA Project.☆11Updated this week
- FC-KAN: Function Combinations in Kolmogorov-Arnold Networks☆40Updated this week
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Modify Entropy Based Sampling to work with Mac Silicon via MLX☆49Nov 6, 2024Updated last year
- ☆23Nov 6, 2022Updated 3 years ago
- A high-efficiency text embedding and reranking model based on RWKV architecture.☆22Sep 12, 2026Updated last week
- ☆16Dec 11, 2025Updated 9 months ago
- ☆121Mar 18, 2026Updated 6 months ago
- SETOL: SemiEmpirical Theory of (Deep) Learning☆29Updated this week
- The official implementation of HybridNorm: Towards Stable and Efficient Transformer Training via Hybrid Normalization☆19Mar 7, 2025Updated last year
- Pretraining and inference code for a large-scale depth-recurrent language model☆967Dec 29, 2025Updated 8 months ago
- NeuMeta transforms neural networks by allowing a single model to adapt on the fly to different sizes, generating the right weights when n…☆45Nov 8, 2024Updated last year
- Bare Metal GPUs on DigitalOcean Gradient AI • AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- look how they massacred my boy☆63Oct 16, 2024Updated last year
- MetaLadder: Ascending Mathematical Solution Quality via Analogical-Problem Reasoning Transfer (EMNLP 2025)☆12Apr 18, 2025Updated last year
- Code for RATIONALYST: Pre-training Process-Supervision for Improving Reasoning https://arxiv.org/pdf/2410.01044☆36Oct 3, 2024Updated last year
- ☆18Apr 19, 2024Updated 2 years ago
- ☆16May 15, 2021Updated 5 years ago
- ☆13Jun 12, 2024Updated 2 years ago
- Code for paper Almost-Orthogonal Layers for Efficient General-Purpose Lipschitz Networks☆13Aug 9, 2022Updated 4 years ago
- ☆15Mar 20, 2025Updated last year
- Code accompanying the paper "Generalized Interpolating Discrete Diffusion"☆124Jun 9, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- The simplest, fastest repository for training/finetuning medium-sized xLSTMs.☆41May 24, 2024Updated 2 years ago
- Work in progress.☆82Nov 25, 2025Updated 9 months ago
- Attention Kernels for Symmetric Power Transformers☆131Sep 25, 2025Updated 11 months ago
- [NeurIPS 2025 Spotlight] TPA: Tensor ProducT ATTenTion Transformer (https://arxiv.org/abs/2501.06425)☆463Sep 4, 2026Updated 2 weeks ago
- slowly building a set of infinite riddle generators for data-hungry methods☆14Nov 15, 2022Updated 3 years ago
- ☆19Sep 1, 2025Updated last year
- Implementation of Not All Contexts Are Equal: Teaching LLMs Credibility-aware Generation. Paper: https://arxiv.org/abs/2404.06809☆22Oct 22, 2024Updated last year