Evolution Pretraining Fully in Int Formats
☆182Feb 25, 2026Updated 6 months ago
Alternatives and similar repositories for nano-egg
Users that are interested in nano-egg are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Jax Codebase for Evolutionary Strategies at the Hyperscale☆372Feb 27, 2026Updated 6 months ago
- Official implementation for "How Should We Meta-Learn Reinforcement Learning Algorithms?"☆23Sep 7, 2025Updated last year
- Neuroevolution Community☆16Nov 17, 2025Updated 9 months ago
- Official implementation of DiscoGen, for "Procedural Generation of Algorithm Discovery Tasks in Machine Learning"☆50Updated this week
- Evolutionary strategies finetuning library for LLMs☆25Jun 29, 2026Updated 2 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- A collection of lightweight interpretability scripts to understand how LLMs think☆92Mar 18, 2026Updated 5 months ago
- Official Codebase: LT2: Linear-Time Looped Transformers.☆59Jul 27, 2026Updated last month
- From-scratch PyTorch reproduction of DreamerV4 (Hafner et al., 2025): masked-autoencoder tokenizer, block-causal flow-matching dynamics w…☆38Updated this week
- Internal Coherence Maximization (ICM): A Label-Free, Unsupervised Training Framework for LLMs☆27Sep 5, 2025Updated last year
- Experiments Notebook of "Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism"☆17Apr 30, 2025Updated last year
- Generic building-block toolbox for training neural networks with adaptive and recursive execution. It provides reusable components to con…☆27Jun 29, 2026Updated 2 months ago
- nanoGPT using Equinox☆15Mar 3, 2023Updated 3 years ago
- Example for a Monty-enabled RLM in DSPy☆20Feb 16, 2026Updated 6 months ago
- GCRL in JAX. Official repository for LEO (ICML 2026).☆30Jun 20, 2026Updated 2 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Implementation and explorations into PopuLoRA, Co-Evolving LLM Populations for Reasoning Self-Play☆18Updated this week
- A simple, performant and scalable JAX-based world modeling codebase.☆163Jan 15, 2026Updated 7 months ago
- Official implementation of the paper "Next Embedding Prediction Makes World Models Stronger"☆39Aug 24, 2026Updated 2 weeks ago
- An imaginary extension of rotary position embeddings for long-context language models☆33Dec 9, 2025Updated 8 months ago
- Fast reinforcement learning 💨☆29Jul 15, 2025Updated last year
- Global CoT Analysis: Initial attempts to uncover patterns across many chains of thought☆20Feb 10, 2026Updated 6 months ago
- Marketplace ML experiment - training without backprop☆28Sep 9, 2025Updated 11 months ago
- TEVO: evolve LM motifs cheaply, then validate them in downstream train.py loops.☆19Apr 18, 2026Updated 4 months ago
- Softened ROSA QKV Operators for Training Next-Generation LLM Models☆39Aug 5, 2026Updated last month
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- explore token trajectory trees on instruct and base models☆158May 29, 2025Updated last year
- Linear Attention for Efficient Bidirectional Sequence Modeling☆18May 13, 2025Updated last year
- [ICML 2026] NITP: Next Implicit Token Prediction for LLM Pre-training☆35May 26, 2026Updated 3 months ago
- CIFAR-10 speedrun: Trains to 94% accuracy in 1.98 seconds on a single NVIDIA A100 GPU.☆80Jul 30, 2026Updated last month
- Cost aware hyperparameter tuning algorithm☆191Jun 27, 2024Updated 2 years ago
- Cuda kernels for leveraging LLM sparsity to improve throughput and decrease the memory requirements during inference and training.☆258Jun 29, 2026Updated 2 months ago
- Scalable Option Learning☆25Aug 9, 2026Updated 3 weeks ago
- Reference implementation of "Self-Attention at Constant Cost per Token via Symmetry-Aware Taylor Approximation" (Heinsen and Kozachkov, 2…☆36Feb 21, 2026Updated 6 months ago
- Gumbel-Softmax post-training quantization for LLMs (1–3 bit scalar, INT/GGUF-compatible).☆36Jul 11, 2026Updated last month
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆24Mar 23, 2026Updated 5 months ago
- Efficient optimizers☆340Aug 29, 2026Updated last week
- Official JAX implementation of End-to-End Test-Time Training for Long Context☆698Feb 15, 2026Updated 6 months ago
- PyTorch Code for Energy-Based Transformers paper -- generalizable reasoning and scalable learning☆658Apr 21, 2026Updated 4 months ago
- End-to-end formally verified solvers for the Maxwell and perfectly hyperbolic Maxwell equations in 1D, 2D, and 3D.☆43Jul 20, 2026Updated last month
- FrontierSmith, a new system that uses AI to synthesize open-ended coding problems at scale☆52May 30, 2026Updated 3 months ago
- Parse and Manipulate R Code☆33Aug 6, 2026Updated last month