Evolution Pretraining Fully in Int Formats
☆181Feb 25, 2026Updated 5 months ago
Alternatives and similar repositories for nano-egg
Users that are interested in nano-egg are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Jax Codebase for Evolutionary Strategies at the Hyperscale☆369Feb 27, 2026Updated 5 months ago
- Official implementation for "How Should We Meta-Learn Reinforcement Learning Algorithms?"☆23Sep 7, 2025Updated 11 months ago
- This repo contains the source code for the paper "Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning"☆379Jun 26, 2026Updated last month
- Neuroevolution Community☆16Nov 17, 2025Updated 9 months ago
- Official implementation of DiscoGen, for "Procedural Generation of Algorithm Discovery Tasks in Machine Learning"☆50Aug 4, 2026Updated 2 weeks ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- A collection of lightweight interpretability scripts to understand how LLMs think☆91Mar 18, 2026Updated 5 months ago
- Internal Coherence Maximization (ICM): A Label-Free, Unsupervised Training Framework for LLMs☆27Sep 5, 2025Updated 11 months ago
- Experiments Notebook of "Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism"☆16Apr 30, 2025Updated last year
- Generic building-block toolbox for training neural networks with adaptive and recursive execution. It provides reusable components to con…☆27Jun 29, 2026Updated last month
- nanoGPT using Equinox☆15Mar 3, 2023Updated 3 years ago
- smolLM with Entropix sampler on pytorch☆148Oct 31, 2024Updated last year
- Example for a Monty-enabled RLM in DSPy☆20Feb 16, 2026Updated 6 months ago
- GCRL in JAX. Official repository for LEO (ICML 2026).☆29Jun 20, 2026Updated last month
- Implementation and explorations into PopuLoRA, Co-Evolving LLM Populations for Reasoning Self-Play☆17Updated this week
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- A simple, performant and scalable JAX-based world modeling codebase.☆162Jan 15, 2026Updated 7 months ago
- Official implementation of the paper "Next Embedding Prediction Makes World Models Stronger"☆37Mar 5, 2026Updated 5 months ago
- [ICLR26] Beyond Real: Imaginary Extension of Rotary Position Embeddings for Long-Context LLMs☆33Dec 9, 2025Updated 8 months ago
- Fast reinforcement learning 💨☆29Jul 15, 2025Updated last year
- Global CoT Analysis: Initial attempts to uncover patterns across many chains of thought☆20Feb 10, 2026Updated 6 months ago
- Marketplace ML experiment - training without backprop☆28Sep 9, 2025Updated 11 months ago
- TEVO: evolve LM motifs cheaply, then validate them in downstream train.py loops.☆19Apr 18, 2026Updated 4 months ago
- Softened ROSA QKV Operators for Training Next-Generation LLM Models☆39Aug 5, 2026Updated 2 weeks ago
- explore token trajectory trees on instruct and base models☆157May 29, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [ICML 2026] NITP: Next Implicit Token Prediction for LLM Pre-training☆35May 26, 2026Updated 2 months ago
- Cost aware hyperparameter tuning algorithm☆191Jun 27, 2024Updated 2 years ago
- Exploring techniques to generate diverse conventions in multi-agent settings☆16Nov 14, 2023Updated 2 years ago
- Cuda kernels for leveraging LLM sparsity to improve throughput and decrease the memory requirements during inference and training.☆257Jun 29, 2026Updated last month
- Reference implementation of "Self-Attention at Constant Cost per Token via Symmetry-Aware Taylor Approximation" (Heinsen and Kozachkov, 2…☆36Feb 21, 2026Updated 5 months ago
- A codebase for pretraining multi-billion-scale sparse GPTs.☆27Feb 9, 2026Updated 6 months ago
- Lottery Tickets in Evolutionary Optimization (Lange & Sprekeler, ICML 2023)☆17Jun 2, 2023Updated 3 years ago
- ☆24Mar 23, 2026Updated 4 months ago
- Efficient optimizers☆339Jul 25, 2026Updated 3 weeks ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Official JAX implementation of End-to-End Test-Time Training for Long Context☆675Feb 15, 2026Updated 6 months ago
- Implementation of Poly-attention, a higher-order self-attention proposed by Chakrabarti et al. of Columbia☆55Jul 22, 2026Updated 3 weeks ago
- PyTorch Code for Energy-Based Transformers paper -- generalizable reasoning and scalable learning☆648Apr 21, 2026Updated 3 months ago
- FrontierSmith, a new system that uses AI to synthesize open-ended coding problems at scale☆50May 30, 2026Updated 2 months ago
- End-to-end formally verified solvers for the Maxwell and perfectly hyperbolic Maxwell equations in 1D, 2D, and 3D.☆42Jul 20, 2026Updated 3 weeks ago
- Parse and Manipulate R Code☆33Aug 6, 2026Updated last week
- Unified Implementations of Offline Reinforcement Learning Algorithms☆228Dec 19, 2025Updated 7 months ago