Implementation for POET and POET-X for LLM pretraining
☆41Jun 9, 2026Updated 4 months ago
Alternatives and similar repositories for poet
Users that are interested in poet are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆39Sep 25, 2026Updated 2 weeks ago
- Official repository of PEFT-Arena: Understanding Parameter-Efficient Finetuning from a Stability-Plasticity Perspective☆29Jun 13, 2026Updated 3 months ago
- Implementation of <Orthogonal Model Merging>☆37Aug 5, 2026Updated 2 months ago
- Implementation of <Symbolic Graphics Programming with Large Language Models>☆39Sep 14, 2025Updated last year
- Implementation of <Model Merging with Functional Dual Anchors>☆47Nov 23, 2025Updated 10 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Repository of <FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models> in EMNLP 2026 Findings☆76Aug 24, 2026Updated last month
- My attempt to improve the speed of the newton schulz algorithm, starting from the dion implementation.☆42Apr 30, 2026Updated 5 months ago
- [ECCV 2026] Official implementation of "TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning"☆26Feb 8, 2026Updated 8 months ago
- SimKO: Simple Pass@K Policy Optimization☆31Oct 24, 2025Updated 11 months ago
- [ICLR 2025] Implementation of Accelerating Auto-regressive Text-to-Image Generation with Training-free Speculative Jacobi Decoding☆54Apr 21, 2025Updated last year
- [NeurIPS 2024] DN-4DGS: Denoised Deformable Network with Temporal-Spatial Aggregation for Dynamic Scene Rendering☆12Oct 22, 2024Updated last year
- [ICLR 2026] GRAPE: Group Representational Position Encoding (https://arxiv.org/abs/2512.07805)☆120Jun 15, 2026Updated 3 months ago
- MuZero for Combinatorial Action Spaces: open-source codebase for MA-Gumbel-AlphaZero, MA-Sampled-AlphaZero, MA-Gumbel-MuZero and MA-Sampl…☆24Jan 22, 2024Updated 2 years ago
- DreamSmooth: Improving Model-Based RL with Reward Smoothing (ICLR 2024)☆13May 6, 2024Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Easy no-frills Jax implementations of common abstractions for simple diffusion models.☆12Feb 23, 2026Updated 7 months ago
- A minimal re-implementation of orthogonal fine-tuning (OFT), a diffusion method, for LLMs. Based on nanoGPT and minLoRA.☆14Nov 17, 2023Updated 2 years ago
- The official repository for SkyLadder: Better and Faster Pretraining via Context Window Scheduling☆44Dec 29, 2025Updated 9 months ago
- ICML2025-Inductive Gradient Adjustment for Spectral Bias in Implicit Neural Representations☆47May 31, 2025Updated last year
- Benchmarking Optimizers for LLM Pretraining☆61May 3, 2026Updated 5 months ago
- [NeurIPS 2023] MoVie: Visual Model-Based Policy Adaptation for View Generalization☆12Sep 22, 2023Updated 3 years ago
- This repository is an open source implementation of the MuonClip strategy from the KIMI K2 Model from Moonshot AI☆29Nov 7, 2025Updated 11 months ago
- Code for CVPR2018 "Iterative Learning with Open-set Noisy Labels"☆12Mar 12, 2021Updated 5 years ago
- Official implementation of "Continual Learning by Modeling Intra-Class Variation" (MOCA). [TMLR 2023]☆16Mar 3, 2023Updated 3 years ago
- Simple, predictable pricing with DigitalOcean hosting • AdAlways know what you'll pay with monthly caps and flat pricing. Enterprise-grade infrastructure trusted by 600k+ customers.
- Code for "Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining"☆30Oct 14, 2025Updated 11 months ago
- MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models☆458Feb 1, 2024Updated 2 years ago
- ☆114Jul 24, 2025Updated last year
- Understanding deep networks and large models.☆30Jan 23, 2026Updated 8 months ago
- Puzzle Generator; Einstein's Riddle, Zebra Puzzle and Blood Donation Puzzle Solver. For non-commercial use only!☆20Mar 4, 2023Updated 3 years ago
- The official code of "Mano: Restriking Manifold Optimization for LLM Training".☆25Jun 1, 2026Updated 4 months ago
- Multi-Agent Verification: Scaling Test-Time Compute with Multiple Verifiers☆33Mar 1, 2025Updated last year
- [ICML‘25] Official code for paper "Occult: Optimizing Collaborative Communication across Experts for Accelerated Parallel MoE Training an…☆14Apr 17, 2025Updated last year
- [ICLR 2026] M2-Miner: Multi-Agent Enhanced MCTS for Mobile GUI Agent Data Mining☆55Apr 22, 2026Updated 5 months ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Official code for ICML 2024 paper Reinformer: Max-Return Sequence Modeling for offline RL☆49Oct 16, 2024Updated last year
- A latest curated list of resources on implicit neural representations.☆18Apr 18, 2025Updated last year
- [ICML 2026] Less Is More: Training-Free Sparse Attention with Global Locality for Efficient Reasoning☆36Sep 12, 2025Updated last year
- Implementation of Iterative Machine Teaching algorithm with PyTorch.☆10Aug 27, 2023Updated 3 years ago
- ☆12May 19, 2025Updated last year
- ☆29Jan 14, 2025Updated last year
- Kinetics: Rethinking Test-Time Scaling Laws☆87Jul 11, 2025Updated last year