Implementation for POET and POET-X for LLM pretraining
☆41Jun 9, 2026Updated 3 months ago
Alternatives and similar repositories for poet
Users that are interested in poet are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆39Jul 2, 2026Updated 2 months ago
- Official repository of PEFT-Arena: Understanding Parameter-Efficient Finetuning from a Stability-Plasticity Perspective☆29Jun 13, 2026Updated 3 months ago
- Implementation of <Model Merging with Functional Dual Anchors>☆47Nov 23, 2025Updated 9 months ago
- Stable and Efficient Reinforcement Learning for Trillion-Parameter LLMs☆157Sep 10, 2026Updated last week
- Repository of <FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models> in EMNLP 2026 Findings☆76Aug 24, 2026Updated 3 weeks ago
- Managed Database hosting by DigitalOcean • AdPostgreSQL, MySQL, MongoDB, Kafka, Valkey, and OpenSearch available. Automatically scale up storage and focus on building your apps.
- Implementation for <Orthogonal Over-Parameterized Training> in CVPR'21.☆22Jul 16, 2021Updated 5 years ago
- ☆32Jul 14, 2025Updated last year
- My attempt to improve the speed of the newton schulz algorithm, starting from the dion implementation.☆42Apr 30, 2026Updated 4 months ago
- [ECCV 2026] Official implementation of "TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning"☆25Feb 8, 2026Updated 7 months ago
- M^3PC: Test-Time Model Predictive Control for Pretrained Masked Trajectory Model, ICLR 2025☆19Mar 17, 2025Updated last year
- SimKO: Simple Pass@K Policy Optimization☆31Oct 24, 2025Updated 10 months ago
- [ICLR 2025] Implementation of Accelerating Auto-regressive Text-to-Image Generation with Training-free Speculative Jacobi Decoding☆54Apr 21, 2025Updated last year
- [NeurIPS 2024] DN-4DGS: Denoised Deformable Network with Temporal-Spatial Aggregation for Dynamic Scene Rendering☆12Oct 22, 2024Updated last year
- [ICLR 2026] GRAPE: Group Representational Position Encoding (https://arxiv.org/abs/2512.07805)☆120Jun 15, 2026Updated 3 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- DreamSmooth: Improving Model-Based RL with Reward Smoothing (ICLR 2024)☆13May 6, 2024Updated 2 years ago
- Score-Guided Planning☆16Jun 29, 2026Updated 2 months ago
- A minimal re-implementation of orthogonal fine-tuning (OFT), a diffusion method, for LLMs. Based on nanoGPT and minLoRA.☆14Nov 17, 2023Updated 2 years ago
- Muon in Int8 Precision Made Possible☆20Jun 18, 2026Updated 3 months ago
- The official repository for SkyLadder: Better and Faster Pretraining via Context Window Scheduling☆43Dec 29, 2025Updated 8 months ago
- Benchmarking Optimizers for LLM Pretraining☆61May 3, 2026Updated 4 months ago
- Code for CVPR2018 "Iterative Learning with Open-set Noisy Labels"☆12Mar 12, 2021Updated 5 years ago
- Official implementation of "Continual Learning by Modeling Intra-Class Variation" (MOCA). [TMLR 2023]☆16Mar 3, 2023Updated 3 years ago
- Code for "Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining"☆30Oct 14, 2025Updated 11 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- Understanding deep networks and large models.☆30Jan 23, 2026Updated 7 months ago
- Model-based Offline Policy Optimization re-implement all by pytorch☆44Sep 13, 2023Updated 3 years ago
- Multi-Agent Verification: Scaling Test-Time Compute with Multiple Verifiers☆33Mar 1, 2025Updated last year
- [ICML‘25] Official code for paper "Occult: Optimizing Collaborative Communication across Experts for Accelerated Parallel MoE Training an…☆14Apr 17, 2025Updated last year
- [ICLR 2026] M2-Miner: Multi-Agent Enhanced MCTS for Mobile GUI Agent Data Mining☆55Apr 22, 2026Updated 4 months ago
- [ICML 2026] Less Is More: Training-Free Sparse Attention with Global Locality for Efficient Reasoning☆36Sep 12, 2025Updated last year
- Implementation of Iterative Machine Teaching algorithm with PyTorch.☆10Aug 27, 2023Updated 3 years ago
- ☆12May 19, 2025Updated last year
- ☆29Jan 14, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Kinetics: Rethinking Test-Time Scaling Laws☆87Jul 11, 2025Updated last year
- ☆34Dec 29, 2025Updated 8 months ago
- [ICLR 2025 & COLM 2025] Official PyTorch implementation of the Forgetting Transformer and Adaptive Computation Pruning☆154Feb 25, 2026Updated 6 months ago
- [CVPR2025] MetricGrids: Arbitrary Nonlinear Approximation with Elementary Metric Grids based Implicit Neural Representation☆16Jun 1, 2025Updated last year
- ☆10Jun 28, 2023Updated 3 years ago
- 🤗 PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.☆18Apr 11, 2024Updated 2 years ago
- ☆13Jan 7, 2025Updated last year