Research without Re-search: Maximal Update Parametrization Yields Accurate Loss Prediction across Scales
☆33Jul 17, 2023Updated 3 years ago
Alternatives and similar repositories for Mu-scaling
Users that are interested in Mu-scaling are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆18Sep 5, 2024Updated 2 years ago
- 智源研究院、旋玑智能、面壁智能发布的具身智能操作系统☆25Apr 8, 2026Updated 5 months ago
- Open Source Implementation of Dual Modality MAGVIT2 Tokenizer☆26Nov 26, 2024Updated last year
- ☆112Jul 15, 2025Updated last year
- Implementation of OpenAI paper with Simple Noise Scale on Fastai V2☆19Apr 16, 2021Updated 5 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Is In-Context Learning Sufficient for Instruction Following in LLMs? [ICLR 2025]☆34Jan 23, 2025Updated last year
- 微信小程序-答题练习☆12Feb 4, 2023Updated 3 years ago
- Implementation for the paper "Fictitious Synthetic Data Can Improve LLM Factuality via Prerequisite Learning"☆11Jan 10, 2025Updated last year
- The official repository for SkyLadder: Better and Faster Pretraining via Context Window Scheduling☆43Dec 29, 2025Updated 8 months ago
- Fluid Language Model Benchmarking☆29Sep 16, 2025Updated last year
- ☆13Jun 4, 2024Updated 2 years ago
- ☆14Mar 2, 2025Updated last year
- Codebase for Instruction Following without Instruction Tuning☆36Sep 24, 2024Updated 2 years ago
- Combining SOAP and MUON☆25Feb 11, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Masked Structural Growth for 2x Faster Language Model Pre-training☆26Apr 28, 2024Updated 2 years ago
- Crop Yield Prediction using various ML approaches - Random-Forest Regressor, Gradient-Boosting Regressor, Decision-Tree Regressor, Suppo…☆11Jul 13, 2023Updated 3 years ago
- Code for paper 'Are We Falling in a Middle-Intelligence Trap? An Analysis and Mitigation of the Reversal Curse'☆14Aug 2, 2024Updated 2 years ago
- Self-Supervised Alignment with Mutual Information☆20May 24, 2024Updated 2 years ago
- personal settings for linux tools, including zsh, vim, tmux, pip.☆11Dec 2, 2019Updated 6 years ago
- Longitudinal Evaluation of LLMs via Data Compression☆32May 29, 2024Updated 2 years ago
- ☆38Oct 4, 2025Updated 11 months ago
- Code for the paper "Function-Space Learning Rates"☆23Jun 3, 2025Updated last year
- [NAACL'24] Self-data filtering of LLM instruction-tuning data using a novel perplexity-based difficulty score, without using any other mo…☆420Jun 25, 2025Updated last year
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Stick-breaking attention☆64Jul 1, 2025Updated last year
- Pytorch implementation of "Oscillation-Reduced MXFP4 Training for Vision Transformers" on DeiT Model Pre-training☆41May 4, 2026Updated 4 months ago
- ☆47Jun 11, 2025Updated last year
- ☆65Jun 12, 2025Updated last year
- [ICLR 2026] M2-Miner: Multi-Agent Enhanced MCTS for Mobile GUI Agent Data Mining☆55Apr 22, 2026Updated 5 months ago
- A python wrapper for Stanford CoreNLP, simple and customizable.☆13Oct 26, 2021Updated 4 years ago
- Code for NeurIPS 2024 Spotlight: "Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations"☆94Oct 30, 2024Updated last year
- Minimal (400 LOC) implementation Maximum (multi-node, FSDP) GPT training☆132Apr 17, 2024Updated 2 years ago
- GoldFinch and other hybrid transformer components☆46Jul 20, 2024Updated 2 years ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- ☆14Jul 21, 2022Updated 4 years ago
- Rice Yield CNN is a model to estimate the rice yield based on RGB image of rice canopy at harvest. The model is developed based on more t…☆13Nov 8, 2021Updated 4 years ago
- ☆11Feb 2, 2023Updated 3 years ago
- [ICLR 2026] BARREL: Boundary-Aware Reasoning for Factual and Reliable LRMs☆18May 21, 2025Updated last year
- Public repository for content related to the the Plotline project.☆15Apr 21, 2026Updated 5 months ago
- Implementation of NAACL 2024 Outstanding Paper "LM-Infinite: Simple On-the-Fly Length Generalization for Large Language Models"☆153Mar 13, 2025Updated last year
- Spectral Sphere Optimizer☆133Mar 23, 2026Updated 6 months ago