Welcome to the 'In Context Learning Theory' Reading Group
☆31Nov 8, 2024Updated last year
Alternatives and similar repositories for Awesome_Large_Foundation_Model_Theory
Users that are interested in Awesome_Large_Foundation_Model_Theory are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Welcome to the Awesome Feature Learning in Deep Learning Thoery Reading Group! This repository serves as a collaborative platform for sch…☆210Apr 13, 2026Updated 5 months ago
- SLTrain: a sparse plus low-rank approach for parameter and memory efficient pretraining (NeurIPS 2024)☆39Nov 1, 2024Updated last year
- This repo contains papers, books, tutorials and resources on Riemannian optimization.☆66Sep 11, 2026Updated 2 weeks ago
- Implementation of "RaanA: A Fast, Flexible, and Data-Efficient Post-Training Quantization Algorithm"☆18Apr 11, 2025Updated last year
- [KDD 2023] code for "Test accuracy vs. generalization gap: model selection in NLP without accessing training or testing data" https://arx…☆12Oct 17, 2022Updated 3 years ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- [NeurIPS 2023] Code release for "Going Beyond Linear Mode Connectivity: The Layerwise Linear Feature Connectivity"☆19Oct 19, 2023Updated 2 years ago
- Code to generate figures of paper "When do spectral gradient updates help in deep learning?"☆17Dec 3, 2025Updated 9 months ago
- ☆28Feb 20, 2026Updated 7 months ago
- [ICLR 2025 Spotlight] Code release for "Sharpness-Aware Minimization Efficiently Selects Flatter Minima Late In Training"☆20Feb 20, 2025Updated last year
- [NeurIPS 2025] Official implementation for our paper "Scaling Diffusion Transformers Efficiently via μP".☆101Nov 2, 2025Updated 10 months ago
- ☆18Jan 17, 2024Updated 2 years ago
- [ICML 2024] Code release for "On the Emergence of Cross-Task Linearity in Pretraining-Finetuning Paradigm"☆11Feb 20, 2025Updated last year
- A toy eval suite for tracing generalization dynamics of LM pre-training☆22May 19, 2026Updated 4 months ago
- [NeurIPS 2023 Spotlight] Temperature Balancing, Layer-wise Weight Analysis, and Neural Network Training☆37Apr 7, 2025Updated last year
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- The official implementation of A Unified Game-Theoretic Interpretation of Adversarial Robustness.☆22Jun 9, 2022Updated 4 years ago
- This is the repo for constructing a comprehensive and rigorous evaluation framework for LLM calibration.☆14Apr 9, 2024Updated 2 years ago
- An Elegant Library for Bayesian Deep Learning in PyTorch☆27Dec 19, 2022Updated 3 years ago
- Implementation of the Regularized Nonlinear Acceleration algorithm☆13Oct 4, 2018Updated 7 years ago
- Find context neurons in Pythia models.☆13Jun 13, 2023Updated 3 years ago
- ☆27Apr 11, 2023Updated 3 years ago
- AnchorAttention: Improved attention for LLMs long-context training☆216Jan 15, 2025Updated last year
- SimPER: A Minimalist Approach to Preference Alignment without Hyperparameters (ICLR 2025)☆17Aug 22, 2025Updated last year
- ☆13Feb 2, 2022Updated 4 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- ☆122Feb 25, 2025Updated last year
- SLURM Tutorial☆28May 29, 2025Updated last year
- The official implementation of the paper "Large Scale Knowledge Washing"☆10Jun 12, 2024Updated 2 years ago
- PyTorch implementation of the paper "Discovering and Explaining the Representation Bottleneck of DNNs" (ICLR 2022 Oral)☆37Oct 30, 2024Updated last year
- A brief and partial summary of RLHF algorithms.☆154Mar 4, 2025Updated last year
- Scaling Sparse Fine-Tuning to Large Language Models☆20Jan 31, 2024Updated 2 years ago
- ☆15May 2, 2026Updated 4 months ago
- The official implementation of the paper "Self-Updatable Large Language Models by Integrating Context into Model Parameters"☆16May 18, 2025Updated last year
- ☆20Oct 3, 2019Updated 6 years ago
- Proton VPN Special Offer - Get 70% off • AdSpecial partner offer. Trusted by over 100 million users worldwide. Tested, Approved and Recommended by Experts.
- This is a list of peer-reviewed representative papers on deep learning dynamics (optimization dynamics of neural networks). The success o…☆305Apr 10, 2024Updated 2 years ago
- ☆12Jul 4, 2024Updated 2 years ago
- Code for lin-RFM used for sparse recovery tasks☆17Mar 13, 2025Updated last year
- Code for Paper: Learning Implicit Representation for Reconstructing Articulated Objects☆27Jun 5, 2024Updated 2 years ago
- Code for the paper: Why Transformers Need Adam: A Hessian Perspective☆66Mar 11, 2025Updated last year
- mHC-lite: You Don’t Need 20 Sinkhorn-Knopp Iterations☆94Jan 12, 2026Updated 8 months ago
- ☆17Mar 23, 2020Updated 6 years ago