☆26Feb 20, 2026Updated 6 months ago
Alternatives and similar repositories for icl-dynamics
Users that are interested in icl-dynamics are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Official Code for What Makes and Breaks Safety Fine-tuning? A Mechanistic Study (NeurIPS 2024)☆11Oct 31, 2024Updated last year
- ☆15Sep 29, 2022Updated 3 years ago
- Official Repository of Paper "Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs"☆15Sep 25, 2025Updated 11 months ago
- [ICML 2024] Code release for "On the Emergence of Cross-Task Linearity in Pretraining-Finetuning Paradigm"☆11Feb 20, 2025Updated last year
- This repository contains the code used for the experiments in the paper "Fine-Tuning Enhances Existing Mechanisms: A Case Study on Entity…☆32Oct 27, 2025Updated 10 months ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Welcome to the 'In Context Learning Theory' Reading Group☆31Nov 8, 2024Updated last year
- Official implementation of "Reasoning by Superposition: A Theoretical Perspective on Chain of Continuous Thought" (NeurIPS 2025)☆44Oct 8, 2025Updated 10 months ago
- 6,080-param transformer achieving 100% accuracy on 10-digit addition. Trained from scratch in 10 minutes.☆22Feb 19, 2026Updated 6 months ago
- The official implementation of A Unified Game-Theoretic Interpretation of Adversarial Robustness.☆22Jun 9, 2022Updated 4 years ago
- Minimum Description Length probing for neural network representations☆20Jan 28, 2025Updated last year
- Code for paper "Robustness of Bayesian Neural Networks to Gradient-Based Attacks"☆17Feb 26, 2024Updated 2 years ago
- [ICLR 2025 Spotlight] Code release for "Sharpness-Aware Minimization Efficiently Selects Flatter Minima Late In Training"☆20Feb 20, 2025Updated last year
- Official code release for Delta Activations: A Representation for Finetuned Large Language Models☆21Sep 5, 2025Updated 11 months ago
- A benchmark for mechanistic discovery of circuits in Transformers☆18Dec 15, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- How do transformer LMs encode relations?☆60Feb 24, 2024Updated 2 years ago
- The official code of FineRMoE.☆23Mar 17, 2026Updated 5 months ago
- [ICCV 2023] Black Box Few-Shot Adaptation for Vision-Language models☆27May 14, 2024Updated 2 years ago
- PyTorch and NNsight implementation of AtP* (Kramar et al 2024, DeepMind)☆21Jan 19, 2025Updated last year
- Code for "Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining"☆30Oct 14, 2025Updated 10 months ago
- [NeurIPS2024] Fast T2T: Optimization Consistency Speeds Up Diffusion-Based Training-to-Testing Solving for Combinatorial Optimization; [N…☆22Jul 2, 2025Updated last year
- Some thoughts about writing scientific papers☆23Nov 8, 2024Updated last year
- Code for "Evidence of Learned Look-Ahead in a Chess-Playing Neural Network"☆31Jun 4, 2024Updated 2 years ago
- ☆35Jul 5, 2023Updated 3 years ago
- Deploy open-source AI quickly and easily - Special Bonus Offer • AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Code accompanying the ICML'24 paper "Feature Contamination: Neural Networks Learn Uncorrelated Features and Fail to Generalize"☆22Feb 13, 2025Updated last year
- ☆23Mar 11, 2026Updated 5 months ago
- Multi-Layer Sparse Autoencoders (ICLR 2025)☆30Feb 6, 2026Updated 6 months ago
- Supporting code for the blog post on modular manifolds.☆129Sep 26, 2025Updated 11 months ago
- 📄🕸️ Generalizing Cross-Document Event Coreference Resolution Across Multiple Corpora☆10May 25, 2022Updated 4 years ago
- Mechanistic Interpretability for Transformer Models☆54Jun 1, 2022Updated 4 years ago
- ☆35Nov 30, 2025Updated 9 months ago
- FeatureAlignment = Alignment + Mechanistic Interpretability☆35Mar 8, 2025Updated last year
- ⚓️ Repository for the "Thought Anchors: Which LLM Reasoning Steps Matter?" paper.☆142Oct 27, 2025Updated 10 months ago
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- ☆15Apr 15, 2026Updated 4 months ago
- [CVPR 2024] Friendly Sharpness-Aware Minimization☆37Oct 29, 2024Updated last year
- ☆23Jun 2, 2026Updated 2 months ago
- BH hackathon☆14Apr 4, 2024Updated 2 years ago
- ☆10Jan 20, 2023Updated 3 years ago
- A remote Scala code evaluator☆14May 16, 2023Updated 3 years ago
- A toolkit that provides a range of model diffing techniques including a UI to visualize them interactively.☆82Jul 20, 2026Updated last month