☆47Oct 11, 2023Updated 2 years ago
Alternatives and similar repositories for transformer_in_transformer
Users that are interested in transformer_in_transformer are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Q-Probe: A Lightweight Approach to Reward Maximization for Language Models☆40Jun 10, 2024Updated 2 years ago
- Code for our ACL '23 paper titled "Grokking of Hierarchical Structure in Vanilla Transformers"☆26Oct 8, 2023Updated 2 years ago
- [NeurIPS 2022] Your Transformer May Not be as Powerful as You Expect (official implementation)☆35Aug 6, 2023Updated 3 years ago
- ☆53Jun 10, 2024Updated 2 years ago
- Code for the paper "Distinguishing the Knowable from the Unknowable with Language Models"☆11Jul 18, 2026Updated last month
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆10Jun 28, 2022Updated 4 years ago
- Repository for the code and dataset for the paper: "Have LLMs Advanced enough? Towards Harder Problem Solving Benchmarks For Large Langu…☆39Dec 18, 2023Updated 2 years ago
- Universal Neurons in GPT2 Language Models☆30May 28, 2024Updated 2 years ago
- Study of Pre-Trained Positional Embeddings☆16Nov 6, 2020Updated 5 years ago
- The offcial repository for 'CharacterBERT and Self-Teaching for Improving the Robustness of Dense Retrievers on Queries with Typos', SIGI…☆16May 4, 2022Updated 4 years ago
- ☆15Jul 24, 2022Updated 4 years ago
- T-Projection is a method to perform high-quality Annotation Projection of Sequence Labeling datasets.☆13Nov 21, 2023Updated 2 years ago
- ☆12Jul 4, 2024Updated 2 years ago
- 🧮 Algebraic Positional Encodings.☆21Jun 5, 2026Updated 2 months ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆11May 24, 2024Updated 2 years ago
- ☆19Jun 10, 2024Updated 2 years ago
- [NeurIPS 2023] MeZO: Fine-Tuning Language Models with Just Forward Passes. https://arxiv.org/abs/2305.17333☆1,173Jan 11, 2024Updated 2 years ago
- Memory Mosaics are networks of associative memories working in concert to achieve a prediction task.☆64Aug 3, 2026Updated 2 weeks ago
- Code for Fast-weight Product Key Memory (FwPKM)☆21Mar 18, 2026Updated 5 months ago
- DeltaProduct is a new linear recurrent neural network architecture that uses products of generalized Householder matrices as state-transi…☆17Oct 13, 2025Updated 10 months ago
- Why Do We Need Weight Decay in Modern Deep Learning? [NeurIPS 2024]☆73Sep 25, 2024Updated last year
- Implementation of average- and worst-case robust flatness measures for adversarial training.☆15Nov 5, 2021Updated 4 years ago
- ☆83Apr 16, 2024Updated 2 years ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Experiments for "A Closer Look at In-Context Learning under Distribution Shifts"☆18May 29, 2023Updated 3 years ago
- Code for 'The Linear Representation Hypothesis and the Geometry of Large Language Models' (ICML 2024)☆127Feb 11, 2025Updated last year
- Align, a general text alignment function☆15Dec 7, 2023Updated 2 years ago
- TART: A plug-and-play Transformer module for task-agnostic reasoning☆201Jun 22, 2023Updated 3 years ago
- ☆12Nov 22, 2024Updated last year
- This repository is for the paper "Is BERT Blind? Exploring the Effect of Vision-and-Language Pretraining on Visual Language Understanding…☆21Nov 2, 2023Updated 2 years ago
- This repository contains source codes for SoftCTC. Original paper can be found here: https://arxiv.org/abs/2212.02135☆19Mar 7, 2023Updated 3 years ago
- Discretized Integrated Gradients for Explaining Language Models (EMNLP 2021)☆27Mar 26, 2022Updated 4 years ago
- MDL Complexity computations and experiments from the paper "Revisiting complexity and the bias-variance tradeoff".☆18Jun 12, 2023Updated 3 years ago
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- A modern look at the relationship between sharpness and generalization [ICML 2023]☆44Sep 11, 2023Updated 2 years ago
- ☆16Updated this week
- Official code for the paper "Attention as a Hypernetwork"☆59Feb 24, 2026Updated 5 months ago
- Efficient empirical NTKs in PyTorch☆22Jun 13, 2022Updated 4 years ago
- The Full Spectrum of Deepnet Hessians at Scale: Dynamics with SGD Training and Sample Size☆19May 19, 2019Updated 7 years ago
- Repository for reproducing `Model-Based Robust Deep Learning`☆17Jan 22, 2021Updated 5 years ago
- ☆14Nov 15, 2022Updated 3 years ago