☆260Jun 6, 2025Updated last year
Alternatives and similar repositories for meliad
Users that are interested in meliad are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Implementation of Memorizing Transformers (ICLR 2022), attention net augmented with indexing and retrieval of memories using approximate …☆646Jul 17, 2023Updated 3 years ago
- Implementation of Block Recurrent Transformer - Pytorch☆226Aug 20, 2024Updated last year
- The official Languini Kitchen repository☆14May 6, 2024Updated 2 years ago
- ☆54Jan 19, 2023Updated 3 years ago
- Source-to-Source Debuggable Derivatives in Pure Python☆15Jan 23, 2024Updated 2 years ago
- End-to-end encrypted cloud storage - Proton Drive • AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- Sequence modeling with Mega.☆302Jan 28, 2023Updated 3 years ago
- Demonstration that finetuning RoPE model on larger sequences than the pre-trained model adapts the model context limit☆62Jun 21, 2023Updated 3 years ago
- ☆13Aug 23, 2024Updated last year
- FlexAttention w/ FlashAttention3 Support☆27Oct 5, 2024Updated last year
- An implementation of local windowed attention for language modeling☆503Jul 16, 2025Updated last year
- playing with gpt4☆13Mar 17, 2023Updated 3 years ago
- Convenient Text-to-Text Training for Transformers☆18Dec 10, 2021Updated 4 years ago
- ☆23Oct 15, 2022Updated 3 years ago
- HyPe: Better Pre-trained Language Model Fine-tuning with Hidden Representation Perturbation [ACL 2023]☆14Jul 11, 2023Updated 3 years ago
- End-to-end encrypted email - Proton Mail • AdSpecial offer: 40% Off Yearly / 80% Off First Month. All Proton services are open source and independently audited for security.
- Understand and test language model architectures on synthetic tasks.☆283Mar 22, 2026Updated 4 months ago
- Convolutions for Sequence Modeling☆916Jun 13, 2024Updated 2 years ago
- Sequence Modeling with Structured State Spaces☆69Aug 2, 2022Updated 4 years ago
- Large Context Attention☆774Oct 13, 2025Updated 10 months ago
- Fine-Tuning Pre-trained Transformers into Decaying Fast Weights☆20Oct 9, 2022Updated 3 years ago
- ☆10Dec 17, 2020Updated 5 years ago