Code repository for the paper "MrT5: Dynamic Token Merging for Efficient Byte-level Language Models."
☆59Sep 25, 2025Updated 10 months ago
Alternatives and similar repositories for mrt5
Users that are interested in mrt5 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- PathPiece tokenizer☆14Nov 10, 2024Updated last year
- MARs: Multi-view Attention Regularizations for Patch-based Feature Recognition of Space Terrain☆11Nov 20, 2024Updated last year
- [ICML 2024] VQDNA: Unleashing the Power of Vector Quantization for Multi-Species Genomic Sequence Modeling☆10Sep 22, 2024Updated last year
- ☆11Nov 18, 2024Updated last year
- Official repository for the paper "Approximating Two-Layer Feedforward Networks for Efficient Transformers"☆39Jun 11, 2025Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Code for the paper "Greed is All You Need: An Evaluation of Tokenizer Inference Methods"☆15Nov 26, 2024Updated last year
- ModuleFormer is a MoE-based architecture that includes two different types of experts: stick-breaking attention heads and feedforward exp…☆225Sep 18, 2025Updated 10 months ago
- Python package for Natural Language Processing (NLP), focused on low-resource languages spoken in Mexico.☆24Sep 4, 2025Updated 11 months ago
- ☆60Nov 18, 2025Updated 8 months ago
- [CVPR'25] MergeVQ: A Unified Framework for Visual Generation and Representation with Token Merging and Quantization☆51Jul 22, 2025Updated last year
- [NeurIPS 2024] Image Understanding Makes for A Good Tokenizer for Image Generation☆21Dec 17, 2024Updated last year
- ☆18Jun 12, 2023Updated 3 years ago
- Efficient encoder-decoder architecture for small language models (≤1B parameters) with cross-architecture knowledge distillation and visi…☆32Feb 7, 2025Updated last year
- Flash-Linear-Attention models beyond language☆21Aug 28, 2025Updated 11 months ago
- Virtual machines for every use case on DigitalOcean • AdGet dependable uptime with 99.99% SLA, simple security tools, and predictable monthly pricing with DigitalOcean's virtual machines, called Droplets.
- Tool to perform paired evaluation of automatic systems☆13Oct 20, 2021Updated 4 years ago
- Landing repository for the paper "Predicting the Order of Upcoming Tokens Improves Language Modeling"☆48May 13, 2026Updated 3 months ago
- ☆45Jul 5, 2022Updated 4 years ago
- [ICLR 2025 Oral] "Your Mixture-of-Experts LLM Is Secretly an Embedding Model For Free"☆93Oct 15, 2024Updated last year
- Landing repository for the paper "Softpick: No Attention Sink, No Massive Activations with Rectified Softmax"☆92Sep 12, 2025Updated 11 months ago
- A high-efficiency text embedding and reranking model based on RWKV architecture.☆20Jan 10, 2026Updated 7 months ago
- Fuel innovation and advance language models with HomoScriptor: A vibrant, community-driven dataset for fine-tuning large language models.☆18Oct 14, 2023Updated 2 years ago
- Implementation of MambaFormer in Pytorch ++ Zeta from the paper: "Can Mamba Learn How to Learn? A Comparative Study on In-Context Learnin…☆23Aug 3, 2026Updated last week
- Linear Attention for Efficient Bidirectional Sequence Modeling☆18May 13, 2025Updated last year
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- The offcial repository for 'CharacterBERT and Self-Teaching for Improving the Robustness of Dense Retrievers on Queries with Typos', SIGI…☆16May 4, 2022Updated 4 years ago
- ☆13Aug 19, 2024Updated last year
- ☆21Mar 1, 2023Updated 3 years ago
- Speed up Transformers With Spectrum-Preserving Token Merging☆57Apr 30, 2026Updated 3 months ago
- ☆15Mar 20, 2025Updated last year
- Code for the paper "BPE stays on SCRIPT", "Which Pieces Does Unigram Tokenization Really Need?" and MinGram☆19Updated this week
- new optimizer☆20Aug 4, 2024Updated 2 years ago
- triple-encoders is a library for contextualizing distributed Sentence Transformers representations.☆15Sep 3, 2024Updated last year
- ☆25Oct 13, 2024Updated last year
- Deploy on Railway without the complexity - Free Credits Offer • AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- SCT: An Efficient Self-Supervised Cross-View Training For Sentence Embedding (TACL)☆16Jul 27, 2024Updated 2 years ago
- Official codebase for our paper "Do Language Models Use Their Depth Efficiently?"☆29Jun 25, 2025Updated last year
- [ICLR 2025] Official Pytorch Implementation of "Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN" by Pengxia…☆30Jul 24, 2025Updated last year
- A Shared Nearest Neighbors clustering implementation. This code is basically a wrapper of sklearn DBSCAN, implementing the neighborhood s…☆16Jan 10, 2022Updated 4 years ago
- Implementation of Kronecker Attention in Pytorch☆20Sep 12, 2020Updated 5 years ago
- Official implementation of "Data Mixture Inference: What do BPE tokenizers reveal about their training data?"☆23May 15, 2025Updated last year
- Official repository of the paper "JIST: Joint Image and Sequence Training for Sequential Visual Place Recognition"☆24Dec 15, 2023Updated 2 years ago