Code repository for the paper "MrT5: Dynamic Token Merging for Efficient Byte-level Language Models."
☆59Sep 25, 2025Updated 11 months ago
Alternatives and similar repositories for mrt5
Users that are interested in mrt5 are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- PathPiece tokenizer☆14Nov 10, 2024Updated last year
- Digital texts in Prakrit☆11Sep 14, 2025Updated last year
- MARs: Multi-view Attention Regularizations for Patch-based Feature Recognition of Space Terrain☆11Nov 20, 2024Updated last year
- ☆11Nov 18, 2024Updated last year
- Official repository for the paper "Approximating Two-Layer Feedforward Networks for Efficient Transformers"☆39Jun 11, 2025Updated last year
- GPUs on demand by Runpod - Special Offer Available • AdRun AI, ML, and HPC workloads on powerful cloud GPUs—without limits or wasted spend. Deploy GPUs in under a minute and pay by the second.
- ModuleFormer is a MoE-based architecture that includes two different types of experts: stick-breaking attention heads and feedforward exp…☆225Sep 18, 2025Updated last year
- Python package for Natural Language Processing (NLP), focused on low-resource languages spoken in Mexico.☆24Sep 4, 2025Updated last year
- This repo contains VPR models that have been fine-tuned for indoor usage.☆16May 15, 2024Updated 2 years ago
- ☆60Nov 18, 2025Updated 10 months ago
- [NeurIPS 2024] Image Understanding Makes for A Good Tokenizer for Image Generation☆21Dec 17, 2024Updated last year
- ☆18Jun 12, 2023Updated 3 years ago
- Official repository for BMVC 2022 paper: Global Proxy-based Hard Mining for Visual Place Recognition☆18Mar 7, 2023Updated 3 years ago
- Efficient encoder-decoder architecture for small language models (≤1B parameters) with cross-architecture knowledge distillation and visi…☆33Feb 7, 2025Updated last year
- Flash-Linear-Attention models beyond language☆21Aug 28, 2025Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- Landing repository for the paper "Predicting the Order of Upcoming Tokens Improves Language Modeling"☆48May 13, 2026Updated 4 months ago
- [ICLR 2025 Oral] "Your Mixture-of-Experts LLM Is Secretly an Embedding Model For Free"☆93Oct 15, 2024Updated last year
- Landing repository for the paper "Softpick: No Attention Sink, No Massive Activations with Rectified Softmax"☆92Sep 12, 2025Updated last year
- ☆16Dec 9, 2023Updated 2 years ago
- A high-efficiency text embedding and reranking model based on RWKV architecture.☆22Sep 12, 2026Updated last week
- Fuel innovation and advance language models with HomoScriptor: A vibrant, community-driven dataset for fine-tuning large language models.☆18Oct 14, 2023Updated 2 years ago
- Implementation of MambaFormer in Pytorch ++ Zeta from the paper: "Can Mamba Learn How to Learn? A Comparative Study on In-Context Learnin…☆23Aug 29, 2026Updated 3 weeks ago
- A latent-variable model for learning bilingual word embedding mappings☆19Feb 11, 2019Updated 7 years ago
- This repository contains all code and data for the Inside Out Visual Place Recognition task☆23Nov 24, 2021Updated 4 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- An NLP pipeline for Hebrew☆43Jun 16, 2025Updated last year
- ☆13Aug 19, 2024Updated 2 years ago
- ☆22Mar 1, 2023Updated 3 years ago
- Speed up Transformers With Spectrum-Preserving Token Merging☆59Apr 30, 2026Updated 4 months ago
- ☆15Mar 20, 2025Updated last year
- Code for the paper "BPE stays on SCRIPT", "Which Pieces Does Unigram Tokenization Really Need?" and MinGram☆22Updated this week
- Adding new tasks to T0 without catastrophic forgetting☆33Oct 20, 2022Updated 3 years ago
- new optimizer☆20Aug 4, 2024Updated 2 years ago
- triple-encoders is a library for contextualizing distributed Sentence Transformers representations.☆15Sep 3, 2024Updated 2 years ago
- Wordpress hosting with auto-scaling - Free Trial Offer • AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- ☆25Oct 13, 2024Updated last year
- SCT: An Efficient Self-Supervised Cross-View Training For Sentence Embedding (TACL)☆16Jul 27, 2024Updated 2 years ago
- Official codebase for our paper "Do Language Models Use Their Depth Efficiently?"☆29Jun 25, 2025Updated last year
- Experiments for efforts to train a new and improved t5☆76Apr 15, 2024Updated 2 years ago
- [ICLR 2025 & COLM 2025] Official PyTorch implementation of the Forgetting Transformer and Adaptive Computation Pruning☆154Feb 25, 2026Updated 6 months ago
- [ICLR 2025] Official Pytorch Implementation of "Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN" by Pengxia…☆30Jul 24, 2025Updated last year
- Statewide Visual Geolocalization in the Wild (ECCV 2024)☆75Dec 2, 2024Updated last year