[ICLR 2026 π₯] Dr.LLM: Dynamic Layer Routing in LLMs
β55Apr 24, 2026Updated 2 months ago
Alternatives and similar repositories for dr-llm
Users that are interested in dr-llm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- WorldCache: Content-Aware Caching for Accelerated Video World Modelsβ21Jun 28, 2026Updated 3 weeks ago
- The official implementation of HybridNorm: Towards Stable and Efficient Transformer Training via Hybrid Normalizationβ19Mar 7, 2025Updated last year
- [NAACL 2025 π₯] CAMEL-Bench is an Arabic benchmark for evaluating multimodal models across eight domains with 29,000 questions.β38Apr 17, 2025Updated last year
- β19Jul 24, 2023Updated 2 years ago
- transformer layers behavior as paintersπ§βπ¨β15May 6, 2025Updated last year
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [ICCV2025] Hierarchical Visual Prompt Learning for Continual Video Instance Segmentationβ14Feb 18, 2026Updated 5 months ago
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Modelsβ19Jan 21, 2026Updated 6 months ago
- β17Jul 24, 2023Updated 2 years ago
- Code and data for "Timo: Towards Better Temporal Reasoning for Language Models" (COLM 2024)β26Oct 23, 2024Updated last year
- [NAACL'25 π SAC Award] Official code for "Advancing MoE Efficiency: A Collaboration-Constrained Routing (C2R) Strategy for Better Expertβ¦β16Feb 4, 2025Updated last year
- Official Repository of Paper "Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs"β15Sep 25, 2025Updated 9 months ago
- Source code of "Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers" EMNLP 2025β17Jan 12, 2026Updated 6 months ago
- Official repository of paper titled "D3Former: Debiased Dual Distilled Transformer for Incremental Learning".β25Jul 10, 2023Updated 3 years ago
- xRouter: Training Cost-Aware LLMs Orchestration System via Reinforcement Learningβ31Jun 2, 2026Updated last month
- End-to-end encrypted cloud storage - Proton Drive β’ AdSpecial offer: 40% Off Yearly / 80% Off First Month. Protect your most important files, photos, and documents from prying eyes.
- [arxiv: 2512.19673] Bottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal Policiesβ60Feb 6, 2026Updated 5 months ago
- [NeurIPS 2024] BLAST: Block Level Adaptive Structured Matrix for Efficient Deep Neural Network Inferenceβ18Nov 6, 2024Updated last year
- [CVPR 2025 π₯]A Large Multimodal Model for Pixel-Level Visual Grounding in Videosβ104Apr 14, 2025Updated last year
- β26May 19, 2026Updated 2 months ago
- [NAACL 2022] "Learning to Win Lottery Tickets in BERT Transfer via Task-agnostic Mask Training", Yuanxin Liu, Fandong Meng, Zheng Lin, Peβ¦β15Oct 18, 2022Updated 3 years ago
- Official repo of dataset-decomposition paper [NeurIPS 2024]β21Jan 8, 2025Updated last year
- Composed Video Retrievalβ62May 2, 2024Updated 2 years ago
- [CVPR'26 Demo] Mobile-O: Unified Multimodal Understanding and Generation on Mobile Deviceβ153Apr 13, 2026Updated 3 months ago
- Source code and data for ADEPT: A DEbiasing PrompT Framework (AAAI-23).β15Dec 13, 2024Updated last year
- Deploy open-source AI quickly and easily - Special Bonus Offer β’ AdRunpod Hub is built for open source. One-click deployment and autoscaling endpoints without provisioning your own infrastructure.
- Experiments for "A Closer Look at In-Context Learning under Distribution Shifts"β18May 29, 2023Updated 3 years ago
- [ACL 2025 π₯] A Comprehensive Multi-Domain Benchmark for Arabic OCR and Document Understandingβ76May 24, 2025Updated last year
- β42Nov 9, 2023Updated 2 years ago
- Official code of the paper "VideoMolmo: Spatio-Temporal Grounding meets Pointing"β56Jul 5, 2025Updated last year
- πΌ Official implementation of Dynamic Data Mixing Maximizes Instruction Tuning for Mixture-of-Expertsβ41Sep 29, 2024Updated last year
- [ICLR 2024] Official code for the paper "LLM Blueprint: Enabling Text-to-Image Generation with Complex and Detailed Prompts"β85May 18, 2024Updated 2 years ago
- SCARA Robot Sim - Twincat 3, Python 3, CoppeliaSim(V-Rep)β14Jan 1, 2023Updated 3 years ago
- [ICLR 2025] Official Pytorch Implementation of "Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN" by Pengxiaβ¦β30Jul 24, 2025Updated 11 months ago
- [SIGIR '25] This is the code repo for our SIGIR '25 paper: Enhancing the Patent Matching Capability of Large Language Models via Memory Gβ¦β19Apr 22, 2025Updated last year
- Deploy to Railway using AI coding agents - Free Credits Offer β’ AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ImageNet-12k subset of ImageNet-21k (fall11)β23Jun 13, 2023Updated 3 years ago
- A Benchmark and Agentic Framework for Omni-Modal Reasoning and Tool Use in Long Videosβ21Jun 20, 2026Updated last month
- Code for ProTrix: Building Models for Planning and Reasoning over Tables with Sentence Contextβ17Nov 15, 2024Updated last year
- Official code release for Delta Activations: A Representation for Finetuned Large Language Modelsβ20Sep 5, 2025Updated 10 months ago
- code for Learning the Unlearned: Mitigating Feature Suppression in Contrastive Learningβ20Jul 16, 2024Updated 2 years ago
- Source codes for paper "BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity".β19Jan 10, 2026Updated 6 months ago
- ML model trained on data from Bayut.com to predict housing prices in Dubaiβ17Aug 21, 2025Updated 11 months ago