[ICLR 2026 π₯] Dr.LLM: Dynamic Layer Routing in LLMs
β57Apr 24, 2026Updated 3 months ago
Alternatives and similar repositories for dr-llm
Users that are interested in dr-llm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The official implementation of HybridNorm: Towards Stable and Efficient Transformer Training via Hybrid Normalizationβ19Mar 7, 2025Updated last year
- [NAACL 2025 π₯] CAMEL-Bench is an Arabic benchmark for evaluating multimodal models across eight domains with 29,000 questions.β38Apr 17, 2025Updated last year
- β19Jul 24, 2023Updated 3 years ago
- transformer layers behavior as paintersπ§βπ¨β15May 6, 2025Updated last year
- [ICCV2025] Hierarchical Visual Prompt Learning for Continual Video Instance Segmentationβ14Feb 18, 2026Updated 5 months ago
- Deploy on Railway without the complexity - Free Credits Offer β’ AdConnect your repo and Railway handles the rest with instant previews. Quickly provision container image services, databases, and storage volumes.
- Video-CoM: Interactive Video Reasoning via Chain of Manipulationsβ23Jun 17, 2026Updated last month
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Modelsβ19Jan 21, 2026Updated 6 months ago
- Code and data for "Timo: Towards Better Temporal Reasoning for Language Models" (COLM 2024)β26Oct 23, 2024Updated last year
- Official Repository of Paper "Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs"β15Sep 25, 2025Updated 10 months ago
- Source code of "Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers" EMNLP 2025β17Jan 12, 2026Updated 6 months ago
- Official repository of paper titled "D3Former: Debiased Dual Distilled Transformer for Incremental Learning".β25Jul 10, 2023Updated 3 years ago
- xRouter: Training Cost-Aware LLMs Orchestration System via Reinforcement Learningβ33Jun 2, 2026Updated 2 months ago
- [arxiv: 2512.19673] Bottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal Policiesβ60Feb 6, 2026Updated 6 months ago
- [NeurIPS 2024] BLAST: Block Level Adaptive Structured Matrix for Efficient Deep Neural Network Inferenceβ18Nov 6, 2024Updated last year
- Wordpress hosting with auto-scaling - Free Trial Offer β’ AdFully Managed hosting for WordPress and WooCommerce businesses that need reliable, auto-scalable performance. Cloudways SafeUpdates now available.
- β26Jul 24, 2026Updated 2 weeks ago
- [NAACL 2022] "Learning to Win Lottery Tickets in BERT Transfer via Task-agnostic Mask Training", Yuanxin Liu, Fandong Meng, Zheng Lin, Peβ¦β15Oct 18, 2022Updated 3 years ago
- [EMNLP 2025] TokenSkip: Controllable Chain-of-Thought Compression in LLMsβ227Nov 30, 2025Updated 8 months ago
- Official repo of dataset-decomposition paper [NeurIPS 2024]β21Jan 8, 2025Updated last year
- Composed Video Retrievalβ62May 2, 2024Updated 2 years ago
- [NAACL'25 π SAC Award] Official code for "Advancing MoE Efficiency: A Collaboration-Constrained Routing (C2R) Strategy for Better Expertβ¦β16Feb 4, 2025Updated last year
- Source code of "C-SEO Bench: Does Conversational SEO Work?" NeurIPS D&B 2025β18Sep 28, 2025Updated 10 months ago
- Source code and data for ADEPT: A DEbiasing PrompT Framework (AAAI-23).β15Dec 13, 2024Updated last year
- Source code of "Calibrating Large Language Models Using Their Generations Only", ACL2024β22Nov 20, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- [ICLR 2025] When Attention Sink Emerges in Language Models: An Empirical View (Spotlight)β165Jul 8, 2025Updated last year
- β11May 9, 2023Updated 3 years ago
- ABB 140 Robot Draws a Given Pictureβ14Oct 17, 2020Updated 5 years ago
- Experiments for "A Closer Look at In-Context Learning under Distribution Shifts"β18May 29, 2023Updated 3 years ago
- [ACL 2025 π₯] A Comprehensive Multi-Domain Benchmark for Arabic OCR and Document Understandingβ77Updated this week
- Official code of the paper "VideoMolmo: Spatio-Temporal Grounding meets Pointing"β56Jul 5, 2025Updated last year
- Official repository for "Stylized Adversarial Training" (TPAMI 2022)β11Dec 30, 2022Updated 3 years ago
- [ICLR 2024] Official code for the paper "LLM Blueprint: Enabling Text-to-Image Generation with Complex and Detailed Prompts"β85May 18, 2024Updated 2 years ago
- SCARA Robot Sim - Twincat 3, Python 3, CoppeliaSim(V-Rep)β14Jan 1, 2023Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Code for COLM 2026 Paper "Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs"β26Jul 1, 2026Updated last month
- [EMNLP'23] ClimateGPT: a specialized LLM for conversations related to Climate Change and Sustainability topics in both English and Arabiβ¦β79Sep 24, 2024Updated last year
- [SIGIR '25] This is the code repo for our SIGIR '25 paper: Enhancing the Patent Matching Capability of Large Language Models via Memory Gβ¦β19Apr 22, 2025Updated last year
- β45Jan 30, 2026Updated 6 months ago
- β23Dec 17, 2024Updated last year
- β38Jul 24, 2023Updated 3 years ago
- Official code for the paper Towards Fully Exploiting LLM Internal States to Enhance Knowledge Boundary Perception. The code is based on tβ¦β22Aug 5, 2025Updated last year