[ICLR 2026 π₯] Dr.LLM: Dynamic Layer Routing in LLMs
β58Apr 24, 2026Updated 5 months ago
Alternatives and similar repositories for dr-llm
Users that are interested in dr-llm are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- The official implementation of HybridNorm: Towards Stable and Efficient Transformer Training via Hybrid Normalizationβ19Mar 7, 2025Updated last year
- [NAACL 2025 π₯] CAMEL-Bench is an Arabic benchmark for evaluating multimodal models across eight domains with 29,000 questions.β38Apr 17, 2025Updated last year
- Video-CoM: Interactive Video Reasoning via Chain of Manipulationsβ23Sep 5, 2026Updated last month
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Modelsβ19Sep 5, 2026Updated last month
- β17Jul 24, 2023Updated 3 years ago
- Managed hosting for WordPress and PHP on Cloudways β’ AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Code and data for "Timo: Towards Better Temporal Reasoning for Language Models" (COLM 2024)β26Oct 23, 2024Updated last year
- [CVPR 2025 π₯]A Large Multimodal Model for Pixel-Level Visual Grounding in Videosβ106Sep 5, 2026Updated last month
- Official Repository of Paper "Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs"β15Sep 25, 2025Updated last year
- Source code of "Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers" EMNLP 2025β18Jan 12, 2026Updated 8 months ago
- VideoMathQA is a benchmark designed to evaluate mathematical reasoning in real-world educational videosβ25Sep 5, 2026Updated last month
- xRouter: Training Cost-Aware LLMs Orchestration System via Reinforcement Learningβ34Jun 2, 2026Updated 4 months ago
- [EMNLP 2026 Oral] Bottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal Policiesβ60Feb 6, 2026Updated 8 months ago
- Search engine results page scraperβ13Dec 19, 2018Updated 7 years ago
- β26Jul 24, 2026Updated 2 months ago
- Bare Metal GPUs on DigitalOcean Gradient AI β’ AdPurpose-built for serious AI teams training foundational models, running large-scale inference, and pushing the boundaries of what's possible.
- Better coding experience for Flaskβ16Aug 11, 2026Updated last month
- [NAACL 2022] "Learning to Win Lottery Tickets in BERT Transfer via Task-agnostic Mask Training", Yuanxin Liu, Fandong Meng, Zheng Lin, Peβ¦β15Oct 18, 2022Updated 3 years ago
- [EMNLP 2025] TokenSkip: Controllable Chain-of-Thought Compression in LLMsβ226Nov 30, 2025Updated 10 months ago
- Composed Video Retrievalβ62May 2, 2024Updated 2 years ago
- Official repo of dataset-decomposition paper [NeurIPS 2024]β21Sep 11, 2026Updated 3 weeks ago
- [CVPR'26 Demo] Mobile-O: Unified Multimodal Understanding and Generation on Mobile Deviceβ159Apr 13, 2026Updated 5 months ago
- [NAACL'25 π SAC Award] Official code for "Advancing MoE Efficiency: A Collaboration-Constrained Routing (C2R) Strategy for Better Expertβ¦β17Feb 4, 2025Updated last year
- Source code of "C-SEO Bench: Does Conversational SEO Work?" NeurIPS D&B 2025β20Sep 28, 2025Updated last year
- Source code of "Calibrating Large Language Models Using Their Generations Only", ACL2024β23Nov 20, 2024Updated last year
- AI Agents on DigitalOcean Gradient AI Platform β’ AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- [ICLR 2025] When Attention Sink Emerges in Language Models: An Empirical View (Spotlight)β164Jul 8, 2025Updated last year
- β11May 9, 2023Updated 3 years ago
- Experiments for "A Closer Look at In-Context Learning under Distribution Shifts"β17May 29, 2023Updated 3 years ago
- [ACL 2025 π₯] A Comprehensive Multi-Domain Benchmark for Arabic OCR and Document Understandingβ79Aug 10, 2026Updated 2 months ago
- Official code of the paper "VideoMolmo: Spatio-Temporal Grounding meets Pointing"β57Jul 5, 2025Updated last year
- β42Nov 9, 2023Updated 2 years ago
- πΌ Official implementation of Dynamic Data Mixing Maximizes Instruction Tuning for Mixture-of-Expertsβ40Sep 29, 2024Updated 2 years ago
- [ICLR 2024] Official code for the paper "LLM Blueprint: Enabling Text-to-Image Generation with Complex and Detailed Prompts"β85May 18, 2024Updated 2 years ago
- [EMNLP'23] ClimateGPT: a specialized LLM for conversations related to Climate Change and Sustainability topics in both English and Arabiβ¦β80Sep 24, 2024Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient β’ AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- LaTeXDataHub is an open-source platform dedicated to the sharing and contribution of real-world LaTeX image datasets and their annotationβ¦β12Aug 13, 2024Updated 2 years ago
- ImageNet-12k subset of ImageNet-21k (fall11)β23Jun 13, 2023Updated 3 years ago
- β48Jan 30, 2026Updated 8 months ago
- β23Dec 17, 2024Updated last year
- Code for COLM 2026 Paper "Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs"β34Jul 1, 2026Updated 3 months ago
- β37Jul 24, 2023Updated 3 years ago
- Official code for the paper Towards Fully Exploiting LLM Internal States to Enhance Knowledge Boundary Perception. The code is based on tβ¦β21Aug 5, 2025Updated last year