[Findings of EMNLP 2024] AdaMoE: Token-Adaptive Routing with Null Experts for Mixture-of-Experts Language Models
☆20Oct 2, 2024Updated 2 years ago
Alternatives and similar repositories for AdaMoE
Users that are interested in AdaMoE are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- ☆40May 20, 2025Updated last year
- LoPA: Scaling dLLM Inference via Lookahead Parallel Decoding☆40Aug 3, 2026Updated 2 months ago
- [ICML 2023] Structural Re-weighting Improves Graph Domain Adaptation (StruRW)☆23Jun 20, 2023Updated 3 years ago
- [CVPR 2026] Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight☆95Jun 5, 2026Updated 4 months ago
- The open-source materials for paper "Sparsing Law: Towards Large Language Models with Greater Activation Sparsity".☆32Nov 12, 2024Updated last year
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- ☆62Apr 9, 2026Updated 6 months ago
- Repo for paper "CODIS: Benchmarking Context-Dependent Visual Comprehension for Multimodal Large Language Models".☆13Oct 14, 2024Updated last year
- Official PyTorch Implementation of EMoE: Unlocking Emergent Modularity in Large Language Models [main conference @ NAACL2024]☆40May 28, 2024Updated 2 years ago
- One Network, Many Masks: Towards More Parameter-Efficient Transfer Learning☆40Jul 1, 2023Updated 3 years ago
- ☆20Jan 10, 2025Updated last year
- [COLM 2025] "C3PO: Critical-Layer, Core-Expert, Collaborative Pathway Optimization for Test-Time Expert Re-Mixing"☆21Apr 9, 2025Updated last year
- Non-IID Transfer Learning on Graphs☆13Jul 4, 2023Updated 3 years ago
- 深度学习课程自己所做答案☆10Apr 23, 2018Updated 8 years ago
- ☆18Apr 21, 2024Updated 2 years ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- [CoRL 2026] World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis☆147Sep 16, 2026Updated 3 weeks ago
- The official implementation of InterBERT☆11Oct 18, 2022Updated 3 years ago
- [ICLR 2023] "Graph Domain Adaptation via Theory-Grounded Spectral Regularization" by Yuning You, Tianlong Chen, Zhangyang Wang, Yang Shen☆25Feb 27, 2023Updated 3 years ago
- Neural image compression models optimized for Mask R-CNN from paper "Boosting Neural Image Compression for Machines Using Latent Space Ma…☆11Aug 16, 2022Updated 4 years ago
- Searching a High Performance Feature Extractor for Text Recognition Network. TPAMI 2022☆13Nov 25, 2022Updated 3 years ago
- [Findings of EMNLP 2022] Code of paper Generative Prompt Tuning for Relation Classification. https://arxiv.org/abs/2210.12435☆20May 7, 2023Updated 3 years ago
- Implementation of the paper End-to-end Learning of Deterministic Decision Trees☆17May 19, 2022Updated 4 years ago
- Scripts for KGIRNet model for ESWC☆10Jul 6, 2023Updated 3 years ago
- ICML'20: SIGUA: Forgetting May Make Learning with Noisy Labels More Robust☆17Dec 14, 2020Updated 5 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆15Mar 31, 2022Updated 4 years ago
- PyTorch implementation of paper "Evolving Parameterized Prompt Memory for Continual Learning" in AAAI 2024 (Oral).☆13Apr 15, 2024Updated 2 years ago
- code for our BMVC 2021 paper "HCV: Hierarchy-Consistency Verification for Incremental Implicitly-Refined Classification"☆15Oct 28, 2022Updated 3 years ago
- Official implementation for "LOVECon: Text-driven Training-free Long Video Editing with ControlNet"☆43Oct 26, 2023Updated 2 years ago
- NVFP4 Flash-Attention 4 on BlackWell☆68Sep 26, 2026Updated last week
- ☆42Jun 9, 2025Updated last year
- The extented code of layered conceptual image compression. Journal submitted.☆15Aug 29, 2022Updated 4 years ago
- ☆10Aug 31, 2023Updated 3 years ago
- ☆20Dec 8, 2024Updated last year
- GPU virtual machines on DigitalOcean Gradient AI • AdGet to production fast with high-performance AMD and NVIDIA GPUs you can spin up in seconds. The definition of operational simplicity.
- DSTC9 Submission☆16Apr 12, 2021Updated 5 years ago
- FocusLLM: Scaling LLM’s Context by Parallel Decoding☆45Dec 8, 2024Updated last year
- Implementation of 'Attention-guided Feature Fusion for Small Object Detection'☆14Dec 21, 2023Updated 2 years ago
- ☆16Jun 4, 2024Updated 2 years ago
- ☆11Nov 11, 2018Updated 7 years ago
- ☆21Oct 31, 2022Updated 3 years ago
- Source code of "What Makes Graph Neural Networks Miscalibrated?" (NeurIPS 2022)☆24Jun 9, 2025Updated last year