[ICLR 2025] CAMEx: Curvature-Aware Merging of Experts
☆24Mar 1, 2025Updated last year
Alternatives and similar repositories for CAMEx
Users that are interested in CAMEx are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- Code for paper "Towards Efficient Pareto Set Approximation via Weight-Ensembling Mixture of Experts"☆11Jul 30, 2026Updated last month
- Code for the paper "Rethinking Importance Weighting for Deep Learning under Distribution Shift".☆32Apr 9, 2021Updated 5 years ago
- Public code release for the paper "Reawakening knowledge: Anticipatory recovery from catastrophic interference via structured training"☆11Oct 27, 2025Updated 10 months ago
- ☆12Jun 24, 2021Updated 5 years ago
- ☆17Apr 30, 2026Updated 4 months ago
- Managed hosting for WordPress and PHP on Cloudways • AdManaged hosting for WordPress, Magento, Laravel, or PHP apps, on multiple cloud providers. Deploy in minutes on Cloudways by DigitalOcean.
- Adaptive gradient descent without descent☆54Oct 12, 2021Updated 4 years ago
- [MedIA 2026] No Modality Left Behind: Adapting to Missing Modalities via Knowledge Distillation for Brain Tumor Segmentation.☆31Aug 19, 2026Updated last month
- Joint Distillation for Fast Likelihood Evaluation and Sampling in Flow-based Models☆23Mar 5, 2026Updated 6 months ago
- Official implementation of "One-Prompt Strikes Back: Sparse Mixture of Experts for Prompt-based Continual Learning " (ICLR 2026)☆17Mar 29, 2026Updated 5 months ago
- Here I gather promising research directions to make DNNs interpretable☆17Apr 11, 2024Updated 2 years ago
- XL-VLMs: General Repository for eXplainable Large Vision Language Models☆52Sep 8, 2025Updated last year
- MEDREASON-R1: Learning to Reason for CT Diagnosis with Reinforcement Learning and Local Zoom☆16Oct 10, 2025Updated 11 months ago
- An implementation of various tensor-based decomposition for NN & RNN parameters☆18Jun 4, 2018Updated 8 years ago
- MNIST experiment from Tensorizing neural networks (Novikov et al. 2015)☆14Oct 22, 2019Updated 6 years ago
- Deploy to Railway using AI coding agents - Free Credits Offer • AdUse Claude Code, Codex, OpenCode, and more. Autonomous software development now has the infrastructure to match with Railway.
- ☆18Dec 9, 2020Updated 5 years ago
- Official PyTorch codes for "Enhancing Diffusion Models with Text-Encoder Reinforcement Learning", ECCV2024☆59Aug 13, 2024Updated 2 years ago
- MAP: Low-compute Model Merging with Amortized Pareto Fronts via Quadratic Approximation☆18Sep 2, 2024Updated 2 years ago
- ☆25Nov 23, 2021Updated 4 years ago
- This is the public github for our paper "Transformer with a Mixture of Gaussian Keys"☆29Aug 13, 2022Updated 4 years ago
- u-MPS implementation and experimentation code used in the paper Tensor Networks for Probabilistic Sequence Modeling (https://arxiv.org/ab…☆19Jul 2, 2020Updated 6 years ago
- ☆22Oct 14, 2021Updated 4 years ago
- A strong baseline for liveness detection. The source code could be used for similar tasks, such as face anti-spoofing or detecting fake v…☆23Nov 29, 2022Updated 3 years ago
- End-to-end training of Retrieval-Augmented LMs (REALM, RAG)☆23Nov 22, 2023Updated 2 years ago
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- An online multiplayer board game similar to Catan☆37Sep 23, 2024Updated last year
- ☆16Jun 3, 2026Updated 3 months ago
- PyTorch implementation of Drifting Models by Kaiming He et al.☆22Feb 6, 2026Updated 7 months ago
- Apply CP, Tucker, TT/TR, HT to compress neural networks. Train from scratch.☆17Nov 26, 2020Updated 5 years ago
- Code for the paper "Tensor Networks for Maching Learning"☆17Nov 7, 2019Updated 6 years ago
- Neural Sentiment Analyzer for Modern Hebrew☆21Nov 21, 2022Updated 3 years ago
- Preserving Generalization of Language Models in Few-shot Continual Relation Extraction (EMNLP2024)☆17Nov 21, 2024Updated last year
- [ICLR'25] R-Sparse: Rank-Aware Activation Sparsity for Efficient LLM Inference☆21Apr 28, 2025Updated last year
- An efficient distillation method for flow matching models☆31Feb 1, 2026Updated 7 months ago
- Managed Kubernetes at scale on DigitalOcean • AdDigitalOcean Kubernetes includes the control plane, bandwidth allowance, container registry, automatic updates, and more for free.
- ☆27Jan 25, 2021Updated 5 years ago
- [CVPR2024] Official Codes for "Adversarial Score Distillation: When score distillation meets GAN"☆38Apr 24, 2025Updated last year
- ☆24May 2, 2026Updated 4 months ago
- Jax implementation of VIT-VQGAN☆10Jan 25, 2024Updated 2 years ago
- Official implementation of "Modeling Multi-Task Model Merging as Adaptive Projective Gradient Descent".☆23May 23, 2025Updated last year
- ICLR 2026☆29May 13, 2026Updated 4 months ago
- Official Repository for Task-Circuit Quantization☆28Jun 1, 2025Updated last year