A framework for steering MoE models by detecting and controlling behavior-linked experts.
☆39Sep 12, 2025Updated last year
Alternatives and similar repositories for SteerMoE
Users that are interested in SteerMoE are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.
Sorting:
- [NAACL 2022] GlobEnc: Quantifying Global Token Attribution by Incorporating the Whole Encoder Layer in Transformers☆24May 16, 2023Updated 3 years ago
- [COLM 2025] "C3PO: Critical-Layer, Core-Expert, Collaborative Pathway Optimization for Test-Time Expert Re-Mixing"☆21Apr 9, 2025Updated last year
- [CVPR Findings 2026] HoliSafe: Holistic Safety Benchmarking and Modeling for Vision-Language Model☆17Mar 8, 2026Updated 6 months ago
- [COLM 2025] JailDAM: Jailbreak Detection with Adaptive Memory for Vision-Language Model☆26Nov 25, 2025Updated 9 months ago
- ☆13Jan 14, 2025Updated last year
- Open source password manager - Proton Pass • AdSecurely store, share, and autofill your credentials with Proton Pass, the end-to-end encrypted password manager trusted by millions.
- Official Codebase of the ACL 2026 Oral paper "Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contra…☆28Jun 25, 2026Updated 2 months ago
- ☆14Feb 12, 2024Updated 2 years ago
- [EMNLP 2025 Findings] Familiarity-aware Evidence Compression for Retrieval Augmented Generation☆15Aug 20, 2025Updated last year
- [ICLR 2025] RaSA: Rank-Sharing Low-Rank Adaptation☆10May 19, 2025Updated last year
- [ICML 2024] Generalizing Knowledge Graph Embedding with Universal Orthogonal Parameterization☆16May 12, 2024Updated 2 years ago
- ☆14Sep 22, 2025Updated last year
- Rad-cGAN v1.0: Radar-based precipitation nowcasting model with conditional Generative Adversarial Networks for multiple dam domains☆11Jul 22, 2022Updated 4 years ago
- [ICLR 2025 Spotlight] Code release for "Sharpness-Aware Minimization Efficiently Selects Flatter Minima Late In Training"☆20Feb 20, 2025Updated last year
- [ICML 2026] An Evaluation Suite for Chain-of-Thought Controllability☆54Mar 10, 2026Updated 6 months ago
- AI Agents on DigitalOcean Gradient AI Platform • AdBuild production-ready AI agents using customizable tools or access multiple LLMs through a single endpoint. Create custom knowledge bases or connect external data.
- Dialog2Flow: convert your dialogs to flows. This repository accompanies the paper "Dialog2Flow: Pre-training Soft-Contrastive Sentence Em…☆20Jul 1, 2025Updated last year
- ☆15Mar 20, 2025Updated last year
- Simple Python Socket-based Split Learning technique using PyTorch☆14Mar 13, 2020Updated 6 years ago
- This is the official code for the paper "Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable".☆35Mar 11, 2025Updated last year
- Code for paper "Concrete Subspace Learning based Interference Elimination for Multi-task Model Fusion"☆14Mar 28, 2024Updated 2 years ago
- The implementation of ACL 2026 paper "Rethinking entropy interventions in rlvr: An entropy change perspective"☆27Jul 19, 2026Updated 2 months ago
- Supporting code for the EMNLP 2019 paper "Answers Unite! Unsupervised Metrics for Reinforced Summarization Models"☆14Jun 12, 2023Updated 3 years ago
- ☆15Aug 17, 2023Updated 3 years ago
- Code for "Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMs"☆19Nov 6, 2025Updated 10 months ago
- Serverless GPU API endpoints on Runpod - Get Bonus Credits • AdSkip the infrastructure headaches. Auto-scaling, pay-as-you-go, no-ops approach lets you focus on innovating your application.
- Code for Augment & Reduce, a scalable stochastic algorithm for large categorical distributions☆10May 16, 2018Updated 8 years ago
- Developer project for getting basic API integrations working in under 5 minutes☆11May 22, 2026Updated 4 months ago
- ☆43Jul 3, 2026Updated 2 months ago
- [NAACL 2025] A Closer Look into Mixture-of-Experts in Large Language Models☆60Feb 7, 2025Updated last year
- NeurIPS 2025☆16Feb 4, 2026Updated 7 months ago
- A toolkit for embedding text datasets with sparse autoencoders☆31Mar 24, 2026Updated 6 months ago
- Code for "DynaGuard: A Dynamic Guardrail Model With User-Defined Policies."☆25Nov 3, 2025Updated 10 months ago
- Official Code Repository for OmniRetrieval☆35Jun 1, 2026Updated 3 months ago
- codes and plots for "Active-Dormant Attention Heads: Mechanistically Demystifying Extreme-Token Phenomena in LLMs"☆11Dec 30, 2024Updated last year
- 1-Click AI Models by DigitalOcean Gradient • AdDeploy popular AI models on DigitalOcean Gradient GPU virtual machines with just a single click. Zero configuration with optimized deployments.
- Official Repo of Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents☆99Jun 2, 2026Updated 3 months ago
- A repository for LotteryFL re-implementation and experiments☆13Dec 18, 2020Updated 5 years ago
- ☆10Aug 26, 2022Updated 4 years ago
- Jupyter notebooks from our weekly (or so) hackathons☆11Dec 3, 2024Updated last year
- Code for "Exponential Family Estimation via Adversarial Dynamics Embedding" (NeurIPS 2019)☆14Nov 26, 2019Updated 6 years ago
- This is a pip package implementing Reinforcement Learning algorithms in non-stationary environments supported by the OpenAI Gym toolkit.☆16Jun 28, 2024Updated 2 years ago
- Use ChatGPT to write README, based on your code. This repo's readme is written by this tool. So if you think this readme sucks, literally…☆23Jul 26, 2024Updated 2 years ago