nancui0000 / DiSuSumDetView on GitHub
End-to-end RLHF detoxification pipeline implementing the InstructGPT architecture: LoRA SFT on DialogSum, custom reward model from synthetic preferences, and multi-objective PPO balancing toxicity, faithfulness, and quality. Achieves 60% toxicity reduction with zero ROUGE-L degradation
50Apr 13, 2026Updated 3 months ago

Alternatives and similar repositories for DiSuSumDet

Users that are interested in DiSuSumDet are comparing it to the libraries listed below. We may earn a commission when you buy through links labeled 'Ad' on this page.

Sorting:

Are these results useful?