cuba6112 avatar

unsloth-dpo

Direct Preference Optimization (DPO) for aligning models with preference data without separate rewar

提供方 cuba6112|开源