unsloth-dpo
Direct Preference Optimization (DPO) for aligning models with preference data without separate rewar
2401 个 Skills
Direct Preference Optimization (DPO) for aligning models with preference data without separate rewar
Supervised fine-tuning using SFTTrainer, instruction formatting, and multi-turn dataset preparation
Instructions for updating the Python backend and leaderboard database.
This skill should be used when you need to refresh the Dashboard with current system status, pending
Update Nexus system files from upstream repository. Load when user says "update nexus", "sync nexus"
Updates the Monarch Initiative publications page with latest data from Google Scholar. Use this when
Programmatically update marathon-ralph state file using deterministic jq commands. Use this instead
Use before committing staged changes when you need to verify all related documentation is current -
Updates person records in the auntruth genealogy database (3,004 people across 10 lineages) while ma
Use core utilities from @openapi-lsp/core for URI/URL and JSON Pointer handling. Apply when working
Parse URLs in CSV files and extract query parameters as new columns. Use when working with CSV files
Search and query Goodreads library from CSV export. Use when the user asks about books, TBR (to-be-r