openclawrl
OpenClaw-RL (arXiv:2603.10165) trains a personal agent to satisfy the preferences of its user. It learns from their conversations: the user's next message shows whether a reply was accepted, and the model is updated while the agent stays in service.
|
Evolves |
model weights |
|
Signal |
none sent by the agent; Reef reads its recorded traffic |
|
Loss family |
openclawrl |
|
Package |
recipes/openclawrl/ |
|
Processor |
computed feedback |
|
Needs |
GPUs, plus a second served model for the PRM judge |
|
Example |
recipes/openclawrl/examples/openclawrl/ |
What it does #
The learning signal comes from the conversation itself: when the user accepts a reply and moves on, the turn counts as positive, and when the user complains, it counts as negative. Reef reads this from the traffic it already records, so the agent only needs to use Reef as its inference endpoint. It does not send reports, set session headers, or run behind a proxy.
How Reef implements it #
The processor rebuilds sessions from the recorded traffic. A chat agent resends the transcript on every request, so each request can be matched to its session by trace. When a session's next user message arrives, the finished turn is judged by a PRM on a private worker. The PRM votes on whether the message shows acceptance, and on acceptance it also proposes a hindsight hint, a short instruction that would have produced this reply if the user had given it up front. Judged turns are batched for training directly.
OpenClaw-RL trains with two signals, reinforcement learning and on-policy distillation (OPD). The RL term is a PPO clipped surrogate on the sampled tokens, with the turn's raw reward as its advantage. The OPD term pulls the policy toward a teacher on the top-K token candidates the engine recorded at generation time. The teacher is the frozen base model conditioned on one of the accepted hindsight hints, and the objective selects which hint by how well the teacher's top-K overlaps the policy's. The weights of the two terms are --openclawrl-w-rl and --openclawrl-w-opd in training.slime_flags, both 1.0 by default.
Configuration #
batch_size16judged turns per training step. Must equal the driver's --global-batch-size.session_ttl_s900.0idle window before a session expiresprm_url""the PRM's sglang serverprm_tokenizer_path""required whenever prm_url is setprm_m3judge votes per turnprm_temperature0.6judge sampling temperatureprm_max_tokens8192judge generation budgetprm_timeout_s120.0per-judge timeoutprm_concurrency8turns judged at onceprm_max_hint_candidates3hint candidates collected per turnprm_record_file""judged population, one JSON line per batch; empty disables itmax_staleness0accepted lag between the producing and serving version. Env REEF_MAX_STALENESS.The values above are the recipe's defaults. The example's serve.yaml overrides three of them. It points prm_url and prm_tokenizer_path at the stack's own PRM engine, raises prm_timeout_s to 3600 because a thinking model can spend minutes on a single vote, and sets prm_record_file so every batch leaves one line with the reward split and the judge counters.
Run the example #
The example reproduces the paper's personal-agent experiment. A simulated student brings 72 GSM8K homework problems to a Hermes agent with one session per problem. Each session is scored on whether the agent's first solution reply satisfies the student's preferences.
docker build -f docker/Dockerfile.reef -t reef-openclawrl .
hf download Qwen/Qwen3-4B-Thinking-2507 --local-dir ~/models/Qwen3-4B-Thinking-2507 # policy and PRM
hf download Qwen/Qwen3-32B --local-dir ~/models/Qwen3-32B # the student
bash recipes/openclawrl/examples/openclawrl/run.shrun.sh builds the simulated student's image, starts the stack that serve.yaml describes, and runs the 72 sessions in order. The stack uses seven GPUs: four for the tensor-parallel actor, one for the rollout engine, one for the PRM, and one for the Qwen3-32B student model.
Results #
The example's README records one run over the first 36 sessions of the stream.
A session passes when the agent's first solution reply already matches the student's preferences, with the work shown and the correct answer. The student wants homework that does not look AI-written, so a reply that uses bold text or a bullet or numbered list draws a complaint. The bold rate and the list rate measure this habit over the run, as the fraction of the last ten sessions whose first reply still contains bold text or a list. From the curves we see both fall as training goes on, and the run reaches the paper's adaptation criterion (three passed sessions in a row) at session 14.
The demo above replays two sessions from the same run. In session 1 the student rejects a formatted reply, reef keeps the training going, and by session 16 the first reply passes directly.