Hugging Face and Liquid AI have released an open stack for reinforcement learning directly inside coding agents, including Claude Code, Codex and OpenCode. Developers can connect training without modifying the agents’ source code by directing model requests through a proxy that preserves each agent’s native communication format and records the data needed for learning.
Hugging Face and Liquid AI report LFM2.5-2.6B gains across four environments
In Hugging Face and Liquid AI’s main experiment, LFM2.5-2.6B’s overall success rate rose from 42.2% to 54.2% on pass@1, which measures the share of tasks solved on the first attempt. The teams recorded improvements in all four evaluated agent environments.
The teams’ results also showed fewer tool calls. The model trained across multiple environments needed 31% fewer calls than the base model, with that comparison restricted to tasks both versions successfully solved. The model trained exclusively in OpenCode reduced tool calls by 11%.
According to Hugging Face and Liquid AI, training in a single environment produced a different pattern of results: the OpenCode-only model achieved a 58% success rate when evaluated in OpenCode. The multi-environment model performed better in Claude Code and Codex. When running in Claude Code, the OpenCode-only model also used more tool calls and tokens than the base version.
Connecting through the model endpoint and translating API formats
The capture proxy sits between the agent and the vLLM inference server. Developers configure the agent to use the proxy’s address as its model endpoint, enabling training capture without changing the agent’s source code. The agent continues sending requests and receiving responses in its native format.
The proxy identifies one of four supported API formats using the request’s path, headers and body. Converters taken from NVIDIA Polar translate the input into Chat Completions format. Once an answer has been generated, the proxy translates it back into the format expected by the calling agent.
Training records preserve exact tokens and branching runs
Archiving the output text alone is not enough for policy-gradient updates. The proxy stores the exact IDs of tokens generated by the model, along with their processed log probabilities. Saving the text first and tokenizing it later could produce different token IDs for the same visible content.
Another part of the capture process handles runs involving retries, subagents or context compaction. Each model call is stored as a node, with connections determined by the longest exact shared token prefix. Each path from the root to the end of a branch becomes a training sequence.
The proxy samples from the model’s full distribution, with top_p set to 1.0 and top_k truncation disabled. Hugging Face and Liquid AI report that setting top_p to 1.0 moved the measured importance-sampling ratio from 0.985–0.993 to 0.9984–0.9999. A value near 1 indicates alignment between the probabilities used to generate training runs and those used during training itself.
Eight attempts per task, with rewards for fewer tool calls
Hugging Face and Liquid AI conducted their experiments on SmolDataEnvs, a collection of 1,000 data-analysis tasks derived from Kaggle notebooks. Both training variants used Async GRPO in TRL, running for 1,000 steps on two H100 GPUs.
The variants differed in their choice of environment. In the first, every attempt ran in OpenCode. In the second, each group of attempts used one of four environments: OpenCode, Claude Code, Codex or Mini-SWE-Agent.
In the teams’ training setup, each GRPO group contained eight attempts at the same task. A correct answer received a reward of 1, while an incorrect answer received 0. Successful solutions could also earn a bonus of up to 0.1 for using fewer tool calls; this supplied an additional learning signal even in groups where all eight attempts succeeded.
Reinforcement learning versus supervised fine-tuning
Alongside reinforcement learning, the teams tested supervised fine-tuning. They ran Qwen3.8-27B in all four environments and selected 3,189 successful solution trajectories from its runs. They then used that data to fine-tune LFM2.5-2.6B.
In that comparison, Hugging Face and Liquid AI reported a pass@1 score of 54.6% for multi-environment reinforcement learning. OpenCode-only supervised fine-tuning reached 47.5%, while supervised fine-tuning across multiple environments reached 43.1%.
The teams also recorded a regression in Mini-SWE-Agent for the model supervised-fine-tuned across multiple environments. Its success rate fell from the base model’s 62.1% to 45.2%, offsetting gains achieved in the other three environments.
Tools, data and seven checkpoints are public
The released stack includes the OpenEnv capture proxy and FineEnvs scripts for defining environments and launching training. The TRL trainer integration and SmolDataEnvs task suite are also available. The release includes supervised fine-tuning data and all seven trained checkpoints.



