Hugging Face and Liquid AI bring model training to coding agents without changing their code

Hugging Face and Liquid AI bring model training to coding agents without changing their code

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
3. 10. 2026
4 minutes reading · 1 views
Listen to the article
Audio version of the article
Hugging Face and Liquid AI bring model training to coding agents without changing their code

Hugging Face and Liquid AI have released an open stack for reinforcement learning directly inside coding agents, including Claude Code, Codex and OpenCode. Developers can connect training without modifying the agents’ source code by directing model requests through a proxy that preserves each agent’s native communication format and records the data needed for learning.

Hugging Face and Liquid AI report LFM2.5-2.6B gains across four environments

In Hugging Face and Liquid AI’s main experiment, LFM2.5-2.6B’s overall success rate rose from 42.2% to 54.2% on pass@1, which measures the share of tasks solved on the first attempt. The teams recorded improvements in all four evaluated agent environments.

The teams’ results also showed fewer tool calls. The model trained across multiple environments needed 31% fewer calls than the base model, with that comparison restricted to tasks both versions successfully solved. The model trained exclusively in OpenCode reduced tool calls by 11%.

According to Hugging Face and Liquid AI, training in a single environment produced a different pattern of results: the OpenCode-only model achieved a 58% success rate when evaluated in OpenCode. The multi-environment model performed better in Claude Code and Codex. When running in Claude Code, the OpenCode-only model also used more tool calls and tokens than the base version.

Connecting through the model endpoint and translating API formats

The capture proxy sits between the agent and the vLLM inference server. Developers configure the agent to use the proxy’s address as its model endpoint, enabling training capture without changing the agent’s source code. The agent continues sending requests and receiving responses in its native format.

The proxy identifies one of four supported API formats using the request’s path, headers and body. Converters taken from NVIDIA Polar translate the input into Chat Completions format. Once an answer has been generated, the proxy translates it back into the format expected by the calling agent.

Training records preserve exact tokens and branching runs

Archiving the output text alone is not enough for policy-gradient updates. The proxy stores the exact IDs of tokens generated by the model, along with their processed log probabilities. Saving the text first and tokenizing it later could produce different token IDs for the same visible content.

Another part of the capture process handles runs involving retries, subagents or context compaction. Each model call is stored as a node, with connections determined by the longest exact shared token prefix. Each path from the root to the end of a branch becomes a training sequence.

The proxy samples from the model’s full distribution, with top_p set to 1.0 and top_k truncation disabled. Hugging Face and Liquid AI report that setting top_p to 1.0 moved the measured importance-sampling ratio from 0.985–0.993 to 0.9984–0.9999. A value near 1 indicates alignment between the probabilities used to generate training runs and those used during training itself.

Eight attempts per task, with rewards for fewer tool calls

Hugging Face and Liquid AI conducted their experiments on SmolDataEnvs, a collection of 1,000 data-analysis tasks derived from Kaggle notebooks. Both training variants used Async GRPO in TRL, running for 1,000 steps on two H100 GPUs.

The variants differed in their choice of environment. In the first, every attempt ran in OpenCode. In the second, each group of attempts used one of four environments: OpenCode, Claude Code, Codex or Mini-SWE-Agent.

In the teams’ training setup, each GRPO group contained eight attempts at the same task. A correct answer received a reward of 1, while an incorrect answer received 0. Successful solutions could also earn a bonus of up to 0.1 for using fewer tool calls; this supplied an additional learning signal even in groups where all eight attempts succeeded.

Reinforcement learning versus supervised fine-tuning

Alongside reinforcement learning, the teams tested supervised fine-tuning. They ran Qwen3.8-27B in all four environments and selected 3,189 successful solution trajectories from its runs. They then used that data to fine-tune LFM2.5-2.6B.

In that comparison, Hugging Face and Liquid AI reported a pass@1 score of 54.6% for multi-environment reinforcement learning. OpenCode-only supervised fine-tuning reached 47.5%, while supervised fine-tuning across multiple environments reached 43.1%.

The teams also recorded a regression in Mini-SWE-Agent for the model supervised-fine-tuned across multiple environments. Its success rate fell from the base model’s 62.1% to 45.2%, offsetting gains achieved in the other three environments.

Tools, data and seven checkpoints are public

The released stack includes the OpenEnv capture proxy and FineEnvs scripts for defining environments and launching training. The TRL trainer integration and SmolDataEnvs task suite are also available. The release includes supervised fine-tuning data and all seven trained checkpoints.

Advertisement

Content created with help from UpTier.

SEO and GEO on autopilot. UpTier’s multi-agent systems write and optimize content for search engines and AI answers.

Discover UpTier ↗

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Cohere Embed 5 pairs Pro indexing with Fast searchCohere Embed 5 pairs Pro indexing with Fast search
The two Embed 5 variants share a vector space, allowing developers to index documents with Pro and process queries with Fast without rebuilding the index. The family supports multimodal inputs, more than 100 languages and a 128,000-token context.
3 min read
3. 10. 2026
Qwen-Image-2.1 combines image generation and editing with transparent outputQwen-Image-2.1 combines image generation and editing with transparent output
Alibaba has released a single checkpoint for image generation and editing. Qwen-Image-2.1 supports transparent PNGs and, according to the company, accepts up to 10 reference images. Commercial deployment requires a separate license.
3 min read
3. 10. 2026
CoreWeave launches Forge for training and continuously improving AI agentsCoreWeave launches Forge for training and continuously improving AI agents
Forge connects model training, evaluation and improvement with insights from production. It offers post-training without a dedicated cluster, experiment analysis and isolated environments for agents.
2 min read
3. 10. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok