A benchmark score is usually read as a verdict on the model. A new release from Hugging Face’s FineEnvs team argues that it also depends on the software wrapped around the model. In its own tests, one small model with exactly the same weights solved 62% of a set of data-analysis tasks in one coding tool and 33% in another [1][2].
The team’s response is an open framework that trains models inside unmodified coding tools [1]. Hugging Face chief executive Clément Delangue announced it on X on 5 October [2]. All results below are Hugging Face’s own.
First, some terms
- Harness. The program around a model that runs the loop: which tools it gets, what context it sees, how its replies are parsed and when to stop [1]. Claude Code, Codex, OpenCode and Mini-SWE-Agent are harnesses.
- Reinforcement learning (RL). The model attempts a task, a grader scores the result, and the model is nudged towards what scored well. In “agentic RL” the attempt is multi-step: write code, run it, read the output, try again [1].
- Pass@1. The share of tasks solved on the first graded attempt [3].
- Logprobs. The probability the model gave each word fragment (token) it produced, stored as a logarithm. RL needs the exact tokens and these numbers, not just the final text, to know precisely what to reinforce [1].
Why the harness changes the score
Each harness gives the model different tools, prompts and rules. A model trained in one can call a tool by a name another spells differently, and the call is rejected before it runs [1]. The guide also cites a measurement of GLM-5.2 at 23% in one harness and 52% in another on SWE-bench Pro [1].
The usual fix is to rebuild a harness as a training environment. Delangue says this means “most models get trained in a scaffold nobody actually ships” [2].
How the framework works
The trick, Delangue says, is “a proxy, not a rewrite” [2]:
- The capture proxy (in OpenEnv). The harness thinks it is talking to a normal model API. In fact it talks to a proxy that speaks the four formats coding agents use (OpenAI Chat Completions, OpenAI Responses, Anthropic Messages and Gemini), forwards requests to a vLLM server running the model in training, and records the exact token IDs and logprobs [1][2].
- The trainer (TRL). Hugging Face’s TRL library trains on those recordings with Async GRPO, which compares several attempts at the same task and reinforces the better ones. The guide cites the merged TRL pull request #6947; an earlier one, #6420, added the first version of the integration [1][4][5].
- Tasks and sandboxes (Harbor). Harbor supplies the tasks and runs each attempt in an isolated sandbox (E2B in these runs) [1].
The guide says OpenEnv has validated ten harnesses end to end; Delangue says “10 harnesses run through it today, none modified” [1][2].
What Hugging Face found
The test model was Liquid AI’s 2.6-billion-parameter LFM2.5-2.6B, scored on 250 held-out data-analysis tasks from FineEnvs’ SmolDataEnvs suite under four harnesses [1][3]. One of the guide’s authors is from Liquid AI [1].
- Before training: 62% under Mini-SWE-Agent, 33% under Claude Code. The average across four harnesses was 42.2% [1].
- Trained across four harnesses (OpenCode, Claude Code, Codex, Mini-SWE-Agent): the average rose to 54.2% [1][3]. At release, the model card lists OpenCode 49.6%, Claude Code 48.8%, Codex 53.6% and Mini-SWE-Agent 64.8% [3].
- Trained in OpenCode only: the OpenCode score rose from 34% to 58%, and the overall average reached 52.3% [1]. The guide says the 1.9-point overall gap between the two runs “is within the noise”. The difference lies in where the gains landed. The multi-harness model did better under Claude Code and Codex, while the OpenCode-only model did better in OpenCode [1].
- Fewer tool calls: a small reward bonus for solving tasks in fewer steps helped the multi-harness model make 31% fewer tool calls than the base model, on the tasks both solved [1][3]. The guide notes there was no run without the bonus, so the bonus’s effect cannot be isolated [1].
- Imitation did less well: fine-tuning on a 27B model’s successful attempts reached 47.5% with OpenCode data and 43.1% with all 3,189 attempts from four harnesses, within the noise of the base model. Both are below the RL runs [1]. (Delangue’s post attaches 47.5% to the 3,189-attempt run; the guide does not [1][2].)
Hugging Face has released the proxy, trainer integration, tasks, fine-tuning data, training code and seven models [1][2][3]. The weights carry Liquid AI’s LFM Open License, and the model card says it “is not an official Liquid AI release” [3].
What this does not prove
- That it works for bigger models. Every run used a small model of 2 to 2.6 billion parameters, trained once. Hugging Face says larger runs are still being set up [1].
- That it works in harnesses left out of training. The trained model was scored only in the four harnesses it trained in. When the OpenCode-only model was tested in other harnesses, most of its gain stayed in OpenCode [1].
- That it applies to general coding. The tasks are data analysis. Liquid AI’s own model card does not recommend LFM2.5-2.6B for agentic coding [6].
- That these four are the model’s native harnesses. Liquid AI says it trained LFM2.5 inside Hermes Agent, OpenClaw “and other harnesses” [7]. The guide notes those are not the four used here [1].
- That the gains are independently confirmed. These are single-run, company-reported results, and the best checkpoint was picked on the test set [1][3].
The Bottom Line
Hugging Face’s point is simple: a score describes a model in a particular harness, not the model in general [1]. Its framework lets developers train open models inside the tools people actually use, and its early, small-scale results suggest that training across several harnesses spreads the gains [1][3]. Larger models, unseen harnesses and everyday coding remain untested.
Sources
- Adithya S Kolavi, Joel Niklaus, Sergio Paniego Blanco, Leonie Monigatti, Amine Dirhoussi, Ben Burtenshaw, Lewis Tunstall and Leandro von Werra, “The ultimate guide to multi-harness RL” (Hugging Face / Liquid AI), published 24 September 2026, updated 1 October 2026 (PDF). https://fineenvs-multi-harness-rl.hf.space/the-ultimate-guide-to-multi-harness-rl.pdf. Interactive version: https://huggingface.co/spaces/FineEnvs/multi-harness-rl
- Clément Delangue (@ClementDelangue), X post, 5 October 2026, 15:48 BST (read via the X API). https://x.com/ClementDelangue/status/2107120717980471638
- FineEnvs, “LFM2.5-2.6B-multiharness-RL” (Hugging Face model card). https://huggingface.co/FineEnvs/LFM2.5-2.6B-multiharness-RL
- Hugging Face TRL, pull request #6420, “Async grpo OpenEnv harness rollout” (merged 24 July 2026). https://github.com/huggingface/trl/pull/6420
- Hugging Face TRL, pull request #6947, “AsyncGRPO: train OpenEnv harnesses from validated token captures” (merged 1 October 2026). https://github.com/huggingface/trl/pull/6947
- Liquid AI, “LFM2.5-2.6B” (Hugging Face model card). https://huggingface.co/LiquidAI/LFM2.5-2.6B
- Liquid AI, “LFM2.5-2.6B: Deploy Agents Everywhere” (blog), 4 August 2026. https://www.liquid.ai/blog/lfm2-5-2-6b

