A “harness” is the software wrapped around an AI model to turn it into an agent: the tools it can call, the loops that drive it, and any planners or helper agents. A new preprint asks how much of that machinery still helps. Two companies, meanwhile, are packaging expertise as add-on “skills” for existing agents.
Spotted via himanshu (@himanshustwts) on X and Pasqal (@PasqalQuantum) on X.
A minimal coding agent matched the elaborate ones
Preprint, submitted 30 September; it drew attention on X on 9 October. Kirill Brilliantov, Alejandro Hernández-Cano and Emmanuel Abbé, listed by Apple’s research site as EPFL researchers, with Hernández-Cano’s work done at Apple (as an intern, per arXiv), tested agents that build machine-learning models for Kaggle-style competitions [1][2].
Their abstract says that “under an equal time budget and the same frontier LLM backbone, open-source state-of-the-art harnesses provide no advantages over a single session of a minimal-harness coding agent baseline, pointing to the backbone as the primary driver for performance” [1]. Their own minimal agent, Malena, is one long coding-agent session with a few tools. The paper says: “With GLM 5.2, Malena earns a medal on 62.5% of the competitions, compared with 47.1% for the best external harness tested”, on the authors’ 30-task MLE-bench set with a 24-hour budget [1].
The biggest single gain they measured came from letting the model read, write and run code directly. The single long session was not the cheapest on model costs: at GLM 5.2 its modelled cost was about 6.2 times that of one rival harness, though the authors say hardware dominates overall cost [1]. They list limits: MLE-bench may be in models’ training data, some results lack statistical power, and tuning effort may have been uneven [1]. TSN covered a related finding on 8 October: the same model can score very differently in different harnesses [3].
Pasqal: quantum experiments from a coding agent
Confirmed (company release). Pasqal says its “agent skills are now available” for coding tools “including Claude Code, Cursor, and Codex” [4]. The skills take a researcher from an idea or paper to an experiment specification, a Pulser pulse sequence, emulation, submission to a Pasqal quantum processor through Pasqal Cloud, and analysis. Pasqal’s docs say local emulation needs no account; QPU runs need an account with QPU access, and the scripts ask for approval of the shot count before spending [5].
Pasqal’s July paper (arXiv 2607.25834) found that of 526 theory papers it classified, “258 are assessed as implementable with publicly available Pasqal QPUs today”, and its own caveat is that “execution success is not the same as scientific correctness” [4]. In one test, an expert had to correct the agent’s choice of measurement.
Microsoft: Dynamics 365 skills for Copilot
Confirmed; parts in private preview. Microsoft announced “30 new pre-built CRM skills for Copilot Cowork, now generally available. Autopilot is in private preview, and Code is available through the Frontier program” [6]. The skills draw on “more than 180 generally available tools” across Dynamics 365 Sales, Customer Service and Customer Insights, exposed through MCP server tools. Microsoft says “People remain in control of business actions”, and Autopilot’s recurring automation stays “subject to private-preview limitations” [6].
What this does not prove
- That harnesses are useless. The study covers ML-engineering benchmarks; it is a preprint, and weaker models still benefited from extra structure [1].
- That Apple endorses the finding. It is a research paper by EPFL authors, one of whom worked at Apple [2].
- That Pasqal’s agents do science unaided. Pasqal’s own paper records two failures caught only by experts [4].
- That Autopilot is available. It is in private preview; Code is early access [6].
The Bottom Line
The EPFL–Apple preprint suggests that, for now, the model and a good coding environment matter more than elaborate scaffolding. Pasqal and Microsoft are betting on lighter add-ons instead: skills plugged into agents people already use, with humans still approving the important steps.
Sources
- Kirill Brilliantov, Alejandro Hernández-Cano and Emmanuel Abbé, “How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?”, arXiv:2609.40303, submitted 30 September 2026 (preprint, not peer reviewed). https://arxiv.org/abs/2609.40303
- Apple Machine Learning Research, paper page (lists authors as EPFL; Hernández-Cano “Work done while at Apple”; “published October 2026”). https://machinelearning.apple.com/research/harness-autonomous-ml-engineering
- TSN, “Same Model, Different Score: Training AI Inside the Coding Tools People Actually Use”, 8 October 2026. https://tsnmedia.org/fineenvs-multi-harness-rl-coding-agent-harnesses-training/
- Jaap Kautz, “Turning Ideas into QPU Experiments: How Agentic Workflows Make Quantum Computing More Accessible”, Pasqal blog, 8 October 2026 (company release; cites C. Dalyac et al., arXiv:2607.25834, July 2026, Pasqal’s own preprint). https://www.pasqal.com/blog/turning-ideas-into-qpu-experiments-how-agentic-workflows-make-quantum-computing-more-accessible/ and https://arxiv.org/abs/2607.25834
- Pasqal, “Agentic workflows for neutral-atom experiments”, documentation, last updated 8 October 2026. https://docs.pasqal.com/agentic-toolkit/
- Deva Rajamohan (CVP, Dynamics 365 Customer Experience), “Bring Microsoft Dynamics 365 into the flow of work with 30 new CRM skills”, Microsoft Dynamics 365 blog, 8 October 2026 (company blog; Autopilot private preview). https://www.microsoft.com/en-us/dynamics-365/blog/it-professional/2026/10/08/bring-microsoft-dynamics-365-into-the-flow-of-work-with-30-new-crm-skills/
- himanshu (@himanshustwts), X post, 9 October 2026, 09:45 BST (read via the X API; first seen here). It calls the work a “paper from apple”; the authors are EPFL researchers, one of whom did the work at Apple. https://x.com/himanshustwts/status/2108479042521624931
- Pasqal (@PasqalQuantum), X post, 9 October 2026, 07:58 BST (company post, read via the X API). It says “Today we’re launching” the toolkit; Pasqal’s blog and docs are dated 8 October. https://x.com/PasqalQuantum/status/2108452030369202376

