AI agents are being asked to do longer, messier jobs, and to be graded on them. Three items from 8 October 2026 cover both sides: Anthropic’s account of agents assembling a sky map, and two studies on how far agent scores can be trusted.
Spotted via Anthropic (@AnthropicAI) on X.
Claude Science and the missing UV map
Confirmed (Anthropic’s own post); results are the company’s account. Brice Ménard, “an astrophysicist at Johns Hopkins University and a researcher at Anthropic”, writes that he used Claude Science to produce “the first complete map of the sky in UV light” [1]. Ultraviolet light is blocked by the ozone layer, so it can only be observed from space, and existing surveys left gaps.
He says “Claude orchestrated a team of AI agents”. They searched for public UV surveys, cleaned up NASA’s GALEX data (“some 38,000 separate observations”), cross-calibrated the surveys and merged them [1]. For the sky never observed in UV, Claude learned how UV brightness relates to visible, infrared and radio data, then estimated the missing areas. In Anthropic’s words: “About a third of this map, including much of the galactic plane, was predicted with Claude Science” [1]. Extra layers label each pixel “measured” or “predicted”.
In a test where known data was hidden, the method estimated it “to within about 10% of the real UV measurements”. Claude then added estimates from “more than 100 million individual stars” using Gaia data. The work took “several days” while Ménard worked on other projects [1].
The post is candid about a miss: a calibration artefact “passed two rounds of review by other agents without the problem being caught” until Ménard spotted it [1]. The post does not say what Claude Science costs or who can use it.
TRACE: when the grader, not the agent, moves the score
Preprint. Radhika Gaonkar of Prime Intellect posted TRACE to arXiv on 8 October [2]. It tests whether a change in an agent’s score reflects the agent or the measurement. In a synthetic suite, renaming tools “lowers a scripted agent’s score by 0.250 even though it performs exactly the same operations”, because the checker matched tool names [2].
On 88 new tau2-bench tasks (simulated customer-service conversations), renaming or reformatting “leaves reward unchanged to within ±0.10 for seven of eight agent–change pairs”, while deliberately misleading tool names “lower every agent’s reward by 0.20–0.44”. But “identical reruns flip 15–36% of task outcomes” [2]. Two frontier judges, which the paper names as GPT-6.1 Sol and Claude Opus 5.5, “disagree with each other on 57% of the same records”, largely because one grades procedure rather than outcome [2]. The paper’s point: a single run cannot separate a real change from noise.
Exa’s ATLAS search benchmark
Vendor benchmark. Exa, which sells a search API, is “previewing ATLAS”, a set of 547 “deep and wide research tasks” seeded from its anonymised search demand [3]. Its findings: “no agent run costing less than $1 per task achieved a row F1 over 0.5”; “the best system we tested reaches 0.66 row F1 only with a budget of $8.92 and 18 minutes per task”; and “Even the most expensive search agents miss about 1/3 of the golden results” [3]. Row F1 scores whole table rows, which count only if every cell is right.
Exa also says “Exa defines the cost-performance Pareto frontier” on its own benchmark. The tasks, answer tables and grader are due “in the coming weeks”, so no one has yet reproduced the results [3].
What this does not prove
- That Claude observed the whole UV sky. About a third of the map is predicted, not measured, and the post is Anthropic’s own, not a peer-reviewed paper [1].
- That agents work unsupervised. Agent reviewers missed the artefact; a human caught it [1].
- That agent leaderboards are wrong. TRACE is a preprint covering four agents and two presentation changes; it shows single runs are noisy [2].
- That Exa’s search is best. Those are Exa’s results on Exa’s benchmark, not yet independently checked [3].
The Bottom Line
Agents can now do the slow assembly work scientists put off, but people still have to check it, and the scores used to rank agents are noisier than they look. Treat single-run results and vendor-run benchmarks as starting points, not verdicts.
Sources
- Brice Ménard, “The missing map of the sky”, Anthropic, 8 October 2026 (company post; author is an Anthropic researcher; not peer reviewed). https://www.anthropic.com/research/the-missing-map-of-the-sky
- Radhika Gaonkar (Prime Intellect), “TRACE: Diagnosing Verifier Brittleness in Agentic Evaluation”, arXiv:2610.11678, submitted 8 October 2026 (preprint, not peer reviewed). https://arxiv.org/abs/2610.11678 (code: https://github.com/RGaonkar/trace-verifier-stress-tests)
- Alexander Goldberg, Joshua Ahn and Scott Langille, “ATLAS: Evaluating Agents on Search-Intensive Tasks”, Exa blog, 8 October 2026 (vendor benchmark; Exa sells a search API; tasks not yet released). https://exa.ai/blog/atlas-benchmark
- Anthropic (@AnthropicAI), X post, 8 October 2026, 21:16 BST (company post, read via the X API; first seen here). Its line that the work “would have taken humans weeks” is Anthropic’s framing; the blog says such work “takes weeks of painstaking work” and that this project spread over “several days”. https://x.com/AnthropicAI/status/2108290395599667700

