More of the work of building AI is being handed to AI. Three releases this week tackle different parts of that shift: how to measure whether a model can run experiments well, how to train it without letting it cheat, and how to keep experiments repeatable.
- TasteVal, from P-Zero Research, measures how well models design and interpret their own experiments. Its headline results are the authors’ own and the paper is a preprint that has not been peer reviewed [1][2].
- Karotte, from Preference Model, is an open-source framework for building training environments that resist “reward hacking” (confirmed release; robustness claims are the company’s) [3].
- Zephon, from DatologyAI, is an open-source data loader that feeds a model the same data in the same order however many GPUs a run uses (confirmed release; test results are DatologyAI’s) [4].
What does TasteVal measure?
P-Zero Research defines “research taste” as picking good problems, designing experiments and interpreting the results. TasteVal measures the middle part: given a fixed AI research problem, how well a model chooses and learns from experiments [1][2]. To separate this from coding skill, the model under test only proposes experiments, while a fixed coding agent carries them out on a single high-end GPU [1][2].
The benchmark has eight new tasks drawn from frontier AI work, such as curating training data and training language models. The tasks are kept private so models cannot learn the answers [2]. The human baseline is the best attempt on each task from 24 experts who have recently worked at organisations including OpenAI, Google DeepMind and NVIDIA [1][2].
The key measure is compute efficiency: how much experimental computing a model needs to reach the expert’s score. A model that gets there with half the compute counts as having twice the taste [1].
What did it find?
According to the abstract and the authors’ summary [1][2]:
- The best model, Opus 5.5, beat the expert baseline with a compute multiplier of 2.3 (95% confidence interval 1.15 to 4.37), at roughly 1/30 of the experts’ average cost per attempt.
- Across 20 models released from 2023 to 2026, that multiplier has doubled about every 3.0 months since December 2025 (interval 1.7 to 5.0 months), up from every 14 months before.
- But the models’ final scores, ignoring compute, show no acceleration: they have doubled about every 14.6 months across the whole period.
In other words, models are getting much faster at reaching good results, but not, so far, reaching dramatically better ones.
The authors also run what they call “a naive extrapolation rather than a forecast” through an existing AI forecasting model, which pulls its estimated arrival date for superintelligence forward [1]. TSN treats that as an illustration of the authors’ assumptions, not a prediction.
What are the caveats?
The authors list them themselves [1]. Their results may overstate progress: the tasks are quick and cheap to check, which is exactly what AI labs find easiest to optimise for, and the human baseline does not include top researchers. TasteVal also does not measure choosing which problems are worth working on. The results may understate progress too: the models were given little tuning, and the best one spent about 1/30 as much per attempt as the human experts [1].
Because the paper is a preprint, TSN reports only its abstract and the authors’ public summary.
What is Karotte?
Reinforcement learning (RL) trains a model by rewarding it for succeeding at tasks inside software “environments”. Reward hacking is when the model finds a way to collect the reward without doing the task, for example by tampering with the grader. Preference Model argues this matters more in training than in testing: in a test, a hack skews one score; in training, it gets reinforced until it becomes the model’s habit [3].
Karotte is the framework Preference Model says it has used internally for a year to build such environments, now open-sourced under the MIT licence [3]. By default, anything the model does runs as an unprivileged user inside a sandbox; any processes it starts are stopped before grading; and suspicious files that could crash the grader are rejected [3]. The company says Karotte has been hardened through more than a million evaluation runs and controlled red-teaming [3]. Those are its own claims.
Preference Model says it has built RL environments for several frontier labs over the past year and has $16 million in seed funding led by a16z [3].
What is Zephon?
When researchers test whether a change to training data helps, everything else has to stay fixed. DatologyAI says the data loader often does not: restart a run on a different number of GPUs and most loaders quietly change which data each GPU sees, and in what order [4].
To show the effect, DatologyAI trained the same 1-billion-parameter model on 8, 16, 32 and 64 GPUs. With a standard loader, scores varied with the GPU count by up to 0.64 points on one evaluation suite and 0.82 points on another. With Zephon, the spread was 0.011 and 0.05 points [4]. DatologyAI says the standard loader’s spread is roughly four times the differences other researchers have used to judge data-cleaning methods, so a comparison could end up measuring the GPU count instead of the data [4]. These are the company’s own tests.
Zephon is released under the Apache 2.0 licence [4]. DatologyAI notes that more of its experiments are now designed and run by AI agents, which, like people, “won’t notice when a result depends on the GPU count” [4].
What links the three?
All three are about trusting the numbers when AI helps build AI. TasteVal tries to measure a research skill directly; Karotte tries to make sure a training reward means what it says; Zephon tries to make sure an experiment’s result comes from the change being tested.
What this does not prove
- That AI can now do frontier research. TasteVal’s tasks are small, quick to check and exclude problem selection; its authors say results may overstate progress [1].
- That the 3-month doubling will continue. It is a fitted trend over a short period, in a preprint not yet peer reviewed [2].
- That Karotte stops reward hacking. The robustness claims are Preference Model’s; no independent test is cited [3].
- That Zephon’s gains hold elsewhere. The score spreads come from DatologyAI’s own experiments [4].
The Bottom Line
TasteVal’s authors say the best model now reaches expert-level results on their tasks with 2.3 times less compute, and that this efficiency is doubling fast, while final scores improve more slowly [1][2]. That is a striking claim from a small, unreviewed benchmark, and its own authors list why it may overstate things. Karotte and Zephon are less dramatic but practical: tools to make AI training harder to game and AI experiments easier to repeat [3][4].
Sources
- Ollie Jaffe and Dane Sherburn, “Introducing TasteVal,” P-Zero Research, last updated 6 October 2026. https://pzeroresearch.com/work/tasteval/
- Oliver Jaffe and Dane Sherburn, “TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human Experts,” arXiv:2610.06824 (preprint, not peer reviewed), October 2026. Abstract only. https://arxiv.org/abs/2610.06824
- Jennifer Zhou and Felipe Peter, “Introducing Karotte: A Framework for Building Robust RL Environments,” Preference Model blog, 7 October 2026. https://preferencemodel.com/blog/introducing-karotte/
- DatologyAI, “Zephon: Fast, Flexible, Elastically Deterministic Data Loading,” DatologyAI blog, October 2026. https://www.datologyai.com/blog/zephon
Source note: For TasteVal, only the arXiv abstract and the authors’ own summary page were used, per TSN’s preprint rule; figures from the paper’s body are not reported. The brief’s “Opus 5.5 matches the best expert” is corrected to “exceeds the expert baseline”, the authors’ wording. Licences were checked on GitHub: preferencemodel/karotte (MIT) and datologyai/zephon (Apache-2.0). The brief’s Zephon figure (0.05) is the DCLM Core spread; the FineWeb spread was 0.011. The brief describes Preference Model as “out of stealth”; its post does not say that, so the phrase is dropped.

