Two pieces of research this week make the same broad point from opposite ends: what an AI model can do and what it does do can be far apart.
The first is a new benchmark, Humanity’s Sixth Sense (HSS), from Scale AI and the visual-AI company Elorian. People score 93.1% on its everyday visual questions; the best AI model scores 53.6% [1][2]. The second is a preprint, a paper posted publicly before peer review, which reports that “base” models, ones not specially trained to reason, can reason almost as well as trained versions if their answer simply starts with the right word [3].
Humanity’s Sixth Sense: what does it test?
HSS asks the kind of questions people answer about a scene without thinking: what just happened, what will happen next, whether something will fit, who is in charge, what someone else can see [1]. It has 522 tasks, 288 images and 234 short video clips, across four areas: time and cause, physical and spatial logic, social understanding, and abstract or contextual inference [1].
Every question is open-ended and marked against a rubric written by people. An answer only counts if it meets every point, so a model cannot pass by guessing [1]. One example from Elorian’s post shows a library shelf and asks whether two more white books would fit. The answer is yes, in gaps already there; GPT-6 Astra, Gemini 3.8 Flash and Claude Opus 5.5 each said no on all three attempts [1].
How did the models do?
Results are from the paper’s Table 1, as published by Elorian and on Scale’s leaderboard [1][2]:
- Human baseline: 93.1%.
- GPT-6 Astra (at its highest “max” reasoning setting): 53.6%, the best model.
- GPT-6.1 Sol (max): 46.6%.
- Claude Opus 5.5 (xhigh): 44.6%.
- Median of the 25 models tested, from eight companies: 30.9%.
As a check, the researchers ran GPT-6 Astra again with the image or video removed. It scored 6.6%, roughly what you get from guessing at the question text alone [1].
Why do the models fail?
This is the most useful finding. Across 8,573 labelled failures, 94% traced to seeing or inferring rather than reasoning [1]. In 53% of cases the model misread something in the picture; in 41% it failed to infer something the picture implied but did not show. Only 5% came down to faulty logic [1].
Other patterns [1]:
- Social understanding (beliefs, intentions, who defers to whom) was the weakest area for 21 of the 25 models.
- Video was harder than still images for 23 of 25 models, by 7.3 points on average.
- More thinking does not always help. GPT-6 Astra dropped 14 points on working out what happened just before a scene when its reasoning setting went from high to xhigh, then partly recovered at max.
- Zoom tools help a little. Letting models crop and zoom through agent tools such as Claude Code and Codex lifted the best setup to 59.3% on a 388-task subset.
One caveat on independence: Elorian says it is “building the foundation of visual thinking”, so it has a commercial interest in showing that today’s models see poorly [1]. The results are also tied to particular model versions and settings, and Elorian notes that vendors update models continuously [1].
The preprint: reasoning on cue
The second study, posted to arXiv on 5 October by Sophie L. Wang, Amil Dravid, Rulin Shao, Kevin Farhat, Sewon Min and Alexei A. Efros, has not been peer reviewed. We describe only what its abstract says [3].
Some background. A base model is one trained only to predict the next word on large amounts of text. Many reasoning models are then trained further with reinforcement learning (RL), rewarding them for correct answers. A common assumption is that RL teaches them to reason.
The abstract challenges that. It reports that fixing particular opening words (“token cues”) at the start of a base model’s answer makes its maths and coding performance “competitive with that of its reinforcement learning (RL)-trained counterparts” [3]:
- Starting the answer with a full stop, a blank line and “Okay” raised Olmo-3-7B‘s score on the MATH-500 maths test from 42% to 78%.
- Starting with “Alright,” raised Qwen3-14B from 72% to 87%.
The abstract also says [3]:
- RL mostly makes these cues more likely, and fixing the cues “recovers much of” RL’s gain over the base model.
- The effect comes from training data. By editing the data, the authors turned an arbitrary word, “chicken”, into a working reasoning cue, or removed an existing cue’s effect. A similar edit made “Think duck duck goose” as effective as “Think step by step”.
- Opening words also changed safety behaviour: different cues produced different patterns of refusing or complying, linked to different kinds of training data.
If it holds up, the finding suggests that some of what RL appears to add is already present in the base model, waiting for the right trigger.
What this does not prove
- That AI cannot see. HSS measures one kind of intuitive visual reasoning, on 522 tasks chosen by its creators [1].
- That one model is “best” at vision generally. Published scores carry margins of roughly ±3 to ±4 points and depend on the reasoning settings used [1].
- Anything beyond the preprint’s abstract. It is unreviewed, and we have not drawn on its body [3].
- That the cue trick works on every model or task. The abstract reports results for specific models and benchmarks [3].
- That RL is unnecessary. The abstract says cues recover “much of” RL’s gain, not all of it [3].
The Bottom Line
HSS shows that today’s best model gets barely half of everyday visual questions right, against 93.1% for people, and that the problem is mostly seeing, not thinking [1]. The arXiv preprint, not yet peer reviewed, argues that much reasoning ability is already in base models and that the right opening word can unlock it [3]. Together they are a reminder to ask what a score is really measuring, and under what conditions.
Sources
- Elorian, “Introducing Humanity’s Sixth Sense (HSS),” 8 October 2026 (figures quoted from the Scale AI and Elorian paper “Humanity’s Sixth Sense: Benchmarking Intuitive Visual Reasoning in Multimodal Models”, Guo, Gu, Jang et al.). https://elorian.ai/blog/humanitys-sixth-sense
- Scale Labs, “AI Model Leaderboards & Benchmarks” (Humanity’s Sixth Sense leaderboard), read 8 October 2026. https://labs.scale.com/leaderboard
- Sophie L. Wang, Amil Dravid, Rulin Shao, Kevin Farhat, Sewon Min, Alexei A. Efros, “Base Models Can Reason By Taking a Cue From Training Data,” arXiv:2610.06851 (preprint, not peer reviewed; abstract only used), 5 October 2026. https://arxiv.org/abs/2610.06851

