AI Still Misreads Pictures, and Reasoning May Hinge on One Word: Two New Studies
Scale AI and Elorian's Humanity's Sixth Sense benchmark finds people score 93.1% on everyday visual questions while the best model, GPT-6 Astra, scores 53.6%, mostly because models misread the picture. A separate preprint, not yet peer reviewed, says base models can reason almost as well as RL-trained ones when their answer starts with the right word.
Top stories
Latest
No posts to display
Latest on TSN
Free TSN tools: crypto calculator, Flux dashboard and more.
