Scale AI releases visual-reasoning benchmark; best model scores 53.6% versus humans’ 93.1%
Humanity’s Sixth Sense tests what images and videos imply, not just what they show. Social understanding was the weakest area for 21 of 25 models, though those scores measure agreement with human interpretations.
Scale AI and Elorian released Humanity’s Sixth Sense on October 7, 2026, a benchmark for judging spatial, social, temporal and abstract meaning in visual scenes—not just recognizing objects. Its 522 fixed tasks include 234 silent-video tasks, and the release includes a dataset and evaluation code for testing other models. In Scale’s launch results, GPT-6-astra scored 53.6%, versus 93.1% for human participants; even extensive reasoning and visual tools did not close the gap. Results should be read with limits including automated grading and possible prior exposure to public-web media.
01
Across 25 tested models, the median pass@1 score was 30.9%; GPT-6.1-Sol was second at 46.6%.
02
Social understanding was the weakest domain for 21 of 25 models, averaging 24.4% accuracy versus 34.1% across the other domains.
03
Video lowered scores by an average of 7.3 percentage points compared with still images for 23 models; clips had no audio or transcripts.
Complex visual capabilities do not necessarily translate into the inferences people make at a glance. Scale AI and Elorian’s Humanity’s Sixth Sense benchmark, released October 7, 2026, puts that tension in numbers: Scale reports that the strongest tested model scored 53.6%, even at maximum reasoning effort, against 93.1% for human participants.
Beyond identifying what is in the frame
Humanity’s Sixth Sense, or HSS, asks models to interpret what a scene implies. Scale’s examples include judging whether a vehicle can fit between parked cars and inferring who holds authority in a room from a few seconds of video. Human-written questions probe time, space, social relationships and abstract patterns.
The release contains 522 tasks: 288 image-based and 234 video-based, spanning four main domains and eleven subdomains. Scale says those tasks survived three independent review rounds from an initial pool of 3,466. The release includes a dataset and evaluation code, with prompts and grading instructions for testing additional models.
The video material totals 17.6 hours. The median clip lasts 76 seconds, while the longest runs 28 minutes. All confirmatory metrics use a frozen, versioned set of the 522 tasks, fixing the test material for that release.
Long reasoning chains, weak social readings
The leaderboard covers eight vendors. GPT-6.1-Sol ranked second at 46.6%, followed by Claude-Opus-5.5 at 44.6% and Gemini-3.8-Flash at 41.6%. Most tested models scored below 40%, so the shortfall was not confined to a few low-performing systems.
The weakness was not evenly distributed. Social understanding was the lowest-scoring domain for 21 of 25 models. Accuracy there averaged 24.4%, compared with 34.1% across the other three domains. Video also proved harder than still images for 23 models, with an average decline of 7.3 percentage points.
These were not quick-answer runs. Scale evaluated models with high reasoning effort and the maximum permitted output length. Across the 25 models, reasoning averaged 4,000 tokens—the chunks of text models process—per task. Scale says models frequently overthought these questions without reaching the correct answer.
The researchers also tested a tool-using setup that dynamically manipulated visual inputs. Their paper says this narrowed the human-model gap but did not close it. Closer examination helped, but the reported improvement still left models below human performance on the benchmark.
What a correct answer means here
Scale’s leaderboard calls its main measure pass@1: the fraction of attempts judged correct, averaged across tasks. Each model gets three attempts per task, and the score averages those attempts rather than selecting the best answer. A response that never commits to an answer fails every grading criterion, even if it uses its entire output budget reasoning.
A separate rubric score awards partial credit, distinguishing a near-miss from a total miss. It matches pass@1 on the 71.5% of tasks with one grading criterion. The leaderboard also reports uncertainty intervals based on resampling tasks, reflecting sensitivity to which questions the benchmark contains.
An automated language-model judge, Claude-Opus-5, grades responses against reference answers and task-specific criteria. Scale flags several boundaries on interpreting the results:
Video is silent: clips have neither audio nor transcripts, excluding questions that depend on speech or off-screen sound.
Grading has judgment calls: automated judges can carry biases, and answers about people’s beliefs measure agreement with annotators, not objective truth.
Prior exposure remains possible: public-web images and clips may have appeared in training data, although the questions were newly written.
Reader comments
Newest comments first. Replies stay oldest first.