Introducing Humanity's Sixth Sense (HSS)

A benchmark measuring the implicit visual reasoning and intuitive understanding we draw from a scene. Built in partnership with Scale AI.

Six examples from HSS. How would you answer?

  1. A long run of library shelving filled with bound law journals, seen at an angle down the row.
    “Law library books” by Janet Lindenmuth, CC BY 2.0. Resized for the web; otherwise unaltered.

    Would two more white books fit on this shelf?

    Yes — in the gaps already there.

    In the paper’s evaluation, GPT 6 Astra, Gemini 3.8 Flash and Claude Opus 5.5 each answered “no, there is too little space” on all three attempts. Fig. 10 · affordance & feasibility

  2. A painted mountain landscape: a rider and a walking figure with two dogs on a track beside a stone bridge, a hilltop temple beyond.
    Gebirgslandschaft (Mountain landscape) by Paul Bril (c. 1553–1626). Public domain, via Wikimedia Commons. Resized for the web.

    From which side of the frame did the traveler with the two dogs enter?

    The bottom edge — you can tell by tracing the path back.

    In the paper’s evaluation, GPT 6 Astra, Gemini 3.8 Flash and Claude Opus 5.5 each answered “the lower-left side” on all three attempts. Fig. 11 · retrodiction

  3. Two teams of draught horses in harness pulling ploughs across a stubble field, with spectators behind.
    Horse Ploughing (8) by Bernard Spragg, 2010. Public domain, via PICRYL.

    Are these chains bearing the plough’s drag weight?

    Yes — they are taut and loaded.

    In the paper’s evaluation, GPT 6 Astra, Gemini 3.8 Flash and Claude Opus 5.5 each answered “no, they hang slack” on all three attempts. Fig. 12 · mechanistic causality

  4. A white single-engine light aircraft parked in a hangar, with a larger military transport aircraft and several people behind it.

    Are the people around the aircraft visitors, or crew preparing it?

    Crew — you can tell from the open access panels.

    In the paper’s evaluation, GPT 6 Astra, Gemini 3.8 Flash and Claude Opus 5.5 each answered “visitors examining it” on all three attempts. Fig. 13 · social role, norm & power

  5. A 1920s six-cylinder in-line aircraft engine on a museum stand, propeller hub at the left.
    Walter IV aircraft engine, Museum of Aviation, Košice by ZemplinTemplar, 2010. CC BY-SA 4.0, via Wikimedia Commons.

    If the shaft spins, does the cage around it spin too?

    No — it is bolted to the engine casing.

    In the paper’s evaluation, GPT 6 Astra, Gemini 3.8 Flash and Claude Opus 5.5 each answered “yes, it turns with the shaft” on all three attempts. Fig. 13 · change & consequence

  6. A striped flag on a pole against an open sky.

    Which way is the wind blowing, from the striped flag?

    Toward the camera.

    In the paper’s evaluation, GPT 6 Astra, Gemini 3.8 Flash and Claude Opus 5.5 each answered “from right to left across the frame” on all three attempts. Fig. 9 · hidden & invisible

These questions are among the 522 in Humanity's Sixth Sense (HSS), a new benchmark we built in partnership with Scale AI. HSS measures how a model interprets and reasons through information a visual scene implies but doesn’t explicitly show. When a person looks at something, they can draw countless conclusions without significant thought. As we grow into adults we develop a sort of sixth sense, allowing us to conceptualize real-world scenes, including the spatial layout, context, and structural relationships beyond isolated objects.

This cognitive development grows gradually from infancy through childhood, driven by specialized brain regions and a framework known as “scene grammar.” By their early teens, most people can accurately sketch out a scene, even with limited visual cues and anticipate what will happen next. This is done largely by filling in the blanks and predicting actions based on past experiences. In theory, it seems like frontier models should excel in compiling complex information from a myriad of perspectives, leading to accurate predictions. Most models currently lack fundamental visual skills which would allow them to operate natively with visually-dependent tasks.

The whole picture

Model research has largely focused on text-based problems utilizing code, math, books, and exam questions. Most “visual” benchmarks test the same text-oriented flows. HSS is unique in that it tests simple visual reasoning, across four areas:

  • Temporal & causal dynamics What happened before the scene, and what happens next.
  • Physical & spatial logic What fits, what’s in reach, and what someone else can see from where they are standing.
  • Social understanding Intent, attention, belonging.
  • Abstract & contextual inference The patterns and rules a scene implies.

The benchmark is 522 tasks in all: 288 images and 234 short video clips, 17.6 hours of video between them.

What HSS measures

Frontier models write production code and pass professional exams. Given they also edit images and can identify objects in videos, it can seem like these models understand visual tasks better than they can. Most visual benchmarks reinforce that impression.

Humans are capable of visual reasoning well beyond what most visual benchmarks measure. We can usually tell what happened a moment before a photo was taken, whether a box will fit through a door, or why a dog took off running. Every question in HSS is open-ended, and answers are graded against rubrics written by humans. An answer only counts if it meets every point on the rubric, so a model can't pass by guessing.

Leaderboard

Pass@1 (%) on HSS, from Table 1 of the paper. These are the figures as published, not live numbers — they reflect the model versions and reasoning-effort settings used in that evaluation. Model names are written out; the paper and the live leaderboard use the API identifiers, so GPT 6 Astra appears there as GPT-6-astra. For the current picture, see Scale’s leaderboard.
ModelEffortpass@1 ImageVideo TimeSpaceMindContext
Human baseline—93.194.191.997.094.989.189.2
GPT 6 Astramax53.6 ±4.155.651.158.253.646.155.6
GPT 6.1 Solmax46.6 ±3.845.547.955.245.836.447.4
Claude Opus 5.5xhigh44.6 ±3.851.336.544.044.934.855.6
Gemini 3.8 Flashhigh41.6 ±3.747.234.843.337.540.048.4
Claude Fable 5.1xhigh40.8 ±3.644.736.043.341.134.244.1
Gemini 3.7 Flashhigh39.8 ±3.745.433.040.540.935.841.5
Muse Spark 1.3max37.4 ±3.544.828.239.638.125.845.8
Claude Fable 5max34.5 ±3.737.231.338.137.130.030.4
Qwen 3.8 Maxmax33.0 ±3.638.326.238.135.221.034.7
Gemini 3.5 Flashhigh32.6 ±3.435.129.635.631.631.831.4
Gemini 3.6 Flashhigh31.9 ±3.432.531.237.129.732.428.4
GPT 6 Solmax31.2 ±3.431.830.538.128.622.436.3
Qwen 3.8 Flashmax30.9 ±3.337.222.937.929.218.637.1
Claude Opus 5max30.5 ±3.532.527.931.335.022.430.1
GPT 5.6 Solmax30.0 ±3.328.531.938.626.327.328.1
GLM 5.3 Flashmax26.8 ±3.030.222.631.326.318.231.0
Kimi K3max25.5 ±3.129.121.233.127.415.523.2
MiMo v2.6 Flashhigh25.4 ±2.929.220.630.129.612.126.7
MiMo v2.6 Prohigh25.2 ±2.930.418.628.329.514.525.7
Claude Sonnet 5max24.8 ±3.229.518.929.127.514.225.8
Muse Spark 1.2xhigh24.5 ±3.127.520.823.624.221.229.7
Claude Opus 4.8max23.5 ±3.227.218.927.126.514.223.5
GPT 5.6 Terramax22.7 ±3.224.120.928.421.019.121.9
GPT 5.6 Lunamax21.6 ±3.223.519.426.924.613.318.6
GPT 6 Lunamax21.0 ±3.023.318.226.924.19.420.6
GPT 6 Astra, media removedmax6.6 ±1.52.411.79.55.54.86.5

Twenty-five models from eight companies. Most score below 40%, and the median model scores 30.9%. The last row is the paper’s blind control — the strongest model, run again with the image or video taken away. It scores 6.6%, which is the floor you get from guessing at the question text alone.

Where models fail

HSS tasks are reasonably easy for the average person, with a baseline score of 93.1%.

Across the 8,573 individual failures the team labeled, 94% trace to seeing or inferring rather than to reasoning. The model misread something in the image in 53% of cases, and in a further 41% it failed to infer something the image implied but did not show. Only 5% came down to faulty logic.

The two most common single mistakes were (1) missing a detail that decides the answer (21% of failures) and (2) misidentifying an object, person, or role (20%). Failures were consistent across the 24 models that analysis covers. Struggling significantly with nuance, perspective, object variations, lighting effects, and causality. On tasks that five or more models failed, a median of 80% failed for the same reason.

Stacked bar chart of eleven subdomains, each split into perception, latent and reasoning failures. Perception dominates affordance and feasibility at 91 percent; latent dominates theory of mind at 74 percent; alien viewpoint is the only subdomain where reasoning is substantial, at 32 percent.
Almost nothing is a reasoning failure. Where a subdomain turns on what the scene shows, models misread it; where it turns on what the scene implies, they miss it. Alien viewpoint is the exception — a third of its failures are models reasoning wrongly from facts they had right. Redrawn from Figure 3 of the paper.

The weakest domain is social understanding — reading beliefs, intentions, and who defers to whom. It is the lowest-scoring domain for 21 of the 25 models, averaging 24.4% against 34.1% across the other three. Video is harder than still images for 23 of the 25 models, by 7.3 points on average.

Increasing reasoning effort helps overall, but on some skills it makes things worse. The paper reports a 14-point drop for GPT 6 Astra on retrodiction — working out what happened just before the moment shown — when reasoning effort is scaled from high to xhigh; the score then partly recovers at max. Best scores often came at a middle setting. Increasing reasoning effort on several models frequently led to the model overthinking or focusing on incorrect evidence.

Agent tools such as Claude Code and Codex modestly improved scores by letting a model crop, zoom, and take a second look. On the 388-task subset used for the agentic evaluation, the best setup reached 59.3%. A second look cut failures from 944 to 759 by catching missed cues and misread objects. What a closer look does not fix is depth: errors from reading a 2D overlap as 3D alignment were essentially unchanged, 102 to 101. Agents sometimes talked themselves out of answers they had gotten right the first time.

Humans noted that they mainly answered the questions intuitively. Typically spending more time explaining their reasoning than coming up with the correct answer. Many human failure cases could be traced to sloppiness, misunderstanding the question, and arithmetic errors. Author’s note: definitely never happens to me.

Is the horse walking towards you, or away?

Via r/opticalillusions.

Elorian is building the foundation of visual thinking. HSS is a preliminary step in identifying and understanding how to solve the complex gaps in frontier models. HSS highlights how today's models ability to see and interpret visuals is highly limited and not a quirk of one lab's model.

Read the paper The full leaderboard The dataset

Figures, questions and reference answers in this post are quoted from Humanity’s Sixth Sense: Benchmarking Intuitive Visual Reasoning in Multimodal Models, by Guo, Gu, Jang et al. — a joint paper from Scale AI and Elorian. Model results are as reported there, for the model versions and reasoning-effort settings used in that evaluation. Scores on this benchmark will move as vendors ship new versions; the leaderboard above is the current picture.