Distinguishing Performance From Competence in Evaluations of Humanlike Abstract Reasoning

Claas Beger1,*, Ryan Yi1, Shuhao Fu1, Kaleda Denton1, Arseny Moskvichev2, Sarah W. Tsai3, Sivasankaran Rajamanickam3, Melanie Mitchell1
1 Santa Fe Institute, 2 Sandia National Laboratories, 3 Advanced Micro Devices Inc.
* Correspondence: claasbeger@santafe.edu
NeurIPS 2026 (Evaluations and Datasets Track)

Abstract

AI reasoning models have exceeded human performance on the ARC-AGI-1 benchmark, but does that mean state-of-the-art models have the underlying competence—humanlike abstract reasoning—the benchmark was designed to test? Here we investigate the abstraction abilities of AI models using the closely related but simpler ConceptARC benchmark. Our evaluations vary input modality (textual vs. visual), use of external Python tools, and reasoning effort. Beyond output accuracy, we evaluate the natural-language rules that models generate to explain their solutions, enabling us to assess whether models recognize the abstractions that ConceptARC was designed to elicit. We show that the best models' rules are frequently based on less abstract, more domain-specific concepts, capturing intended abstractions considerably less often than humans. In the visual modality, AI models' output accuracy drops sharply; however, our rule-level analysis reveals that a substantial share of their rules capture intended abstractions, even as the models struggle to apply these concepts to generate correct solutions. In short, we show that using performance (accuracy) alone can substantially overestimate AI competence in textual modalities while underestimating it in visual modalities—an illustration of the risk of mistaking performance for competence.

Key findings

We evaluate o3, o4-mini, Claude Sonnet 4, Gemini 2.5 Pro, GPT-4o, Llama 4 Scout and Qwen 2.5 VL 72B on the 480 tasks of ConceptARC (16 spatial and semantic concepts, 30 tasks each), with textual and visual inputs, low and medium reasoning effort, and with and without Python tools. We compare their outputs and rules with those of human participants.

75.6%
o3 accuracy with textual inputs, vs. 73% for humans
27%
of o3's correct textual outputs rest on unintended or incorrect rules
29.2%
o3 accuracy with visual inputs
28%
of o3's incorrect visual outputs still come with a correct-intended rule

o3 results are for medium reasoning effort with Python tools.

Accuracy and rule correctness diverge

We classify the natural-language rule behind each answer as correct-intended (it captures the intended abstraction), correct-unintended (it works on the demonstrations but is an unintended solution), or incorrect. With textual inputs, a sizeable share of the models' correct outputs come with rules that miss the intended abstraction. With visual inputs, output accuracy drops sharply, yet many incorrect outputs come with rules that do capture the intended abstraction.

Stacked bar chart of the percentage of tasks for o3, Claude and Gemini with textual and visual inputs, and for humans, split by correct and incorrect output grid. Each bar is divided into correct-intended, correct-unintended, incorrect and not-classified rules.
Rule classification for tasks with correct and incorrect output grids, as a percentage of the 480 tasks, for o3, Claude Sonnet 4 and Gemini 2.5 Pro (medium effort, with Python tools) with textual and visual inputs, and for human participants. Rules were not collected for humans' incorrect outputs, so these are shown as not classified.

Humans describe objects, models describe grids

ConceptARC builds on core-knowledge priors such as objectness. Most human rules are phrased in terms of objects, while the models' rules more often focus on colours, individual pixels and other low-level features of the grid.

Bar chart of the proportion of rules using objectness terms and grid-specific terms. Humans: about 0.89 objectness and 0.14 grid-specific. Claude Sonnet 4, o3 and Gemini 2.5 Pro: about 0.42 to 0.53 objectness and 0.80 to 0.88 grid-specific.
Proportion of rules that use objectness terms and grid-specific terms, for humans and for Claude Sonnet 4, o3 and Gemini 2.5 Pro with Python tools.

Correct outputs from unintended rules

In each example below the model's output matches the ground truth, but its rule does not capture the intended abstraction.

Example task with training examples, the model's rule, the test input, and a correct model output that matches the ground truth.
Heuristic. The rule picks the colour with the lowest density relative to its bounding box. Generic heuristics like bounding boxes and cell connectivity recur throughout the models' rules.

Dataset viewer

Explore the model and human outputs and rules for each ConceptARC task.

Full data can be downloaded from AIHumanAbstraction/ConceptARC_Rule_Annotations on Hugging Face.

BibTeX

@inproceedings{beger2026distinguishing,
  title     = {Distinguishing Performance From Competence in Evaluations of Humanlike Abstract Reasoning},
  author    = {Beger, Claas and Yi, Ryan and Fu, Shuhao and Denton, Kaleda and Moskvichev, Arseny and Tsai, Sarah W. and Rajamanickam, Sivasankaran and Mitchell, Melanie},
  booktitle = {Advances in Neural Information Processing Systems},
  year      = {2026}
}