PotARCin: Multi-Dimensional Evaluation of Skill Acquisition in Abstract Reasoning Tasks

Claas Beger, Ryan Yi, Melanie Mitchell
Santa Fe Institute
Correspondence: claasbeger@santafe.edu
NeurIPS 2026 (Evaluations and Datasets Track)

Abstract

The Abstraction and Reasoning Corpus (ARC) has become a prominent benchmark for evaluating general abstract reasoning and fluid intelligence in AI models. Yet standard ARC evaluation considers only a single capability: producing the correct output grid for a test input. We argue that this narrow format fails to evaluate the diversity of abilities that genuine abstract skill acquisition should enable. We introduce PotARCin, a benchmark that extends ARC by assessing understanding of a task's underlying abstract rule across five dimensions: Definition, Classification, Constrained Generation, Editing, and Inversion. PotARCin employs programmatic methods to generate new task instances and transform given inputs for a given ARC task, enabling dynamic generative sampling beyond fixed input-output pairs. Across five state-of-the-art models evaluated on the ARC-AGI-1 training set, we observe a 25–52 percentage-point performance gap between standard ARC evaluation and evaluation on PotARCin, and find that multi-dimensional evaluation reorders models that standard accuracy ranks alike. We further investigate effects of generative sampling, difficulty of corruption types, and questions of self-consistency, showing that models frequently contradict their own formalized rule even where they have stated it correctly. We also introduce P-ARC, a held-out hand-crafted test set, on which models achieve 1–8% accuracy across all five dimensions, underscoring the importance of more holistic evaluations of abstract reasoning capabilities.

Five uses of the same rule

Each ARC task defines a small skill: inferring and using the rule behind its demonstrations. Standard ARC evaluation tests one use of it, transforming a test input. PotARCin presents the same demonstrations and asks for four more, all scored automatically with task-specific generator and verifier programs.

Diagram of the five PotARCin dimensions around the standard ARC task: Definition, Classification, Constrained Generation, Editing and Inversion, each illustrated on the same ARC task.
Overview of PotARCin on ARC-AGI-1 task 890034e9, for which GPT-5.4 solved the standard test-grid task but failed all five additional dimensions.

Definition

Write a Python program implementing the rule, checked on the demonstrations, the test input and hundreds of generated examples.

Classification

Decide for five candidate pairs whether each follows the rule; candidates mix valid examples with plausible corruptions.

Constrained generation

Produce a new valid input–output pair that differs from the demonstrations.

Editing

Repair a corrupted pair so that it follows the rule again, while staying close to it.

Inversion

Given an output grid, produce an input that the rule maps to it.

Key findings

We evaluated GPT-5.4, Gemini 3.1 Pro, Claude Opus 4.6, Kimi K2.5 and MiniMax M2.5 on the 400 ARC-AGI-1 training tasks and on P-ARC, with two independent runs each. A task counts as fully solved only if all five dimensions are correct.

25–52 pts
drop from output-grid accuracy (63–89%) to full-task accuracy (21–58%) on ARC-AGI-1
1–8%
full-task accuracy of every model on the held-out P-ARC tasks
21.1%
of wrong answers agree with the model's own correct program, vs. 99.1% of right ones
15 of 20
fully solved tasks fail at least one of 10 resampled instances (GPT-5.4)

Output accuracy does not reflect rule competence

Definition and Classification are the hardest dimensions and Inversion the easiest. Multi-dimensional evaluation also separates models that output-grid accuracy ranks alike: Kimi K2.5 has the lowest output-grid accuracy of the five models, yet nearly doubles MiniMax M2.5's full-task accuracy (38.6% vs. 20.6%). On P-ARC, every level drops sharply.

Grouped bar charts for ARC-AGI-1 and P-ARC showing output-grid accuracy, accuracy on each of the five PotARCin dimensions, and full-task accuracy for five models.
Performance on the standard output-grid task, each PotARCin dimension, and all five together, for 400 ARC-AGI-1 training tasks and 50 P-ARC tasks, pooled over two runs. Whiskers are 95% Wilson intervals.

Models contradict rules they have stated correctly

Because Definition yields an executable program, we can check the model's other answers against its own program. Where that program is correct and the other answer is right, the two agree 99.1% of the time on ARC-AGI-1. Where the other answer is wrong, agreement falls to 21.1% (17.1% on P-ARC): most mistakes are not consequences of a wrong rule, but failures to apply a rule the model has already formalized.

Single samples overestimate robustness

For 20 tasks that GPT-5.4 initially solved in all five dimensions, we sampled 10 new instances per dimension. On 15 of them it fails at least one instance, mostly within the first two additional samples. Probing five different dimensions also exposes more failures than repeating one dimension five times (25.5% vs. 20.2% of trials).

Line chart of accuracy against the number of additional samples for each dimension and for average and hard full-task accuracy, for GPT-5.4 on 20 tasks.
Accuracy of GPT-5.4 as more instances are sampled for 20 ARC-AGI-1 tasks it initially solved in all five dimensions. Hard full-task accuracy counts tasks that pass every sampled instance so far.

P-ARC

P-ARC is a held-out set of 50 hand-crafted ARC-style tasks, which we estimate to lie between ARC-AGI-1 and ARC-AGI-2 in difficulty. Each task comes with a generator and a verifier program, 50 fixed generated examples, three human-made corruptions (erroneous attempts from the feasibility check or errors designed by the task's creator), and a reviewed natural-language rule. The dataset is available on Hugging Face under the MIT license. Four of its tasks:

BibTeX

@inproceedings{beger2026potarcin,
  title         = {PotARCin: Multi-Dimensional Evaluation of Skill Acquisition in Abstract Reasoning Tasks},
  author        = {Beger, Claas and Yi, Ryan and Mitchell, Melanie},
  booktitle     = {Advances in Neural Information Processing Systems},
  year          = {2026},
  eprint        = {2609.27288},
  archivePrefix = {arXiv}
}