PotARCin: Multi-Dimensional Evaluation of Skill Acquisition in Abstract Reasoning Tasks
Abstract
The Abstraction and Reasoning Corpus (ARC) has become a prominent benchmark for evaluating general abstract reasoning and fluid intelligence in AI models. Yet standard ARC evaluation considers only a single capability: producing the correct output grid for a test input. We argue that this narrow format fails to evaluate the diversity of abilities that genuine abstract skill acquisition should enable. We introduce PotARCin, a benchmark that extends ARC by assessing understanding of a task's underlying abstract rule across five dimensions: Definition, Classification, Constrained Generation, Editing, and Inversion. PotARCin employs programmatic methods to generate new task instances and transform given inputs for a given ARC task, enabling dynamic generative sampling beyond fixed input-output pairs. Across five state-of-the-art models evaluated on the ARC-AGI-1 training set, we observe a 25–52 percentage-point performance gap between standard ARC evaluation and evaluation on PotARCin, and find that multi-dimensional evaluation reorders models that standard accuracy ranks alike. We further investigate effects of generative sampling, difficulty of corruption types, and questions of self-consistency, showing that models frequently contradict their own formalized rule even where they have stated it correctly. We also introduce P-ARC, a held-out hand-crafted test set, on which models achieve 1–8% accuracy across all five dimensions, underscoring the importance of more holistic evaluations of abstract reasoning capabilities.
Five uses of the same rule
Each ARC task defines a small skill: inferring and using the rule behind its demonstrations. Standard ARC evaluation tests one use of it, transforming a test input. PotARCin presents the same demonstrations and asks for four more, all scored automatically with task-specific generator and verifier programs.
Definition
Write a Python program implementing the rule, checked on the demonstrations, the test input and hundreds of generated examples.
Classification
Decide for five candidate pairs whether each follows the rule; candidates mix valid examples with plausible corruptions.
Constrained generation
Produce a new valid input–output pair that differs from the demonstrations.
Editing
Repair a corrupted pair so that it follows the rule again, while staying close to it.
Inversion
Given an output grid, produce an input that the rule maps to it.
Key findings
We evaluated GPT-5.4, Gemini 3.1 Pro, Claude Opus 4.6, Kimi K2.5 and MiniMax M2.5 on the 400 ARC-AGI-1 training tasks and on P-ARC, with two independent runs each. A task counts as fully solved only if all five dimensions are correct.
Output accuracy does not reflect rule competence
Definition and Classification are the hardest dimensions and Inversion the easiest. Multi-dimensional evaluation also separates models that output-grid accuracy ranks alike: Kimi K2.5 has the lowest output-grid accuracy of the five models, yet nearly doubles MiniMax M2.5's full-task accuracy (38.6% vs. 20.6%). On P-ARC, every level drops sharply.
Models contradict rules they have stated correctly
Because Definition yields an executable program, we can check the model's other answers against its own program. Where that program is correct and the other answer is right, the two agree 99.1% of the time on ARC-AGI-1. Where the other answer is wrong, agreement falls to 21.1% (17.1% on P-ARC): most mistakes are not consequences of a wrong rule, but failures to apply a rule the model has already formalized.
Single samples overestimate robustness
For 20 tasks that GPT-5.4 initially solved in all five dimensions, we sampled 10 new instances per dimension. On 15 of them it fails at least one instance, mostly within the first two additional samples. Probing five different dimensions also exposes more failures than repeating one dimension five times (25.5% vs. 20.2% of trials).
P-ARC
P-ARC is a held-out set of 50 hand-crafted ARC-style tasks, which we estimate to lie between ARC-AGI-1 and ARC-AGI-2 in difficulty. Each task comes with a generator and a verifier program, 50 fixed generated examples, three human-made corruptions (erroneous attempts from the feasibility check or errors designed by the task's creator), and a reviewed natural-language rule. The dataset is available on Hugging Face under the MIT license. Four of its tasks:
Rule (t10). Every cell has a 180-degree partner: the cell at the opposite position through the grid's centre. For each cell, keep its own colour if it is non-black, otherwise take the colour of its partner. In these puzzles at most one of any partner pair is non-black to begin with, so each pair ends up filled with that single colour on both sides.
Show answer
Rule (t2). Find the 8-connected objects of non-black cells; their leftmost cells all sit in different columns. Order the objects by that leftmost cell (by column). Rotate the colours one step along that order with wrap-around: each object takes the colour of the object before it, and the first takes the colour of the last. Shapes and positions are unchanged.
Show answer
Rule (t9). Several coloured pieces lie scattered on the grid. Exactly one subset of them fits together, by translation only, into a solid rectangle with no gaps and no overlaps. Output that assembled rectangle, each piece keeping its own colour. Where more than one subset could work it is the one covering the largest area; pieces not needed for the assembly are left out.
Show answer
Rule (t47). Two equal halves are separated by a full grey line, and both share one background colour. Each half carries a single other colour as its foreground. Compare the halves position by position: where only the first half has foreground, the output takes the first half's colour; where only the second does, it takes the second half's colour; where both have foreground, or neither does, the output is black.
Show answer
BibTeX
@inproceedings{beger2026potarcin,
title = {PotARCin: Multi-Dimensional Evaluation of Skill Acquisition in Abstract Reasoning Tasks},
author = {Beger, Claas and Yi, Ryan and Mitchell, Melanie},
booktitle = {Advances in Neural Information Processing Systems},
year = {2026},
eprint = {2609.27288},
archivePrefix = {arXiv}
}