Claas Beger
Research Scientist, Stealth Startup
I am a Research Scientist at a stealth startup in New York City. Previously, I was a Graduate Fellow at the Santa Fe Institute, working on multimodal abstract reasoning with Professor Melanie Mitchell. Before that, I completed my Master’s in Computer Science at Cornell University, where I worked with Kevin Ellis, Kilian Weinberger, and Saikat Dutta. My research focuses on human-like artificial intelligence. I believe the most promising path there is to learn from natural intelligence, drawing on both neuroscience and cognitive psychology.
News
| Sep 25, 2026 | Two papers are accepted to NeurIPS 2026: “Distinguishing Performance From Competence in Evaluations of Humanlike Abstract Reasoning” and “PotARCin: Multi-Dimensional Evaluation of Skill Acquisition in Abstract Reasoning Tasks”! |
|---|---|
| Sep 14, 2026 | π-CoT is accepted at TMLR! |
| Jul 01, 2026 | Joined a Stealth Startup as a Research Scientist in New York. |
| Apr 06, 2026 | “Bongards at the Boundary of Perception and Reasoning” accepted as full paper at CogSci 2026. “Cognitive Science-Inspired Evaluation of Large Language Models” accepted as Symposium at CogSci 2026. |
| Apr 06, 2026 | OmniCode is accepted to ACL 2026 Findings! |
Selected Publications
-
In Advances in Neural Information Processing Systems (NeurIPS 2026), 2026TLDR: Extending ARC to five dimensions of rule understanding (definition, classification, constrained generation, editing, inversion) reveals a 25-52 point gap relative to standard ARC evaluation. -
In Advances in Neural Information Processing Systems (NeurIPS 2026), 2026TLDR: Vision-language models can perform well on abstract reasoning tasks without using the intended human-like core knowledge priors, so accuracy alone overstates their competence. -
Transactions on Machine Learning Research (TMLR), 2026Also known as Memento: Note-Taking for Your Future SelfTLDR: Decomposing multi-hop questions into single-step Prolog definitions improves performance on various long-context question datasets. -
In Proceedings of the Annual Meeting of the Cognitive Science Society (CogSci 2026), 2026Full paperTLDR: Through a Bayesian approach combining natural language rules and program synthesis, models can approach human performance on visual pattern recognition puzzles (Bongard problems). -
In Findings of the Association for Computational Linguistics: ACL 2026, 2026TLDR: A multilingual benchmark for LLM-based software engineering agents covering bug fixing, test generation, style fixing, and addressing code reviews. -
In NeurIPS 2025 Workshop on Multimodal Algorithmic Reasoning (MAR), 2025SpotlightTLDR: The o3 model relies on diverse shortcuts and heuristics to solve ARC tasks. Tool use improves accuracy in the visual modality, while increased reasoning effort helps in the text modality.