Claas Beger

Research Scientist, Stealth Startup

prof_pic.jpg

I am a Research Scientist at a stealth startup in New York City. Previously, I was a Graduate Fellow at the Santa Fe Institute, working on multimodal abstract reasoning with Professor Melanie Mitchell. Before that, I completed my Master’s in Computer Science at Cornell University, where I worked with Kevin Ellis, Kilian Weinberger, and Saikat Dutta. My research focuses on human-like artificial intelligence. I believe the most promising path there is to learn from natural intelligence, drawing on both neuroscience and cognitive psychology.

News

Sep 25, 2026 Two papers are accepted to NeurIPS 2026: “Distinguishing Performance From Competence in Evaluations of Humanlike Abstract Reasoning” and “PotARCin: Multi-Dimensional Evaluation of Skill Acquisition in Abstract Reasoning Tasks”!
Sep 14, 2026 π-CoT is accepted at TMLR!
Jul 01, 2026 Joined a Stealth Startup as a Research Scientist in New York.
Apr 06, 2026 “Bongards at the Boundary of Perception and Reasoning” accepted as full paper at CogSci 2026. “Cognitive Science-Inspired Evaluation of Large Language Models” accepted as Symposium at CogSci 2026.
Apr 06, 2026 OmniCode is accepted to ACL 2026 Findings!

Selected Publications

  1. PotARCin_Overview.png
    Claas Beger, Ryan Yi, and Melanie Mitchell
    In Advances in Neural Information Processing Systems (NeurIPS 2026), 2026
    TLDR: Extending ARC to five dimensions of rule understanding (definition, classification, constrained generation, editing, inversion) reveals a 25-52 point gap relative to standard ARC evaluation.
  2. Shortcut_Overview.png
    Claas Beger, Ryan Yi, Shuhao Fu, Arseny Moskvichev, Sarah W. Tsai, Sivasankaran Rajamanickam, and Melanie Mitchell
    In Advances in Neural Information Processing Systems (NeurIPS 2026), 2026
    TLDR: Vision-language models can perform well on abstract reasoning tasks without using the intended human-like core knowledge priors, so accuracy alone overstates their competence.
  3. Memento.jpg
    Chao Wan, Albert Gong, Mihir Mishra, Carl-Leander Henneking, Claas Beger, and Kilian Q. Weinberger
    Transactions on Machine Learning Research (TMLR), 2026
    Also known as Memento: Note-Taking for Your Future Self
    TLDR: Decomposing multi-hop questions into single-step Prolog definitions improves performance on various long-context question datasets.
  4. Bongards.png
    Cassidy Langenfeld, Claas Beger, Gloria Geng, Wasu Top Piriyakulkij, Keya Hu, Yewen Pu, and Kevin Ellis
    In Proceedings of the Annual Meeting of the Cognitive Science Society (CogSci 2026), 2026
    Full paper
    TLDR: Through a Bayesian approach combining natural language rules and program synthesis, models can approach human performance on visual pattern recognition puzzles (Bongard problems).
  5. OmniCode.png
    Atharv Sonwane*, Eng-Shen Tu*, Wei-Chung Lu*, Claas Beger*, Carter Larsen, Debjit Dhar, Simon Alford, Rachel Chen, Ronit Pattanayak, Tuan Anh Dang, Guohao Chen, Gloria Geng, Kevin Ellis, and Saikat Dutta
    In Findings of the Association for Computational Linguistics: ACL 2026, 2026
    TLDR: A multilingual benchmark for LLM-based software engineering agents covering bug fixing, test generation, style fixing, and addressing code reviews.
  6. Chart.png
    Award ribbon
    Claas Beger, Shuhao Fu, Ryan Yi, Arseny Moskvichev, and Melanie Mitchell
    In NeurIPS 2025 Workshop on Multimodal Algorithmic Reasoning (MAR), 2025
    Spotlight
    TLDR: The o3 model relies on diverse shortcuts and heuristics to solve ARC tasks. Tool use improves accuracy in the visual modality, while increased reasoning effort helps in the text modality.