Claas Beger

Research Scientist, Stealth Startup

prof_pic.jpg

I am currently a Research Scientist at a stealth startup in New York City. Previously, I was a Graduate Fellow at the Santa Fe Institute, where I worked on multimodal reasoning with Professor Melanie Mitchell. Prior to that I finished my Master’s in Computer Science at Cornell University, where I worked with Kevin Ellis, Kilian Weinberger and Saikat Dutta. My research interest centers broadly around Human-like Artificial Intelligence. For this purpose, I think it is the most promising direction to look towards Natural Intelligence, both with regard to the brain and psychology/cognition.

news

Jul 01, 2026 Started as a Research Scientist at a stealth startup in New York City.
Apr 06, 2026 “Bongards at the Boundary of Perception and Reasoning” accepted as full paper at CogSci 2026. “Cognitive Science-Inspired Evaluation of Large Language Models” accepted as Symposium at CogSci 2026.
Apr 06, 2026 OmniCode is accepted to ACL 2026 Findings!
Mar 28, 2026 I was awarded a fellowship by Princeton University’s Natural and Artificial Minds (NAM) Initiative
Feb 03, 2026 New paper on using VLMs for visual concept hypothesis formation on Bongard problems out on arxiv! arXiv:2602.03038

selected publications

  1. Bongards.png
    Cassidy Langenfeld, Claas Beger, Gloria Geng, Wasu Top Piriyakulkij, Keya Hu, Yewen Pu, and Kevin Ellis
    In Proceedings of the Annual Meeting of the Cognitive Science Society (CogSci 2026), 2026
    Full paper
    TLDR: Through a Bayesian approach to natural language rules and program synthesis, models can approach human performance on visual pattern recognition puzzles (Bongards).
  2. OmniCode.png
    Atharv Sonwane, Eng-Shen Tu, Wei-Chung Lu, Claas Beger, Carter Larsen, Debjit Dhar, Simon Alford, Rachel Chen, Ronit Pattanayak, Tuan Anh Dang, Guohao Chen, Gloria Geng, Kevin Ellis, and Saikat Dutta
    In Findings of the Association for Computational Linguistics: ACL 2026, 2026
    TLDR: A multilingual programming benchmark for LLM-based software engineering agents consisting of bug-fixing, test-generation, style-fixing and addressing code reviews.
  3. Figure_2_MIRAGE-1.png
    Alex Noviello*Claas Beger*, Jacob Groner, Kevin Ellis, and Weinan Sun
    In NeurIPS 2025 Workshop on Interpreting Cognition in Deep Learning Models (CogInterp), 2025
    TLDR: An architecture of schema learning and iterative application can resolve arbitrarily deep compositional statements.
  4. Memento.jpg
    Chao Wan, Albert Gong, Mihir Mishra, Carl-Leander Henneking, Claas Beger, and Kilian Q. Weinberger
    2025
    Also known as Memento: Note-Taking for Your Future Self
    TLDR: Decomposing multi-hop questions into single-step Prolog definitions improves performance on various long-context question datasets.
  5. Error_Vis.jpg
    Claas Beger, and Saikat Dutta
    In 2025 IEEE/ACM International Workshop on Large Language Models for Code (LLM4Code), 2025
    TLDR: Various language models struggle with dry-execution of simple and advanced code structures (Recursion, Concurrency OOP)
  6. Content_Style_Sentiment_Clustering.jpg
    Carl-Leander Henneking*, and Claas Beger*
    2025
    TLDR: Using improved clustering and a more diverse embedding approach our technique can more accurately compress preference datasets into human-readable constitutions
  7. citegeist-pipeline.jpg
    Award ribbon
    Claas Beger*, and Carl-Leander Henneking*
    2025
    TLDR: Through a multi-step retrieval and summarization pipeline with three definable properties Citegeist can synthesize related work for a given scientific paper
  8. Shortcut_Overview.png
    Claas Beger, Ryan Yi, Shuhao Fu, Arseny Moskvichev, Sarah W. Tsai, Sivasankaran Rajamanickam, and Melanie Mitchell
    2025
    TLDR: Vision-Language models have good performance on abstract reasoning tasks, but do not utilize the intended human-core knowledge priors.