Claas Beger
Research Scientist, Stealth Startup
I am currently a Research Scientist at a stealth startup in New York City. Previously, I was a Graduate Fellow at the Santa Fe Institute, where I worked on multimodal reasoning with Professor Melanie Mitchell. Prior to that I finished my Master’s in Computer Science at Cornell University, where I worked with Kevin Ellis, Kilian Weinberger and Saikat Dutta. My research interest centers broadly around Human-like Artificial Intelligence. For this purpose, I think it is the most promising direction to look towards Natural Intelligence, both with regard to the brain and psychology/cognition.
news
| Jul 01, 2026 | Started as a Research Scientist at a stealth startup in New York City. |
|---|---|
| Apr 06, 2026 | “Bongards at the Boundary of Perception and Reasoning” accepted as full paper at CogSci 2026. “Cognitive Science-Inspired Evaluation of Large Language Models” accepted as Symposium at CogSci 2026. |
| Apr 06, 2026 | OmniCode is accepted to ACL 2026 Findings! |
| Mar 28, 2026 | I was awarded a fellowship by Princeton University’s Natural and Artificial Minds (NAM) Initiative |
| Feb 03, 2026 | New paper on using VLMs for visual concept hypothesis formation on Bongard problems out on arxiv! arXiv:2602.03038 |
selected publications
-
In Proceedings of the Annual Meeting of the Cognitive Science Society (CogSci 2026), 2026Full paperTLDR: Through a Bayesian approach to natural language rules and program synthesis, models can approach human performance on visual pattern recognition puzzles (Bongards). -
In Findings of the Association for Computational Linguistics: ACL 2026, 2026TLDR: A multilingual programming benchmark for LLM-based software engineering agents consisting of bug-fixing, test-generation, style-fixing and addressing code reviews. -
In NeurIPS 2025 Workshop on Interpreting Cognition in Deep Learning Models (CogInterp), 2025TLDR: An architecture of schema learning and iterative application can resolve arbitrarily deep compositional statements. -
2025Also known as Memento: Note-Taking for Your Future SelfTLDR: Decomposing multi-hop questions into single-step Prolog definitions improves performance on various long-context question datasets. -
In 2025 IEEE/ACM International Workshop on Large Language Models for Code (LLM4Code), 2025TLDR: Various language models struggle with dry-execution of simple and advanced code structures (Recursion, Concurrency OOP) -
2025TLDR: Using improved clustering and a more diverse embedding approach our technique can more accurately compress preference datasets into human-readable constitutions -
-
2025TLDR: Vision-Language models have good performance on abstract reasoning tasks, but do not utilize the intended human-core knowledge priors.