Publications
Publications generated by jekyll-scholar
2026
- Cognitive Science-Inspired Evaluation of Large Language ModelsIn Proceedings of the Annual Meeting of the Cognitive Science Society (CogSci 2026), 2026SymposiumTLDR: Symposium on evaluating language models with methods and principles from cognitive science.
-
In Proceedings of the Annual Meeting of the Cognitive Science Society (CogSci 2026), 2026Full paperTLDR: Through a Bayesian approach to natural language rules and program synthesis, models can approach human performance on visual pattern recognition puzzles (Bongards). -
In Findings of the Association for Computational Linguistics: ACL 2026, 2026TLDR: A multilingual programming benchmark for LLM-based software engineering agents consisting of bug-fixing, test-generation, style-fixing and addressing code reviews.
2025
-
In NeurIPS 2025 Workshop on Interpreting Cognition in Deep Learning Models (CogInterp), 2025TLDR: An architecture of schema learning and iterative application can resolve arbitrarily deep compositional statements. - In NeurIPS 2025 Workshop on Regulatable Machine Learning (RegML), 2025TLDR: We refine Inverse Constitutional AI to extract interpretable principles from preference datasets for EU AI Act Article 10 compliance.
-
2025Also known as Memento: Note-Taking for Your Future SelfTLDR: Decomposing multi-hop questions into single-step Prolog definitions improves performance on various long-context question datasets. -
In 2025 IEEE/ACM International Workshop on Large Language Models for Code (LLM4Code), 2025TLDR: Various language models struggle with dry-execution of simple and advanced code structures (Recursion, Concurrency OOP) -
2025TLDR: Using improved clustering and a more diverse embedding approach our technique can more accurately compress preference datasets into human-readable constitutions -
-
2025TLDR: Vision-Language models have good performance on abstract reasoning tasks, but do not utilize the intended human-core knowledge priors. -
2025TLDR: Models are commonly aligned on human preferences, but that does not mean they follow normative goals in their favor. Beyond that, new alignment techniques may be required to enable this. -
2025TLDR: The o3 model uses diverse shortcuts or heuristics to solve ARC tasks. Tool-usage can help output accuracy in visual modality, whereas increased reasoning effort helps for text.