Ax Negar Arabzadeh, Wenjie Ma, Sewon Min, Matei Zaharia 5/7/2026

RAG over Thinking Traces Can Improve Reasoning Tasks

RAG applied to reasoning tasks by retrieving intermediate thinking traces instead of documents, improving math and code generation performance.

Ax Doojin Baek, Gyubin Lee, Junyeob Baek, Hosung Lee, Sungjin Ahn 5/7/2026

Learning to Theorize the World from Observation

Theory-building world models inspired by developmental cognition: learning internal representations of world dynamics from observation.

Ax Thanh Dat Hoang, Thanh Trung Huynh, Matthias Weidlich, Thanh Tam Nguyen, Tong Chen, Hongzhi Yin, Quoc Viet Hung Nguyen 5/7/2026

FINER-SQL: Boosting Small Language Models for Text-to-SQL

FINER-SQL boosts small language models for text-to-SQL tasks, enabling efficient on-premise deployment with improved reasoning and instruction following.

Ax John Yang, Kilian Lieret, Jeffrey Ma, Parth Thakkar, Dmitrii Pedchenko, Sten Sootla, Emily McMilin, Pengcheng Yin, Rui Hou, Gabriel Synnaeve, Diyi Yang, Ofir Press 5/7/2026

ProgramBench: Can Language Models Rebuild Programs From Scratch?

Benchmark testing language model agents' ability to build complete software projects from scratch with minimal human oversight.

Ax Sebastian Wind, Tri-Thien Nguyen, Jeta Sopa, Mahshad Lotfinia, Sebastian Bickelhaup, Michael Uder, Harald K\"ostler, Gerhard Wellein, Sven Nebelung, Daniel Truhn, Andreas Maier, Soroosh Tayebi Arasteh 5/7/2026

Safety and accuracy follow different scaling laws in clinical large language models

SaFE-Scale framework showing clinical LLM safety and accuracy follow different scaling laws; safety doesn't necessarily improve with model size or context.

Ax Ivaxi Sheth, Jan Wehner, Sahar Abdelnabi, Ruta Binkyte, Mario Fritz 5/7/2026

Safety Must Precede the Deployment of Open-Ended AI

Position paper on safety requirements for open-ended AI agents with autonomous behavior generation and self-evolution capabilities before deployment.