Ax Jinyuan Li, Langlin Huang, Chengsong Huang, Shaoyang Xu, Donghong Cai, Yuyi Yang, Wenxuan Zhang, Jiaxin Huang 5/18/2026

Process Rewards with Learned Reliability

Distributional process reward model predicting step-level success probability and reliability for reasoning tasks.

Ax Fengfei Yu, Ruijia Niu, Dongxia Wu, Yian Ma, Rose Yu 5/18/2026

Calibrating LLMs with Semantic-level Reward

Calibration method for LLMs using semantic-level rewards to improve uncertainty estimation in high-stakes tasks.

Ax Nils Feldhus, Tanja Baeumel, Elena Golimblevskaia, Qianli Wang, Van Bach Nguyen, Aaron Louis Eidt, Christopher Ebert, Wojciech Samek, Jing Yang, Vera Schmitt, Sebastian M\"oller, Simon Ostermann 5/18/2026

Judge Circuits

Causal analysis of format inconsistencies in LLM-as-judge scoring using PEAP to investigate internal mechanisms.

Ax Augusto B. Corr\^ea, Andr\'e G. Pereira, Jendrik Seipp 5/18/2026

Property-Guided LLM Program Synthesis for Planning

Uses property-guided synthesis to reduce LLM inference costs in program synthesis for planning, guiding generation with formal properties.

Ax Nathan Roll, Jill Kries, Laura Gwilliams, Cory Shain 5/18/2026

Artificial Aphasias in Lesioned Language Models

Technique to characterize language model organization by lesioning parameters, inspired by neuroscience aphasia studies.

Ax Yuantu Zhu, Zheyan Li, Dai Shi, Luke Thompson, Oliver Nash, Jose Miguel Lara Rangel, Siran Li, Bingguang Chen, Rongchan Zhu, Qi Meng, Hao Ni 5/18/2026

SPDEBench: An Extensive Benchmark for Learning Stochastic PDEs

Comprehensive benchmark for learning surrogate models of stochastic PDEs with complex spatio-temporal dynamics.

Ax Prabhat Nagarajan, Martha White, Marlos C. Machado 5/18/2026

Deep Double Q-learning

Extends double Q-learning to deep RL, improving target bootstrap decoupling in value function estimation.

Ax Parth Asawa, Alan Zhu, Abigail O'Neill, Matei Zaharia, Alexandros G. Dimakis, Joseph E. Gonzalez 5/18/2026

How to Train Your Advisor: Steering Black-Box LLMs with Advisor Models

Method to train small open-weight advisor models that generate dynamic prompts to improve black-box LLM performance. Demonstrates 27.4% improvement on GPT-5.2 tax tasks.

Ax Jamison Meindl, Yunsheng Tian, Tony Cui, Veronika Thost, Zhang-Wei Hong, Jie Chen, Wojciech Matusik, Mina Konakovi\'c Lukovi\'c 5/18/2026

SemanticOpt: Towards LLM-Based Semantic Black-Box Optimization

LLM-based semantic optimization approach for expensive black-box problems incorporating domain knowledge and heuristics.