Ax Seyed Mohammad Asghari, Chris Chute, Vikranth Dwaracherla, Xiuyuan Lu, Mehdi Jafarnia, Victor Minden, Zheng Wen, Benjamin Van Roy 3/19/2026

Efficient Exploration at Scale

Online learning algorithm improving data efficiency of RLHF by incrementally updating reward and language models during preference learning.

Ax Siqi Pei, Liang Tang, Tiaonan Duan, Long Chen, Shuxian Li, Kaer Huang, Yanzhe Jing, Yiqiang Yan, Bo Zhang, Chenghao Jiang, Borui Zhang, Jiwen Lu 3/19/2026

AdaZoom-GUI: Adaptive Zoom-based GUI Grounding with Instruction Refinement

AdaZoom-GUI improves vision-language models for GUI grounding by using adaptive zoom to handle high-resolution screenshots and ambiguous instructions for UI automation.

Ax Rui Xiao, Sanghwan Kim, Yongqin Xian, Zeynep Akata, Stephan Alaniz 3/19/2026

FINER: MLLMs Hallucinate under Fine-grained Negative Queries

Benchmark and analysis of hallucinations in multimodal LLMs with fine-grained negative queries covering multi-object, multi-attribute, and multi-relation scenarios.

Ax Yihong Chen, Quanming Yao 3/19/2026

Attention Sinks Induce Gradient Sinks

Analysis showing attention sinks in transformers induce gradient concentration during backpropagation under causal masking, affecting training dynamics.

Ax Luca Hinkamp, Simon Kl\"uttermann, Emmanuel M\"uller 3/19/2026

RangeAD: Fast On-Model Anomaly Detection

On-model anomaly detection method that leverages primary model representations to detect distributional shifts without separate AD models.

Ax Lintang Sutawika, Aditya Bharat Soni, Bharath Sriraam R R, Apurva Gandhi, Taha Yassine, Sanidhya Vijayvargiya, Yuchen Li, Xuhui Zhou, Yilin Zhang, Leander Melroy Maben, Graham Neubig 3/19/2026

CodeScout: An Effective Recipe for Reinforcement Learning of Code Search Agents

Reinforcement learning method for training code search agents to localize relevant files, classes, and functions in large repositories as prerequisite for coding tasks.

Ax Dharshan Kumaran, Arthur Conmy, Federico Barbero, Simon Osindero, Viorica Patraucean, Petar Velickovic 3/19/2026

How do LLMs Compute Verbal Confidence

arXiv paper proposing GeCO, time-unconditional flow matching framework for adaptive robotic control using diffusion models.

Ax Alexander D. Goldie, Zilin Wang, Adrian Hayler, Deepak Nathani, Edan Toledo, Ken Thampiratwong, Aleksandra Kalisz, Michael Beukman, Alistair Letcher, Shashank Reddy, Clarisse Wibault, Theo Wolf, Charles O'Neill, Uljad Berdica, Nicholas Roberts, Saeed Rahmani, Hannah Erlebach, Roberta Raileanu, Shimon Whiteson, Jakob N. Foerster 3/19/2026

Procedural Generation of Algorithm Discovery Tasks in Machine Learning

arXiv paper investigating how LLMs compute verbal confidence scores and whether they're generated just-in-time or cached during inference.