Ax Jehyeok Yeon, Isha Chaudhary, Gagandeep Singh 5/14/2026

Quantitative Certification of Agentic Tool Selection

Framework for certifying tool selection in LLM-based agentic systems, evaluating robustness against adversarial tool pools and deployment scenarios.

Ax Luca Belli, Kate H. Bentley, Will Alexander, Emily Ward, Matt Hawrilenko, Kelly Johnston, Mill Brown, Adam M. Chekroud 5/14/2026

VERA-MH Concept Paper

Automated safety evaluation framework for mental health AI chatbots using clinician-informed rubric and multi-agent validation approach.

Ax Dario Shariatian, Alain Durmus, Umut Simsekli, Stefano Peluchetti 5/14/2026

Latent-Augmented Discrete Diffusion Models

Latent-augmented discrete diffusion model with auxiliary channel for improved few-step language generation and cross-token dependencies.

Ax John Yang, Kilian Lieret, Joyce Yang, Carlos E. Jimenez, Muhtasham Oblokulov, Aryan Siddiqui, Ofir Press, Ludwig Schmidt, Diyi Yang 5/14/2026

CodeClash: Benchmarking Goal-Oriented Software Engineering

Benchmark evaluating whether LLMs can iteratively develop code toward high-level goals beyond isolated task completion.

Ax Yifan Zhang, Yifeng Liu, Mengdi Wang, Quanquan Gu 5/14/2026

Deep Delta Learning

Deep Delta Learning introduces residual update rule for transformer layers enabling selective content rewriting in deep networks.

Ax Or Ordentlich, Yury Polyanskiy 5/14/2026

High-Rate Quantized Matrix Multiplication I

Research on quantized matrix multiplication optimization for efficient LLM deployment with weight and activation quantization.

Ax Tomas Ruiz, Zhen Qin, Yifan Zhang, Xuyang Shen, Yiran Zhong, Mengdi Wang 5/14/2026

FlashSampling: Fast and Memory-Efficient Exact Sampling

Optimized sampling primitive that fuses categorical sampling into LM-head computation, reducing memory traffic for large-vocabulary decoding.

Ax Kai-Wei Chang, Wei-Chih Chen, En-Pei Hu, Hung-yi Lee, James Glass 5/14/2026

TiCo: Time-Controllable Spoken Dialogue Model

Spoken dialogue model with controllable response duration for voice assistants and interactive agents.

Ax Vladimir Zaigrajew, Micha{\l} Piechota, Gaspar Sekula, Pawe{\l} Gelar, Przemys{\l}aw Biecek 5/14/2026

LINE: LLM-based Iterative Neuron Explanations for Vision Models

LINE: Training-free iterative method using LLMs to explain individual neurons in vision models, improving interpretability beyond predefined concept vocabularies.

Ax Priyal Deep, Shane Emmons, Amy Fox, Kyle Bacon, Kelley McAllister, Peter Ortiz, Krisztian Flautner 5/14/2026

Evaluation of Prompt Injection Defenses in Large Language Models

Evaluation of prompt injection defenses across 20,000 adaptive attacks on nine defense configurations; only output filtering remained unbroken.