Ax Jielin Qiu, Jianguo Zhang, Zixiang Chen, Liangwei Yang, Ming Zhu, Juntao Tan, Haolin Chen, Wenting Zhao, Rithesh Murthy, Roshan Ram, Akshara Prabhakar, Shelby Heinecke, Caiming, Xiong, Silvio Savarese, Huan Wang 3/2/2026

AudioCapBench: Quick Evaluation on Audio Captioning across Sound, Music, and Speech

AudioCapBench: benchmark for evaluating audio captioning of multimodal LLMs across sound, music, speech with 1,000 samples and LLM-as-Judge evaluation.

Ax Hariz Yet, Nguyen Thanh Tam, Mao V. Ngo, Lim Yi Shen, Lin Wei, Jihong Park, Binbin Chen, Tony Q. S. Quek 3/2/2026

SLA-Aware Distributed LLM Inference Across Device-RAN-Cloud

System design for distributed LLM inference across device, RAN-edge, and cloud tiers with latency constraints for 5G embodied AI applications.

Ax Dongxu Zhang, Yiding Sun, Pengcheng Li, Yumou Liu, Hongqiang Lin, Haoran Xu, Xiaoxuan Mu, Liang Lin, Wenbiao Yan, Ning Yang, Chaowei Fang, Juanjuan Zhao, Jihua Zhu, Conghui He, Cheng Tan 3/2/2026

PointCoT: A Multi-modal Benchmark for Explicit 3D Geometric Reasoning

Multimodal benchmark evaluating MLLMs on explicit 3D geometric reasoning with point clouds, exposing geometric hallucinations.

Ax Oscar Hill, Mateo Espinosa Zarlenga, Mateja Jamnik 3/2/2026

Hierarchical Concept-based Interpretable Models

Hierarchical concept embedding models improving neural network interpretability through human-readable concept representations.

Ax Daniel Yang, Samuel Stante, Florian Redhardt, Lena Libon, Parnian Kassraie, Ido Hakimi, Barna P\'asztor, Andreas Krause 3/2/2026

RewardUQ: A Unified Framework for Uncertainty-Aware Reward Models

RewardUQ: Framework for uncertainty quantification in reward models used to align LLMs with human preferences, reducing annotation costs.

Ax Haritz Puerto, Haonan Li, Xudong Han, Timothy Baldwin, Iryna Gurevych 3/2/2026

Controllable Reasoning Models Are Private Thinkers

Method for training reasoning models to follow instructions in reasoning traces to prevent unintended leakage of private information in AI agents processing sensitive user data.

Ax Ali Behrouz, Zeman Li, Yuan Deng, Peilin Zhong, Meisam Razaviyayn, Vahab Mirrokni 3/2/2026

Memory Caching: RNNs with Growing Memory

Exploration of recurrent architectures with growing memory as subquadratic alternatives to Transformers for sequence modeling.

Ax Weinan Dai, Hanlin Wu, Qiying Yu, Huan-ang Gao, Jiahao Li, Chengquan Jiang, Weiqiang Lou, Yufan Song, Hongli Yu, Jiaze Chen, Wei-Ying Ma, Ya-Qin Zhang, Jingjing Liu, Mingxuan Wang, Xin Liu, Hao Zhou 3/2/2026

CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation

CUDA Agent system using large-scale agentic RL to generate optimized GPU kernels, bridging gap between LLMs and compiler-based systems.

Ax Jenny Y. Huang, Leshem Choshen, Ramon Astudillo, Tamara Broderick, Jacob Andreas 3/2/2026

Do LLMs Benefit From Their Own Words?

Study comparing standard multi-turn prompting with user-turn-only prompting to determine if LLMs benefit from their own prior responses.