Ax Yi Yu, Guangquan Hu, Chenghuang Shen, Xingyan Liu, Jing Gu, Hangyi Sun, Junzhuo Ma, Weiting Liu, Jianfeng Liu, Mingyue Pu, Yu Wang, Zhengdong Xiao, Rui Xie, Longjiu Luo, Qianrong Wang, Gurong Cui, Honglin Qiao, Wenlian Lu 3/31/2026

CirrusBench: Evaluating LLM-based Agents Beyond Correctness in Real-World Cloud Service Environments

CirrusBench benchmark evaluates LLM-based agents in real-world cloud service environments beyond correctness, measuring robustness and efficiency in high-complexity customer-assistant interactions.

Ax Yash Savani, Branislav Kveton, Yuchen Liu, Yilin Wang, Jing Shi, Subhojyoti Mukherjee, Nikos Vlassis, Krishna Kumar Singh 3/31/2026

Stepwise Credit Assignment for GRPO on Flow-Matching Models

Stepwise credit assignment method for reinforcement learning on flow-matching generative models, differentiating early vs. late diffusion steps.

Ax Jinyi Han, Ying Huang, Ying Liao, Zishang Jiang, Xikun Lu, Haiquan Zhao, Xinyi Wang, Guanghao Zhou, Sihang Jiang, Jiaqing Liang, Weikang Zhou, Zeye Sun, Fei Yu, Yanghua Xiao 3/31/2026

Your Models Have Thought Enough: Training Large Reasoning Models to Stop Overthinking

Trains large reasoning models to stop reasoning early by detecting sufficient evidence accumulation, reducing computational costs while maintaining performance.

Ax Yuanqi Du, Botao Yu, Tianyu Liu, Tony Shen, Junwu Chen, Jan G. Rittig, Kunyang Sun, Yikun Zhang, Aarti Krishnan, Yu Zhang, Daniel Rosen, Rosali Pirone, Zhangde Song, Bo Zhou, Cassandra Masschelein, Yingze Wang, Haorui Wang, Haojun Jia, Chao Zhang, Hongyu Zhao, Martin Ester, Nir Hacohen, Teresa Head-Gordon, Carla P. Gomes, Huan Sun, Chenru Duan, Philippe Schwaller, Wengong Jin 3/31/2026

Accelerating Scientific Discovery with Autonomous Goal-evolving Agents

Scientific Autonomous Goal-evolving Agent automates objective function design to guide scientific discovery agents beyond manually-specified quantitative proxies.

Ax Mia Hopman, Jannes Elstner, Maria Avramidou, Amritanshu Prasad, David Lindner 3/31/2026

Evaluating and Understanding Scheming Propensity in LLM Agents

Evaluates when LLM agents exhibit scheming behavior by decomposing incentives into agent and environmental factors to understand misalignment risks in autonomous systems.

Ax Daattavya Aggarwal, Oisin Kim, Carl Henrik Ek, Challenger Mishra 3/31/2026

Discovering mathematical concepts through a multi-agent system

Multi-agent LLM system for automated mathematical discovery that generates conjectures, attempts proofs, and uses feedback to evolve a distribution of mathematical concepts.

Ax Jiang Liu, John Martabano Landy, Yao Xuan, Swamy Muddu, Nhat Le, Munaf Sahaf, Luc Kien Hang, Rupinder Khandpour, Kevin De Angeli, Chang Yang, Shouyuan Chen, Shiblee Sadik, Anirudh Agrawal, Djordje Gligorijevic, Jingzheng Qin, Peggy Yao, Alireza Vahdatpour 3/31/2026

Design Once, Deploy at Scale: Template-Driven ML Development for Large Model Ecosystems

Template-driven ML development approach for scaling ML model ecosystems in computational advertising platforms.

Ax Canfer Akbulut, Rasmi Elasmar, Abhishek Roy, Anthony Payne, Priyanka Suresh, Lujain Ibrahim, Seliem El-Sayed, Charvi Rastogi, Ashyana Kachra, Will Hawkins, Kristian Lum, Laura Weidinger 3/31/2026

Evaluating Language Models for Harmful Manipulation

Framework for evaluating harmful AI manipulation through human-AI interaction studies across policy, finance, and health domains.

Ax Qiao Yuan, Sheng-Uei Guan, Pin Ni, Tianlun Luo, Ka Lok Man, Prudence Wong, Victor Chang 3/31/2026

Continual Graph Learning: A Survey

Survey of continual graph learning methods for incrementally learning from streaming graph data without catastrophic forgetting.

Ax Weiwei Gu, Suresh Kondepudi, Anmol Gupta, Lixiao Huang, Nakul Gopalan 3/31/2026

Continual Robot Skill and Task Learning via Dialogue

Framework for robots to continually learn tasks and skills through dialogue interactions, maintaining a skill library with LLM support.