Ax Hanna Foerster, Tom Blanchard, Kristina Nikoli\'c, Ilia Shumailov, Cheng Zhang, Robert Mullins, Nicolas Papernot, Florian Tram\`er, Yiren Zhao 3/10/2026

CaMeLs Can Use Computers Too: System-level Security for Computer Use Agents

Security architecture for Computer Use Agents using strict isolation between trusted task planning and untrusted environment observations to defend against prompt injection.

Ax Zixuan Huang, Xin Xia, Yuxi Ren, Jianbin Zheng, Xuefeng Xiao, Hongyan Xie, Li Huaqiu, Songshi Liang, Zhongxiang Dai, Fuzhen Zhuang, Jianxin Li, Yikun Ban, Deqing Wang 3/10/2026

Real-Time Aligned Reward Model beyond Semantics

Framework addressing reward overoptimization in RLHF by moving beyond semantic information to capture true human intent and prevent policy exploitation.

Ax Jihwan Oh, Murad Aghazada, Yooju Shin, Se-Young Yun, Taehyeon Kim 3/10/2026

MERIT Feedback Elicits Better Bargaining in LLM Negotiators

Framework and benchmark (AgoraBench) for improving LLM negotiation capabilities through utility-focused feedback, spanning nine complex bargaining scenarios.

Ax Xiangyi Li, Wenbo Chen, Yimin Liu, Shenghan Zheng, Xiaokun Chen, Yifeng He, Yubo Li, Bingran You, Haotian Shen, Jiankai Sun, Shuyi Wang, Binxu Li, Qunhong Zeng, Di Wang, Xuandong Zhao, Yuanli Wang, Roey Ben Chaim, Zonglin Di, Yipeng Gao, Junwei He, Yizhuo He, Liqiang Jing, Luyang Kong, Xin Lan, Jiachen Li, Songlin Li, Yijiang Li, Yueqian Lin, Xinyi Liu, Xuanqing Liu, Haoran Lyu, Ze Ma, Bowei Wang, Runhui Wang, Tianyu Wang, Wengao Ye, Yue Zhang, Hanwen Xing, Yiqi Xue, Steven Dillmann, Han-chung Lee 3/10/2026

SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

SkillsBench: benchmark with 86 tasks across 11 domains measuring effectiveness of agent skills with curated and self-generated skill evaluations.

Ax Lve Meng (University of Science,Technology of China, Zhongguancun Academy), Weilong Zhao (Universit\'e Paris Cit\'e), Yanzhi Zhang (Zhongguancun Academy), Haoxiang Guan (Zhongguancun Academy), Jiyan He (Zhongguancun Academy) 3/10/2026

Can a Lightweight Automated AI Pipeline Solve Research-Level Mathematical Problems?

Investigation of lightweight automated AI pipelines for research-level mathematics using LLMs beyond competition benchmarks to practical applications.

Ax Xiaoxuan Wang, Han Zhang, Haixin Wang, Yidan Shi, Ruoyan Li, Kaiqiao Han, Chenyi Tong, Haoran Deng, Renliang Sun, Alexander Taylor, Yanqiao Zhu, Jason Cong, Yizhou Sun, Wei Wang 3/10/2026

ARLArena: A Unified Framework for Stable Agentic Reinforcement Learning

ARLArena: unified framework for stable agentic reinforcement learning addressing training instability and collapse in multi-step interactive tasks.

Ax Maxwell A. Xu, Harish Haresamudram, Catherine W. Liu, Patrick Langer, Jathurshan Pradeepkumar, Wanting Mao, Sunita J. Ferns, Aradhana Verma, Jimeng Sun, Paul Schmiedmayer, Xin Liu, Daniel McDuff, Emily B. Fox, James M. Rehg 3/10/2026

How Well Do Multimodal Models Reason on ECG Signals?

Evaluation of multimodal LLM reasoning on ECG signals with focus on verifying validity of clinical reasoning traces beyond proxy metrics.

Ax Zhiyu Ni, Yifeng Xiao, Zheng Liang 3/10/2026

Agentified Assessment of Logical Reasoning Agents

Framework for benchmarking logical reasoning agents using an assessor agent to enforce budgets, parse outputs, and record structured failure types with reproducible evaluation.

Ax Reza Refaei Afshar, Joaquin Vanschoren, Uzay Kaymak, Rui Zhang, Yaoxin Wu, Wen Song, Yingqian Zhang 3/10/2026

Automated Reinforcement Learning: An Overview

Overview of automated reinforcement learning automating problem modeling, algorithm selection, and hyperparameter tuning.

Ax Yan Zhuang, Qi Liu, Haoyang Bi, Zhenya Huang, Weizhe Huang, Jiatong Li, Junhao Yu, Zirui Liu, Zirui Hu, Yuting Hong, Zachary A. Pardos, Haiping Ma, Mengxiao Zhu, Shijin Wang, Enhong Chen 3/10/2026

Survey of Computerized Adaptive Testing: A Machine Learning Perspective

Survey of computerized adaptive testing systems using machine learning for personalized and efficient assessment.

Ax Connor Douglas, Foster Provost, Arun Sundararajan 3/10/2026

The Illusion of Collusion

Research on whether competing multi-armed bandit agents exhibit collusion in repeated games, testing algorithmic behavior without explicit coordination.