Ax Borui Zhang, Bo Zhang, Bo Wang, Wenzhao Zheng, Yuhao Cheng, Liang Tang, Yiqiang Yan, Jie Zhou, Jiwen Lu 5/8/2026

BAMI: Training-Free Bias Mitigation in GUI Grounding

Training-free bias mitigation method for GUI grounding in agents, improving performance on complex screen interaction tasks.

Ax Yidan Sun, Viktor Schlegel, Srinivasan Nandakumar, Iqra Zahid, Yuping Wu, Yulong Wu, Hao Li, Jie Zhang, Warren Del-Pinto, Goran Nenadic, Siew Kei Lam, Anil Anthony Bharath 5/8/2026

SynBench: A Benchmark for Differentially Private Text Generation

Benchmark for evaluating differentially private text generation methods using LLMs, enabling secure sharing of sensitive datasets across institutions.

Ax Sihan Hu, Xiansheng Cai, Yuan Huang, Zhiyuan Yao, Linfeng Zhang, Pan Zhang, Youjin Deng, Kun Chen 5/8/2026

Emergent Slow Thinking in LLMs as Inverse Tree Freezing

Research on how reinforcement learning enables LLMs to develop multi-step reasoning capabilities, using statistical physics framework to explain emergence of slow thinking.

Ax Martin Odersky, Yaoyu Zhao, Yichen Xu, Oliver Bra\v{c}evac, Cao Nguyen Pham 5/8/2026

Tracking Capabilities for Safer Agents

Safety harness using capability-safe Scala 3 language to restrict agent tool calls and prevent information leakage, unintended side effects, and prompt injection attacks.

Ax Bowen Ye, Rang Li, Qibin Yang, Yuanxin Liu, Linli Yao, Hanglong Lv, Zhihui Xie, Chenxin An, Lei Li, Lingpeng Kong, Qi Liu, Zhifang Sui, Tong Yang 5/8/2026

Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents

Claw-Eval benchmark suite with 300 human-verified tasks across 9 categories for trustworthy evaluation of autonomous LLM agents in real-world software environments.

Ax Wentao Zhang, Zhe Zhao, Haibin Wen, Yingcheng Wu, Cankun Guo, Ming Yin, Bo An, Mengdi Wang 5/8/2026

Autogenesis: A Self-Evolving Agent Protocol

Autogenesis Protocol for self-evolving LLM-based agent systems, addressing lifecycle management, version tracking, and safe updates to enable modular agent composition.

Ax Theodore Papamarkou, Pierre Alquier, Matthias Bauer, Wray Buntine, Andrew Davison, Gintare Karolina Dziugaite, Maurizio Filippone, Andrew Y. K. Foong, Vincent Fortuin, Dimitris Fouskakis, Jes Frellsen, Eyke H\"ullermeier, Theofanis Karaletsos, Mohammad Emtiyaz Khan, Nikita Kotelevskii, Salem Lahlou, Yingzhen Li, Fang Liu, Clare Lyle, Thomas M\"ollenhoff, Konstantina Palla, Maxim Panov, Yusuf Sale, Kajetan Schweighofer, Artem Shelmanov, Siddharth Swaroop, Martin Trapp, Willem Waegeman, Andrew Gordon Wilson, Alexey Zaytsev 5/8/2026

Position: agentic AI orchestration should be Bayes-consistent

Position paper arguing agentic AI control layers should be Bayes-consistent for tool/expert selection under uncertainty.