Ax Yaolun Zhang, Ruohui Wang, Jiahao Wang, Yepeng Tang, Xuanyu Zheng, Haonan Duan, Hao Lu, Hanming Deng, Lewei Lu 3/25/2026

EVA: Efficient Reinforcement Learning for End-to-End Video Agent

EVA: Reinforcement learning method for video understanding agents using multimodal LLMs. Adaptive frame sampling and reasoning without manual workflows.

Ax Miao Yu, Siyuan Fu, Moayad Aloqaily, Zhenhong Zhou, Safa Otoum, Xing fan, Kun Wang, Yufei Guo, Qingsong Wen 3/25/2026

SafeSeek: Universal Attribution of Safety Circuits in Language Models

SafeSeek framework for universal attribution of safety circuits in LLMs using mechanistic interpretability to understand alignment, jailbreak, and backdoor behaviors.

Ax Shaid Hasan, Breenice Lee, Sujan Sarker, Tariq Iqbal 3/25/2026

A Multimodal Framework for Human-Multi-Agent Interaction

Multimodal framework for human-multi-agent interaction integrating perception, embodied expression, and coordinated decision-making in shared physical spaces.

Ax Mehmet Caner, Agostino Capponi, Nathan Sun, Jonathan Y. Tan 3/25/2026

Designing Agentic AI-Based Screening for Portfolio Investment

Agentic AI platform for portfolio investment screening using LLM agents for fundamental analysis and sentiment analysis with deliberation mechanism for buy/sell signals.

Ax Yuntong Zhang, Zhiyuan Pan, Imam Nur Bani Yusuf, Haifeng Ruan, Ridwan Shariffdeen, Abhik Roychoudhury 3/25/2026

Code Review Agent Benchmark

Introduces benchmark dataset and evaluation framework for code review agents, addressing code quality assurance as AI-generated code scales.

Ax Haoran Yuan, Weigang Yi, Zhenyu Zhang, Wendi Chen, Yuchen Mo, Jiashi Yin, Xinzhuo Li, Xiangyu Zeng, Chuan Wen, Cewu Lu, Katherine Driggs-Campbell, Ismini Lourentzou 3/25/2026

VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs

Proposes VTAM, extending video-action models for embodied AI with tactile sensing for contact-rich physical interactions beyond vision-only approaches.

Ax Ufaq Khan, Umair Nawaz, L D M S S Teja, Numaan Saeed, Muhammad Bilal, Yutong Xie, Mohammad Yaqub, Muhammad Haris Khan 3/25/2026

MedObvious: Exposing the Medical Moravec's Paradox in VLMs via Clinical Triage

Evaluates Vision Language Models' ability to perform pre-diagnostic sanity checks in medical imaging, identifying gaps between fluent text generation and safe visual understanding.

Ax Nan Huo, Xiaohan Xu, Jinyang Li, Per Jacobsson, Shipei Lin, Bowen Qin, Binyuan Hui, Xiaolong Li, Ge Qu, Shuzheng Si, Linheng Han, Edward Alexander, Xintong Zhu, Rui Qin, Ruihan Yu, Yiyao Jin, Feige Zhou, Weihao Zhong, Yun Chen, Hongyu Liu, Chenhao Ma, Fatma Ozcan, Yannis Papakonstantinou, Reynold Cheng 3/25/2026

BIRD-INTERACT: Re-imagining Text-to-SQL Evaluation for Large Language Models via Lens of Dynamic Interactions

BIRD-INTERACT benchmark evaluating LLMs on multi-turn text-to-SQL tasks with dynamic interactions and error handling.

Ax Raj Ghugare, Roger Creus Castanyer, Catherine Ji, Kathryn Wantlin, Jin Schofield, Karthik Narasimhan, Benjamin Eysenbach 3/25/2026

BuilderBench: The Building Blocks of Intelligent Agents

BuilderBench benchmark for evaluating AI agents' ability to learn through exploration and interaction beyond training data patterns.