Ax Rick Chen, Joseph Ternasky, Afriyie Samuel Kwesi, Ben Griffin, Aaron Ontoyin Yin, Zakari Salifu, Kelvin Amoaba, Xianling Mu, Fuat Alican, Yigit Ihlamur 5/7/2026

VCBench: Benchmarking LLMs in Venture Capital

VCBench: First benchmark for predicting founder success in venture capital using LLMs, with sparse signals and uncertain outcomes exceeding market index performance.

Ax Wenyue Hua, Tianyi Peng, Chi Wang, Jiaxin Pei, Ian Kaufman, Bryan Lim, Chandler Fang 5/7/2026

Quantifying Trust: Financial Risk Management for Trustworthy AI Agents

Framework for quantifying trust in autonomous AI agents through end-to-end operational outcomes rather than model-internal properties, relevant to deployed agents with financial risk.

Ax Saad Alqithami 5/7/2026

Soft Tournament Equilibrium

Soft Tournament Equilibrium: Framework for evaluating non-transitive LLM-based agents using set-valued cores instead of linear rankings for cyclic competitive domains.

Ax Tu Trinh, Mohamed Elfeki, Guangze Luo, Kelvin Luu, Nathan Hunt, Ernesto Hernandez, Nandan Marwaha, Yannis Yiming He, Charles Wang, Fernando Carabedo, Alessa Castillo, Bing Liu 5/7/2026

HiL-Bench (Human-in-Loop Benchmark): Do Agents Know When to Ask for Help?

HiL-Bench evaluates whether coding agents know when to ask for help versus act autonomously on incomplete specifications, addressing judgment gaps in frontier agents.

Ax Mehul Bafna, Siddhant anand Jadhav, David Sweet 5/7/2026

Taking the GP Out of the Loop

Scaling Bayesian optimization to many observations by replacing Gaussian process surrogates with more scalable alternatives.

Ax Yi Ru Wang, Carter Ung, Christopher Tan, Grant Tannert, Jiafei Duan, Josephine Li, Anh Le, Rishabh Oswal, Markus Grotz, Wilbert Pumacay, Yuquan Deng, Ranjay Krishna, Dieter Fox, Siddhartha Srinivasa 5/7/2026

RoboEval: Where Robotic Manipulation Meets Structured and Scalable Evaluation

RoboEval: structured evaluation framework and benchmark for robotic manipulation with behavioral and outcome metrics beyond binary success.

Ax Friederike Groschupp, Daniele Lain, Aritra Dhar, Lara Magdalena Lazier, Srdjan \v{C}apkun 5/7/2026

Can LLMs Make (Personalized) Access Control Decisions?

Study evaluating LLM capability for personalized access control decisions in applications and agent-based systems to reduce user cognitive burden.

Ax Ignacio Heredia, \'Alvaro L\'opez Garc\'ia, Fernando Aguilar G\'omez, Diego Aguirre, Caterina Alarc\'on Mar\'in, Khadijeh Alibabaei, Lisana Berberi, Miguel Caballer, Amanda Calatrava, Pedro Castro, Alessandro Costantini, Mario David, Jaime D\'iez Stefan Dlugolinsky, Borja Esteban Sanchis, Giacinto Donvito, Leonhard Duda, Sa\'ul Fernandez, Andr\'es Heredia Canales, Valentin Kozlov, Sergio Langarita, Jo\~ao Machado, Germ\'an Molt\'o, Daniel San Mart\'in, Martin \v{S}eleng, Giang Nguyen, Marcin P{\l}\'ociennik, Marta Obreg\'on Ruiz, Susana Rebolledo Ruiz, Vicente Rodriguez, Judith S\'ainz-Pardo D\'iaz, Viet Tran 5/7/2026

AI4EOSC: a Federated Cloud Platform for Artificial Intelligence in Scientific Research

Open-source federated platform for operationalizing AI/ML lifecycle in scientific research with FAIR principles and MLOps tooling.

Ax Mingzhe Lu, Yiwen Wang, Yanbing Liu, Qi You, Chong Liu, Ruize Qin, Haoyu Dong, Wenyu Zhang, Jiarui Zhang, Yue Hu, Yunpeng Li 5/7/2026

LitVISTA: A Benchmark for Narrative Orchestration in Literary Text

Benchmark dataset (LitVISTA) for evaluating LLM narrative structure and story arcs in literary text generation versus human narratives.