Ax Yunhe Li, Hao Shi, Wenhao Liu, Mengzhe Ruan, Hanxu Hou, Zhongxiang Dai, Shuang Qiu, Linqi Song 29d ago

DemoPSD: Disagreement-Modulated Policy Self-Distillation

DemoPSD method improves LLM reasoning via self-distillation with disagreement modulation to reduce overfitting and improve cross-domain generalization.

Ax Hana Chockler, David A. Kelly, Daniel Kroening, Youcheng Sun 29d ago

Causal Explanations for Image Classifiers

Black-box method for computing image classifier explanations using formal causal theory and actual causality definitions.

Ax Raj Ghugare, Roger Creus Castanyer, Catherine Ji, Kathryn Wantlin, Jin Schofield, Karthik Narasimhan, Benjamin Eysenbach 29d ago

BuilderBench: The Building Blocks of Intelligent Agents

BuilderBench benchmark for developing AI agents that learn through interaction and exploration rather than mimicry alone.

Ax Anirudh Ajith, Amanpreet Singh, Jay DeYoung, Nadav Kunievsky, Austin C. Kozlowski, Oyvind Tafjord, James Evans, Daniel S. Weld, Tom Hope, Doug Downey 29d ago

PreScience: A Dataset and Benchmark for Scientific Forecasting

PreScience dataset and benchmark for forecasting scientific advances using 98K AI papers with citations and author histories.

Ax Xiaoxi Li, Wenxiang Jiao, Jiarui Jin, Haoxuan Li, Hao Wang, Shijian Wang, Guanting Dong, Jiajie Jin, Yinuo Wang, Yuan Lu, Ji-Rong Wen, Zhicheng Dou, Zhouchen Lin 29d ago

OmniGAIA: Towards Native Omni-Modal AI Agents

OmniGAIA benchmark evaluating omni-modal AI agents with vision, audio, and language integration for complex reasoning and tool usage.

Ax Giona Fieni, Joschua W\"uthrich, Marc-Philippe Neumann, Christopher H. Onder 29d ago

Learning-based Multi-agent Race Strategies in Formula 1

Reinforcement learning approach for multi-agent Formula 1 race strategy optimization, modeling energy, tire degradation, and competitor behavior.

Ax Drew Prinster, Clara Fannjiang, Ji Won Park, Kyunghyun Cho, Anqi Liu, Suchi Saria, Samuel Stanton 29d ago

Conformal Policy Control

Conformal policy control method using safe reference policies to regulate untested agent policies, balancing exploration and safety constraints.

Ax Pengyu Zhu, Lijun Li, Yaxing Lyu, Qianxin Luo, Jingyi Yang, Yi Liu, Tingfeng Hui, Xinyu Yuan, Li Sun, Sen Su, Jing Shao 29d ago

A Unified Framework for the Evaluation of LLM Agentic Capabilities

Unified evaluation framework for LLM agentic capabilities that separates model capability from benchmark implementation choices for fair cross-benchmark comparison.

Ax Xinbao Qiao, Xianglong Du, Wei Liu, Jingqi Zhang, Peihua Mai, Meng Zhang, Yan Pang 29d ago

When Sample Selection Bias Precipitates Model Collapse

Research on model collapse from recursive training on synthetic data and how sample selection bias affects model verification in low-resource regimes.

Ax Tingyang Chen, Shuo Lu, Kang Zhao, Weicheng Meng, Hanlin Teng, Tianhao Li, Chao Li, Xule Liu, Jian Liang, Zhizhong Zhang, Yuan Xie, Heng Qu, Kun Shao, Jian Luan 29d ago

HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry

HarnessX: foundry for composable, adaptive agent harnesses combining prompts, tools, memory, and control flow with systematic evolution from execution traces.