Ax Yin Cheng, Liao Zhou, Xiyu Liang, Dihao Luo, Tewei Lee, Kailun Zheng, Weiwei Zhang, Mingchen Cai, Jian Dong, Andy Zhang 3/31/2026

Let the Agent Steer: Closed-Loop Ranking Optimization via Influence Exchange

Sortify: Agent-based closed-loop ranking optimization for recommendation systems treating ranking as influence allocation problem with online-offline metric calibration.

Ax Swarna Kamal Paul, Shubhendu Sharma, Nitin Sareen 3/31/2026

GAAMA: Graph Augmented Associative Memory for Agents

GAAMA: Graph-based long-term memory system for AI agents maintaining personalized, multi-session behavior through associative memory structure instead of flat retrieval.

Ax Camilo Chac\'on Sartori, Jos\'e H. Garc\'ia, Andrei Voicu Tomut, Christian Blum 3/31/2026

GEAKG: Generative Executable Algorithm Knowledge Graphs

GEAKG: Knowledge graphs framework for representing procedural algorithm knowledge as executable, learnable structures for domain-agnostic problem solving.

Ax Yoonho Lee, Roshen Nair, Qizheng Zhang, Kangwook Lee, Omar Khattab, Chelsea Finn 3/31/2026

Meta-Harness: End-to-End Optimization of Model Harnesses

arXiv paper: Meta-Harness, automated optimization system searching over LLM harness code to improve system performance beyond model weights.

Ax Siyuan Ma, Bo Gao, Zikai Xiao, Hailong Wang, Xinlei Yu, Rui Qian, Jiayu Qian, Luqi Gong, Yang Liu 3/31/2026

CoT2-Meta: Budgeted Metacognitive Control for Test-Time Reasoning

arXiv paper: CoT2-Meta, training-free metacognitive reasoning framework combining chain-of-thought with meta-level control over reasoning trajectories.

Ax Fangda Ye, Yuxin Hu, Pengxiang Zhu, Yibo Li, Ziqi Jin, Yao Xiao, Yibo Wang, Lei Wang, Zhen Zhang, Lu Wang, Yue Deng, Bin Wang, Yifan Zhang, Liangcai Su, Xinyu Wang, He Zhao, Chen Wei, Qiang Ren, Bryan Hooi, An Bo, Shuicheng Yan, Lidong Bing 3/31/2026

MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome

MiroEval benchmarks multimodal deep research agents by evaluating both research process and outcomes, addressing limitations of existing benchmarks with real-world query complexity and multimodal coverage.

Ax Hongtao Wu, Boyun Zheng, Dingjie Song, Yu Jiang, Jianfeng Gao, Lei Xing, Lichao Sun, Yixuan Yuan 3/31/2026

Towards a Medical AI Scientist

Medical AI Scientist: autonomous system generating medical hypotheses, conducting experiments, and drafting manuscripts grounded in clinical evidence.

Ax Songjun Tu, Chengdong Xu, Qichao Zhang, Yaocheng Zhang, Xiangyuan Lan, Linjing Li, Dongbin Zhao 3/31/2026

Dynamic Dual-Granularity Skill Bank for Agentic RL

D2Skill: dynamic dual-granularity skill bank organizing reusable experience into task and step skills for reinforcement learning agents.

Ax Yadi Cao, Sicheng Lai, Jiahe Huang, Yang Zhang, Zach Lawrence, Rohan Bhakta, Izzy F. Thomas, Mingyun Cao, Chung-Hao Tsai, Zihao Zhou, Yidong Zhao, Hao Liu, Alessandro Marinoni, Alexey Arefiev, Rose Yu 3/31/2026

SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMs

SimulCost: cost-aware benchmark for LLM agents in physics simulations accounting for tool-use costs like simulation time and experimental resources.