Ax Lingfeng Li, Yunlong Lu, Yuefei Zhang, Jingyu Yao, Yixin Zhu, KeYuan Cheng, Yongyi Wang, Qirui Zheng, Xionghui Yang, Wenxin Li 2/17/2026

BotzoneBench: Scalable LLM Evaluation via Graded AI Anchors

BotzoneBench framework for evaluating LLMs in interactive strategic decision-making environments using graded AI anchors instead of LLM-vs-LLM tournaments.

Ax Zerui Cheng, Jiashuo Liu, Chunjie Wu, Jianzhu Yao, Pramod Viswanath, Ge Zhang, Wenhao Huang 2/17/2026

VeRA: Verified Reasoning Data Augmentation at Scale

VeRA framework for verified reasoning data augmentation at scale, converting benchmark problems into executable specifications for robust AI evaluation.

Ax Javier Mar\'in 2/17/2026

A Geometric Taxonomy of Hallucinations in LLMs

Taxonomy categorizing LLM hallucinations into three geometric types: unfaithfulness, confabulation, and factual error with distinct embedding signatures.

Ax Roham Koohestani, Ali Al-Kaswan, Jonathan Katzy, Maliheh Izadi 2/17/2026

AST-PAC: AST-guided Membership Inference for Code

AST-PAC applies AST-guided membership inference attacks to detect unauthorized code usage in Code LLMs trained on restricted datasets.

Ax Darren Li, Meiqi Chen, Chenze Shao, Fandong Meng, Jie Zhou 2/17/2026

General learned delegation by clones

SELFCEST: Agentic RL training enabling base models to spawn parallel clones for efficient test-time compute allocation.

Ax William Waites 2/17/2026

Artificial Organisations

Proposes organizational/institutional design patterns for multi-agent AI systems to achieve reliability through compartmentalization and adversarial review.

Ax Yifan Ding, Yuhui Shi, Zhiyan Li, Zilong Wang, Yifeng Gao, Yajun Yang, Mengjie Yang, Yixiu Liang, Xipeng Qiu, Xuanjing Huang, Xingjun Ma, Yu-Gang Jiang, Guoyu Wang 2/17/2026

Mirror: A Multi-Agent System for AI-Assisted Ethics Review

Multi-agent system using LLMs to assist with ethics review and research governance decision-making.

Ax Chen Yang, Guangyue Peng, Jiaying Zhu, Ran Le, Ruixiang Feng, Tao Zhang, Xiyun Xu, Yang Song, Yiming Jia, Yuntao Wen, Yunzhi Xu, Zekai Wang, Zhenwei An, Zhicong Sun, Zongchao Chen 2/17/2026

Nanbeige4.1-3B: A Small General Model that Reasons, Aligns, and Acts

Open-source 3B parameter language model achieving agentic behavior, code generation, and reasoning through reward modeling.

Ax Anhao Zhao, Ziyang Chen, Junlong Tong, Yingqi Fan, Fanghua Ye, Shuhao Li, Yunpu Ma, Wenjie Li, Xiaoyu Shen 2/17/2026

On-Policy Supervised Fine-Tuning for Efficient Reasoning

Proposes on-policy supervised fine-tuning as simpler alternative to RL for training reasoning models with better efficiency.