Ax Hanrong Zhang, Shicheng Fan, Henry Peng Zou, Yankai Chen, Zhenting Wang, Jiayu Zhou, Chengze Li, Wei-Chieh Huang, Yifei Yao, Kening Zheng, Xue Liu, Xiaoxiao Li, Philip S. Yu 4/14/2026

CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification

CoEvoSkills framework enables LLM agents to self-evolve structured multi-file skill artifacts through co-evolutionary verification without manual authoring.

Ax Anushree Sinha, Srivaths Ranganathan, Debanshu Das, Abhishek Dharmaratnakar 4/14/2026

Beyond Fluency: Toward Reliable Trajectories in Agentic IR

Position paper on failure modes in agentic IR systems, analyzing error cascades in multi-step reason-act-observe workflows despite linguistic fluency.

Ax Mohamed Elfeki, Tu Trinh, Kelvin Luu, Guangze Luo, Nathan Hunt, Ernesto Montoya, Nandan Marwaha, Yannis He, Charles Wang, Fernando Crabedo, Alessa Castilo, Bing Liu 4/14/2026

HiL-Bench (Human-in-Loop Benchmark): Do Agents Know When to Ask for Help?

HiL-Bench evaluates whether coding agents know when to request help with incomplete specifications, exposing judgment gaps in frontier models.

Ax Charlie F. Ruan, Yucheng Qin, Akaash R. Parthasarathy, Xun Zhou, Ruihang Lai, Hongyi Jin, Yixin Dong, Bohan Hou, Meng-Shiun Yu, Yiyan Zhai, Sudeep Agarwal, Hangrui Cao, Siyuan Feng, Tianqi Chen 4/14/2026

WebLLM: A High-Performance In-Browser LLM Inference Engine

WebLLM inference engine enabling high-performance LLM execution directly in web browsers for on-device deployment without server GPUs.

Ax Kisu Yang, Yoonna Jang, Hwanseok Jang, Kenneth Choi, Isabelle Augenstein, Heuiseok Lim 4/14/2026

Reliable Evaluation Protocol for Low-Precision Retrieval

Protocol for reliable evaluation of low-precision retrieval systems, addressing spurious ties and variability in relevance scoring with reduced numerical precision.

Ax Wenhong Zhu, Ruobing Xie, Rui Wang, Xingwu Sun, Di Wang, Pengfei Liu 4/14/2026

Proximal Supervised Fine-Tuning

Proximal SFT: supervised fine-tuning method using trust-region constraints to prevent capability deterioration when adapting foundation models to new tasks.