Ax Zichen Liu, Anya Sims, Keyu Duan, Changyu Chen, Simon Yu, Xiangxin Zhou, Haotian Xu, Shaopan Xiong, Bo Liu, Chenmien Tan, Chuen Yang Beh, Weixun Wang, Hao Zhu, Weiyan Shi, Diyi Yang, Michael Shieh, Yee Whye Teh, Wee Sun Lee, Min Lin 3/3/2026

GEM: A Gym for Agentic LLMs

GEM is an open-source environment simulator for training agentic LLMs through experience-based learning, analogous to OpenAI Gym.

Ax Ali Hatamizadeh, Syeda Nahida Akter, Shrimai Prabhumoye, Jan Kautz, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro, Yejin Choi 3/3/2026

RLP: Reinforcement as a Pretraining Objective

arXiv paper proposing RLP, reinforcement learning as a pretraining objective for reasoning models instead of post-training only.

Ax Runzhe Zhan, Yafu Li, Zhi Wang, Xiaoye Qu, Dongrui Liu, Jing Shao, Derek F. Wong, Yu Cheng 3/3/2026

ExGRPO: Learning to Reason from Experience

arXiv paper on ExGRPO, reinforcement learning from verifiable rewards to improve LLM reasoning with experience reuse.

Ax Xinzhe Huang, Wenjing Hu, Tianhang Zheng, Kedong Xiu, Xiaojun Jia, Di Wang, Zhan Qin, Kui Ren 3/3/2026

Untargeted Jailbreak Attack

arXiv paper on untargeted jailbreak attacks against LLMs that optimize adversarial suffixes without fixed target responses.

Ax Junxi Yan, Zixi Wei, Qingyao Ai, Yiqun Liu, Jingtao Zhan 3/3/2026

What Scales in Cross-Entropy Scaling Law?

Analysis of cross-entropy scaling law breakdown at very large LLM scales and investigation of contributing factors.

Ax Seungeun Rho, Aaron Trinh, Danfei Xu, Sehoon Ha 3/3/2026

Reference Grounded Skill Discovery

Reference-Grounded Skill Discovery algorithm for scaling unsupervised skill learning to high-dimensional agent control.

Ax Perry Dong, Chongyi Zheng, Chelsea Finn, Dorsa Sadigh, Benjamin Eysenbach 3/3/2026

Value Flows

Value Flows framework extends distributional RL by modeling return distributions with continuous flow-based methods.

Ax Wenjie Ma, Andrei Cojocaru, Neel Kolhe, Bradley Louie, Robin Said Sharif, Haihan Zhang, Vincent Zhuang, Matei Zaharia, Sewon Min 3/3/2026

Reliable Fine-Grained Evaluation of Natural Language Math Proofs

Systematic methodology for developing reliable fine-grained evaluators of LLM-generated natural language math proofs.