Ax Tianyu Li, Dongchen Han, Zixuan Cao, Haofeng Huang, Mengyu Zhou, Ming Chen, Erchao Zhao, Xiaoxi Jiang, Guanjun Jiang, Gao Huang 5/22/2026

SiameseNorm: Breaking the Barrier to Reconciling Pre/Post-Norm

SiameseNorm architecture reconciles pre-norm and post-norm trade-offs in Transformers for improved stability and capacity.

Ax Zhixia Zhang, Zixuan Huang, Gongxun Li, Huaiyang Wang, Chengyi Yuan, Xin Xia, Deqing Wang, Fuzhen Zhuang, Shuai Ma, Ning Ding, Yaodong Yang, Jianxin Li, Yikun Ban 5/22/2026

Heterogeneous Agent Collaborative Reinforcement Learning

arXiv paper on Heterogeneous Agent Collaborative RL where agents share verified rollouts during training for mutual improvement.

Ax Pavel Golikov, Evgenii Opryshko, Gennady Pekhimenko, Mark C. Jeffrey 5/22/2026

Robust Reasoning Benchmark

arXiv paper introducing Robust Reasoning Benchmark testing 8 LLMs against textual perturbations on AIME math problems.

Ax Amrut Nadgir, Vijay Balasubramanian, Pratik Chaudhari 5/22/2026

How does Chain of Thought decompose complex tasks?

Analysis of how chain-of-thought reasoning decomposes complex tasks by breaking classification into smaller subtasks to reduce error.

Ax Yuanda Xu, Hejian Sang, Zhengze Zhou, Ran He, Zhipeng Wang, Alborz Geramifard 5/22/2026

TIP: Token Importance in On-Policy Distillation

Research on token importance in on-policy knowledge distillation for LLMs, identifying which token positions provide useful learning signals during training.

Ax Zhiming Huang, Jamie Morgenstern, Aaron Roth, Claire Jie Zhang 5/22/2026

Instance-Adaptive Online Multicalibration

Develops online multicalibration algorithm adapting between benign and worst-case sequences via dyadic grid refinement with optimal convergence rates.

Ax Yuxiang Chen, Dingli Liang, Yihang Chen, Ziqin Gong, Chenyang Le, Zhaokai Wang, Jiachen Zhu, Lingyu Yang, Jianghao Lin, Weinan Zhang, Jun Wang 5/22/2026

Holder Policy Optimisation

Introduces Holder Policy Optimisation improving on GRPO by adaptively aggregating token-level probabilities for trajectory-level advantages in LLM training.

Ax Yunshu Wu, Jiayi Cheng, Longxuan Yu, Partha Thakuria, Rob Brekelmans, Evangelos E. Papalexakis, Greg Ver Steeg 5/22/2026

Discrete Stochastic Localization for Non-autoregressive Generation

Proposes Discrete Stochastic Localization for non-autoregressive generation using continuous diffusion with unit-sphere embeddings, improving over masked discrete diffusion models.

Ax Krish Sharma, Omar Naim, Soumadeep Saha, Vinija Jain, Aman Chadha, Nicholas Asher 5/22/2026

TAPIOCA: Why Task- Aware Pruning Improves OOD model Capability

Investigates task-aware layer pruning effects on model generalization. Finds pruning improves out-of-distribution accuracy while harming in-distribution performance across polynomial tasks and LLMs.

Ax Youngin Kim, Ray Sun, Inho Kim, Bumsoo Park, Hyun Oh Song 5/22/2026

Identifiable Token Correspondence for World Models

Introduces Identifiable Token Correspondence to improve temporal consistency in token-based transformer world models for long-horizon visual RL.

Ax Egor Shvetsov, Aleksandr Serkov, Shokorov Viacheslav, Redko Dmitry, Vladislav Goloshchapov, Evgeny Burnaev 5/22/2026

Bug or Feature$^2$: Weight Drift, Activation Sparsity and Spikes

Analyzes weight drift and activation sparsity in neural network training dynamics caused by interactions between standard losses and activation functions.

Ax Muhammad Umer, Muhammad Ahmed Mohsin, Ahsan Bilal, Arslan Chaudhry, Andreas Haupt, Sanmi Koyejo, Emily Fox, John M. Cioffi 5/22/2026

General Preference Reinforcement Learning

Proposes General Preference Reinforcement Learning to unify online RL for verifiable tasks with preference optimization for open-ended generation in LLM alignment.