Ax Shuang Ma, Chon Lam Lao, Zhiying Xu, Zhuang Wang, Ziming Mao, Delong Meng, Jia Zhen, Jun Wu, Ion Stoica, Yida Wang, Yang Zhou 4/23/2026

UCCL-Zip: Lossless Compression Supercharged GPU Communication

UCCL-Zip integrates lossless compression into GPU communication primitives for LLM training without numerical errors or convergence degradation.

Ax Yunke Ao, Le Chen, Bruce D. Lee, Assefa S. Wahd, Aline Czarnobai, Philipp F\"urnstahl, Bernhard Sch\"olkopf, Andreas Krause 4/23/2026

Bounded Ratio Reinforcement Learning

Bounded Ratio Reinforcement Learning framework bridges theory-practice gap in PPO by formalizing trust region methods with bounded ratio constraints.

Ax SLAM Labs, :, Oleksiy Ostapenko, Raymond Li, Torsten Scholak, Alireza Mousavi-Hosseini, Aman Tiwari, Denis Kocetkov, Joel Lamy Poirier, Kelechi Ogueji, Nanda H Krishna, Rafael Pardinas, Sathwik Tejaswi Madhusudhan, Shruthan Radhakrishna, Srinivas Sunkara, Valerie Becaert 4/23/2026

Super Apriel: One Checkpoint, Many Speeds

Super Apriel: 15B-parameter supernet supporting multiple attention mechanisms switchable at inference time without reloading.

Ax Zeyu Shen, Peter Henderson 4/23/2026

Temporally Extended Mixture-of-Experts Models

Proposes temporally extended mixture-of-experts layers using reinforcement learning options framework to improve GPU memory efficiency during inference.

Ax Luke Bailey, Kaiyue Wen, Kefan Dong, Tatsunori Hashimoto, Tengyu Ma 4/23/2026

Scaling Self-Play with Self-Guidance

Self-play scaling approach for LLMs with self-guidance to overcome learning plateaus and reward hacking.

Ax Andrey Vasilyev, Yikai Wang, Xiaocheng Li, Guanting Chen 4/23/2026

Calibrating conditional risk

Methods for estimating expected loss of prediction models conditional on input features in classification and regression settings.

Ax Alessandro Morosini, Matea Gjika, Tomaso Poggio, Pierfrancesco Beneventano 4/23/2026

Too Sharp, Too Sure: When Calibration Follows Curvature

Study of relationship between model calibration and loss surface curvature in neural networks, showing calibration emerges during training on vision tasks.