Ax Yunshu Wu, Jiayi Cheng, Longxuan Yu, Partha Thakuria, Rob Brekelmans, Evangelos E. Papalexakis, Greg Ver Steeg 5/14/2026

Discrete Stochastic Localization for Non-autoregressive Generation

Discrete Stochastic Localization framework for non-autoregressive sequence generation using continuous diffusion with unit-sphere embeddings.

Ax Kaiyang Li, Shaobo Han, Qing Su, Shihao Ji 5/14/2026

Bayesian Model Merging

Bayesian approach to merging multiple task-specific expert models without retraining, using strong anchor models as inductive bias.

Ax Siyuan Liu (IIIS, Tsinghua University), Tinghong Chen (College of AI, Tsinghua University,Shanghai Qi Zhi Institute), Xinghan Li (IIIS, Tsinghua University), Yifei Wang (Amazon AGI SF Lab), Jingzhao Zhang (IIIS, Tsinghua University,Shanghai Qi Zhi Institute) 5/14/2026

Data Difficulty and the Generalization--Extrapolation Tradeoff in LLM Fine-Tuning

Systematic empirical and theoretical analysis of how data difficulty affects generalization and extrapolation in LLM fine-tuning.

Ax Changhao Li, Rushi Qiang, Jiawei Huang, Chenxiao Gao, Chao Zhang, Niao He, Bo Dai 5/14/2026

Revisiting DAgger in the Era of LLM-Agents

Study of DAgger algorithm for training long-horizon LLM agents, addressing covariate shift in multi-turn interactions with teacher supervision.

Ax Celine Lee, Jing Nathan Yan, Chen Liang, Jiaxin Shi, Yin Zhang, Jeremiah Liu, Pengcheng Yin, Fernando Pereira, Ed Chi, Derek Cheng, Alexander M. Rush, Ruoxi Wang 5/14/2026

The Efficiency Gap in Byte Modeling

Study of efficiency gap in byte-level language modeling comparing byte-level and masked diffusion approaches to traditional subword tokenization.

Ax Yanggan Gu, Shuo Cai, Zihao Wang, Wenjun Wang, Yuanyi Wang, Pengkai Wang, Sirui Huang, Su Lu, Jianmin Wu, Hongxia Yang 5/14/2026

FeatCal: Feature Calibration for Post-Merging Models

FeatCal addresses performance gaps in merged models by analyzing feature drift between merged and expert models, proposing methods to calibrate features during model merging.

Ax Zihan Guan, Qiao Jin, Guangzhi Xiong, Fangyuan Chen, Mengxuan Hu, Qingyu Chen, Yifan Peng, Zhiyong Lu, Anil Vullikanti 5/14/2026

Large Language Models Lack Temporal Awareness of Medical Knowledge

Evaluation showing LLMs lack temporal awareness in medical knowledge because benchmarks are atemporal while medical knowledge continuously evolves.

Ax Aaditya L. Kachhadiya 5/14/2026

Local Inverse Geometry Can Be Amortized

Amortized learned surrogate (Deceptron) for solving nonlinear inverse problems by encoding curvature information into reusable reverse operator.

Ax Huiqi Deng, Yibo Li, Quanshi Zhang, Peng Zhang, Hongbin Pei, Xia Hu 5/14/2026

Understanding Generalization through Decision Pattern Shift

Decision Pattern Shift framework analyzing how deep neural network internal decision mechanisms evolve from training to test data for understanding generalization failures.