Ax Yue Wang, Qizhou Wang, Zizhuo Zhang, Gang Niu, Bo Han, Masashi Sugiyama 5/18/2026

What Is Preference Optimization Doing, and Why?

Analysis of optimization dynamics comparing preference optimization methods like DPO and PPO for LLM alignment.

Ax Yixuan Even Xu, John Kirchenbauer, Yash Savani, Asher Trockman, Alexander Robey, Tom Goldstein, Fei Fang, J. Zico Kolter 5/18/2026

Antidistillation Fingerprinting

Fingerprinting method to detect when student LLMs trained on teacher model outputs via distillation

Ax Nick Alonso, Tomas Figliolia, Beren Millidge 5/18/2026

Online Vector Quantized Attention

Vector quantized attention layer balancing efficiency and performance for long context language models

Ax Junxiong Wang, Fengxiang Bie, Jisen Li, Zhongzhu Zhou, Zelei Shao, Yubo Wang, Yinghui Liu, Qingyang Wu, Avner May, Sri Yanamandra, Ce Zhang, Tri Dao, Percy Liang, Ben Athiwaratkun, Shuaiwen Leon Song, Chenfeng Xu, Xiaoxia Wu 5/18/2026

When RL Meets Adaptive Speculative Training: A Unified Training-Serving System

Unified system combining reinforcement learning with adaptive speculative decoding for efficient LLM serving

Ax Uzay Macar, Li Yang, Atticus Wang, Peter Wallich, Emmanuel Ameisen, Jack Lindsey 5/18/2026

Mechanisms of Introspective Awareness

Study of how LLMs detect and identify injected steering vectors in residual streams, revealing introspective awareness mechanisms.

Ax Zigeng Chen, Gongfan Fang, Xinyin Ma, Ruonan Yu, Xinchao Wang 5/18/2026

DMax: Aggressive Parallel Decoding for dLLMs

DMax paradigm for efficient parallel decoding in diffusion language models via progressive self-refinement, reducing error accumulation.

Ax Andrey Bocharnikov, Ivan Ermakov, Denis Kuznedelev, Vyacheslav Zhdanovskiy, Yegor Yershov 5/18/2026

KV Cache Offloading for Context-Intensive Tasks

Research on KV cache offloading to reduce memory and latency bottlenecks for long-context LLM inference tasks.

Ax Zhaokun Wang, Jinyu Guo, Jingwen Pu, Hongli Pu, Meng Yang, Xunlei Chen, Jie Ou, Wenyi Li, Guangchun Luo, Wenhong Tian 5/18/2026

CAP: Controllable Alignment Prompting for Unlearning in LLMs

Research on unlearning sensitive information from LLMs via controllable alignment prompting without modifying weights, addressing regulatory compliance.

Ax Zhaoyuan Su, Olatunji Ruwase, Karthik Ganesan, Aurick Qiao, Samyam Rajbhandari, Juncheng Yang, Yue Cheng, Yuxiong He 5/18/2026

MoE-Prefill: Zero Redundancy Overheads in MoE Prefill Serving

System for serving mixture-of-experts LLMs on prefill-only workloads without redundant distributed execution overhead.

Ax Thomas Walker, T. Mitchell Roddenberry, Ahmed Imtiaz Humayun, Randall Balestriero, Richard Baraniuk 5/18/2026

The Geometric Structure of Models Learning Sparse Data

Theoretical analysis of how models learn sparse data regimes where the manifold hypothesis does not apply.

Ax Christopher Lohse, Anish Dhir, Amadou Ba, Bradley Eck, Marco Ruffini, Jonas Wahl 5/18/2026

PRIM: Meta-Learned Bayesian Root Cause Analysis

Meta-learning approach framing root cause analysis in complex systems as Bayesian inference over synthetic causal model priors.

Ax William Lehn-Schi{\o}ler, Magnus Ruud Kj{\ae}r, Rahul Thapa, Magnus Guldberg Pedersen, Anton Mosquera Storgaard, Nick Williams, Radu Gatej, Tue Lehn-Schi{\o}ler, S\'andor Beniczky, Sadasivan Puthusserypady, James Zou, Lars Kai Hansen 5/18/2026

Mechanistic Interpretability of EEG Foundation Models via Sparse Autoencoders

Applies sparse autoencoders to extract interpretable features from EEG foundation models for clinical applications.

Ax Jerem\'ias Figueiredo Paschmann, Juan Kaplan, Francisco Nattero, Santiago Barron, Juan Wisznia, Luciano del Corro 5/18/2026

Active Learners as Efficient PRP Rerankers

Uses active learning to efficiently rerank LLM pairwise preference judgments, treating it as robust top-K recovery rather than sorting.