Ax Alex Chen, Renato Geh, Aditya Grover, Guy Van den Broeck, Daniel Israel 5/15/2026

The Pitfalls of KV Cache Compression

Study identifying pitfalls in KV cache compression for LLMs in realistic multi-instruction scenarios with practical implications.

Ax Yihong Wu, Liheng Ma, Lei Ding, Muzhi Li, Xinyu Wang, Kejia Chen, Zhan Su, Zhanguang Zhang, Chenyang Huang, Yingxue Zhang, Mark Coates, Jian-Yun Nie 5/15/2026

It Takes Two: Your GRPO Is Secretly DPO

Analysis showing GRPO reinforcement learning algorithm for LLM post-training is equivalent to DPO with group-level baselines.

Ax Ning Yang, Hengyu Zhong, Haijun Zhang, Randall Berry 5/15/2026

Vision-LLMs for Spatiotemporal Traffic Forecasting

Vision-LLM approach for spatiotemporal traffic forecasting combining visual understanding of grid-based traffic data with language model capabilities.

Ax Robert Joseph George, Carson Eisenach, Udaya Ghai, Dominique Perrault-Joncas, Anima Anandkumar, Dean Foster 5/15/2026

BRIDGE: Building Representations In Domain Guided Program Synthesis

BRIDGE framework for structured prompting of LLMs to generate code with formal verification in proof assistants like Lean, handling multiple coupled domains.

Ax Jingkun Liu, Yisong Yue, Max Welling, Yue Song 5/15/2026

Krause Synchronization Transformers

Krause Attention mechanism addressing representation collapse and attention sink phenomena in transformers through principled bounded-confidence dynamics.

Ax Pascal Jr Tikeng Notsawo, Guillaume Dumas, Guillaume Rabusseau 5/15/2026

Grokking Finite-Dimensional Algebra

Study of grokking phenomenon (sudden generalization) in neural networks learning finite-dimensional algebra operations, extending prior work on group operations.

Ax Th\'eo Vincent, Kevin Gerhardt, Yogesh Tripathi, Habib Maraqten, Adam White, Martha White, Jan Peters, Carlo D'Eramo 5/15/2026

Gradient Iterated Temporal-Difference Learning

Gradient Iterated Temporal-Difference Learning addresses divergence issues in TD learning with semi-gradient updates.

Ax Kun Zhang, Jiaqi Sun, Yiqing Li, Ignavier Ng, Namrata Deka, Shaoan Xie 5/15/2026

SEDGE: Structural Extrapolated Data Generation

SEDGE framework for generating structured data beyond training distribution with conditions for reliable extrapolation.

Ax Junyu Guo, Shangding Gu, Ming Jin, Costas Spanos, Javad Lavaei 5/15/2026

LLMs Should Express Uncertainty Explicitly

Research on training LLMs to explicitly express uncertainty signals within responses during reasoning or at answer time.

Ax Yixian Xu, Yusong Wang, Shengjie Luo, Kaiyuan Gao, Tianyu He, Di He, Chang Liu 5/15/2026

Quotient-Space Diffusion Models

Quotient-Space Diffusion Models leverage symmetry in generative tasks for faster 3D molecule structure generation.

Ax Mohamed Ali Souibgui, Jan Fostier, Rodrigo Abad\'ia-Heredia, Bohdan Denysenko, Christian Marschke, Igor Peric 5/15/2026

LayerBoost: Layer-Aware Attention Reduction for Efficient LLMs

LayerBoost proposes layer-aware attention reduction to improve efficient LLM inference by replacing softmax attention selectively across layers.