Ax Ajay Jaiswal, Lauren Hannah, Han-Byul Kim, Duc Hoang, Mehrdad Farajtabar, Minsik Cho 5/8/2026

TIDE: Every Layer Knows the Token Beneath the Context

Research on improving LLM architecture by using token indices at every layer instead of once at input, addressing rare token training and position awareness issues.

Ax Yifan Tang, Qiquan Wang, In\'es Garc\'ia-Redondo, Anthea Monod 5/8/2026

Topological Signatures of Grokking

Topological analysis of grokking phenomenon in neural networks using persistent homology on embedding matrices.

Ax Hongcan Guo, Qinyu Zhao, Yian Zhao, Shen Nie, Rui Zhu, Qiushan Guo, Feng Wang, Tao Yang, Hengshuang Zhao, Guoqiang Wei, Yan Zeng 5/8/2026

Continuous Latent Diffusion Language Model

Hierarchical latent diffusion language model for text generation using non-autoregressive approach with improved efficiency and semantic modeling.

Ax Apurva Gandhi, Satyaki Chakraborty, Xiangjun Wang, Aviral Kumar, Graham Neubig 5/8/2026

Recursive Agent Optimization

Reinforcement learning approach enabling agents to recursively spawn and delegate sub-tasks for divide-and-conquer inference-time scaling.