Ax Ningkang Peng, Qianfeng Yu, Xiaoqian Peng, Linjing Qian, Yafei Liu, Canran Xiao, Xinyu Lu, Tingyu Lu, Zhichao Zheng, Yanhui Gu 3/18/2026

How to Achieve Prototypical Birth and Death for OOD Detection?

Prototype-based OOD detection method with dynamic prototype count adaptation based on category complexity.

Ax Andres Potapczynski, Ravi Kiran Selvam, Tatiana Konstantinova, Shankar Ramasubramanian, Malcolm Wolff, Kin G. Olivares, Ruijun Ma, Mengfei Cao, Michael W. Mahoney, Andrew Gordon Wilson, Boris N. Oreshkin, Dmitry Efimov 3/18/2026

Time-Aware Prior Fitted Networks for Zero-Shot Forecasting with Exogenous Variables

Zero-shot forecasting method for time series with exogenous variables using prior-fitted networks.

Ax Hanxian Huang, Igor Fedorov, Andrey Gromov, Bernard Beckerman, Naveen Suda, David Eriksson, Maximilian Balandat, Rylan Conway, Patrick Huber, Chinnadhurai Sankar, Ayushi Dalmia, Zechun Liu, Lemeng Wu, Tarek Elgamal, Adithya Sagar, Vikas Chandra, Raghuraman Krishnamoorthi 3/18/2026

MobileLLM-Flash: Latency-Guided On-Device LLM Design for Industry Scale

Hardware-in-the-loop architecture search methodology for designing efficient on-device LLMs with real-time latency constraints for mobile deployment.

Ax Swadesh Jana, Cansu Sancaktar, Tom\'a\v{s} Dani\v{s}, Georg Martius, Antonio Orvieto, Pavel Kolev 3/18/2026

GASP: Guided Asymmetric Self-Play For Coding LLMs

Proposes guided asymmetric self-play method for post-training coding LLMs with better problem selection to improve model capabilities.

Ax Xiaolong Han, Ferrante Neri, Zijian Jiang, Fang Wu, Yanfang Ye, Lu Yin, Zehong Wang 3/18/2026

W2T: LoRA Weights Already Know What They Can Do

Analyzes whether LoRA checkpoint weights encode task performance information readable without running the base model, enabling efficient adapter analysis.

Ax Christina Baek, Ricardo Pio Monti, David Schwab, Amro Abbas, Rishabh Adiga, Cody Blakeney, Maximilian B\"other, Paul Burstein, Aldo Gael Carranza, Alvin Deng, Parth Doshi, Vineeth Dorna, Alex Fang, Tony Jiang, Siddharth Joshi, Brett W. Larsen, Jason Chan Lee, Katherine L. Mentzer, Luke Merrick, Haakon Mongstad, Fan Pan, Anshuman Suri, Darren Teh, Jason Telanoff, Jack Urbanek, Zhengping Wang, Josh Wills, Haoli Yin, Aditi Raghunathan, J. Zico Kolter, Bogdan Gaza, Ari Morcos, Matthew Leavitt, Pratyush Maini 3/18/2026

The Finetuner's Fallacy: When to Pretrain with Your Finetuning Data

Study of specialized pretraining strategy using domain data during pretraining to improve finetuning performance and reduce forgetting.

Ax Hoang Phan, Quang H. Nguyen, Hung T. Q. Le, Xiusi Chen, Heng Ji, Khoa D. Doan 3/18/2026

Decoding the Critique Mechanism in Large Reasoning Models

Study of how large reasoning models use backtracking and self-verification to detect and correct errors in complex logical reasoning tasks.