Ax Marc Becker, Lennart Schneider, Martin Binder, Lars Kotthoff, Bernd Bischl 4/1/2026

mlr3mbo: Bayesian Optimization in R

mlr3mbo: modular R toolbox for Bayesian optimization supporting single/multi-objective, parallelization, and custom algorithm construction.

Ax Tim R. Davidson, Benoit Seguin, Enrico Bacis, Cesar Ilharco, Hamza Harkous 4/1/2026

Reasoning-Driven Synthetic Data Generation and Evaluation

Reasoning-driven approach for generating synthetic multi-modal training data without manual prompts, addressing scarcity of specialized AI training datasets.

Ax Xue Jiang, Tianyu Zhang, Ge Li, Mengyang Liu, Taozhi Chen, Zhenhua Xu, Binhua Li, Wenpin Jiao, Zhi Jin, Yongbin Li, Yihong Dong 4/1/2026

Think Anywhere in Code Generation

Proposes adaptive reasoning allocation during code generation for LLMs, addressing limitations of upfront thinking approaches in handling code complexity.

Ax Minhyuk Seo, Seongwon Cho, Minjae Lee, Diganta Misra, Hyeonbeom Choi, Seon Joo Kim, Jonghyun Choi 4/1/2026

GenOL: Generating Diverse Examples for Name-only Online Learning

GenOL framework for online learning with only concept names (name-only setup) enabling real-time adaptation to data distribution shifts in continual learning scenarios.

Ax Zehua Pei, Ying Zhang, Hui-Ling Zhen, Tao Yuan, Xianzhi Yu, Zhenhua Dong, Sinno Jialin Pan, Mingxuan Yuan, Bei Yu 4/1/2026

PreMoE: Proactive Inference for Efficient Mixture-of-Experts

Training-free framework for compiling sparse Mixture-of-Experts variants with predicted expert utility metric for deployment optimization.

Ax Da Chang, Yongxiang Liu, Ganzhao Yuan 4/1/2026

On the Convergence of Muon and Beyond

Theoretical convergence analysis of Muon optimizer for matrix-structured parameters in neural network training.

Ax Ahmed A. Elhag, Arun Raja, Alex Morehead, Samuel M. Blau, Hongtao Zhao, Christian Tyrchan, Eva Nittinger, Garrett M. Morris, Michael M. Bronstein 4/1/2026

Learning Inter-Atomic Potentials without Explicit Equivariance

Transformer-based inter-atomic potential model for molecular simulations without explicit equivariance constraints.

Ax Mingzhi Chen, Taiming Lu, Jiachen Zhu, Mingjie Sun, Zhuang Liu 4/1/2026

Stronger Normalization-Free Transformers

Research on normalization-free transformer architectures using Dynamic Tanh as alternative to standard normalization layers.

Ax Yuze Wang, Yujia Tong, Xuan Liu, Junhao Dong 4/1/2026

Sparsity-Aware Unlearning for Large Language Models

Addresses machine unlearning for sparse LLMs to remove memorized sensitive information while maintaining model sparsification benefits for efficient deployment.

Ax Jie Xiao, Meng Chen, Qingnan Ren, Jingwei Song, Jiaqi Huang, Yangshen Deng, Chris Tong, Wanyi Chen, Suli Wang, Ziqian Bi, Shuo Lu, Yiqun Duan, Xu Wang, Rymon Yu, Ween Yang, Lynn Ai, Eric Yang, Bill Shi 4/1/2026

ECHO-2: A Large-Scale Distributed Rollout Framework for Cost-Efficient Reinforcement Learning

ECHO-2 is distributed RL framework for LLM post-training via reinforcement learning, optimizing cost-efficiency of rollout generation across distributed resources.

Ax Amin Oji, Paul Fieguth 4/1/2026

Joint Embedding Variational Bayes

VJE introduces reconstruction-free latent-variable framework for self-supervised learning using symmetric conditional ELBO on paired embeddings.