Ax Thomas Bartz-Beielstein 4/16/2026

Optimization with SpotOptim

Python package for surrogate-model-based optimization using Kriging, Expected Improvement, and multi-objective extensions for expensive functions.

Ax Youling Huang, Guanqiao Chen, Junchi Yao, Lu Wang, Fangkai Yang, Chao Du, ChenZhuo Zhao, Pu Zhao, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang 4/16/2026

Beyond State Consistency: Behavior Consistency in Text-Based World Models

arXiv paper on behavior consistency in text-based world models for evaluating agent planning and offline evaluation beyond single-step metrics.

Ax Adam Lahouari, Shen Ai, Jihye Han, Jillian Hoffstadt, Philipp Hoellmer, Charlotte Infante, Pulkita Jain, Sangram Kadam, Maya M. Martirossyan, Amara McCune, Hypatia Newton, Shlok J. Paul, Willmor Pena, Jonathan Raghoonanan, Sumon Sahu, Oliver Tan, Andrea Vergara, Jutta Rogal, Mark E. Tuckerman 4/16/2026

MolCryst-MLIPs: A Machine-Learned Interatomic Potentials Database for Molecular Crystals

MolCryst-MLIPs: Open database of machine-learned interatomic potentials for molecular crystals. MACE models for nine systems with automated ML pipeline.

Ax Zijian Gao, Wangwang Jia, Xingxing Zhang, Pengfei Qian, Tao Sun, Bo Ding, Yong Dou, Huaimin Wang, Kele Xu 4/16/2026

MAny: Merge Anything for Multimodal Continual Instruction Tuning

MAny: Method for multimodal continual instruction tuning of MLLMs. Addresses catastrophic forgetting via parameter merging across perception and reasoning spaces.

Ax Gerg\H{o} Szalay, Gergely Zsolt Kov\'acs, S\'andor Teleki, Bal\'azs Pint\'er, Tibor Gregorics 4/16/2026

Neural architectures for resolving references in program code

Neural architectures for resolving code references and indirect indexing. Proposes seq2seq models for decompilation tasks with synthetic benchmarks.

Ax Yuanda Xu, Hejian Sang, Zhengze Zhou, Ran He, Zhipeng Wang, Alborz Geramifard 4/16/2026

TIP: Token Importance in On-Policy Distillation

Token Importance in on-policy knowledge distillation for LLMs. Identifies which token positions provide useful learning signals during student training on teacher supervision.

Ax Sumeet Ramesh Motwani, Daniel Nichols, Charles London, Peggy Li, Fabio Pizzati, Acer Blake, Hasan Hammoud, Tavish McDonald, Akshat Naik, Alesia Ivanova, Vignesh Baskaran, Ivan Laptev, Ruben Glatt, Tal Ben-Nun, Philip Torr, Natasha Jaques, Ameya Prabhu, Brian Bartoldson, Bhavya Kailkhura, Christian Schroeder de Witt 4/16/2026

LongCoT: Benchmarking Long-Horizon Chain-of-Thought Reasoning

LongCoT: benchmark of 2,500 expert-designed problems measuring long-horizon chain-of-thought reasoning across chemistry, math, CS, chess, and logic.

Ax Maksim Ivanov, Abhijay Rana, Gokul Prabhakaran 4/16/2026

Can Coding Agents Be General Agents?

Case study evaluating coding agents on business process automation tasks in ERP systems, identifying capability gaps beyond software engineering.

Ax Dikshant Kukreja (IIIT Delhi, India), Kshitij Sah (IIIT Delhi, India), Gautam Gupta (IIIT Delhi, India), Avinash Anand (Singapore Institute of Technology), Rajiv Ratn Shah (IIIT Delhi, India), Zhengkui Wang (Singapore Institute of Technology), Aik Beng Ng (NVIDIA), Erik Cambria (Nanyang Technological University) 4/16/2026

Better and Worse with Scale: How Contextual Entrainment Diverges with Model Size

Scaling laws for contextual entrainment showing larger language models simultaneously improve at ignoring false claims but worsen at ignoring irrelevant tokens.

Ax Hongyi Jin, Bohan Hou, Guanjie Wang, Ruihang Lai, Jinqi Chen, Zihao Ye, Yaxing Cai, Yixin Dong, Xinhao Cheng, Zhihao Zhang, Yilong Zhao, Yingyi Huang, Lijie Yang, Jinchen Jiang, Gabriele Oliaro, Jianan Ji, Xupeng Miao, Vinod Grover, Todd C. Mowry, Zhihao Jia, Tianqi Chen 4/16/2026

Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel

Event Tensor compiler framework unifying dynamic megakernel abstraction to improve LLM inference performance by eliminating kernel launch overhead and enabling inter-kernel parallelism.

Ax Joel Niklaus, Atsuki Yamaguchi, Michal \v{S}tef\'anik, Guilherme Penedo, Hynek Kydl\'i\v{c}ek, Elie Bakouch, Lewis Tunstall, Edward Emanuel Beeching, Thibaud Frere, Colin Raffel, Leandro von Werra, Thomas Wolf 4/16/2026

How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data

Systematic study of synthetic data generation for LLM pretraining, testing rephrasing strategies, generator models, and source data across one trillion tokens to identify optimal design choices.