Ax Sviatoslav Lushnei, Dmytro Shumskyi, Severyn Shykula, Ernesto Jimenez-Ruiz, Artur d'Avila Garcez 2/17/2026

Large Language Models as Oracles for Ontology Alignment

Using LLMs as oracles for ontology alignment with human-in-the-loop approaches to improve mapping quality for large ontologies.

Ax Caorui Li, Yu Chen, Yiyan Ji, Jin Xu, Zhenyu Cui, Shihao Li, Yuanxing Zhang, Wentao Wang, Zhenghao Song, Dingling Zhang, Ying He, Haoxiang Liu, Yuxuan Wang, Qiufeng Wang, Jiafu Tang, Zhenhe Wu, Jiehui Luo, Zhiyu Pan, Weihao Xie, Chenchen Zhang, Zhaohui Wang, Jiayi Tian, Yanghai Wang, Zhe Cao, Minxin Dai, Ke Wang, Runzhe Wen, Yinghao Ma, Yaning Pan, Sungkyun Chang, Termeh Taheri, Haiwen Xia, Christos Plachouras, Emmanouil Benetos, Yizhi Li, Ge Zhang, Jian Yang, Tianhao Peng, Zili Wang, Minghao Liu, Junran Peng, Zhaoxiang Zhang, Jiaheng Liu 2/17/2026

OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs

OmniVideoBench: evaluation benchmark for multimodal LLMs on audio-visual understanding tasks with comprehensive synergistic reasoning assessment.

Ax Shiqi Zhang, Xinbei Ma, Yunqing Xu, Zouying Cao, Pengrui Lu, Haobo Yuan, Tiancheng Shen, Zhuosheng Zhang, Hai Zhao, Ming-Hsuan Yang 2/17/2026

ParaCook: On Time-Efficient Planning for Multi-Agent Systems

ParaCook benchmark for evaluating time-efficient collaborative planning in multi-agent systems using LLMs for long-horizon reasoning.

Ax Minwei Kong, Ao Qu, Xiaotong Guo, Wenbin Ouyang, Chonghe Jiang, Han Zheng, Yining Ma, Dingyi Zhuang, Yuhan Tang, Junyi Li, Shenhao Wang, Haris Koutsopoulos, Hai Wang, Cathy Wu, Jinhua Zhao 2/17/2026

AlphaOPT: Formulating Optimization Programs with Self-Improving LLM Experience Library

AlphaOPT uses LLMs with self-improving experience libraries to automate optimization problem formulation from natural language into mathematical models and solver code.

Ax Ricardo Vinuesa, Steven L. Brunton, Gianmarco Mengaldo 2/17/2026

Explainable AI: Learning from the Learners

Perspective on explainable AI combined with causal reasoning for extracting insights from foundation models.

Ax Hyejun Jeong, Amir Houmansadr, Shlomo Zilberstein, Eugene Bagdasarian 2/17/2026

Persuasion Propagation in LLM Agents

Study on persuasion propagation: how belief-level intervention affects downstream behavior in LLM agents executing long-horizon tasks.

Ax Alisia Lupidi, Bhavul Gauri, Thomas Simon Foster, Bassel Al Omari, Despoina Magka, Alberto Pepe, Alexis Audran-Reiss, Muna Aghamelu, Nicolas Baldwin, Lucia Cipolina-Kun, Jean-Christophe Gagnon-Audet, Chee Hau Leow, Sandra Lefdal, Hossam Mossalam, Abhinav Moudgil, Saba Nazir, Emanuel Tewolde, Isabel Urrego, Jordi Armengol Estape, Amar Budhiraja, Gaurav Chaurasia, Abhishek Charnalia, Derek Dunfield, Karen Hambardzumyan, Daniel Izcovich, Martin Josifoski, Ishita Mediratta, Kelvin Niu, Parth Pathak, Michael Shvartsman, Edan Toledo, Anton Protopopov, Roberta Raileanu, Alexander Miller, Tatiana Shavrina, Jakob Foerster, Yoram Bachrach 2/17/2026

AIRS-Bench: a Suite of Tasks for Frontier AI Research Science Agents

AIRS-Bench: benchmark of 20 ML research tasks for evaluating AI agent capabilities across language modeling, mathematics, bioinformatics, and time series forecasting.

Ax John Muchovej, Amanda Royka, Shane Lee, Julian Jara-Ettinger 2/17/2026

GPT-4o Lacks Core Features of Theory of Mind

Tests whether GPT-4o possesses Theory of Mind via causal model evaluation, finding it lacks core ToM representations.

Ax Tianyu Chen, Shuai Lu, Shan Lu, Yeyun Gong, Chenyuan Yang, Xuheng Li, Md Rakib Hossain Misu, Hao Yu, Nan Duan, Peng Cheng, Fan Yang, Shuvendu K Lahiri, Tao Xie, Lidong Zhou 2/17/2026

Automated Proof Generation for Rust Code via Self-Evolution

SAFE framework automates formal proof generation for Rust code using LLMs via self-evolution to overcome proof data scarcity.

Ax Federico Errica, Henrik Christiansen, Viktor Zaverkin, Mathias Niepert, Francesco Alesiani 2/17/2026

Adaptive Width Neural Networks

Technique for learning neural network layer width during training without manual hyperparameter tuning or architecture search.

Ax Xianrui Zhong, Bowen Jin, Siru Ouyang, Yanzhen Shen, Qiao Jin, Yin Fang, Zhiyong Lu, Jiawei Han 2/17/2026

Benchmarking Retrieval-Augmented Generation for Chemistry

Benchmark for retrieval-augmented generation in chemistry domain with curated evaluation datasets and domain-specific corpora.