Ax Hongbo Zhang, Yue Yang, Jianhao Yan, Guangsheng Bao, Yue Zhang, Yue Zhang 2/13/2026

Detecting RLVR Training Data via Structural Convergence of Reasoning

Detection method for training data used in RLVR (reinforcement learning with verifiable rewards) by analyzing structural convergence of reasoning to identify benchmark contamination.

Ax Yordan Yordanov, Matteo Forasassi, Bayar Menzat, Ruizhi Wang, Chang Qi, Markus Kaltenberger, Amine M'Charrak, Tommaso Salvatori, Thomas Lukasiewicz 2/13/2026

Prototype Transformer: Towards Language Model Architectures Interpretable by Design

Prototype Transformer (ProtoT) introduces interpretable-by-design LM architecture using prototypes to make reasoning explicit and reduce opacity, hallucination, and deception risks.

Ax Nenad Toma\v{s}ev, Matija Franklin, Simon Osindero 2/13/2026

Intelligent AI Delegation

Framework for AI agents to dynamically decompose complex tasks into sub-components and delegate to other agents/humans while adapting to environmental changes and handling failures robustly.

Ax Taian Guo, Haiyang Shen, Junyu Luo, Zhongshi Xing, Hanchun Lian, Jinsheng Huang, Binqi Chen, Luchen Liu, Yun Ma, Ming Zhang 2/13/2026

MEME: Modeling the Evolutionary Modes of Financial Markets

LLM-based approach for quantitative finance using logic-oriented paradigm to model financial market movements beyond asset-centric or market-centric methods.

Ax Romain Froger, Pierre Andrews, Matteo Bettini, Amar Budhiraja, Ricardo Silveira Cabral, Virginie Do, Emilien Garreau, Jean-Baptiste Gaya, Hugo Lauren\c{c}on, Maxime Lecanu, Kunal Malkan, Dheeraj Mekala, Pierre M\'enard, Gerard Moreno-Torres Bertran, Ulyana Piterbarg, Mikhail Plekhanov, Mathieu Rita, Andrey Rusakov, Vladislav Vorotilov, Mengjue Wang, Ian Yu, Amine Benhalloum, Gr\'egoire Mialon, Thomas Scialom 2/13/2026

Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous Environments

Gaia2 benchmark for evaluating LLM agents in realistic asynchronous environments with temporal constraints, dynamic events, and multi-agent collaboration.

Ax Hanno Ackermann, Hong Cai, Mohsen Ghafoorian, Amirhossein Habibian 2/13/2026

HLA: Hadamard Linear Attention

Hadamard Linear Attention mechanism reducing computational cost of standard quadratic attention in transformers using kernel functions.

Ax John Muchovej, Amanda Royka, Shane Lee, Julian Jara-Ettinger 2/13/2026

GPT-4o Lacks Core Features of Theory of Mind

Evaluation framework testing whether GPT-4o possesses theory of mind through causal models of mental states, finding core features lacking.

Ax Nicholas Lee, Lutfi Eren Erdogan, Chris Joseph John, Surya Krishnapillai, Michael W. Mahoney, Kurt Keutzer, Amir Gholami 2/13/2026

Agentic Test-Time Scaling for WebAgents

CATTS technique for dynamically allocating compute in multi-step web agents using test-time scaling to improve agentic task performance and reliability.

Ax Zhendong Huang, Hengjie Cao, Fang Dong, Ruijun Huang, Mengyi Chen, Yifeng Yang, Xin Zhang, Anrui Chen, Mingzhi Dong, Yujiang Wang, Jinlong Hou, Qin Lv, Robert P. Dick, Yuan Cheng, Fan Yang, Tun Lu, Li Shang 2/13/2026

Spectra: Rethinking Optimizers for LLMs Under Spectral Anisotropy

Research on optimizer design for LLM training, analyzing spectral anisotropy in gradient signals to improve learning of contextual information.