Ax Perry Dong, Alexander Swerdlow, Dorsa Sadigh, Chelsea Finn 4/22/2026

FASTER: Value-Guided Sampling for Fast RL

FASTER: Value-guided sampling method reducing computational cost of test-time scaling in diffusion-based RL policies.

Ax Joongwon Kim, Wannan Yang, Kelvin Niu, Hongming Zhang, Yun Zhu, Eryk Helenowski, Ruan Silva, Zhengxing Chen, Srinivasan Iyer, Manzil Zaheer, Daniel Fried, Hannaneh Hajishirzi, Sanjeev Arora, Gabriel Synnaeve, Ruslan Salakhutdinov, Anirudh Goyal 4/22/2026

Scaling Test-Time Compute for Agentic Coding

Method for scaling test-time compute in agentic coding systems using trajectory ranking and value estimates for long-horizon tasks.

Ax Marti\~no R\'ios-Garc\'ia, Nawaf Alampara, Chandan Gupta, Indrajeet Mandal, Sajid Mannan, Ali Asghar Aghajani, N. M. Anoop Krishnan, Kevin Maik Jablonka 4/22/2026

AI scientists produce results without reasoning scientifically

Evaluation of LLM-based scientific agents across 8 domains with 25k runs assessing whether they follow scientific reasoning norms.

Ax Rania Elbadry, Sarfraz Ahmad, Ahmed Heakl, Dani Bouch, Momina Ahsan, Muhra AlMahri, Marwa Elsaid khalil, Yuxia Wang, Salem Lahlou, Sophia Ananiadou, Veselin Stoyanov, Jimin Huang, Xueqing Peng, Preslav Nakov, Zhuohan Xie 4/22/2026

SAHM: A Benchmark for Arabic Financial and Shari'ah-Compliant Reasoning

SAHM: benchmark and instruction-tuning dataset for Arabic financial NLP and Sharia-compliant reasoning with 14,380 examples.

Ax Zhenghua Ma, G Abarajithan, Dimitrios Danopoulos, Olivia Weng, Francesco Restuccia, Ryan Kastner 4/22/2026

Design Rules for Extreme-Edge Scientific Computing on AI Engines

Design principles for deploying machine learning models on extreme-edge AI hardware with latency/throughput constraints using spatial dataflow.

Ax Witold Wydma\'nski, Marek \'Smieja 4/22/2026

AutoNFS: Automatic Neural Feature Selection

AutoNFS: Automatic neural feature selection method for high-dimensional tabular data that detects optimal feature count without retraining.