Ax Sunshine Jiang, John Marangola, David Zhang, Raghuram Kowdeed, Ruiyang Luo, Nitish Dashora, Richard Li, Pulkit Agrawal, Zhang-Wei Hong 7/13/2026

Prompt-Driven Exploration

Method using LLMs and vision-language-action models to improve exploration in reinforcement learning by generating diverse policy perturbations via natural language prompts.

Ax Iris Xu, Sunshine Jiang, John Marangola, Nitish Dashora, Richard Li, Thomas Liu, Zexue He, Yuheng Zhi, Alex Pentland, Pulkit Agrawal, Zhang-Wei Hong 7/13/2026

Learning More from Less: Reinforcement Learning from Hindsight

Improves sample efficiency in vision-language-action model RL post-training by learning from hindsight on sparse-reward manipulation tasks.

Ax Alexander Tian, Aditya Ghai, Sanjit Neelam, Zaal Vasania, Akshay Mishra 7/13/2026

COBS: Cumulant Order Block Sparse Attention

Analyzes block sparse attention in LLMs as KV cache optimization, studying DeepSeek's Native Sparse Attention block selection mechanism.

Ax Jayadeva, Madhur Aswani 7/13/2026

All you need is SAMPAT

SAMPAT: three-layer interpretable neural architecture using multivariate polynomials for scientific data analysis with provable continuous function learning.

Ax Kris Atallah (New York University, New York, USA) 7/13/2026

LionVote: Per-Layer Learning Rate Adaptation for Lion

Per-layer learning rate adaptation mechanism for Lion optimizer addressing 2.6-2.8x scaling disparity across attention, MLP, and normalization layers.

Ax Pedro P. Santos, F\'abio Vital, Alberto Sardinha, Francisco S. Melo 7/13/2026

Risk-Aware General-Utility Markov Decision Processes

Risk-aware Markov Decision Processes framework enabling agents to optimize risk measures of objective value distributions while trading off expected performance.

Ax Foundation Model Team 7/13/2026

Mach-Mind-4-Flash Technical Report

Mach-Mind-4-Flash: 35B-parameter Mixture-of-Experts agentic model with 3B activated parameters achieving 100B-class performance through post-training and agentic RL.