Ax Iris Xu, Sunshine Jiang, John Marangola, Nitish Dashora, Richard Li, Thomas Liu, Zexue He, Yuheng Zhi, Alex Pentland, Pulkit Agrawal, Zhang-Wei Hong 16d ago

Learning More from Less: Reinforcement Learning from Hindsight

Improves sample efficiency in vision-language-action model RL post-training by learning from hindsight on sparse-reward manipulation tasks.

Ax Alexander Tian, Aditya Ghai, Sanjit Neelam, Zaal Vasania, Akshay Mishra 16d ago

COBS: Cumulant Order Block Sparse Attention

Analyzes block sparse attention in LLMs as KV cache optimization, studying DeepSeek's Native Sparse Attention block selection mechanism.

Ax Jayadeva, Madhur Aswani 16d ago

All you need is SAMPAT

SAMPAT: three-layer interpretable neural architecture using multivariate polynomials for scientific data analysis with provable continuous function learning.

Ax Kris Atallah (New York University, New York, USA) 16d ago

LionVote: Per-Layer Learning Rate Adaptation for Lion

Per-layer learning rate adaptation mechanism for Lion optimizer addressing 2.6-2.8x scaling disparity across attention, MLP, and normalization layers.

Ax Pedro P. Santos, F\'abio Vital, Alberto Sardinha, Francisco S. Melo 16d ago

Risk-Aware General-Utility Markov Decision Processes

Risk-aware Markov Decision Processes framework enabling agents to optimize risk measures of objective value distributions while trading off expected performance.

Ax Foundation Model Team 16d ago

Mach-Mind-4-Flash Technical Report

Mach-Mind-4-Flash: 35B-parameter Mixture-of-Experts agentic model with 3B activated parameters achieving 100B-class performance through post-training and agentic RL.

Ax Hyunjin Seo, Hyeon Hwang, Gyubok Lee, Jay Shin, Jimin Park, Taesoo Kim, Sanghoon Lee, Hongjoon Ahn, Sungjun Han, Sangwon Jung 16d ago

TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology

TheBioCollection unifies scattered biological databases and resources into a cohesive pretraining corpus for biology-focused large language models.

Ax Michael Murray, Daphne Chen, Simran Bagaria, Dean Fortier, Tess Hellebrekers, Galen Mullins, Harshavardhan Gajarla, Oier Mees, Maya Cakmak, Andrey Kolobov 16d ago

FlowDAgger: Human-in-the-Loop Adaptation of Generative Robot Policies in Latent Space

FlowDAgger enables human-in-the-loop adaptation of pretrained generative robot policies using diffusion models, allowing rapid real-world deployment without large-scale retraining.

Ax Pulkit Madan, Sanjay Haresh, Reza Ebrahimi, Sunny Panchal, Apratim Bhattacharyya, Roland Memisevic 16d ago

On Locality and Length Generalization in Visual Reasoning

Analysis of local sequential vision models versus global models, investigating computational benefits for visual reasoning tasks.