Ax Mark Braverman, Roi Livni, Yishay Mansour, Shay Moran, Kobbi Nissim 6/26/2026

Learning from Equivalence Queries, Revisited

Learning framework for evolving models through user interaction queries; theoretical foundations for deployed systems.

Ax Mykola Vysotskyi, Runqi Lin, Grzegorz Biziel, Michal Zakrzewski, Sebastian Montagna, Damian Rynczak, Shreyansh Padarha, Kumail Alhamoud, Zihao Fu, William Lugoloobi, Kai Rawal, Hanna Yershova, Xander Davies, Taras Rumezhak, Guohao Li, Fazl Barez, Baoyuan Wu, Arkadiusz Drohomirecki, Yarin Gal, Chris Russell, Christopher Summerfield, Adam Mahdi, Volodymyr Karpiv, Philip Torr, Adel Bibi 6/26/2026

Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments

Benchmark for evaluating AI agent capabilities across diverse environments beyond common applications, addressing limitations of saturated performance on existing benchmarks.

Ax You Zuo (ALMAnaCH), Kim Gerdes (LISN), Eric Villemonte de La Clergerie (ALMAnaCH), Beno\^it Sagot (ALMAnaCH) 6/26/2026

Patent Representation Learning via Self-supervision

Self-supervised contrastive learning method for patent document representation, optimizing dropout and temperature settings.

Ax Ilia Kulikov, Chenxi Whitehouse, Tianhao Wu, Yixin Nie, Swarnadeep Saha, Eryk Helenowski, Weizhe Yuan, Olga Golovneva, Jack Lanchantin, Yoram Bachrach, Jakob Foerster, Xian Li, Han Fang, Sainbayar Sukhbaatar, Jason Weston 6/26/2026

Autodata: An agentic data scientist to create high quality synthetic data

arXiv paper introducing Autodata, an AI agent that acts as data scientist to generate synthetic training data via Agentic Self-Instruct.

HN sebg 6/25/2026

Scaling Laws, Carefully

Technical analysis of scaling laws in deep learning, their empirical foundations, and optimal compute allocation strategies.