HN hpcaitech 5/21/2026

New API Models Launched

DeepSeek V4 Pro and V4 Flash models now available via HPC-AI.COM API with competitive pricing for LLM applications.

Ax Yingwei Li, Xin Huang, Yang Liu, Yang Fu, Alex Zihao Zhu, Chen Song, Junwen Yao, Anant Subramanian, Hao Xiang, Weijing Shi, Yuliang Zou, Tom Hoddes, Zhaoqi Leng, Govind Thattai, Dragomir Anguelov, Mingxing Tan 5/21/2026

STELLAR: Scaling 3D Perception Large Models for Autonomous Driving

Study on scaling 3D perception models for autonomous driving, analyzing impact of model scale on multi-sensor fusion and spatial understanding.

Ax Yifeng He, Ethan Wang, Jicheng Wang, Xuanxin Ouyang, Hao Chen 5/21/2026

Code Generation by Differential Test Time Scaling

DiffCodeGen: test-time scaling method for code generation using coverage-guided approach with reduced token consumption.

Ax Parsa Mazaheri, Kasra Mazaheri 5/21/2026

AgentAtlas: Beyond Outcome Leaderboards for LLM Agents

AgentAtlas provides unified benchmark beyond single-metric leaderboards for evaluating LLM agents across multiple dimensions including task success, tool validity, safety, and robustness.

Ax Slim Barkallah, Luke Bailey, Kaiyue Wen, Mohammed Abouzaid, Tengyu Ma 5/21/2026

Pseudo-Formalization for Automatic Proof Verification

Pseudo-Formalization method for automatic proof verification translating informal proofs into formal languages to enable AI mathematical reasoning evaluation.

Ax Hamza Golubovic, Matthew Shen, Genevera I. Allen, Tarek M. Zikry 5/21/2026

Group-Aware Matrix Estimation and Latent Subspace Recovery

Matrix completion method for heterogeneous data with multiple group memberships, preserving subgroup-specific variation in recommendation and neuroscience applications.

Ax Mengyang Liu, Taozhi Chen, Zhenhua Xu, Xue Jiang, Yihong Dong 5/21/2026

Multi-agent Collaboration with State Management

Framework for multi-agent collaboration addressing concurrent codebase edits through state management to prevent silent conflicts and integration failures.

Ax Seungone Kim, Dongkeun Yoon, Kiril Gashteovski, Juyoung Suk, Jinheon Baek, Pranjal Aggarwal, Ian Wu, Viktor Zaverkin, Spase Petkoski, Daniel R. Schrider, Ilija Dukovski, Francesco Santini, Biljana Mitreska, Yong Jeong, Kyeongha Kwon, Young Min Sim, Dragana Manasova, Arthur Porto, Biljana Mojsoska, Makoto Takamoto, Marko Shuntov, Ruoqi Liu, Hyunjoo Jenny Lee, Niyazi Ulas Din\c{c}, Yehhyun Jo, Sunkyu Han, Chungwoo Lee, Huishan Li, Esther H. R. Tsai, Ergun Simsek, Khushboo Shafi, Yeonseung Chung, Jihye Park, Aleksandar Shulevski, Henrik Christiansen, Yoosang Son, Elly Knight, Amanda Montoya, Jeongyoun Ahn, Christian Langkammer, Heera Moon, Changwon Yoon, Nikola Stikov, Mooseok Jang, Edward Choi, Junhan Kim, Yeon Sik Jung, Woo Youn Kim, Jae Kyoung Kim, Ishraq Md Anjum, Hyun Uk Kim, Drew Bridges, Carolin Lawrence, Xiang Yue, Alice Oh, Akari Asai, Sean Welleck, Graham Neubig 5/21/2026

On the limits and opportunities of AI reviewers: Reviewing the reviews of Nature-family papers with 45 expert scientists

Study of AI peer reviewers on Nature papers with 45 expert scientists; evaluates capabilities, limitations, and credibility.

Ax Connor Pedersen, Dong H. Ahn, Michel Migdal, Collin Neale, Nik Konyuchenko 5/21/2026

Instant GPU Efficiency Visibility at Fleet Scale

GPU efficiency metric (OFU) for AI workloads on HPC systems using performance counters; no application instrumentation needed.