Ax Zhanming Shen, Jintao Tong, Shaotian Yan, Chen Shen, Hao Chen, Wentao Ye, Xiaomeng Hu, Rui Miao, Haobo Wang, Junbo Zhao, Gang Chen, Jieping Ye 29d ago

Purified OPSD: On-Policy Self-Distillation Without Losing How to Think

Purified OPSD improves on-policy self-distillation for long chain-of-thought reasoning by preserving reflective capabilities while providing token-level supervision.

Ax Xiangchen Cheng, Yunwei Jiang, Jianwen Sun, Zizhen Li, Chuanhao Li, Xiangcheng Cao, Yihao Liu, Fanrui Zhang, Li Jin, Kaipeng Zhang 29d ago

AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents

AgenticSTS is a testbed for evaluating long-horizon LLM agents with bounded memory contracts, isolating effects of memory components on agent decision-making.

Ax Xianhui Meng, Zirui Song, Yuchen Zhang, Li Zhang, Yongxuan Lv, Xiuying Chen, Kun Wang, Yan Luo, Kai Chen, Hangjun Ye, Long Chen, Jun Liu, Xiaoshuai Hao 29d ago

Text-Driven 3D Indoor Scene Synthesis in Non-Manhattan Environments

SPG-Layout uses LLMs for text-driven 3D indoor scene synthesis in non-Manhattan environments by modeling non-orthogonal spatial relationships.

Ax Zhilin Wang, Han Song, Runzhe Zhan, Jusen Du, Jiacheng Chen, Tianle Li, Qingyu Yin, Yulun Wu, Zhennan Shen, Tong Zhu, Yanshu Li, Guanjie Chen, Derek F. Wong, Yafu Li, Yu Cheng, Yang Yang 29d ago

EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments

EvoPolicyGym evaluates autonomous agents' ability to improve policies through feedback in interactive environments with controlled evaluation methodology.

Ax Mona Schirmer, Metod Jazbec, Alexander Timans, Christian Naesseth, Maja Waldron, Eric Nalisnick 29d ago

Online Safety Monitoring for LLMs

Real-time safety monitoring method for LLMs using verifier signals with risk-calibrated thresholding to detect unsafe outputs at deployment.

Ax Josh Hills, Ida Caspary, Asa Cooper Stickland 29d ago

Distributed Attacks in Persistent-State AI Control

Studies attack surface of persistent-state AI coding agents shipping code iteratively across sessions; introduces Iterative VibeCoding for AI control.

Ax Firoz Shaik, Mateus Pican\c{c}o Lima Gomes, Tanvir Aumi, Jingci Wang, Milos Milunovic, Filip Basara, Ivana Jovanovic, Vishwas Suryanarayanan, Neha Nandan Kenkare, Weiyao Xie, Zhipeng Han, Zheng Zhang, Waleed Shahid, Jay Rathi, Russell Scherer, Thong Q. Nguyen, Michael Bentley, Tamara Stankovic, Rasika Chakravarthy, Vishal Chowdhary 29d ago

Office Comprehension Benchmark

Introduces Office Comprehension Benchmark for evaluating LLMs on Word, Excel, PowerPoint comprehension across native file formats and variants.

Ax Esra D\"onmez, Agnieszka Falenska 29d ago

Structuring the Space of Sociotechnical Alignment

Framework for specifying sociotechnical alignment of AI systems, addressing gap between technical and normative aspects of socially desirable behavior.

Ax Kathan Shah 29d ago

Token Geometry

Research on gradient geometry of embedding tables in language models; introduces Ember optimizer for efficient finetuning and pretraining with minimal memory.

Ax Jiatong Li, Samuel Yeh, Sharon Li 29d ago

Multi-Head Recurrent Memory Agents

Multi-Head Recurrent Memory Agents: Addresses reliability degradation in LLM agents with long contexts by decomposing memory capture and retention.

Ax Ege Onur Taga, Yilin Zhuang, M. Emrullah Ildiz, Petros Mol, Abhimanyu Das, Karthik Duraisamy, Samet Oymak 29d ago

Evolutionary Feature Engineering for Structured Data

EFE: Framework using LLM-based evolutionary search to discover preprocessing transformations for structured data as composable Python programs.