Ax Dexter Hadley 23d ago

CANONIC: Governance Is Compilation

CANONIC: system that applies compiler-like governance to LLM-generated content, using formal grammars to admit/reject artifacts into evidence ledgers at scale.

Ax Nishan Pantha, Pranath Reddy Kumbam, Sajil Awale, Pushwitha Krishnappa, Muthukumaran Ramasubramanian, Nidhi Jha, Emily Foshee, Ankur Kumar, Rachel Slank, Ashkbiz Danehkar, Rahul Ramachandran 23d ago

Scientific Code Search at Scale: A Multi-Domain Dataset and Benchmark

Multi-domain scientific code search benchmark with 5,264 curated repositories across scientific computing domains to evaluate code discovery tools.

Ax Edwin H. Wintermute, Harmon Bhasin, Christina M. Agapakis, Dianzhuo Wang, Evan Seeyave, Arjun Banerjee, Daniel Fulop, Matthew C. Watson, Adam J. Meyer, Sandrine Boissel, Jens H. Kuhn, Rishi Jain, Noah D. Taylor, Helena Shomar, Patrick M. Boyle, Kenny Workman 23d ago

Evaluating calibrated refusal and safe usefulness in dual-use biology settings

BioSecBench-Refusal: benchmark for evaluating AI agent safety in biology tasks, pairing 61 routine tasks with 46 red-team scenarios for biosecurity assessment.

Ax Bo Huang, Fengxiang Li, Hao Xu, Haoyang Huang, Hongyi Fu, Jinhua Hao, Kun Yuan, Minglei Zhang, Pengcheng Xu, Shiyang Liu, Wenhao Zhuang, Yuze Shi, Zongxian Feng, Chao Wang, Cheng He, Chongling Rao, Deyu Cao, Fan Yang, Gang Xiong, Haochen Liu, Jiabao Li, Jian Liang, Jinghui Jia, Jingwen Chang, Jun Du, Junyu Shi, Min Li, Mingqi Wu, Qiang Gao, Shangpeng Yan, Shaotong Qi, Shu Xu, Shuo Zhou, Tiankuo Xu, Tong Zheng, Weilun Zhao, Xiancheng Meng, Xianda Sun, Xiaoyu Jiang, Xunhao Jia, Yao Xia, Yimeng Xu, Yinghan Cui, Yingpeng Chen, Yiwen Ning, Yong Wang, Yuxuan Sun, Zhongsheng Liu, Ming Sun, Cheng Luo, Chen Yang, Han Li, Kun Gai 23d ago

KAT-Coder-V2.5 Technical Report

KAT-Coder-V2.5: agentic coding model trained autonomously in executable repositories. Introduces AutoBuilder for sandbox reconstruction and end-to-end agentic post-training framework.

Ax Giulia Lanzillotta, Mandana Samiei, Doina Precup, Razvan Pascanu, Claire Vernade 23d ago

To Retain or to Adapt? Generalizing Continual Learning

Continual learning research challenging retention-centered assumptions and prioritizing real-time adaptation in non-stationary environments.