Ax Charles Ye, Jasmine Cui, Dylan Hadfield-Menell 3/16/2026

Prompt Injection as Role Confusion

Analysis of prompt injection vulnerabilities traced to role confusion in LLMs, with novel detection probes and mitigation approaches.

Ax Brian Zhang, Deepti Guntur, Zhiyang Zuo, Abhinav Sharma, Shreyas Chaudhari, Wenlong Zhao, Franck Dernoncourt, Puneet Mathur, Ryan Rossi, Nedim Lipka 3/16/2026

Test-Time Strategies for More Efficient and Accurate Agentic RAG

Test-time optimization strategies for agentic RAG systems to reduce inefficient retrieval and improve accuracy on complex multi-hop questions.

Ax Zheda Mai, Ke Zhang, Fu-En Wang, Zixiao Ken Wang, Albert Y. C. Chen, Lu Xia, Min Sun, Wei-Lun Chao, Cheng-Hao Kuo 3/16/2026

Revisiting Model Stitching In the Foundation Model Era

Research on connecting layers across Vision Foundation Models (CLIP, etc.) to measure representational compatibility via stitching.

Ax Alexander K Taylor, Junyi Zhang, Ethan Ji, Vigyan Sahai, Haikang Deng, Yuanzhou Chen, Yifan Yuan, Di Wu, Jia-Chen Gu, Kai-Wei Chang, Nanyun Peng, Amit Sahai, Wei Wang 3/16/2026

TaoBench: Do Automated Theorem Prover LLMs Generalize Beyond MathLib?

TaoBench evaluates generalization of automated theorem prover LLMs beyond MathLib using novel definitional frameworks.

Ax Yichen Zhang, Da Peng, Zonghao Guo, Zijian Zhang, Xuesong Yang, Tong Sun, Shichu Sun, Yidan Zhang, Yanghao Li, Haiyan Zhao, Wang Xu, Qi Shi, Yangang Sun, Chi Chen, Shuo Wang, Yukun Yan, Xu Han, Qiang Ma, Wei Ke, Liang Wang, Zhiyuan Liu, Maosong Sun 3/16/2026

Cheers: Decoupling Patch Details from Semantic Representations Enables Unified Multimodal Comprehension and Generation

arXiv paper presenting Cheers, unified multimodal model decoupling patch details from semantics for image comprehension and generation.