Ax Giulio Frey, Kawin Ethayarajh 5/18/2026

Mecha-nudges for Machines

Studies mecha-nudging: how internet environments can be systematically modified to influence AI agent decisions without degrading human decision-making.

Ax Aleksandr Bowkis, Marie Davidsen Buhl, Jacob Pfau, Geoffrey Irving 5/18/2026

Automated alignment is harder than you think

Argues that automating alignment research with AI agents risks producing misleading safety assessments even without deliberate deception due to fundamental misunderstandings.

Ax Moein Hasani, Hamidreza Shahidi, Trace Levinson, Yuan Zhong, Guanghua Shu, Vinesh Gudla, Tejaswi Tenneti 5/18/2026

A Cascaded Generative Approach for e-Commerce Recommendations

Cascaded generative approach for personalized e-commerce recommendations that assembles dynamic storefronts with semantic cohesion across placements.

Ax Ziyu Liu, Zeyi Sun, Yuhang Zang, Wei Li, Pan Zhang, Xiaoyi Dong, Yuanjun Xiong, Dahua Lin, Jiaqi Wang 5/18/2026

RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition

RAR framework combines CLIP's broad recognition with MLLMs' fine-grained classification via retrieval and ranking for improved visual recognition.

Ax Yue Liu, Xiaoxin He, Miao Xiong, Jinlan Fu, Shumin Deng, Yingwei Ma, Jiaheng Zhang, Bryan Hooi 5/18/2026

FlipAttack: Jailbreak LLMs via Flipping

FlipAttack exploits left-to-right text comprehension weakness in black-box LLMs using left-side noise to disguise harmful prompts.

Ax ChonLam Lao, Jiaqi Gao, Jiamin Cao, Zhipeng Zhang, Pengcheng Zhang, Jiangfei Duan, Zhilong Zheng, Yu Guan, Yichi Xu, Yong Li, Zhengping Qian, Aditya Akella, Minlan Yu, Ennan Zhai, Dennis Cai, Jingren Zhou 5/18/2026

TrainMover: An Interruption-Resilient Runtime for ML Training

TrainMover runtime enables interruption-resilient LLM training using elastic/standby machines with minimal downtime and zero memory overhead.

Ax Yash Akhauri, Ahmed F AbouElhamayed, Yifei Gao, Chi-Chih Chang, Sameh Gobriel, Nilesh Jain, Mohamed S. Abdelfattah 5/18/2026

TokenButler: Token Importance is Predictable

TokenButler predicts token importance in KV-Cache to identify critical tokens and reduce memory/computation bottlenecks in LLM inference.

Ax Andrea W Wen-Yi, Unso Eun Seo Jo, David Mimno 5/18/2026

Do Chinese models speak Chinese languages?

Comparative analysis of multilingual capabilities in Chinese open-weight LLMs versus US/European models, examining pre-training data curation strategies.

Ax Yuyang Xu, Renjun Hu, Haochao Ying, Jian Wu, Xing Shi, Wei Lin 5/18/2026

Large Language Models Could Be Rote Learners

Study investigates benchmark contamination in LLM evaluation, showing models can achieve inflated performance through rote learning of test data.

Ax Prabhat Nagarajan, Martha White, Marlos C. Machado 5/18/2026

Deep Double Q-learning

arXiv paper introducing Double DQN improvements to address maximization bias in deep reinforcement learning with decoupled action-selection and evaluation.