Ax Aoxiong Zeng, Yuxin Yang, Xiangquan Yang 22d ago

Online Data Selection Is Implicit Alignment

Study showing online data selection during fine-tuning acts as implicit alignment mechanism for LLM behavioral preferences.

Ax Maximilian S. Ernst (Max Planck School of Cognition, Center for Lifespan Psychology Max Planck Institute for Human Development, Machine Learning Group Technische Universit\"at Berlin), Lorenz Linhardt (Machine Learning Group Technische Universit\"at Berlin, Berlin Institute for the Foundations of Learning and Data), Aaron Peikert (Center for Lifespan Psychology Max Planck Institute for Human Development), Oliver Eberle (Machine Learning Group Technische Universit\"at Berlin, Berlin Institute for the Foundations of Learning and Data) 22d ago

Distributed Sparse Interventions in Language Models

Studies causal interventions in language model components to understand task behavior, extending beyond global activation-space steering.

Ax Marcus Williams, Hannah Sheahan, Cameron Raymond, Tomek Korbak, Deng Pan, Peilin Yang, Leon Maksin, Ningyi Xie, Phillip Guo, Ian Kivlichan, Micah Carroll 22d ago

Predicting LLM Safety Before Release by Simulating Deployment

Method for pre-deployment safety evaluation of LLMs by simulating realistic deployments from de-identified conversations to assess failure rates.

Ax Lo\"ic Cabannes, Pierre-Emmanuel Mazar\'e, Gergely Szilvasy, Matthijs Douze, Maria Lomeli, Ilze Amanda Auzina, Justin Carpentier, Gabriel Synnaeve, Herv\'e J\'egou 22d ago

Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity

Research introducing Sparse Delta Memory architecture for scaling linear RNNs; improves long-context recall while reducing FLOPs.

Ax Bojie Li, Noah Shi 22d ago

RLVP: Penalize the Path, Reward the Outcome

Research on reinforcement learning for real-world agents with irreversible interactions; proposes penalizing unsafe paths while rewarding outcomes.

Ax Zeyuan Ding, Wenhai Liu, Yang Xu, Jiayu Hu, Yinda Chen, Yi Zhang, Yong Dai, Jian Tang, Xiaozhu Ju 22d ago

Pelican-VLA 0.5: Attending Before Acting Benefits Generalization

Vision-language-action model integrating perception, future prediction, and action planning with attention-based generalization without task-specific tuning.