Ax Alexandra Souly, Robert Kirk, Jacob Merizian, Abby D'Cruz, Xander Davies 4/2/2026

UK AISI Alignment Evaluation Case-Study

UK AISI evaluates whether frontier LLMs sabotage safety research when used as coding assistants, testing alignment and intended behavior in real-world deployment scenarios.

Ax Jiaqi Liu, Zipeng Ling, Shi Qiu, Yanqing Liu, Siwei Han, Peng Xia, Haoqin Tu, Zeyu Zheng, Cihang Xie, Charles Fleming, Mingyu Ding, Huaxiu Yao 4/2/2026

OmniMem: Autoresearch-Guided Discovery of Lifelong Multimodal Agent Memory

OmniMem uses autoresearch to design lifelong multimodal memory systems for extended-horizon AI agents, automating exploration of architecture and retrieval strategies.

Ax Saeid Jamshidi, Foutse Khomh, Arghavan Moradi Dakhel, Amin Nikanjam, Mohammad Hamdaqa, Kawser Wazed Nafi 4/2/2026

Adversarial Moral Stress Testing of Large Language Models

Evaluates ethical robustness of LLMs under sustained adversarial multi-turn interactions using moral stress testing beyond single-round safety benchmarks.

Ax Esakkivel Esakkiraja, Sai Rajeswar, Denis Akhiyarov, Rajagopal Venkatesaramani 4/2/2026

Therefore I am. I Think

Analyzes whether LLM reasoning models make decisions before or after chain-of-thought reasoning using linear probes on activation patterns for tool-calling decisions.

Ax Zhe Yang, Shulin Tian, Kairui Hu, Shuai Liu, Hoang-Nhat Nguyen, Yichi Zhang, Zujin Guo, Mengying Yu, Zinan Zhang, Jingkang Yang, Chen Change Loy, Ziwei Liu 4/2/2026

HippoCamp: Benchmarking Contextual Agents on Personal Computers

HippoCamp benchmark evaluates multimodal AI agents on personal computer file management tasks requiring context-aware reasoning in user-centric environments.

Ax Eftychia Makri, Nikolaos Nakis, Laura Sisson, Gigi Minsky, Leandros Tassiulas, Vahid Satarifard, Nicholas A. Christakis 4/2/2026

Benchmark for Assessing Olfactory Perception of Large Language Models

Introduces Olfactory Perception benchmark with 1,010 questions to assess LLM reasoning capabilities about smell across odor classification and related tasks.

Ax Jaeik Kim, Woojin Kim, Jihwan Hong, Yejoon Lee, Sieun Hyeon, Mintaek Lim, Yunseok Han, Dogeun Kim, Hoeun Lee, Hyunggeun Kim, Jaeyoung Do 4/2/2026

Dynin-Omni: Omnimodal Unified Large Diffusion Language Model

Dynin-Omni: First masked-diffusion omnimodal foundation model unifying text, image, speech, and video in single architecture.

Ax Gabriel U. Talasso, Meghdad Kurmanji, Allan M. de Souza, Nicholas D. Lane, Leandro A. Villas 4/2/2026

Task-Centric Personalized Federated Fine-Tuning of Language Models

Personalized federated learning approach for fine-tuning language models on heterogeneous distributed tasks while maintaining individual client performance.

Ax Leonardo Medrano Sandonas, David Balcells, Anton Bochkarev, Jacqueline M. Cole, Volker L. Deringer, Werner Dobrautz, Adrian Ehrenhofer, Thorben Frank, Pascal Friederich, Rico Friedrich, Janine George, Luca Ghiringhelli, Alejandra Hinostroza Caldas, Veronika Juraskova, Hannes Kneiding, Yury Lysogorskiy, Johannes T. Margraf, Hanna T\"urk, Anatole von Lilienfeld, Milica Todorovi\'c, Alexandre Tkatchenko, Mariana Rossi, Gianaurelio Cuniberti 4/2/2026

Perspective: Towards sustainable exploration of chemical spaces with machine learning

Perspective on sustainability challenges in AI-driven molecular discovery pipelines including computation and data generation costs.

Ax Patrice Bechard, Orlando Marquez Ayala, Emily Chen, Jordan Skelton, Sagar Davasam, Srinivas Sunkara, Vikas Yadav, Sai Rajeswar 4/2/2026

Terminal Agents Suffice for Enterprise Automation

Analysis showing terminal agents suffice for enterprise automation tasks compared to complex tool-augmented or web-based agentic systems.