Ax Ning Yang, Hai Lin, Yibo Liu, Baoliang Tian, Guoqing Liu, Haijun Zhang 3/3/2026

Token-Importance Guided Direct Preference Optimization

Token-Importance Guided DPO: Enhanced direct preference optimization for LLM alignment using token-level importance weighting beyond standard DPO.

Ax Mikhail Terekhov, Zhen Ning David Liu, Caglar Gulcehre, Samuel Albanie 3/3/2026

Control Tax: The Price of Keeping AI in Check

Control Tax: Framework for measuring operational and financial costs of integrating AI control mechanisms into agentic AI pipelines for high-stakes applications.

Ax Cong Chen, Omer Karaduman, Xu Kuang 3/3/2026

Behavioral Generative Agents for Energy Operations

Study of generative agents for modeling consumer behavior in energy operations, examining their role in operational decision-making and uncertainty handling.

Ax Linhao Luo, Zicheng Zhao, Junnan Liu, Zhangchi Qiu, Junnan Dong, Serge Panev, Chen Gong, Thuy-Trang Vu, Gholamreza Haffari, Dinh Phung, Alan Wee-Chung Liew, Shirui Pan 3/3/2026

G-reasoner: Foundation Models for Unified Reasoning over Graph-structured Knowledge

G-reasoner foundation model enables unified reasoning over graph-structured knowledge, combining LLM capabilities with structured knowledge integration.

Ax Hanane Nour Moussa, Patrick Queiroz Da Silva, Daniel Adu-Ampratwum, Alyson East, Zitong Lu, Nikki Puccetti, Mingyi Xue, Huan Sun, Bodhisattwa Prasad Majumder, Sachin Kumar 3/3/2026

ScholarEval: Research Idea Evaluation Grounded in Literature

ScholarEval: retrieval-augmented evaluation framework assessing research ideas on soundness and contribution using literature grounding.

Ax Annan Li, Chufan Wu, Zengle Ge, Yee Hin Chong, Zhinan Hou, Lizhe Cao, Cheng Ju, Jianmin Wu, Huaiming Li, Haobo Zhang, Shenghao Feng, Mo Zhao, Fengzhi Qiu, Rui Yang, Mengmeng Zhang, Wenyi Zhu, Yingying Sun, Quan Sun, Shunhao Yan, Danyu Liu, Dawei Yin, Dou Shen 3/3/2026

The FM Agent

FM Agent: multi-agent framework combining LLM reasoning with large-scale evolutionary search for complex scientific and engineering discovery tasks.

Ax Elinor Poole-Dayan, Jiayi Wu, Taylor Sorensen, Jiaxin Pei, Michiel A. Bakker 3/3/2026

Benchmarking Overton Pluralism in LLMs

OVERTONBENCH framework measures viewpoint diversity in LLM outputs through set coverage metrics, validated against 1208-person human study across 8 models.