Ax Johannes Schmitt, Tim Gehrunger, Jasper Dekoninck, Gergely B\'erczi, Uri Kreitner, Liam Price, David Holmes 16d ago

ProofCouncil: An LLM Agent for Solving Open Mathematical Problems

LLM agent with author-critic architecture for solving open mathematical problems through agentic workflow inspired by mathematical practice.

Ax Jiayu Yao, Yiwei Wang, Anmeng Zhang, Zhe Sun, Songsong Wang, Lingrui Mei, Yuyao Ge, Shenghua Liu 16d ago

Multimodal Reward Hacking in Reinforcement Learning

Study of reward hacking in multimodal LLM reinforcement learning across VQA and safety tasks with varying reward designs and model scales.

Ax Ian Colbert, Eashan Dash, Pablo Monteagudo-Lago, Juan Amboage, Srinidhi N, Giuseppe Franco, Nicholas J. Fraser, Arun Ramachandran 16d ago

Signed Symmetric Quantization for Few-Bit Integers

Quantization technique for few-bit integer representation addressing asymmetric clipping issues in signed symmetric quantizers.

Ax Hyunjin Seo, Hyeon Hwang, Gyubok Lee, Jay Shin, Jimin Park, Taesoo Kim, Sanghoon Lee, Hongjoon Ahn, Sungjun Han, Sangwon Jung 16d ago

TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology

Unified corpus for biology domain combining heterogeneous biological databases and resources for pretraining specialized large language models.

Ax Sunshine Jiang, John Marangola, David Zhang, Raghuram Kowdeed, Ruiyang Luo, Nitish Dashora, Richard Li, Pulkit Agrawal, Zhang-Wei Hong 16d ago

Prompt-Driven Exploration

Method using LLMs and VLAs as exploration guidance in reinforcement learning to escape weak policies through natural language prompts.

Ax Yuri Ishitoya, Jeremy Siburian, Masashi Hamaya, Kuniaki Saito, Cristian C. Beltran-Hernandez, Mai Nishimura 16d ago

CLAP: Direct VLM-to-VLA Adaptation via Language-Action Grounding

Research on adapting pretrained vision-language models to vision-language-action models for robotics with minimal architectural changes to preserve VLM contributions.

Ax Viraaji Mothukuri, Reza M. Parizi 16d ago

The Patchwork Problem in LLM-Generated Code

Identifies structural incoherence in LLM-generated code where locally valid patches fail globally due to missing configs, imports, or authentication guards.