Ax Masum Hasan, Junjie Zhao, Ehsan Hoque 29d ago

HAL: Inducing Human-likeness in LLMs with Alignment

HAL framework aligns LLMs to conversational human-likeness using interpretable, data-driven methods rather than relying solely on scale or broad supervised training.

Ax Xiaoxi Li, Wenxiang Jiao, Jiarui Jin, Haoxuan Li, Hao Wang, Shijian Wang, Guanting Dong, Jiajie Jin, Yinuo Wang, Yuan Lu, Ji-Rong Wen, Zhicheng Dou, Zhouchen Lin 29d ago

OmniGAIA: Towards Native Omni-Modal AI Agents

OmniGAIA is a benchmark for evaluating omni-modal AI agents with vision, audio, and language capabilities for complex reasoning and tool usage tasks.

Ax Drew Prinster, Clara Fannjiang, Ji Won Park, Kyunghyun Cho, Anqi Liu, Suchi Saria, Samuel Stanton 29d ago

Conformal Policy Control

Conformal Policy Control uses safe reference policies to regulate untested agent behaviors, balancing exploration and safety constraints in high-stakes environments.

Ax Eden Saig, Tamar Garbuz, Ariel D. Procaccia, Inbal Talgam-Cohen, Jamie Tucker-Foltz 29d ago

Adaptive Contracts for Cost-Effective AI Delegation

Adaptive contracts framework for cost-effective AI delegation balancing evaluation noise and costs in pay-for-performance tasks.

HN vaishcodescape 29d ago

Using AI Agents with Databases

Data-Spear: autonomous SQL agent for PostgreSQL that plans queries, verifies results, and cites sources.

HN robinhouston 29d ago

The Ramanujan Challenge for AI

Benchmark dataset for evaluating AI mathematical reasoning using formulas for mathematical constants. Research evaluation tool.

HN Gabrieliam42 29d ago

AI Agent Qubitz

Local-first AI agent using 7B-35B GGUF models with specialized harness for instruction following and tool orchestration. Open source project.