HN deadalus 4/24/2026

Grok Voice Think Fast 1.0

xAI announces grok-voice-think-fast-1.0, voice agent API for customer support and enterprise workflows.

Ax Haebin Seong, Li Yin, Haoran Zhang 4/24/2026

The Last Harness You'll Ever Build

arXiv paper on AI agents handling complex domain-specific workflows, addressing harness engineering challenges for enterprise automation tasks.

Ax Jindi Guo, Xi Fang, Chaozheng Huang 4/24/2026

Can MLLMs "Read" What is Missing?

Benchmark evaluating multimodal LLM ability to reconstruct masked text from visual context in documents and webpages without explicit prompts.

Ax Donggyu Lee, Hyeok Yun, Jungwon Kim, Junsik Min, Sungwon Park, Sangyoon Park, Jihee Kim 4/24/2026

Ideological Bias in LLMs' Economic Causal Reasoning

Systematic evaluation of ideological bias in LLM economic causal reasoning using extended benchmark with ideology-contested intervention cases.

Ax Jian Cui, Zhiyuan Ren, Desheng Weng, Yongqi Zhao, Gong Wenbin, Yu Lei, Zhenning Dong 4/24/2026

ReaGeo: Reasoning-Enhanced End-to-End Geocoding with LLMs

End-to-end geocoding framework using LLMs with geohash sequence reformulation to overcome limitations of traditional multi-stage retrieval approaches.

Ax Fridolin Linder, Thomas Leeper, Daniel Haimovich, Niek Tax, Lorenzo Perini, Milan Vojnovic 4/24/2026

Unbiased Prevalence Estimation with Multicalibrated LLMs

Method using multicalibration to correct LLM prevalence estimation errors across population shifts, addressing diagnostic accuracy in varying contexts.

Ax Maximilian Westermann, Ben Griffin, Aaron Ontoyin Yin, Zakari Salifu, Yagiz Ihlamur, Kelvin Amoaba, Joseph Ternasky, Fuat Alican, Yigit Ihlamur 4/24/2026

CoFEE: Reasoning Control for LLM-Based Feature Discovery

CoFEE provides structured reasoning control for LLM-based feature discovery from unstructured data while avoiding leakage and confounding signals.

Ax Guangxiang Zhao, Qilong Shi, Xusen Xiao, Xiangzheng Zhang, Tong Yang, Lin Sun 4/24/2026

Thinking with Reasoning Skills: Fewer Tokens, More Accuracy

Method to distill reusable reasoning skills from LLM deliberation and retrieve them at inference time, reducing token consumption while improving accuracy.

Ax Nathanael Jo, Zoe De Simone, Mitchell Gordon, Ashia Wilson 4/24/2026

Alignment has a Fantasia Problem

Examines misalignment when LLMs treat incomplete prompts as full intent expressions, arguing behavioral research shows users engage with systems before goals solidify.