HN gagarwal123 5/7/2026

Ask HN: Real life autonomous AI Agents

Discussion thread asking for real-world autonomous AI agent deployments and use cases, distinguishing between true agents and workflow automation.

HN galsapir 5/7/2026

The Comparator in Clinical AI

Critical analysis of OpenAI o1's clinical diagnostic performance claims, examining evaluation methodology and comparator bias.

HN rcarmo 5/7/2026

Notes on GPT 5.x Model Regressions

Analysis of GPT-5.5 code generation regressions where the model makes unrequested changes and improvements to unrelated code.

LB github.com via gcv 5/7/2026

The Agent Harness Framework

Flue is a TypeScript framework for building AI agents with a built-in harness, designed as a headless, programmable alternative to Claude Code.

Ax Masafumi Oyamada, Kunihiro Takeoka, Kosuke Akimoto, Ryoma Obara, Masafumi Enomoto, Haochen Zhang, Daichi Haraguchi, Takuya Tamura 5/7/2026

cotomi Act: Learning to Automate Work by Watching You

Cotomi Act browser agent learning from user behavior observation with adaptive execution achieving 80.4% task success via multi-step execution.

Ax Zirui Tang, Xuanhe Zhou, Yumou Liu, Linchun Li, Weizheng Wang, Hongzhang Huang, Jun Zhou, Jiachen Song, Shaoli Yu, Jinqi Wang, Zihang Zhou, Hongyi Zhou, Yuting Lv, Jinyang Li, Jiashuo Liu, Ruoyu Chen, Chunwei Liu, GuoLiang Li, Jihua Kang, Fan Wu 5/7/2026

Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies

Benchmark for evaluating AI agents on workspace tasks requiring reasoning over file dependencies in realistic work environments.

Ax Robert Gieselmann, Henrike von Huelsen, Mihai Samson, Marie-Christine Meyer, Dariusz Piotrowski, Oleksandr Radomskyi, Justin Okamoto, Turan Gojayev, Michael Painter, Gavin Brown, Federico Pecora, Jeremy L. Wyatt 5/7/2026

Self-Improvement for Fast, High-Quality Plan Generation

Self-improvement approach for generating high-quality plans in sub-exponential time using transformer-based generative models.