Ax George Pu, Michael S. Lee, Udari Madhushani Sehwag, David J. Lee, Bryan Zhu, Yash Maurya, Mohit Raghavendra, Yuan Xue, Samuel Marc Denton 3/23/2026

LHAW: Controllable Underspecification for Long-Horizon Tasks

Framework for managing ambiguity in long-horizon workflow agents. Task-agnostic approach for curating and measuring impact of underspecified instructions on agent execution.

Ax Andrew Seohwan Yu, Mohsen Hariri, Kunio Nakamura, Mingrui Yang, Xiaojuan Li, Vipin Chaudhary 3/23/2026

Medical Image Spatial Grounding with Semantic Sampling

Study of vision language models for spatial grounding in 3D medical imaging. Examines VLM performance across imaging modalities and slice directions.

Ax Yihao Zhang, Zeming Wei, Xiaokun Luan, Chengcan Wu, Zhixin Zhang, Jiangrong Wu, Haolin Wu, Huanran Chen, Jun Sun, Meng Sun 3/23/2026

ClawWorm: Self-Propagating Attacks Across LLM Agent Ecosystems

Security research on ClawWorm, self-propagating attacks across multi-agent LLM ecosystems. First study of attack propagation in interconnected agent systems like OpenClaw.

Ax Chun-Jui Wang, Jian-Ting Guo, Hung Guei, Chung-Chin Shih, Ti-Rong Wu, I-Chen Wu 3/23/2026

Evaluating Game Difficulty in Tetris Block Puzzle

Research using Stochastic Gumbel AlphaZero to evaluate game difficulty in Tetris Block Puzzle variants. Applies game-playing AI as evaluation metric.

LB tombedor.dev by wils124 3/22/2026

Is Local the Future of AI?

Analysis of local open-source AI models as alternative to datacenter-dependent systems. Discusses performance parity with frontier models within 6 months.

HN flippyhead 3/22/2026

ClawMem

ClawMem: On-device memory system for Claude Code and AI agents with retrieval-augmented search, MCP server, no cloud dependencies. Hybrid retrieval architecture.

HN azhenley 3/22/2026

Teaching Claude to QA a mobile app

Building QA system for mobile app using Claude API. Demonstrates LLM application with content filtering issues.