Show HN: Captain Claw local AI agent, 29 tools, multi-session, DAG orchestration
Captain Claw: local AI agent runtime with 29 tools, multi-session DAG orchestration, SQLite persistence, supporting OpenAI/Anthropic/Gemini/Ollama.
Captain Claw: local AI agent runtime with 29 tools, multi-session DAG orchestration, SQLite persistence, supporting OpenAI/Anthropic/Gemini/Ollama.
Rocky AI: Checkly's GA AI agent integrated into SaaS product; lessons on building AI agents beyond chat interfaces for monitoring workflows.
YSA: containerized sandbox for running Claude CLI on production codebases with network isolation, TLS inspection, and data loss prevention.
Thought Canvas: mind-mapping interface for exploring ideas with AI instead of linear chat format.
Tool for constraining AI assistants to understand codebase architecture and prevent bad architectural decisions during development.
Semantic code search tool supporting 35 languages with call graph understanding, built in Rust with MCP support.
Open-source tool automatically improving and explaining LLM prompt optimizations with educational guidance.
Marketplace offering pre-built AI assistant personas with playbooks, tool integrations and deployment guides.
Experiment platform allowing users to use LLMs to rewrite frontend code, exploring malleable software concepts.
Analysis of how agentic AI systems like Claude Code are displacing specialized legal tech through flexibility and generalization.
Open-source AI desktop character with memory, personality and emotion systems that track user behavior and conversations.
AI agent rewrote LGPL chardet library as MIT-licensed drop-in replacement with improved performance and accuracy.
Framework giving AI coding agents persistent memory management (archivist role) to reduce token waste and codebase rediscovery.
GPT-5.4 release announcement for professional work, including coding capabilities and agentic workflows across ChatGPT, API, and Codex.
Research on reasoning model chain-of-thought control limitations in frontier models and their implications for AI agent safety oversight.
GPT-5.4 Thinking system card detailing safety mitigations for reasoning model, including cybersecurity safeguards similar to previous GPT-5 series models.
Open-source Next.js platform for document management with RAG, embeddings, and predictive analysis using modular architecture and RBAC.
Analysis of AI agent failures in production including behavioral drift, hallucination, and lack of accountability infrastructure with proposed solutions.
Single-file HTML multi-agent AI workspace with no backend, emphasizing local execution, data privacy, and AI sovereignty.
Framework for using LLMs beyond conversational assistants for cognitive auditing and asymmetric execution philosophy.
Autonomous coding agent system that monitors project boards, spawns agents for tasks, provides CI feedback and PR management without human supervision.
Local LLM runtime enabling training and inference on Apple Neural Engine (NPU) without CoreML or GPU, runs offline on 2B+ Apple devices.
Personalized coding education platform with 24/7 AI teacher that adapts to student pace and goals.
Local document indexing tool for AI agents supporting PDF, DOCX, Markdown via CLI/MCP protocol with privacy-first design.
Empirical LLM model comparison data and performance statistics from Strix testing with observations on different models.
Analysis of open-source relicensing challenges and case study of chardet using AI-assisted code rewriting.
NeuroPareto: multi-objective optimization architecture for high-dimensional search with Bayesian uncertainty estimation and calibrated acquisition.
Analyzes gradient issues in RLVR methods (GRPO variants) for LLM reasoning tasks, proposing improvements to reinforcement learning with verifiable rewards.
TIME: new benchmark suite for evaluating time series foundation models, addressing data quality and task formulation limitations in existing benchmarks.
JPmHC proposes orthogonal hyper-connections to maintain identity mapping in residual networks, addressing training instability and scalability issues.
AOT-SFT: adversarial dataset and training method to improve robustness of multimodal large language models against visually complex scenes.
Analyzes temperature parameter selection in knowledge distillation, examining dependencies on optimizer and teacher pretraining.
Studies RL algorithms for MDPs with exogenous dynamics where only subset of state variables are directly controlled.
Reviews reward function design for reinforcement learning in autonomous driving contexts with conflicting objectives.
Proposes LoCo-RLHF framework for RLHF with heterogeneous human feedback using contextual information to improve LLM alignment.
Uses MLLMs to improve text-to-image models' ability to generate images with rich object interactions, addressing dataset limitations for rare interactions.
Benchmark evaluating LLM knowledge accuracy on UK government public health information, testing domain-specific reliability.
EG-MRSI architecture integrating metacognition, emotion-based motivation, and recursive self-modification for AI agent self-improvement.
Framework for out-of-distribution detection and segmentation in multimodal safety-critical applications using outlier synthesis.
TADA framework for targeted diffusion-based image augmentation that selectively generates synthetic data to improve classifier generalization efficiently.
Study of in-context learning biases in LLMs through supervised learning lens, proposing decision boundary adjustment for classification calibration.
Model predictive control framework combining Q-learning guidance and Stein variational inference with RL-informed policy priors.
Fast Equivariant Imaging framework for unsupervised deep network training without ground-truth data using Lagrangian optimization and denoisers.
Interpretability study of in-context learning mechanisms in LLMs using off-by-one addition task with circuit analysis.
ObfusQAte framework and ObfusQA benchmark to evaluate LLM robustness on obfuscated factual question-answering tasks.
Python package for physics simulations of quantum dot devices, addresses ML dataset collection challenges for quantum device calibration and operation.
Action-prompted video segmentation framework for embodied AI that handles label noise and multimodal inconsistencies in object interaction segmentation.
Analysis of best-of-N ensemble selection for LLMs using majority voting at infinite limit, with adaptive generation scheme to reduce inference cost.
Geometric framework modeling LLM reasoning as flows in representation space to study how models internalize logical structure.
ToMCLIP: method for topological alignment of vision-language embedding space in multilingual contrastive models.