Show HN: Signal Core - contracts, approvals, and receipts for OpenClaw agents
Governance layer for AI agents with approvals and receipts. Limited documentation; appears to be early-stage proprietary software.
Governance layer for AI agents with approvals and receipts. Limited documentation; appears to be early-stage proprietary software.
Unified YAML tool definition format supporting MCP, OpenAI, and Anthropic agent specifications. Developer tool.
Guide for building and maintaining personal AI agents. Limited details provided.
Personal AI system with encrypted persona vaults and permission layer for agent access to sensitive data. Open source implementation.
Compliance frameworks (SoC 2, ISO 27001, HIPAA) for AI agents in production environments.
Open-source multi-provider AI coding agent CLI supporting Claude, GPT, and local models with custom tool support.
Open-source terminal coding agent supporting multiple LLM backends (Ollama, OpenAI, Anthropic, Deepseek). Includes 17 tools for file operations, bash, web search, and task management.
Rust implementation of TurboQuant quantization for GGUF models. Addresses compound quantization errors from applying KV cache quantization to already-quantized weights, fixing issues like language mixing and hallucination.
HN discussion about client taking over development with Claude Code after developer introduced LLM tools.
Flemma: Full LLM chat client built as Neovim filetype plugin for in-editor AI workload management with prompt engineering.
HN discussion questioning whether massive foundation models are cost-efficient versus specialized smaller models for developer tools.
Crossmind-CLI fetches data from X, Reddit, HN, GitHub and compresses output for LLM efficiency, reducing X responses from 3900 to 666 tokens.
Discussion of vector search limitations for LLM context retrieval. Title only; no content provided.
Clanker CLI tool for DevOps automation using AI agents.
Browserbeam is a browser API optimized for AI agents to interact with web pages more efficiently, reducing token waste and improving page understanding.
Calx tracks and compiles corrections humans provide to AI agents to improve performance, analyzing patterns in agent mistakes across 82K LOC.
Case study: generated 60-page ERP knowledge base using AI in 24 hours.
Using Gemini API documentation MCP and agent skills to improve coding agent performance.
ChartDiff: 8,541 annotated chart-pair benchmark for evaluating cross-chart comparative reasoning in LLMs and vision systems.
Theoretical framework using category theory to formally define, compare, and analyze artificial general intelligence approaches.
Study documenting emergent social organization patterns in hierarchical multi-agent AI systems using thermodynamic and evolutionary frameworks.
World-Action Model: Action-regularized world model jointly predicting visual observations and actions for improved representation learning in RL.
Mimosa: Evolving multi-agent framework automatically synthesizing and refining task-specific LLM workflows for autonomous scientific research through experimental feedback.
Large-scale experiment: Self-organizing multi-agent LLM systems across 25,000 tasks spontaneously develop specialized roles outperforming designed hierarchies.
WebVoyager evaluation audit identifying shortcomings in AI agent evaluation practices, proposing robust and transparent methodologies for web agents.
Theoretical perspective arguing AI research should shift from individual models toward multi-agent systems for innovation and scientific discovery.
PAR²-RAG: LLM system combining planned retrieval with active reasoning for multi-hop question answering, adapting queries based on intermediate evidence.
GISTBench: Benchmark evaluating LLM ability to understand user interests from interaction histories in recommendation systems via evidence-based verification.
SciVisAgentBench: Comprehensive benchmark for evaluating LLM-based scientific visualization agents on multi-step data analysis tasks.
REFINE: Locally deployable multi-turn LLM system providing interactive, personalized formative feedback at scale for learning environments.
LLMs used to develop databases on viruses and marine toxins for medical countermeasure research using ChatGPT and Grok.
SimMOF: LLM-based AI agent automating metal-organic framework simulations by handling workflow construction, parameter selection, and tool interoperability for computational chemistry.
Webscraper: Framework using multimodal LLMs to autonomously navigate dynamic web applications and extract content via tool invocation.
AEC-Bench: Multimodal benchmark evaluating AI agents on architecture, engineering, construction tasks requiring drawing understanding and cross-sheet reasoning.
Analysis of how routing-style meta prompts affect LLM internal states and activate sparse computation, testing sparsity-certainty hypothesis.
Xuanwu VL-2B: Industrial multimodal foundation model addressing fine-grained visual perception and long-tail noise for content moderation.
Reliability science framework for evaluating long-horizon LLM agents beyond pass@1, introducing metrics for consistent success across task duration.
Mechanistic study of grokking in modular arithmetic examining global structural evolution driving sudden generalization in neural networks.
PSPA-Bench: Personalized benchmark for smartphone GUI agents capturing user-specific workflows and preferences in mobile task automation.
Nomad: Autonomous system for exploring document/database corpora and discovering insights without human query constraints using agentic exploration.
BenchScope tool measuring effective dimensionality of benchmark scores, revealing significant redundancy across 22 benchmarks and 8,400+ evaluations.
Methods for generating human-interpretable explanations for tree ensemble model predictions to improve trustworthiness.
Study of LLM-generated prior authorization letters showing strong clinical content but weak administrative compliance and formatting.
Analysis showing ELT-Bench underestimates agent capabilities; AI agents demonstrate higher success on data pipeline construction when benchmark quality issues are corrected.
Metric for evaluating explanation quality in attribution methods using graph-based structural compactness analysis.
PRoSFI: Method using formal intermediaries to generate formally verifiable step-by-step logic reasoning in LLMs with structured validation.
FlowPIE: Framework for scientific idea generation coupling literature retrieval and generation as co-evolving process for autonomous research.
ASI-Evolve: Agentic framework enabling AI systems to autonomously conduct AI research through learn-design-experiment-analyze cycles.
Framework for analyzing and compiling agent conversation traces with structured content like tool calls, reasoning blocks, and sub-agent invocations.
Research on how LLM use affects metacognitive accuracy and task performance, challenging the Dunning-Kruger amplification narrative.