Show HN: DriftGuard – response drift detection for LangGraph agents
DriftGuard detects when LangGraph agents drift from intended domain without ground-truth labels.
DriftGuard detects when LangGraph agents drift from intended domain without ground-truth labels.
AMD contributes GPU support to tiny-vLLM for improved LLM inference performance.
PEAK framework uses aspect-oriented programming to automate data recording for deep reinforcement learning agents in game development playtesting.
TUI tool orchestrating long-running coding agents through full engineering workflow with deterministic phase gates.
Open-source MCP server/CLI/API for unified email, SMS, and phone integration with LLM providers and AI agents. Production-tested voicebot tooling.
Anime-style UI for visualizing AI coding agents reviewing each other's code.
Exfault: autonomous AI agents for Android app penetration testing using static/dynamic analysis with adb, jadx, apktool and reverse engineering tools.
Standalone Playwright CLI tool for AI agents to control real Chromium browsers with structured commands, achieving 82% task success at $0.22/task.
ZUSE agent discovers causal laws in cellular automata using deterministic policy-driven loops without LLMs, achieving 69x compression of period-15 oscillator.
Wolli framework creates self-building long-running agents that customize themselves around assigned purposes with integrations.
UGC Agent: AI tool discovering Instagram/TikTok creators, enriching profiles, scoring against ICP, and exporting leads with contact information.
PostgreSQL-compatible databases for AI infrastructure scaling; 83% of leaders expect data infrastructure to fail within 24 months without upgrades.
Media chain's attempt to replace 47 newspapers with AI-generated content failed.
Model Context Protocol standardizes how AI applications connect to external tools and data sources, addressing LLMs' knowledge cutoff limitation.
Sonar generates local-first codebase briefings for non-technical stakeholders without frontier model subscriptions.
Author shares workplace observations on AI utility for software development, noting developers accomplishing previously infeasible tasks at scale.
Fognitix: autonomous browser with parallel AI agents executing up to five concurrent tasks simultaneously; vision-native interaction with synthesized results.
AI code generation enables premature optimization strategies previously impractical; developer shares experience with tech debt resolution through AI-assisted refactoring.
Screenmind runs vision model locally on screenshots for privacy-first alternative to Microsoft Recall, supporting search and chat over screen history.
Telnyx Voice API enables routing inbound phone calls to AI agents via webhook-driven infrastructure with programmatic control.
Formally verified distributed object-capability OS with README written for multiple LLM model tiers.
Reference MCP tool enabling AI agents to search and reference past session data across multiple Claude instances for decision tracking and context retrieval.
Programming language concept where LLM serves as CPU, treating language execution through AI.
Analysis of enterprise AI agent strategies from Salesforce, Microsoft, IBM; discusses evaluating which agents justify building.
DeepSeek V4 LLM announced for mid-July release with dynamic pricing structure based on peak/off-peak hours.
Frontier Code is an AI coding benchmark for evaluating LLM coding capabilities.
Analysis of how agentic software development has reduced development costs and timelines; prototype $30k→$1.8k, market leader $2.6M→$180k.
Discussion of AI startup ecosystem dynamics, funding culture, and prevalence of thin wrapper companies and agent orchestration layers.
Agent-swarm: open-source alternative to Claude Tags enabling control over models, context, and memory instead of blackbox approaches.
Open-source pack of 25 executable skills for AI coding agents with step-by-step workflows to improve agent safety and reliability.
Open benchmark for prompt-injection detectors measuring both attack catch-rate and false positives across thresholds; model-agnostic and reproducible.
Brain.md: open agent-agnostic persistent memory layer storing project knowledge as markdown; lives in repo, CLI-based, travels across agents.
Private Whisper: open-source voice-to-text tool running on-device with no data storage or training on user audio.
Title only: VibeRaven production workflows for AI coding agents.
Practical guide to security risks in AI coding agents; discusses sandboxing, permissions, and safe deployment patterns for Claude Code and similar tools.
Marmot: context layer for agents capturing institutional knowledge (column definitions, database locations) that humans carry but agents need explicit access to.
Analysis of limited productivity gains from generative AI in games industry despite heavy compute subsidies; questions ROI as costs rise.
Sidequest: context-aware side-channel interface for Pi (Claude) enabling threaded questions without interrupting main conversation; tool-capable.
pdf-struct-chunker: LLM-free Rust tool for semantically-aware PDF chunking that preserves document structure for improved RAG performance.
Claude Code plugin enforcing spec-driven development workflow (PRD→Architecture→Plan→Build) with markdown-backed kanban board and git tracking.
OctoPerf MCP: Model Context Protocol server enabling AI agents to drive load testing without API keys; stateless OAuth 2.1 authentication.
Deno-based framework creating sandboxed Telegram agents from single YAML config; includes research paper reader example with Anthropic API.
Technical deep-dive on structuring context for analytics agents: evolved from complex entity graphs to minimal grounding layer for SQL generation.
Kog Laneformer 2B is a latency-optimized LLM model designed for the Kog inference engine.
Claude Code plugin enabling visual feedback annotation on frontend interfaces that feeds back into Claude sessions via Playwright MCP.
Analysis of position bias in on-policy distillation showing degraded supervision as student rollouts diverge from teacher distribution.
Hartley Neural Operator replacing complex FFT with real-valued Discrete Hartley Transform for solving PDEs with reduced redundancy.
Protocol for model forensics to determine whether concerning LLM behavior reflects actual misalignment versus benign causes.
Neural network method combining topology-informed approaches for flood detection using optical and SAR satellite imagery.
Theoretical study of sample complexity and identifiability conditions for learning ODEs from solution data in scientific machine learning.