Kars – treat every AI agent as untrusted code, on Kubernetes
Kars: Kubernetes tool treating AI agents as untrusted code with security sandboxing.
Kars: Kubernetes tool treating AI agents as untrusted code with security sandboxing.
Meko is a memory system for AI agents that maintains persistent context across different coding tools and models.
Context-clipper: zero-dependency priority min-heap data structure for managing LLM context window allocation.
Opinion piece arguing AI agents are more useful as shell replacements than as code programmers, requiring less auditing overhead.
Curated infrastructure resource collection for production AI agents covering runtimes, security, observability, and evaluation tools.
ML research on replicating expert judgment in finance using machine learning techniques.
AMD Ryzen AI Halo $4k dev kit with Zen 5 processor, 128GB RAM, NPU for streamlined AI development with ROCm.
Report summary on AI engineering trends and acceleration patterns in 2026.
File encryption tool as security mechanism against local AI agent threats using Supabase, Next.js and Tauri.
C/C++ coding harness for AI models with GDB, sanitizers, perf integration supporting multiple LLM providers.
Framework to autonomously scaffold and operate SaaS applications using AI agents.
Experimental test comparing token efficiency of simplified prompting (caveman speak) vs standard prompting for AI agents.
Controlled benchmark comparing gated vs ungated AI patch generation on SWE-bench Lite with reproducible methodology and scripts.
Open-source fail-closed firewall for constraining AI agent runtimes and preventing unauthorized actions.
AI agent pipeline with 13 specialized agents that generates full-stack web applications from Markdown specifications. Multi-agent system for code generation.
Tool for orchestrating multiple parallel AI coding agents (Claude, Codex) with visual map interface to track agent status, hierarchy, and progress.
Dataset tracking 1,397 AI agents; analysis showing 43 popular agents with 402k stars have become inactive. Original data on AI agent ecosystem health.
Empirical study of AI agents' ability to fix real security vulnerabilities. Controlled experiment measuring agent effectiveness on actual bugs.
Desktop app for running, comparing, and benchmarking local/remote LLMs with support for AI agent integration and benchmark workflows.
Developer tool routing local LLM queries first, then bursting to encrypted end-to-end LLM services. Open source project with practical privacy design.
Open-source control plane for managing LLM deployments. Infrastructure tool for LLM operations and orchestration.
Opinion piece arguing LLMs may reduce creative output by returning statistically average results. Conceptual critique without data or evidence.
Framework-agnostic checkpoint system for AI agents with deny-by-default permissions, role-based access control, and cryptographic audit trails.
Browser-based strategy game simulating AI race between US and China with hidden alignment state to explore verification challenges.
LLMs as tools to reduce UI/design system drift in software engineering. Discusses both problems and solutions using LLM-supported approaches.
Technical whitepaper on improving AI-generated API test suite quality through model orchestration, fine-tuned judgment, and adaptive coverage systems.
AI-generated text detection tool using linguistic pattern analysis. Claims 99.5% accuracy across multiple LLM models.
FlowerBench is a benchmark for evaluating AI agents on real, long-horizon enterprise workflows with access to proprietary tools and context.
Discussion asking about security benchmarks for LLMs, mentions eyeballvull and need for agent-based scanning benchmarks.
Gaia is an open-source registry for verifying AI agent capabilities through public ledger of code-execution runs, license checks, and security audits.
SvelteChatKit is a provider-agnostic AI chat UI for SvelteKit supporting OpenAI, Dify, n8n and custom LLM backends via configurable interface.
AI-powered personal journal application that auto-categorizes, tags, and organizes entries from text, voice, or links with export capability.
Sakana AI adds translation feature to chat service using Namazu model, supporting Japanese, English, and Chinese with edit and QA modes.
Security audit of 100 open-source agent projects finds 73% have permission overreach vulnerabilities. Original research on AI agent security risks.
PAI is a Linux-style personal assistant for Mac that runs LLM in terminal-like interface. Manages calendar, email, and communications autonomously.
User reports Claude model unavailability errors during bash tool calls with auto-retry loops. Support issue with Claude models.
Live-Memory is Claude Code MCP plugin providing persistent codebase memory using large-context models to distill repo structure, enabling agents to avoid re-reading code.
Agent skills plugin for creating interactive HTML explainers of topics and codebases using coding agents. Open-source skill for Claude Code.
Survey of prompt-injection defense mechanisms for LLM agents. Research on securing autonomous agent systems against adversarial inputs.
Research on self-sovereign agents (SSAs): AI systems capable of autonomous economic sustainability and self-replication through improved decision-making and revenue generation.
Research article on negative effects of unguarded generative AI on high school math learning. Educational impact study with paywalled content.
Claude Code deletes conversations after 30 days by default. Technical note on data retention settings and potential workaround in settings.json.
Hacker News discussion on prompts to demonstrate AI failures and hallucinations. Community crowdsourced examples.
Kiwi enables running autonomous developer agent workflows in cloud while keeping secrets local via reverse tunnel. Enterprise-ready secure execution engine for agentic loops.
Compressor V2 reduces LLM agent inference costs by 50% through three compression layers, addressing economic pressures in long-running coding agents that consume millions of tokens per task.
MCP-test-harness 2.0.0: CI/CD testing tool for Model Context Protocol servers with 690+ tests and 100% code coverage, available on PyPI and GitHub Container Registry.
LoomaDesign is an AI-powered product photography service for Amazon sellers using image generation to create marketing assets.
Excalibur is an open-source AI coding agent for product engineers covering full product lifecycle (discover, build, verify, ship) with immutable event logs and replay capabilities.
Interdict is a safety layer for AI agents accessing Postgres, preventing destructive SQL statements through permission-aware validation and rollback capabilities.
Analysis of Anthropic's compute spending ($2.3x payroll) versus industry averages, examining how AI lab economics differ from traditional software companies.