AI agent security needs a composition graph, not just an SBOM
Proposes composition graphs as a security approach for AI agents, moving beyond traditional SBOMs for dependency tracking.
Proposes composition graphs as a security approach for AI agents, moving beyond traditional SBOMs for dependency tracking.
Video demonstrating interaction between multiple AI agents.
Local-first MCP self model enabling AI agents to maintain context of user identity and preferences.
Technical article explaining why traditional unit testing fails for LLM applications and proposing alternatives.
Testing framework for identifying failure cases in RAG pipelines before reaching end users.
CUDA profiler tool for monitoring and optimizing production inference workloads.
Unofficial Visual Studio extension implementing Claude Code IDE protocol integration for native debugger support and compiler error sharing.
Tool to share and replicate AI coding environment setups via single command installation.
Discussion of AI token consumption driving increased enterprise cloud computing costs.
arXiv paper on token merging optimization for Segment Anything Model (SAM). Improves Vision Transformer efficiency without retraining.
Analysis of developer motivations for using LLMs in blog post generation.
Open source control plane for AI inference management. Infrastructure for model serving and deployment.
Open source control plane for managing AI model inference. Infrastructure tooling for deployment and serving.
Open source skills for discovering and reviewing conversational AI agent capabilities in codebases, generating improvement recommendations.
Research on content-routed sparse attention for long contexts with bounded error analysis and experiments on SubQ model performance.
Sanity reports on gathering product feedback from AI agents using their MCP server, with 20K+ agents making millions of tool calls.
Static hosting platform designed for AI agents to generate websites, targeting non-technical users building corporate and personal sites.
Open-source MCP server providing page-cited embedded firmware and datasheet references for coding agents like Claude Code.
Cursor acquires Continue, an open-source GitHub Copilot alternative, signaling consolidation in AI developer tools market.
peerd: Open-source AI agent harness running as browser extension with bring-your-own-key model. Early developer preview released on GitHub.
Python library for programmatic video editing, processing and AI workflows with streaming pipeline support.
Voice-based AI mock interview simulator supporting multiple roles/levels. Functional LLM agent application with practical implementation.
Overview of AI Agent Management Platform (AMP) category and features. Descriptive guide to agent platform concepts.
Conceptual comparison of governance vs observability in AI agent systems. Covers agent management concepts.
Tool demonstrating how AI agents interpret and process startup website content for evaluation and understanding.
ECCV 2026 paper on agentic workflow that generates diverse, controllable image interpretations from single text prompts.
CLI and MCP tool for controlling Android devices with fine-grained permissions via accessibility service, designed for automation and agent integration.
Local-first Node.js harness for running AI agents with event API, file storage, and plugin routing system.
AI agent implementation using Unix philosophy with bash, curl, jq and local models, minimizing external dependencies.
Web-based control center for managing multiple AI coding agents (Claude, Gemini, Codex) with real-time monitoring, task assignment, and autopilot features.
Skill pack to prevent coding agents from endorsing bad startup ideas through evaluation prompts.
Open source debugging platform for AI agents and LLMs with error detection and monitoring capabilities, similar to Sentry.
VibeThinker-3B explores verifiable reasoning in small language models. Research on LLM reasoning capabilities and interpretability.
SharpeBench: benchmark for evaluating AI trading agents with robustness against luck and randomness.
Research on LLM reversal curse: models trained on "A is B" fail to learn the reverse relationship "B is A".
Stub entry on company-wide AI agents with no accessible content.
Guidelines for using LLM-based generative AI in open source software contributions.
Aharness framework enforces AI agent workflows as finite state machines on Codex with runtime validation and guardrails.
Defines AI agents, categorizes types, and discusses why most fail to reach production deployment.
Claude Code AI agent successfully reverses CAN bus signals and vehicle protocols using LLM with custom skills.
N8n report on AI agent builder capabilities and market trends in 2026.
Open source Frame.io alternative for creative collaboration with AI agent integration and self-hosting support.
Analysis of whether LLMs will replace traditional software development or remain complementary tools.
Open-source Django SaaS boilerplate with modern frontend stack, positioned for AI-agent applications.
AWS Lambda MicroVMs enable isolated execution for user/AI-generated code with VM-level isolation and fast startup.
Sakana AI releases Fugu Ultra model with multi-agent orchestration matching frontier models without export controls.
Essay on how AI engineering problems increasingly require philosophical frameworks to solve.
Single-binary dashboard for monitoring Claude Code sessions, showing tokens, costs, and tool usage without external infrastructure.
Tool to automatically create clean, deduplicated, refreshed datasets from arXiv, Semantic Scholar, and GitHub given natural language topic descriptions.
Apple's AI compute strategy and potential cloud business model, comparing to SpaceX/Anthropic approach.