SkillSpector
Security scanner detecting vulnerabilities and malicious patterns in AI agent skills before installation, analyzing 26.1% vulnerability rate.
Security scanner detecting vulnerabilities and malicious patterns in AI agent skills before installation, analyzing 26.1% vulnerability rate.
CLI tool using AI agents to find investor contacts and automate outreach emails; user reports 14% reply rate from 43 emails sent.
Open-source local-first cognitive memory using AGM belief revision on property graphs with propagation engine for updating downstream beliefs.
ACM warns that 'vibe coding' with AI skips fundamental software engineering practices and discipline.
Desktop app that orchestrates Claude and Codex agents to autonomously work on Linear board issues in sandboxed workspaces.
Claude Code-based AI coaching system implementing the Intelligence Emotions model with five specialized coaches and local session memory.
Claude and GPT-5.5 agents orchestrate code architecture and building tasks with 80% token reduction, running on existing subscriptions without API costs.
Open-source offline-first sync engine for SQLite/PostgreSQL using Rust with CRDT conflict resolution at column granularity.
Testing framework for LLM agents. Monitors tool calls, arguments, trace, latency—behavior not just output. Catch regressions in CI.
PyTorch tool detecting silent failures in LLM-generated unit tests for vLLM/SGLang kernels. Addresses hallucination in test generation.
Conceptual article on LLM control planes for production AI systems. Covers routing, budgets, privacy, fallbacks, provider management.
arXiv paper abstract about pattern matching mechanisms shared between human and LLM reasoning. Limited content provided.
Discussion on team etiquette when sharing AI-generated code and documentation output. Addresses AI fatigue and integration practices.
Nature Medicine study comparing frontier LLMs (GPT-5.2, Gemini 3.1 Pro, Claude Opus 4.6) to specialized clinical AI tools on medical benchmarks.
Security researcher demonstrates AI-powered vulnerability discovery across Google infrastructure, finding 1,500 APIs and earning $500k in bounties.
GatekeeperAI is self-hosted platform for enterprise governance of third-party and internal AI applications with security scanning.
Survey of harness engineering for AI agents, analyzing production deployment patterns and the layer below model selection.
Discussion of AI agents struggling with UI padding/alignment due to lack of visual feedback loops during generation.
Cortex is agent-native knowledge OS on Markdown with MCP interface for AI agents to read/write/search knowledge graphs from file uploads.
Research measuring LLM impact on N-day vulnerability exploitation and real-world security harm from publicly disclosed bugs.
Tenet Security demonstrates 'Agentjacking' attack hijacking AI coding agents via fake bug reports through Sentry APIs.
WWDC 2026 video on running local agentic AI on Mac using MLX framework. Developer tools for local AI agents.
Discussion of Claude API rate limits and workarounds for multi-agent workflows. Community experiences optimizing LLM agent usage at scale.
36 NPM packages for open infrastructure targeting AI companies in Argentina. Limited details provided.
Optimization techniques for reducing API costs in agentic coding by parallelizing tool calls and minimizing context resubmission. Technical analysis of LLM agent efficiency.
Centralized vector database for ingesting multi-app data to support AI agents via MCP. Addresses context retrieval and memory management for agentic systems.
UW research: AI agents estimate electronics carbon footprints, addressing gap in device sustainability information for consumers.
BitBoard: YC-backed agentic analytics workspace enabling agents and users to collaborate on live dashboards and data analysis.
Analysis arguing sandboxes alone are insufficient for AI agent security; agents can misuse valid credentials despite isolation. Critical infrastructure security for AI agents.
Documentation written for humans poorly serves AI systems consuming it. Explores gap between human and machine comprehension of implicit context.
Analysis of how AI technology impacts different software markets differently, challenging uniform disruption narratives. References NBER paper.
OpenClaw Skill Workshop: system for agents to learn and reuse procedures from repeated work with human review before application.
Distill: self-hosted AI agent requiring physical evidence before task completion, learns reusable skills across sessions with versioned rollback.
Research testing whether 'swearing' in prompts improves LLM reasoning. Null result study on prompt engineering effects on o1 model performance.
MCP server for SharkClean robot vacuums enabling agentic control via Claude. Open-source tool integrating IoT devices with LLM agents through cloud API.
Discussion thread asking for effective prompting techniques to improve Claude and GPT output quality and performance.
Tokenbrook Vale is open-source visualization tool for running multiple Claude AI agent instances in retro pixel art office environment.
Janus MCP server collects browser and terminal context to simplify prompting AI coding agents with interaction history.
SmithersBot open-source AI agent autonomously built and launched a business in 48 hours to solve payment verification for autonomous agents.
VibeClip is open-source AI video editor controlled via natural language chat interface.
StackScope crawler analyzes 40k+ indie product launches to identify technology stacks, hosting, frameworks, and AI builder signals used in real applications.
Essay on ML engineer career transition to AI-native roles, featuring AI agents automating data pipeline scaffolding and model boilerplate with concrete 90-day reskilling plan.
Local web dashboard for Claude Code sessions with live monitoring, history, usage charts, and token cost analysis.
Apodex-1.0 is a deep-research AI agent team that performs science-based verification of its own evidence. Limited technical details provided in excerpt.
Developer used Fable to clone game Soldat and built neural network AI bot trainer from simulated matches, with live dashboard and replay system. Demonstrates LLM code generation capabilities.
TetherDust is a self-hosted open-source AI Analytics Engineer tool. Title-only submission with no technical details.
Tool that converts videos to ANSI pixel art and displays them in Claude Code statusline during agent execution.
Demonstration of automating Instagram engagement using computer vision to overcome DOM obfuscation.
Analysis of Claude Fable 5 achieving top performance across multiple benchmarks (Artificial Analysis, LMArena, SWE-Bench Pro, FrontierCode), discussing acceleration pace of frontier AI models.
Physical AI system for autonomous industrial bioprocess operations using bio-input sensing instead of vision, requiring safety grading mechanisms for autonomous decision loops.