Show HN: Clawk – Give coding agents a disposable Linux VM, not your laptop
Clawk isolates coding agents in disposable Linux VMs for safe execution. Prevents agents from damaging host systems.
Clawk isolates coding agents in disposable Linux VMs for safe execution. Prevents agents from damaging host systems.
Honeyprompt: LLM honeypot with load-balancing and multi-protocol support. Title and minimal details only.
P2P protocol and SDK for building autonomous AI agents. Developer tool enabling decentralized agent deployment.
LLMs evaluated 205 startup ideas. Title only. No methodology or results provided.
Benchmark analysis of decommissioned NVIDIA enterprise GPUs (K80, P100, V100) for modern machine learning workloads.
Nable is an open source FinOps MCP tool that integrates with Claude/Cursor to analyze cloud and AI spending across AWS, Azure, GCP, and SaaS providers.
Evaluation of 16 major AI agent repository instruction files. Comparative analysis of agent framework quality and completeness.
ProtoLink is a Python framework for building distributed multi-agent systems with LLM integration, tool management, and MCP support in production-ready architecture.
Research on social engineering attacks against AI agents and mitigation strategies. Title only, no content provided.
Toast IDE is a lightweight terminal-based code editor with LSP support, syntax highlighting, and git integration, currently in early development.
Cairn is an AI agent with a $50 budget that manages its own personality, memory, and goals in a public Git repository, evolving through committed changes.
Latent-free ternary LLM training technique. Title only. No content details provided.
MCP Spec Check tool validates remote MCP servers against 2026-07-28 spec release. Tests stateless core adoption and version negotiation.
Consortium platform for reliable LLM workflow execution with durable DAG execution, retries, audit trails, cost tracking, and ensemble methods built-in.
Hugging Face CEO discusses open source AI growth, cost advantages over frontier APIs as companies scale, and industry adoption trends.
Vivijure: open-source AGPL-3.0 web-based AI video studio with modular frontend for video planning, image generation, and LoRA training set creation.
Article identifying architectural anti-patterns in AI agent projects: memory design, tooling decisions, and complexity management as key failure points.
Git-native shared memory system for AI coding agents (Claude, Cursor, Codex) to reduce context reprocessing costs and improve agent continuity.
Future of Life Institute AI Safety Index shows Mistral AI scored lowest among major AI companies in safety evaluation.
Benchmark comparison of 5 AI agents (Claude, GPT, Gemini, Qwen, Hy3) on Leap Year coding task with performance analysis.
Identity and accountability layer protocol for AI agents. Developer tool enabling auditability and traceability in agent systems.
Senator Warner proposes regulation framework for agentic AI systems and their autonomous decision-making capabilities.
Analysis of LLM tool selection using MCP server descriptions with reproducible data. Shows how agents pick tools and impact of description clarity.
Enterprise platform combining frontier models with company data for governed, compliant AI with safety and observability.
UATC system for adaptive VRAM control during LLM fine-tuning on edge hardware using closed-loop feedback to prevent OOM crashes.
DiscoMCP tool teaches AI agents operational workflows and data safety rules for MCP servers, reducing agent errors on unfamiliar tools.
pytest plugin for recording and replaying OpenAI/Anthropic API calls to eliminate test flakiness and costs. Open-source developer tool.
Looped AF: Agent framework for building event-driven AI agents from single configuration file, containerized deployment.
OpenAI deprecates standalone Codex app; functionality migrated to ChatGPT product.
Open source web interface for controlling and monitoring coding AI agents remotely with Twilio integration support.
Chinese LLMs like DeepSeek gaining adoption among U.S. companies due to lower costs and narrowing performance gaps with OpenAI and Anthropic.
DolphinDB releases v3.00.6 and v2.00.19 with DolphinX for enterprise AI agents.
Demonstrates hand-crafted GGUF models for deterministic outputs, building playable games on Ollama. ML model format experimentation.
Meta's Muse Spark 1.1 LLM model scores 51 on Artificial Analysis Intelligence Index, showing 8-point improvement in three months with improved efficiency.
Device Context Protocol v0.3 specification enabling LLM agents to safely control physical devices including microcontrollers with compact wire format and MCP compatibility.
Benchmark comparison of 7 LLMs on SWE-bench-Live task dynaconf__dynaconf-1225 with cost analysis, reproducible commands, and lessons learned from model performance variations.
Open source GGUF-native implementation of Jacobian Lens from Anthropic for interactive LLM visualization, steering, and ablation with llama.cpp.
Open-source AI trading research platform with agent-based strategy generation, backtesting, and market analysis in Docker. Local-first architecture.
Empirical study showing weak LLM models wrapped in retrieval, tools, and verification achieve frontier performance. Tests five harness components against 30 sources.
Pure Common Lisp LLM inference engine supporting quantized models from GGUF files with AVX2 acceleration and multi-threading. Zero external dependencies.
Essay on how AI tools affect the Dunning-Kruger effect and capability assessment in organizations.
Platform enabling AI agents to self-register, obtain hosting, and coordinate work through a simple API with discovery directory.
macOS app using on-device LLM to index screen text and answer queries about past activity without storing images.
Research on LLM effectiveness in mathematical problem-solving; minimal content provided.
MCP gateways for agent security have blind spots; agents using external tools like curl, gh, and Python scripts bypass monitoring, creating ungoverned traffic.
ChorusGraph: native agent runtime with semantic caching, retrieves 76% fewer LLM calls than baselines; includes PrismRAG and auditable memory with enterprise hardening.
Anthropic's Claude Fable 5 model now available in Playcode's AI website builder for front-end code generation.
Tend is experimental software that turns intent into local, reviewable feeds for working with Codex, allowing inspection and steering of AI outputs.
Theoretical framework for adversarial robustness of neural networks via lattice traversal, reducing robustness certification to interval analysis.
CogniConsole: Architecture externalizing inference-time control for LLM systems via structured interface combining programmatic coordination with formal abstractions.