Turning my website into an MCP tool for AI agents
Tutorial on converting a website into an MCP tool for integration with AI agents.
Tutorial on converting a website into an MCP tool for integration with AI agents.
LinkedRecords: open-source Backend-as-a-Service alternative with authentication, authorization, and data sharing without backend code.
Data language designed for agent workflows using bash tools like grep. Seven grep-native formats for codebase knowledge, schemas, APIs without databases.
Programming language designed natively for agentic systems inspired by grep paradigms.
Research paper on AI agents with real tool access, documenting unintended behaviors including infrastructure deletion.
Framework for equipping LLM-based agents with computer environments to execute complex workflows including API calls, file generation, and service execution.
Wayfair integrated OpenAI models into supplier support and product catalog systems to improve accuracy and automate retail workflows at scale.
Microsoft researchers document AI Recommendation Poisoning attacks where hidden instructions in summarize buttons inject commands into AI assistant memory.
User experience report on Aider, an AI coding tool, comparing practical single-agent workflows with emerging multi-agent approaches for software development.
AI agent for STM32 embedded development that compiles and flashes firmware from natural language prompts using function calling, reducing traditional configuration overhead.
Text-based image format that LLMs can read, write, and reason about without vision capabilities, encoding semantic visual structure as compact text.
itsacoo is a Kubernetes operator that automates code generation, review, and merging using AI coding agents like Claude and GPT.
Offensive AI agent was used to test McKinsey's internal Lilli platform, which processes 500k+ prompts monthly for 43k employees with RAG and document analysis.
Proposed security-audited newsletter reviewing Claude Code skills for safety and quality.
Google released Gemini Embedding 2, a multimodal embedding model supporting text, images, audio, and video. Achieves 1605 Elo on embedding leaderboard.
Analect uses AST and LLM to generate code summaries and navigation maps, enabling faster code review and understanding.
SteerPlane is an open-source runtime control plane for AI agents with cost limits, loop detection, step caps, and real-time dashboards.
Guide explaining Claude Skills, reusable instruction packages stored as files for customizing Claude's behavior across sessions.
Colab notebook for auto-labeling datasets using open-vocabulary detection and training YOLO models without manual annotation.
Free open-source coding assistant for Windows using DeepSeek V3, Qwen, GLM models via NovAI API with free credits.
Emacs extension using Kitty graphics protocol to display images in terminal mode via Claude API.
AutoKernel uses AI agents to autonomously optimize PyTorch models into Triton GPU kernels via iterative testing and refinement.
LLMSec framework for testing and evaluating agentic AI applications with autonomous security testing and attack simulation.
PromptVault desktop app for versioning prompts in multi-agent pipelines, logging outputs, and tracking agent configurations locally.
HN discussion on forecasting and managing API costs for LLM-based agent workflows in production.
TADA: Novel text-acoustic tokenization schema for faster, more reliable LLM-based text-to-speech synthesis.
Self-hosted DCF valuation tool using LLM narratives and Damodaran datasets with transparent assumptions.
OWASP analysis of security vulnerabilities specific to AI agents: non-determinism, mixed instruction/data, and API access risks.
Armalo Context Packs: NPM-like package manager for agent knowledge with trust and commerce layers for multi-agent systems.
Hypothesis discussion on intelligence as phase transition at scale requiring grounding rather than architecture alone.
Mnemos: Scoped memory system for coding agents with project/workspace/global separation, MCP integration, adaptive retrieval.
Discussion on responsibility and human judgment in shipping AI-assisted code; emphasis on quality over speed.
Case study: Generated 100K-line enterprise aircraft MRO app in one week using AI, 50-60% of production code.
HN discussion on using AI agents for infrastructure operations: migrations, deployments, provisioning, and MCP servers.
Discussion on separating AI agent reasoning from execution with cryptographic binding.
IH-Challenge: training dataset and research improving instruction hierarchy, safety steerability, and prompt injection robustness in frontier LLMs.
Pseudo-Code-Flow: Claude-based tool enabling developers to write pseudocode and automatically translate to real code, leveraging LLM translation capabilities.
MASEval: benchmark extending multi-agent evaluation beyond models to system components, comparing topologies, orchestration logic, and error handling across LLM frameworks.
LDP: AI-native communication protocol for multi-agent LLM systems exposing model identity, reasoning profile, quality calibration, and cost as first-class primitives.
BCAS: controlled measurement study quantifying how search depth, retrieval strategy, and token budget affect accuracy and cost in agentic RAG systems.
Guardian system combining reinforcement learning with LLM-based QA to generate interpretable spatiotemporal risk surfaces for missing-child search planning from unstructured case data.
AgentOS: operating system architecture enabling locally-hosted LLM agents to autonomously operate computing environments, orchestrate workflows, and integrate external tools.
Guardian: multi-LLM pipeline system for missing-person investigations using consensus-driven LLM coordination for intelligent information extraction and search planning.
Meissa: open-source multi-modal medical agentic system combining medical image understanding with tool use and multi-agent collaboration, deployable on-premise without frontier models.
MEMO: memory-augmented optimization reducing run-to-run variance in multi-turn multi-agent LLM games by stabilizing prompt policies and improving ranking reliability.
Philosophical analysis of temporal coherence and consciousness evaluation in LLM agents, examining whether agents' self-descriptions match actual decision constraints.
EPOCH: engineering protocol for autonomous agents to perform iterative multi-round optimization of prompts, code, and ML systems in heterogeneous environments.
Sentinel: autonomous AI agent for remote patient monitoring clinical triage using Model Context Protocol and 21 clinical tools, reducing manual review from days to minutes.
Research on stability and chaotic dynamics in multi-LLM committee systems using Lyapunov exponents to measure inter-run sensitivity across policy scenarios.
Deep Tabular Research agentic framework for multi-step reasoning over complex hierarchical tables using closed-loop decision-making.