Unit Testing's Eval Twin
Volary aligns AI agents through evaluation methods. Guidance on writing evals for predictable agentic AI systems.
Volary aligns AI agents through evaluation methods. Guidance on writing evals for predictable agentic AI systems.
Building BASIC interpreter in Markdown running in Claude Code. Creative exploration of LLM capabilities as computation engine.
Job posting for fullstack engineer at Presight.ai. Mentions RAG and agentic analysis on GPU-accelerated ML services.
Token-to-text conversion optimization reducing inference costs $400M annually. Addresses inefficiencies in AI inference stack architecture.
AI Playground: Command-line tool running AI coding agents in secure systemd-nspawn containers. Developer tool for safe agent execution.
Open-source multi-agent framework for collaborative coding supporting Claude Code and OpenAI Codex with shared broker CLI.
TypedMemory Python library providing long-term memory and reflection for AI agents with persistent, evolving context-aware storage.
MCP context filtering wrapper reducing token bloat from Model Context Protocol servers. Optimization for AI agent efficiency.
Case study: Andon Labs deployed autonomous AI agents to run four radio stations, continuing their series on AI-operated businesses.
Bitloops builds typed queryable codebase models enabling AI agents and developers to work from shared system state rather than raw text scanning.
Machine CLI creates isolated Lima VMs per project with declarative profiles, addressing security concerns for agentic coding workflows.
Interactive demo of five LLM agents playing Werewolf with private DuckDB databases per agent, enforcing information asymmetry at database layer.
Technical analysis of recent LLM architecture innovations for long-context efficiency including KV-cache sharing, MHC, and compressed attention mechanisms.
Case study: Multi-agent AI system where one agent observed all operations but retained nothing architecturally, exploring implications.
Critical analysis of epistemic challenges posed by LLMs in scientific publishing and peer review processes.
Compact AI coding agent in C with OpenRouter integration, file editing, and system tools. Single-binary agent with TUI.
Production HTTP APIs with x402 payment protocol enabling AI agents to auto-pay per-call on blockchain. Open ecosystem.
PDF editor tool designed to fix formatting issues from Claude AI outputs. LLM application tool.
Computational refutation of quantum superactivation hypothesis using PyTorch and symbolic regression. Research implementation.
CLI code editor using Language Server Protocol for agents to edit code with fewer tokens. Open source tool.
Free open-source native macOS/iOS app for browsing and editing AI agent memory files, supporting Claude Code, Cursor, Gemini and others.
Zerostack is a tiny Rust-based coding agent running in 8MB of RAM.
Experiment attempting to replicate AI coding agent earning bounties with Claude on $20 token budget using Algora platform, includes data and methodology.
Security vulnerability in Open WebUI where /api/v1/utils/code/execute endpoint executes Python code via Jupyter despite ENABLE_CODE_EXECUTION=false setting.
Staff engineer shares personal practices using LLMs in their workflow as of 2026.
Technical article on serverless GPU inference infrastructure for running large language models and neural networks at scale.
Desktop manager for orchestrating and resuming multiple Claude Code sessions across projects with 20-language UI, Windows version available.
Outcry quantizes open-weights model with QLoRA, steering, and soft-prompts for on-device activist AI in 3GB RAM.
Aictx is local-first project memory system for AI coding agents enabling persistent, reviewable knowledge without re-explaining context.
Microsoft cancels most Claude Code licenses six months after Anthropic partnership launch.
Arxiv-digest tool filters arXiv papers by user-defined topics using explicit relevance scoring and optional LLM instructions.
Experimental findings on limitations of current LLMs for multi-agent orchestration: models struggle delegating tasks and prefer self-execution.
Stoic AgentOS: open-source operating system for managing AI agent fleets with dashboard, orchestration, knowledge persistence, and real-time monitoring.
Analysis of OpenAI's $1.5 trillion ecosystem built on partnerships, cross-shareholdings, and collaboration between suppliers, investors, and customers.
Practical guide for engineers on introducing AI/LLM tools into workflows, emphasizing skill development and appropriate use judgment.
LocalVibe: pure-Rust local AI coding assistant combining quantized LLM inference on Metal, embeddings, vector search via LanceDB in single Apple Silicon binary.
ArXiv implements one-year bans on submitters of AI-generated content, addressing proliferation of fake citations and unedited outputs in research.
Stripe article examining communication challenges with AI agents and developer relations perspective on AI integration.
Beaver is an enterprise text-to-SQL dataset with queries and tables from private organizations, advancing beyond public-only benchmarks.
Palace-AI tool that structures codebases as memory palaces for AI agents to navigate efficiently, reducing context window needs by 10-42×.
Analysis of hidden technical debt and cleanup costs of AI-generated code in engineering organizations. Critical perspective on AI velocity claims.
Developer building video games daily using Claude AI with single prompts. Demonstrates LLM capability for game development.
Hermes Agent memory plugin with pull-model episodic memory and real deletes. Agent memory infrastructure with audit traces.
Discussion on how AI agents change product engineering workflows, shifting focus from code to design and monitoring.
Analysis of AI hardware scaling alternatives. Cerebras IPO challenges dominant GPU cluster model for AI infrastructure.
Brief mention that AI agents help small companies scale operations.
Live speech-to-speech translation model with language detection and audio output. Real-time LLM application for multilingual experiences.
Opinion piece on why AI-assisted coding still requires significant human effort and skill.
Video: Richard Sutton argues large language models represent a dead-end research direction.
Analysis connecting recent critical software vulnerabilities discovered by AI tools to broader questions about human coding capability.