Deterministic NLP-to-SQL engine using multi-agent architecture over Delta Lake lakehouse. Replaces static BI dashboards with dynamic conversational SQL generation via Streamlit.
Talk on identity and authentication challenges specific to browser automation agents.
Talk on LLM capabilities for generating production-quality enterprise code.
Talk on deploying LLM-powered code review systems at scale.
Talk on intuitive coding techniques and practices.
PDF research paper on using LLMs to automate scientific discovery processes and experimental design.
Talk on building agents with extensive tool access and financial capabilities.
Tool that estimates coding task duration at AI agent speed rather than human speed, using classification and PERT analysis for dispatch planning.
Analysis of why specialist CI agents outperform generalist coding agents like Claude Code for continuous integration debugging and investigation.
Guide on capturing and reusing knowledge from AI agent tasks so subsequent agents learn from previous discoveries rather than starting cold.
Python library for running CPU-bound code in subinterpreters with asyncio, provides InterpreterThreadPoolExecutor.
Working implementation of 15M-parameter quantized Transformer running on 2007 Sony PSP with 64MB RAM, streaming text at 1-2 tokens/sec.
Talk on performance tradeoffs and costs in deploying fast AI agents.
PII anonymization framework for LLM applications, intercepts sensitive data before API calls using multiple detection backends.
Talk title suggests discussion of identity/authentication challenges.
Talk on building production agentic systems using ADE framework.
Talk on safety mechanisms and constraints for autonomous agent systems.
AI-generated Clojure dev-time UI displaying functions as draggable bubbles on canvas with call graph visualization and REPL integration.
Talk on diversity and complexity of coding agent implementations and approaches.
WebGPU support added to llama.cpp for running LLMs in browsers.
Open-source SDK for browser-native AI APIs and WebMCP with TypeScript composable building blocks, no runtime dependencies, vanilla JS/React.
Blog post on using physical systems with dynamics (gyroscopes, springs) as computational substrates for ML tasks like digit classification.
Explores how AI agents might interact with Git version control, considering its quirks and whether agents need better tooling than humans.
Announcement of DAIR Academy session on building visual LLM artifacts and knowledge bases.
Headless CRM designed for AI agents (Claude/Codex) with CLI interface, SQLite backend, minimizes context window and API costs.
Open-source graph database engine using S3 object storage as single source of truth, supports Cypher queries, embeddable or standalone.
LLM-based assistant for searching sensitive records using declarative approach that doesn't expose raw data to model.
AI agent that configures and runs OpenFOAM physics simulations from natural language descriptions, self-healing on failures.
Mistle: Open-source platform for running/automating sandboxed coding agents with Docker. Supports remote providers like E2B. Dashboard included.
Dhrive tool generates native iOS apps from prompts using local AI CLI, enabling developers to build apps without cloud dependencies.
Physics-based civilization simulator with autonomous JEPA agents. Built with Claude Code and Opus 4.6/4.7. Stress test of agentic coding LLM.
Hardware-verified memory routing system for edge AI agents with anomaly detection on Jetson Orin Nano.
Distribution Fine Tuning method improves model output quality through post-training optimization step for better written responses.
Technical analysis of token streaming for AI applications, comparing SSE vs WebSockets for production LLM deployments at Ably.
Assay platform provides validation layer for AI agents handling financial transactions, ensuring safety in production deployments.
Interview discussing AI slop, vibe coding practices, and future of application security with Tanya Janca.
Distribution Fine Tuning technique for improving LLM output quality through post-training optimization.
Project developing Māori language text-to-speech model emphasizing indigenous ownership and alternative to big tech approaches.
Analysis of game-theoretic implications of AI adoption, exploring how widespread AI use may create negative externalities.
Discussion thread collecting reliable AI prompts and techniques practitioners have tested for productive real-world work.
Brief observation about agent behavior during catastrophization. Limited content and technical depth.
Stubble extension provides context infrastructure for AI agents like Claude and Cursor via MCP protocol. Makes work context available to AI tools.
Running fleet of Claude coding agents for automation. Shows scaling from single agent to multi-subscription setup with token optimization.
Canonry is an open-source, agent-first tool for monitoring how AI cites websites and tracking AI search traffic, integrating with GSC and GA.
Analysis of AI agents in software development, examining what autonomous coding tools solve versus remaining human challenges.
Developer discusses reliance on coding agents (Opus 4.5), reduced code review, and changing relationship with manual programming work.
Essay reframing whether LLMs reduce intelligence to examining how they change learning patterns and cognitive processes.
Benchmark comparing Anthropic's Haiku and Sonnet models across three agentic AI tasks, finding the cheaper model performed best.
Explains throughput vs goodput metrics for LLM deployment performance, showing why goodput better reflects real-world behavior.
Analysis of three AI agent archetypes (personal, team, autonomous) and their distinct identity, audit, and credential requirements.