WebGPU feature detection was not enough to run small LLMs on phones
Technical analysis of running small LLMs in browser on phones using WebGPU, testing feature detection and buffer limitations.
Technical analysis of running small LLMs in browser on phones using WebGPU, testing feature detection and buffer limitations.
Fugu research paper on assembling, routing, and coordinating expert AI agents for complex multi-step tasks.
Discussion about future of AI coding tools and LLM capabilities for developers over next four years.
Sakana Fugu API orchestrates multiple AI models dynamically for complex multi-step tasks without vendor lock-in.
Demo exploring why LLMs lack internal imagery, proposing alternative systems that generate and store visual scenes using phonetic methods.
Woltspace is a containerized sandbox platform for interacting with coding agents remotely via Telegram or Slack, supporting Claude Code with persistent memory.
Open source organization's CI/CD disabled by GitHub due to crypto-mining activity by drive-by contributors.
Developer tool that securely provides API credentials to sandboxed AI agents without exposing real values.
GitHub App detecting prompt injection attempts in issues and comments targeting AI agents, including hidden HTML payloads.
Local-first memory layer enabling LLMs and agents to read/write persistent context across sessions via Model Context Protocol.
Whitepaper on using Codex as persistent workspace for long-running projects, managing context and complex workflows beyond single prompts.
Bean is a runtime tool that prevents AI agents from declaring tasks complete until claims are verified, conflicts resolved, and questions documented. Maintains typed claim ledger and runs verification compiler.
ANMA: YAML-based boundary contracts for enforcing architecture rules in cheaper AI coding agents, with benchmarks showing 100% compliance vs 32% without.
PeekAI provides local-first observability and debugging tools for Python AI agents, enabling developers to monitor agent behavior.
macOS tool that itemizes and verifies AI coding spend from Claude Code and Codex logs with granular breakdowns.
Z.ai open-sourced a frontier coding model amid US regulatory restrictions on rival systems.
Samsung Electronics deploys ChatGPT Enterprise and Codex to employees globally, marking one of OpenAI's largest enterprise deployments.
Personal project fine-tuning Qwen 3:0.6B for household question categorization using RAG with metadata-aware vector search.
Self-contained deception detection pipeline fusing text, visual, and audio features from courtroom video for truthfulness prediction.
Developer built LLM-based French tutoring tool replacing paid human tutor, focusing on improved retention through interactive conversation and personalization.
Compass: local-first config layer for AI coding agents with budget caps, guardrails, and safety features. Prevents unsafe commands and unverified code merges.
CivBench: benchmark for evaluating AI agents by tasking them to run civilizations in game environment.
Tunr exposes local development servers to public internet in under 3 seconds with automatic HTTPS, alternative to ngrok built in Go.
EGC: MCP server giving AI coding agents persistent memory across sessions to retain decisions and learnings.
Video: Jonathan Blow discusses fundamental limitations of LLMs in programming tasks.
Apertus: open-weight foundation model (8B/70B) from Swiss AI Initiative. Open data, reproducible training, EU AI Act compliant.
Analysis of security requirements for increasingly capable AI agents in autonomous tasks like cybersecurity and R&D.
Recall: Local session logger that condenses Claude Code interactions into project summaries without external API calls or data transmission.
Conduit: self-hosted Bitcoin Lightning payment infrastructure for autonomous AI agents with spending guardrails.
AuraText: Windows overlay for AI assistance in any text field with contextual prompt generation.
Analysis of LLM industry economics, subsidy structures, and timeline from research to production use in programming.
API platform for comparing and routing requests across multiple LLM providers with unified billing and model switching.
Study comparing ChatGPT diagnostic performance against physicians in medical settings.
Opinion piece expressing concerns about LLM-generated incident reports becoming commonplace in tech operations.
Authorization engine detecting when AI agents are manipulated via prompt injection or anomalous behavior at runtime.
Memory Magico: CLI tool for agent-focused memory, wiki, and deterministic sprint management with validation checkpoints.
Self-hosted AI traffic proxy with token quotas, response caching, and per-user limits via local Docker deployment.
ClaudiOS: Linux distro that boots directly into Claude Code AI agent with full system access, eliminating traditional OS desktop environment.
Case study of agentic AI coding using GLM-5.2 model service: $0.034 cost, 3 minutes, 2 mistakes vs 1 hour human attempt with 4 mistakes.
Developer tool enabling local reproduction of production bugs by isolating I/O from business logic without requiring production databases.
Browser-only privacy tool suite running local AI models via WebGPU: JSON tools, Base64 encoding, QR codes, PDF/image processing without uploads.
Analysis of AI agent benchmark flaws: judge models cannot verify if agents actually accessed claimed information or checked absence claims.
Pokémon Crystal text-based UI benchmark environment for testing AI agent capabilities in turn-based game scenarios.
Jacobi-IDE developer tool for Abaqus UMAT subroutines with AI-assisted diagnosis for computational mechanics simulation.
LLM API gateway offering Opus-equivalent models at 91% lower cost via pay-per-request pricing across frontier models.
Agent framework where server executes pre-approved actions rather than delegating execution to language models for improved safety and control.
Energy-based pricing model for AI inference where prompt efficiency directly reduces computational cost to users.
Google's Open Knowledge Format specification for universal, vendor-neutral knowledge representation as markdown with YAML. Includes agent proof-of-concept for automatic OKF bundle generation.
Framework/methodology for oversight and control of autonomous AI agents. Title-only submission.
Geospatial chatbot tool combining ChatGPT-style interface with map visualization. Limited technical details.