Training an Agentic Router for Optimal Cost-Performance on SWE Tasks
Research on training router models to dispatch software engineering tasks optimally across different LLMs based on cost, latency, and task type.
Research on training router models to dispatch software engineering tasks optimally across different LLMs based on cost, latency, and task type.
Open-source GitHub issue dispatcher that coordinates human-agent teams using OpenClaw for context and local worker execution.
Workspace tool for GitHub/Gmail integration with AI agent for building filters and PR management, similar to Graphite.
Guidance on building long-running AI agents that handle timeout constraints.
Open-source AI tutor tracking handwritten work via webcam to provide hints, available on Mac App Store.
First formally verified polygon intersection algorithm using Claude Opus 4.8 as AI agent, demonstrating one-shot proof generation capability.
Claude Code plugin that generates Python web scrapers, curates news from RSS/URLs, renders personalized HTML newspaper without server.
Building LLM skills with regression tests to verify quality and compare performance, addressing skill reliability testing.
Product analytics platform for AI agents capturing end-to-end runs to identify failures and continuously improve agent performance.
Fademem memory architecture for AI agents to maintain context, preferences, and corrections across sessions without model retraining.
AI agents fail in production despite passing demos. Prompt engineering alone doesn't solve real-world deployment gaps; agents need agile development practices.
Apple approves Poke, a startup enabling AI agents via iMessage, as first third-party agent on Messages for Business platform.
Law professors rated LLM tutoring responses higher than peer answers in blinded evaluation of contract law questions, showing LLM utility in judgment-based domains.
Analysis of vector search limitations for LLM memory with benchmark data demonstrating failures.
Anthropic releases open-source framework using AI to discover software vulnerabilities. Developer tool leveraging LLMs for security.
Technical analysis of LLM-based anti-bot systems deployed by Apple and Fastly, reverse-engineering their implementations.
Personal reflection on LLM usage in work and coding over 9 months, tracking where LLMs work well and limitations.
Analysis claiming 80% of LLM economic value derives from 20% of tokens, suggesting cost optimization opportunities.
Local AI file explorer with RAG, privacy-first approach, LLM chat integration, and multi-tab semantic mapping. New features include light theme and smart music player.
Research on using LLMs to guide runtime parameter optimization for energy-efficient model inference.
Tool for measuring LLM reliability, latency, and cost metrics. Evaluates performance characteristics of language models.
Supabase raises $500M at $10.5B valuation as backend infrastructure for AI apps. Growth driven by AI-assisted coding and developer tool demand.
Subreddits targeted by companies spamming peptide products to manipulate AI chatbot scraping. Strategy exploits source material to influence AI training data.
Article about airlines using AI chatbots to simulate empathy instead of fixing service issues.
Agent Arena framework for causally evaluating AI agents in real-world scenarios.
Anecdotal account of LLM providing misleading suggestions that broke functional code. Title only, limited detail.
Clarity tool traces LLM concepts and maps them back to training data for interpretability.
Open-source AI toolkit designed for e-commerce applications.
Analysis of pricing model variations causing 20x cost differences for equivalent chatbot AI tasks.
Open source self-improving LLM-connected coding environment using web-based chatbots. Runs in browser, saves locally, includes sample apps.
Developer built YouTube transcription tool using AI; achieved SEO indexing after 30 days.
Sequel: secure database connectivity layer for AI agents. Enables safe LLM/agent access to databases.
CLI loop: multi-agent system performing peer-review of code/content. Demonstrates AI agent coordination.
KVarN: Huawei's vLLM KV-cache quantization backend achieving 3-5x capacity and 1.3x throughput improvements. Plug-and-play integration, calibration-free.
Discussion on why LLMs struggle with NoSQL database querying compared to SQL. Explores lack of unified interface across MongoDB, DynamoDB, Cassandra, Redis, Neo4j.
Boxes.dev: cloud-based agentic dev environment enabling Claude Code and Codex agents to run in isolated cloud containers instead of localhost.
Endava redesigns software delivery workflows around AI agents, reshaping enterprise collaboration and accelerating project delivery.
OpenAI improves ChatGPT memory synthesis system to handle staleness, correctness, and scalability across millions of users.
Training-free method revealing token-level triggers control reasoning behavior in hybrid LLMs, enabling intermediate-budget reasoning without fine-tuning.
Outcome-grounded Advantage Reshaping (OAR) improves credit assignment in GRPO for mathematical reasoning tasks by fine-grained token-level rewards.
Theoretical analysis proving success conditioning (rejection sampling, goal-conditioned RL, Decision Transformers) solves a specific optimization problem for policy improvement.
LLMs show limited generalization of procedures across code, graph, and natural language representations.
Small language models adapted via zero/one-shot learning for role classification in human-robot interaction.
Interactive world models with persistent 3D state representation enable spatially consistent video generation.
Constraining safety-critical tokens during fine-tuning preserves LLM alignment on benign downstream tasks.
Safety evaluation scores differ across deployment configurations (ReAct, multi-agent, map-reduce) versus direct API.
Vision foundation models remain stitchable across different objectives and data despite architectural differences.
Attention-based sampler enables parallel decoding in diffusion language models improving inference efficiency.
Policy Split enables dual-mode exploration in LLM reinforcement learning balancing diversity and accuracy.
MANTA benchmark evaluates LLM value alignment on animal welfare across multi-turn adversarial queries.