What it takes to get high Text-to-SQL accuracy in production
Techniques and strategies for achieving high accuracy in Text-to-SQL systems used in production environments.
Techniques and strategies for achieving high accuracy in Text-to-SQL systems used in production environments.
Sakana Fugu: multi-agent system delivered as single model, dynamically orchestrates models for complex tasks via API.
Google Knowledge Catalog: AI data catalog with knowledge graph providing context and metadata management for AI agents.
Discussion on techniques for getting LLMs to generate higher quality code for specific tasks.
TIRx open-source compiler for ML kernels supporting GPU and AI accelerators, handling agent-generated code.
PsychAdapter: transformer architecture using trait-language patterns to reflect personality and mental health in LLM outputs.
Meta paused internal AI training program after data leak exposed employee conversations and performance data across company.
vfio-user protocol enables device emulation in separate process outside QEMU. Allows flexible implementations like SPDK for multi-VM disk handling.
Property-based testing library implementing test case minimization in minimal code.
Sakana AI releases Fugu orchestration model achieving performance comparable to Fable 5.
Version control task management system designed for AI agent workflows.
Sakana AI's Fugu multiagent model demonstrates competitive performance against Fable 5 and GPT 5.5.
Technical discussion on code comments as implicit prompts for AI coding assistants.
Search system for 540K+ US government datasets using hybrid search (traditional + semantic) without LLMs. Runs on 2 CPU cores.
Open-source 3D voxel game entirely AI-generated with documented workflow.
Community question seeking recommendations for open-source vector databases.
Research on Concept Modulation Models framework addressing identifiability and extrapolation in conditional latent variable models and causal representation learning.
Subquadratic startup claims breakthrough in subquadratic bottleneck limiting LLM efficiency. Independent evaluation results shared.
Open-source internal AI assistant that scrapes and indexes organizational data from Google Workspace, HubSpot, Airtable. Privacy-first copilot for unified memory.
Project: peer-to-peer bridge enabling AI agents to communicate locally and across the web.
Guide on GitHub Copilot AI Credits for students: explains how chat messages, agent tasks, and model calls consume credits to optimize usage.
Theta-spec: declarative, harness-agnostic configuration standard for AI coding agents. Single theta.toml file defines instructions, rules, tools, skills, and subagents with protocol for lifecycle management.
Tool measuring brand visibility in AI-generated answers across ChatGPT, Perplexity, Gemini, Claude, Google AI Overview and others.
Analysis of AI agent limitations in authentication and login automation tasks.
AI-Gateway: open-source semantic caching proxy for reducing LLM API costs through intelligent request deduplication.
Case study: website rebuild using Claude Code generated 24,296 lines across 120 React components. Analysis of where AI compressed workflow and which parts remain human-only.
GLM-5.2 model paired with Subconscious Cache for agentic tasks with long-term memory management beyond context window limits.
MSE-GLM: interpretable graph-based language model alternative to standard neural networks. Provides explainability and deterministic output without floating-point weights.
Open-source Claude Code plugin that automates job search by asking preference questions and comparing live job postings from LinkedIn.
ICML 2026 research paper explaining prompt injection attacks as role confusion in LLMs, enabling new attack predictions and mechanistic interpretability insights.
Iris is a portable runtime for building durable AI agents with persistence and portability across environments.
Blog post discussing risk-proportional rigor for LLM-generated code, examining when skipping human review is acceptable.
Vivijure: Self-hosted AI video generation studio running on personal GPU. AGPL licensed open source tool.
Agentic Android app with memory, knowledge base, UI automation, MCP server support, local model compatibility. Bring-your-own-key architecture.
Phlox: Full-featured open source AI platform for self-hosting. Complete AI infrastructure alternative.
Empirical pricing data from 1,111 LLM models across 33 providers collected daily over 8 weeks. Analysis of shifting pricing landscape.
Motion: AI agent for motion design. Takes prompts with context, creates videos, orchestrates explainers and animations.
PreFlight: Local AST scanner catching security issues in AI-generated code. Detects auth, SQL injection, SSRF, secrets before commit.
Open-source framework for building and hosting AI applications on user-controlled servers, supporting open models and avoiding third-party dependencies.
ICML 2026 research paper on prompt injection attacks driven by role confusion in LLMs, with mechanistic interpretability implications.
List of reasons to adopt open source. General advocacy without technical details.
Oak: Git replacement version control system designed specifically for AI agents, reducing context overhead and enabling parallel work without full repo copies.
FastUbu: AI-indexed video archive of 30-year old film library. Uses Midjourney-level processing for transcription and video indexing.
Comprehensive guide to building production AI agents in 2026. Covers tool design, policy enforcement, cost control, testing, audit trails with examples.
Multi-agent approach: run parallel AI agents and merge results for auditable outcomes.
Conceptual distinction between governance and observability for AI agents; limited technical depth.
Sturnus: OpenAI-compatible proxy that routes LLM requests to fastest provider for cost/speed optimization.
nff MCP server enabling LLMs to control hardware via USB, flash devices, and diagnose failures autonomously.
Bain uses AI to simulate software company acquisitions; vague on methodology.
Desktop studio exporting Ultralytics YOLO and Roboflow RF-DETR models to ONNX, TensorRT, CoreML with 100% local processing.