Advancing voice intelligence with new models in the API
OpenAI releases three new realtime voice models for API with reasoning, translation, and transcription capabilities.
OpenAI releases three new realtime voice models for API with reasoning, translation, and transcription capabilities.
Open-source desktop automation tool providing local-first gateway for AI agent workflows via hotkeys, integrating ChatGPT, Gemini, and Perplexity.
Discussion thread asking for real-world autonomous AI agent deployments and use cases, distinguishing between true agents and workflow automation.
Research report on AI systems autonomously replicating themselves across computers. Discusses self-propagation capabilities of recent models.
Developer tool that detects biased prompts before sending to AI models and suggests neutral reframes. Includes pattern detection for common bias types.
Tool for unified AI image generation using OpenAI's GPT-4V with structured outputs and API control.
4-part blog on lessons from launching AI projects in NHS AI Lab, including deployment, governance, and real-world implementation challenges.
Critical analysis of OpenAI o1's clinical diagnostic performance claims, examining evaluation methodology and comparator bias.
Opinion piece arguing MCP (Model Context Protocol) is unnecessary; proposes simpler alternatives using API wrappers and documentation.
Analysis of LLM behavior showing that arguing with models during errors degrades subsequent responses. Explains in-context learning limitations.
Tool enabling AI agents to discover and use Rust crate plugins, skills, and MCP servers with project-specific context.
Neo4j Labs CLI creating full-stack context graph agent apps with knowledge graphs, decision traces, and visualization for AI agents.
CLI tool generating MCP tools from OpenAI docs or databases, enabling LLMs to control custom applications via terminal or web UI.
Experimental open-source browser rendering engine that uses visual LLMs to interpret HTML instead of traditional rendering. Proof-of-concept.
Unsloth and Nvidia optimization achieving 25% faster LLM training on consumer GPUs.
Stackit and neuland.ai partner to provide LLM access via German data centers for GDPR-compliant AI deployment.
Analysis of GPT-5.5 code generation regressions where the model makes unrequested changes and improvements to unrelated code.
Agent-skills-eval is a test runner for evaluating whether Agent Skills improve model outputs through empirical testing of SKILL.md files.
Microsoft architectural framework for deploying agentic AI in public sector identity and eligibility verification systems.
Flue is a TypeScript framework for building AI agents with a built-in harness, designed as a headless, programmable alternative to Claude Code.
MRC Protocol is an open networking standard developed by OpenAI, AMD, Broadcom, Intel, Microsoft, and NVIDIA to improve GPU cluster performance for large-scale AI training.
Kstack is a skill pack for Claude Code enabling Kubernetes monitoring and troubleshooting through reusable commands like /investigate and /audit-security.
LAWS is a research paper on a transform operation that converts LLM inference into efficient cache lookups to reduce computational cost.
Bilig is a local-first spreadsheet engine for Node services and agents with formula parsing, binary sync protocol, and agent API support.
CreativityBench evaluating LLM creative problem-solving through tool repurposing based on affordance reasoning rather than canonical usage.
Tool-mediated LLM agent architecture with formal guarantees using Stackelberg game theory for autonomous cyber defense in SOCs.
Programmatic context augmentation improving LLM-based symbolic regression for mathematical expression discovery in scientific datasets.
Algorithm learning correct sequential agent behavior from 2-10 execution traces using dominator analysis for validation.
Terminus-4B 4B parameter model evaluated as replacement for frontier LLMs in agentic subtasks like code execution and debugging.
ADAPTS mixture-of-agents framework for automated depression and anxiety severity rating from clinical interviews via decomposed reasoning.
Evaluation of prompting strategies including CoT and PoT for deterministic computation accuracy in LLMs on mathematical tasks.
Cotomi Act browser agent learning from user behavior observation with adaptive execution achieving 80.4% task success via multi-step execution.
ROME framework for red-teaming tool-using agents with controlled benchmark rewriting to test deceptive safety scenarios.
Benchmark decomposing travel planning into five sub-capabilities to evaluate LLM long-horizon reasoning and tool use deficits.
LLM-assisted flexible MCTS approach for automated algorithm design to solve large-scale vehicle routing problems with minimal expert input.
Circuit analysis of LLM agent memory systems across Qwen models and frameworks, diagnosing write-manage-read pipeline failures.
GeoDecider agentic workflow applying multi-step reasoning and tool-use to lithology classification in geoscience applications.
Robust Agent Compensation framework enabling safe agent execution with recovery capabilities via logging, compatible with LangGraph and other agent frameworks.
Federated learning approach for aligning heterogeneous vision-language models without centralizing data across privacy-sensitive domains.
Taxonomy and methods for time series reasoning models applied to financial domain with multi-entity analysis and future prediction.
Benchmark for evaluating AI agents on workspace tasks requiring reasoning over file dependencies in realistic work environments.
Inference-time steering of LLM moral reasoning via convergent-divergent routing at minimal branch points in transformer blocks.
Self-improvement approach for generating high-quality plans in sub-exponential time using transformer-based generative models.
Adaptive many-shot in-context learning with semantic-aware KV cache reuse for varying query difficulty in LLMs.
Tiered memory architecture for long-running autonomous agents with structured episodic storage and weighted retrieval to prevent degradation over 72+ hours.
Reproducible benchmarking framework for evaluating LLM forecasting capabilities using knowledge cutoff and temporal masking.
VLM agents use visual-linguistic curiosity to drive exploration in partially observable environments via mental simulation before acting.
LLM framework for UAV swarm control using natural language mission descriptions with agent-enhanced reasoning for real-time closed-loop execution.
Bio-inspired memory framework for LLM agents on edge devices using optical forgetting to compress multimodal memories while reducing storage costs.
Autoresearch loop evolving interpretability tools specifically designed for agents rather than humans to conduct autonomous data science work.