From Features to Actions: Explainability in Traditional and Agentic AI Systems
Research on explainability methods for agentic AI systems that operate over multi-step trajectories, extending beyond single-prediction interpretability.
Research on explainability methods for agentic AI systems that operate over multi-step trajectories, extending beyond single-prediction interpretability.
Sparse video generation framework enables vision-language navigation agents to navigate unknown environments with minimal high-level instructions via beyond-the-view reasoning.
Analysis of GRPO reinforcement learning limitations in LLM reasoning due to implicit advantage symmetry; proposes improvements for exploration and difficulty adaptation.
Autonomous lab combining GPT-5 with Ginkgo Bioworks' automation reduced cell-free protein synthesis costs 40% via closed-loop experimentation.
GPT-5.3-Codex is a Codex-native agent for long-horizon technical work pairing coding performance with general reasoning.
GPT-5.3-Codex system card detailing capabilities of an agentic coding model combining frontier coding performance with reasoning abilities.
Technical guide to embedding Codex agent via App Server using bidirectional JSON-RPC API with streaming, tool use, and diffs.
GeneralVLA: vision-language-action model with knowledge-guided trajectory planning to improve zero-shot generalization in robotic control.
BPDQ: bit-plane decomposition quantization with variable grid for efficient 2-3 bit LLM inference under memory constraints.
Reasoning Cache (RC) algorithm enables LLMs to improve over long horizons via test-time adaptation and RL, improving extrapolation beyond training distribution.
Investigates computational efficiency advantages of diffusion-based language models versus standard approaches.
Novel method for fine-tuning quantized LLMs using evolution strategies instead of backpropagation, enabling high-precision adaptation on discrete, non-differentiable parameter spaces.
TIC-VLA framework for robot navigation that models delayed semantic reasoning in vision-language-action models for real-time control in dynamic environments.
ECHO-2 distributed RL framework for LLM post-training with remote inference workers, addressing cost efficiency and policy coordination challenges.
OpenAI's internal data agent architecture using GPT-5, Codex, and memory systems to reason over large datasets and deliver insights.
Announcement of model retirements: GPT-4o, GPT-4.1, GPT-4.1 mini, and o4-mini retiring from ChatGPT on February 13, 2026.
Taisei Corporation scales ChatGPT Enterprise across HR and global construction operations for talent development.
OpenAI details data protection mechanisms for AI agents, preventing URL-based exfiltration and prompt injection attacks.
TRUSTBANK partners with Recursive to build Choice AI, a multi-agent system using OpenAI models for personalized donation recommendations.
Opinion article debating whether prompt-driven development qualifies as legitimate programming practice.
Technical guide on implementing agentic workflows for automated document processing and data extraction.
Technical deep dive into Codex agent architecture, explaining orchestration of models, tools, prompts, and performance via Responses API.
Technical overview of diffusion transformers including vision transformers and diffusion transformer variants.
Technical case study on scaling PostgreSQL to handle millions of queries per second for ChatGPT using replicas, caching, and workload isolation.
Praktika builds adaptive AI language tutors using GPT-4.1 and GPT-5.2 for personalized lessons and progress tracking.
Data-driven report analyzing worker adoption patterns and use cases of ChatGPT across industries and departments.
Higgsfield uses GPT-4.1, GPT-5, and Sora 2 to generate cinematic social videos from simple user inputs.
Cisco and OpenAI's Codex agent embeds AI into engineering workflows to automate builds, defect fixes, and enable AI-native development.
Explores text diffusion models as emerging alternative paradigm for language model development.
Veo 3.1 video generation model improves consistency and control with vertical video support
Zenken deploys ChatGPT Enterprise company-wide, improving sales performance, preparation time, and proposal success rates.
Netomi scales enterprise AI agents using GPT-4.1 and GPT-5.2 with concurrency, governance, and multi-step reasoning for production workflows.
Tolan builds voice-first AI companion with GPT-5.1 using low-latency responses, real-time context reconstruction, and memory-driven personalities.
OpenAI hardens ChatGPT Atlas against prompt injection attacks using automated red teaming with reinforcement learning for proactive vulnerability discovery.
OpenAI releases GPT-5.2-Codex, an advanced coding model with long-horizon reasoning, large-scale code transformations, and cybersecurity capabilities.
OpenAI releases GPT-5.2-Codex, an advanced coding model with long-horizon reasoning, large-scale code transformations, and cybersecurity capabilities.
Safety documentation for GPT-5.2-Codex covering model-level and product-level mitigations for coding tasks and agent deployment.
ChatGPT app marketplace launched; developers can submit apps for review with updated SDK and guidelines for chat-native integrations.
Overview of enterprise AI adoption, challenges, and potential for funding accessible AI. Discusses organizational AI implementation but lacks technical depth.
Gemma Scope 2 open-source interpretability tools release for Gemma 3 models to study language model behavior
GPT-5.2 achieves state-of-the-art on math/science benchmarks (GPQA Diamond, FrontierMath) with applications to research and mathematical proofs.
GPT-5.2 announced as advanced frontier model with reasoning, long-context, coding, and vision capabilities for agentic workflows.
Safety mitigation documentation for GPT-5.2 model family covering training data sources and safety approaches.
FACTS Benchmark Suite provides systematic evaluation framework for measuring factuality in large language models
OpenAI co-founds Agentic AI Foundation under Linux Foundation, donates AGENTS.md standard for safe, interoperable agentic AI systems.
OpenAI launches certification courses and AI Foundations program to build real-world AI skills and career preparation.
Commonwealth Bank deploys ChatGPT Enterprise to 50,000 employees to build AI fluency and improve customer service and fraud detection.
OpenAI researchers develop 'confessions' training method that teaches models to acknowledge mistakes, improving transparency and trustworthiness.
Mirakl uses AI agents and ChatGPT Enterprise for commerce optimization, improving documentation and customer support with agent-native architecture.
Accenture and OpenAI partnership to deploy agentic AI in enterprises, with ChatGPT Enterprise rollout to tens of thousands of Accenture professionals.