Travelers deploys AI-powered claims countrywide with OpenAI
Travelers Insurance deploys AI-powered claims assistant using OpenAI, achieving 90% customer completion rate for claims processing.
Travelers Insurance deploys AI-powered claims assistant using OpenAI, achieving 90% customer completion rate for claims processing.
Tool converting LLM responses into interactive UI components via Model Context Protocol (MCP).
Open-source local-first development intelligence system that tracks project history for use with AI coding assistants.
Discussion of best practices for LLM development and usage patterns.
AI agentic coding reshaping software development roles and team structures.
CPU-only tool for text summarization, fact-checking, explanation and translation with local model inference.
OpenAI introduces role-specific Codex plugins and tools for analysts, marketers, designers, and other non-developer roles. 20% of users are non-developers.
Handler: Mac IDE adding review layer to AI code generation with per-edit explanations and contextual chat.
Research comparing data efficiency of world models versus LLMs for training.
OpenAIRE 12-week hackathon on AI and open science with multiple skill tracks launching June 2.
Chrome extension that converts webpage screenshots into production-ready Tailwind CSS and HTML code.
Local-first browser sidebar for storing, searching, and chatting with web content using Ollama and RAG.
Multi-model AI workspace with custom knowledge grounding, file upload, and multimodal support.
Research on aligning LLMs without persona-based training approaches.
Tiiny AI Pocket Lab: 300g offline device running local AI models without subscriptions, raised $1M on Kickstarter.
Discussion exploring whether entire jobs and fields could be replaced by AI, comparing perspectives from workers and management.
GitHub integration tool using AI to generate test cases and QA reports from PRs automatically.
Linux kernel module enabling RDMA over USB4/Thunderbolt for distributed AI workloads across consumer hardware. Supports vLLM and tensor-parallel inference.
Uses LLM-powered agents to simulate context-aware user behavior for evaluating recommender systems beyond offline A/B testing.
Framework for 4D reconstruction of dynamic scenes combining visual-inertial sensing with semantic understanding for robotics.
Uses 3D generative models and vision foundation models to generate diverse robot manipulation demonstrations via afford correspondence.
Improves on-policy distillation for LLM alignment with adaptive weighting and token-level credit assignment.
Explores internalizing tool knowledge within LLMs during reasoning to avoid external tool documentation overhead and improve efficiency.
Analyzes temporal routing pathologies in multi-timescale PPO reinforcement learning and proposes representation-based alternatives.
AI agents use type annotations to draft and mechanize mathematical proofs in Isabelle, generalizing from hints.
Unifies reward-based fine-tuning methods for diffusion and flow models under reward score matching framework.
Addresses state drift in vision-language navigation agents using video LLMs to follow natural language instructions in 3D environments.
Deep Interest Mining framework for generating semantic IDs in multimodal recommendation systems with intent-enriched item vocabulary.
FlowPlace: flow matching generative model for chip placement overcoming synthetic data pre-training and sampling time limitations.
Reinforcement learning approach for chip placement learning from expert layouts to achieve expert-quality physical design.
Method for epistemic uncertainty modeling in deep neural networks balancing Bayesian principled estimates with computational efficiency.
STABLEVAL framework for disagreement-aware evaluation of AI systems, modeling latent item correctness and annotator reliability to improve ranking stability.
SemGrad: gradient-based uncertainty quantification method for LLM free-form generation, sampling-free and computationally efficient alternative to existing approaches.
Theoretical framework for steering intermediate representations in generative models, formalizing concept steering through affine concept erasure.
Prune-OPD: efficient on-policy distillation for long-horizon LLM reasoning by pruning diverged trajectories.
Analysis of Cartesian Shortcut vulnerability in vision-reasoning benchmarks where orthogonal grid layouts enable coordinate exploitation.
Self-distillation method with outcome-guided logit steering to improve LLM reasoning through on-policy learning.
On-policy distillation approach leveraging peer trajectories to provide denser token-level supervision for LLM reasoning improvement.
Adversarial attack method for eliciting LLM hallucinations through semantically coherent prompts with constrained optimization.
RISED: pre-deployment evaluation framework for clinical decision-support AI systems across reliability, inclusivity, and operational dimensions.
Contrastive reformulation of GRPO for LLM post-training on reasoning tasks revealing limitations in reward design.
Training-free token pruning method for reducing visual token overhead in vision-language models while preserving pixel grounding.
Analysis of many-shot chain-of-thought in-context learning for reasoning tasks across various LLM architectures and task types.
Autonomous AI agents for scientific discovery in cosmology using LLM-guided code evolution and multi-agent research laboratories.
PBT-Bench: benchmark for evaluating AI agents on property-based testing, measuring semantic invariant derivation and input-generation strategy construction.
Framework for enforcing privacy policies in RAG systems using density estimators to prevent PII leakage.
Research on knowledge distillation with bilevel optimization for imbalanced learning scenarios.
Research paper identifying phase transition in LLM scaling where reasoning and truthfulness shift from anticorrelated to cooperative behavior.
Mirage representation-level auditing framework for certifying visual unlearning in federated learning contexts.
Co-Fusion4D framework for robust 3D object detection in autonomous driving using spatiotemporal fusion.