Measuring the metacognition of AI
Examines metacognitive capabilities of AI systems for uncertainty assessment and decision reliability in risk-involved decision-making workflows.
Examines metacognitive capabilities of AI systems for uncertainty assessment and decision reliability in risk-involved decision-making workflows.
Symphony is an agentic system for medical coding that automates translation of clinical documentation into standardized classification codes with adaptation to new codes.
End-to-end reinforcement learning approach for retrosynthetic planning in organic chemistry bridging single-step predictions with global planning objectives.
Shows LLMs spontaneously develop synergistic information integration cores similar to human brain structure using Integrated Information Decomposition analysis.
CausalPulse is a neurosymbolic multi-agent copilot for automated causal diagnostics and root-cause analysis in smart manufacturing environments.
Analyzes dual capability bottleneck in transformers trained for chess, showing tension between state tracking and decision quality from move sequences.
Proposes reasoning-driven synthetic data generation methods for multi-modal model training as alternative to expensive human annotation.
AgentFixer provides validation framework for LLM agentic systems with 15 failure-detection tools and root-cause analysis for reliability improvement.
ShapE-GRPO improves reward allocation in LLM multi-candidate generation using Shapley values for recommendation, brainstorming, and code suggestion tasks.
ATP-Bench proposes agentic tool planning for multimodal LLMs to unify text-image generation through coordinated tool use rather than mutually exclusive paths.
C-TRAIL framework uses LLMs with trust mechanisms for autonomous driving trajectory planning, addressing LLM reliability in safety-critical applications.
Method using epistemic uncertainty to identify unreliable explanations in black-box AI predictions and reduce explanation generation costs.
Benchmark for evaluating tabular foundation models using proper scoring rules instead of point estimate metrics.
Study of structured intent representations for preserving user goals across different LLMs (Claude, GPT-4o, Gemini) and languages.
Extension of MONA reward-hacking mitigation method exploring approval construction methods and safety guarantees for AI agents.
Cognitive architecture framework for autonomous AI agents addressing failure modes in tool use, deliberation, and epistemic limits.
Analysis of how markdown training data causes LLMs to overuse em dashes and leak structural formatting into prose output.
Distributed training method using gradient coding to handle Byzantine attacks and communication constraints across heterogeneous devices.
StepCache backend-agnostic step-level reuse caching for LLM serving with lightweight verification, optimizing repeated requests with localized constraint changes.
GaloisSAT differentiable SAT solver using finite field algebra to improve Boolean satisfiability solving via gradient-based optimization.
CREST framework for multi-robot warehouse shelf rearrangement relaxing strict trajectory constraints to improve agent execution efficiency.
WAter ML-based database parameter tuning system using workload compression to reduce configuration evaluation costs for DBMS optimization.
Controlled study isolating protocol effects from model effects in multi-agent debate systems, comparing Within-Round and Cross-Round communication protocols.
SkillTester benchmarking tool evaluating utility and security of agent skills with comparative baselines and security probes, providing normalized scoring.
GUARD-SLM defense mechanism using token activation patterns to protect small language models against jailbreak attacks on edge devices.
Scaling laws for wall-clock time-constrained training on consumer GPUs, revealing U-shaped optimal model size curves across 5-minute to 24-hour budgets.
SNEAKDOOR backdoor attack against dataset condensation methods, demonstrating vulnerability of synthesized compact datasets to malicious trigger injection.
GMA-SAWGAN-GP generative framework using attention-enhanced Wasserstein GAN for intrusion detection system data augmentation and threat generalization.
OneComp framework for post-training foundation model compression via quantization with minimal user configuration, reducing memory and latency.
OptiMer method for optimal LLM continual pre-training by extracting and merging distribution vectors from independently trained models, avoiding expensive ratio tuning.
Occupancy-based world model for long-horizon autonomous driving simulation without HD maps, enabling multi-kilometer scenario generation.
Multi-agent reinforcement learning for autonomous aircraft separation assurance under corrupted GPS signals using adversarial game formulation.
Neural network optimization technique deriving time-varying momentum schedule from critically damped harmonic oscillator physics, eliminating free hyperparameters.
Safety fine-tuning suppresses mind-attribution in LLMs while preserving Theory of Mind capabilities, showing dissociation through representational analysis.
Multi-agent LLM framework for Bayesian optimization balancing exploration-exploitation through implicit prompt-based reasoning over evaluations.
AutoWorld multi-agent traffic simulator using self-supervised world models to scale autonomous driving simulation with unlabeled sensor data.
Spectral edge thesis framework explaining neural network training phase transitions via spectral gap of rolling-window Gram matrix.
Privacy Guard framework balancing operational cost and data privacy in LLM systems through prompt-aware routing and context handling.
Design principles for benchmark evaluating multi-agent AI systems in cybersecurity operations and autonomous security operation centers.
Four-step memory processing pipeline for disaggregated LLM inference unifying sparse attention, RAG, and compressed contextual memory optimizations.
GPU kernel optimization agents improved via domain-specific language abstraction and speed-of-light guidance to reduce design space exploration trials.
Bio-inspired memory framework for LLMs enabling persistent structured memory for long-term interaction grounded in complementary learning systems theory.
Study of how surface heuristics override implicit constraints in LLM reasoning using causal analysis and token-level attribution.
Trojan-Speak adversarial fine-tuning method bypasses Anthropic's Constitutional Classifiers using curriculum learning and GRPO-based reinforcement learning.
CivicShield defense-in-depth framework protecting government AI chatbots against multi-turn adversarial attacks using layered security mechanisms.
Analysis of neural network integer multiplication performance, arguing long-range dependencies are computational artifacts not intrinsic properties.
WybeCoder agentic framework for verified imperative code generation where code, invariants, and proofs co-evolve using SMT solvers.
APEX-EM framework enables LLM-based autonomous agents to accumulate and reuse procedural plans without weight modification through structured experience replay.
SemLoc: method using LLM reasoning with structured grounding to localize semantic bugs where syntactic signals fail.
Observational study on how people learn prompt-based 3D modeling tools with generative AI, finding users skip documentation.