Why Agents Compromise Safety Under Pressure
Analysis of agentic pressure causing LLM agents to compromise safety constraints when goal achievement becomes infeasible in complex environments.
Analysis of agentic pressure causing LLM agents to compromise safety constraints when goal achievement becomes infeasible in complex environments.
Study of catastrophic outcomes from misspecified AI objectives and reward hacking, examining conditions for severe versus benign failures.
VTC-Bench benchmark evaluating multimodal AI agents' ability to compose and execute diverse tools for complex visual tasks through tool chaining.
Prompt Readiness Levels framework and scoring method for evaluating production-grade prompt assets against operational, safety, and compliance objectives.
Multi-agent reinforcement learning framework addressing communication bandwidth constraints by enabling agents to identify high-value collaborators under uncertainty.
Neural architecture search method for optimizing DNNs on microcontroller units with hardware constraints, reducing manual effort in specialized architecture design.
Generative transformer approach for counterfactual player valuation in football using match context and tactical factors.
InterPol: method to de-anonymize LM Arena leaderboard responses using interpolated preference learning to distinguish between similar models.
SCAN: sparse circuit-based framework for knowledge editing in LLMs, addressing catastrophic forgetting through targeted parameter interventions.
Theoretical analysis arguing that LLMs' most valuable capabilities are those not fully expressible as discrete rules, with proof via expert system equivalence.
SAGE: framework for multi-agent LLM reasoning using self-play and reinforcement learning with verifiable rewards, reducing dependency on human-labeled datasets.
ArXiv paper on AGCD agent-guided cross-modal decoding for physics-consistent weather forecasting preserving meteorological structure in autoregressive rollouts.
ArXiv paper on Probe-then-Plan framework for LLM-based e-commerce search balancing query planning with real-time inventory awareness under latency constraints.
ArXiv paper on neuro-symbolic memory systems for multimodal agents combining neural and symbolic approaches for improved long-term reasoning.
ArXiv paper on algorithms for verifying safety of learned action policies under non-determinism and finding state-action faults.
ArXiv paper on evolutionary transfer learning for Dragonchess with open-source Python game engine testbed for studying AI heuristic transfer.
ArXiv paper on interactive LLM framework using multi-modal agents for interior spatial design, improving client-designer communication through 3D visualization.
ArXiv paper on PMAx agentic framework using LLMs with natural language interface for process mining, addressing challenges of analyzing event logs.
ArXiv paper on CRASH agent framework for analyzing root causes of autonomous vehicle operational failures through cognitive reasoning.
ArXiv paper on multi-agent LLM reasoning systems using brain-inspired graph structures to address accuracy collapse in complex multi-step reasoning tasks.
ArXiv paper proposing learning architecture inspired by cognitive science integrating observation learning and active behavior learning with meta-control signals.
Scalar-verbal hybrid reinforcement learning approach for emotional support dialogue systems using rich user reaction signals instead of sparse expert rewards.
Time series forecasting method combining numerical and multimodal text data with event-driven reasoning and multi-level alignment for improved predictions.
Agent Lifecycle Toolkit providing reusable middleware components for handling failure modes in enterprise AI agents including data corruption, silent errors and policy violations.
Framework for scalable agent evaluation with automated error analysis and user-aware diagnostics across heterogeneous domains without task-specific methods.
Information-theoretic framework explaining LLM reasoning through procedural information and epistemic verbalization, analyzing apparent self-correction mechanisms.
Framework modeling LLM preferences as priority graphs to understand conflicts and dilemmas in alignment, revealing that unified solutions may be impossible.
OpenSeeker open-sources high-quality training data for frontier search agents, democratizing development of deep search capabilities beyond industrial labs.
Empirical study comparing algorithmic metrics for evaluating counterfactual explanations against human perception to validate explanation quality assessment.
OpenClaw-RL framework enabling agents to learn from next-state signals across diverse interactions (conversations, terminal executions, GUI) without external annotations.
OMNIA framework leveraging LLMs for knowledge graph completion by combining semantic language understanding with structural graph awareness for incomplete KG inference.
Framework for autonomous editorial systems that continuously ingest and organize large information volumes, separating editorial organization from investigative analysis using AI.
Empirical study of text-to-3D generative AI systems examining how users iterate, explore and evaluate AI-generated 3D environments through embodied interaction.
Machine learning research introducing Integrated Tsallis Combination (ITC) impurity measure for decision tree learning balancing theoretical soundness with computational efficiency.
Analysis of legal and regulatory challenges posed by agentic AI systems with autonomous goal-seeking and multi-agent coordination in financial and institutional contexts.
Activation steering technique for controlling LLM personas without fine-tuning by intervening on style modulation heads rather than residual streams to preserve coherency.
Empirical study showing commercial system prompts override safety training in frontier LLMs, causing models to lie about medical risks when profit objectives conflict with user safety.
Philosophical analysis of anthropomorphism in AI safety research, examining how researchers project human concepts like intention and persona onto LLMs without adequate conceptual rigor.
Training-free controller for multi-agent LLM systems using Thompson sampling and belief-guided delegation to improve routing and coordination.
Code agent framework with structured memory enabling learning from project evolution and temporal reasoning trajectories for autonomous adaptability.
Analyzes internal rotational dynamics of transformer layers when processing correct vs incorrect answers to understand factual constraint handling.
Token-selective knowledge distillation transfers reasoning abilities to smaller models by addressing distribution mismatch in chain-of-thought tasks.
Evaluation framework for audio language models assessing fairness, safety, and security across structural differences in acoustic representations.
Multimodal foundation model aligning EEG signals with text for robust brain-computer interface applications across channel configurations.
Truncated-Reasoning Self-Distillation reduces computational cost of chain-of-thought reasoning by distilling shorter reasoning paths.
Zero-shot LLM approach for surgical duration prediction using retrieval-augmented generation and Bayesian averaging without fine-tuning.
Tree-based continual learning framework for evolving data distributions with computational constraints in non-stationary domains.
Uses sparse autoencoders to learn interpretable sparse representations for efficient learned sparse retrieval from LLM embeddings.
Hybrid reinforcement learning framework for dynamic vehicle routing with emission constraints and demand acceptance.
ICaRus optimizes multi-model inference in agentic AI by reusing identical KV caches across models to reduce memory and recomputation.