Mathematical framework for ensuring dependability in distributed collaborative intelligence systems where local agent decisions compose into acceptable global behaviors.
Training reasoning planners with executor-grounded rewards beyond correctness to ensure faithful reasoning traces in LLM reasoning chains.
Language models self-improve through co-evolved discriminative rubrics without external supervision, eliminating reliance on human annotations or proprietary APIs.
Framework for efficient KV-cache handoff between multi-agent LLM systems on edge devices using quantization.
Argues for context-dependent objective optimization in frontier AI systems for open-ended and long-horizon agent tasks.
Framework for automated composition of multi-agent systems with agent recommendation replacing manual workflow creation.
Pluggable agent skill for dynamic retrieval strategy orchestration across heterogeneous tasks in RAG systems.
Conversational AI agent for symptom assessment evaluating LLM performance in real-world patient interviewing scenarios.
Automated red teaming framework for agentic AI systems that generates workflows in hours instead of weeks.
Framework for training search-capable LLM agents using informative and high-difficulty trajectories without industrial-scale resources.
Large-scale analysis showing frontier LLM personalities converge toward similar trait expressions despite different training.
Survey of safety risks, attacks, and defenses in embodied AI systems operating in autonomous domains.
AI interface for content exploration when users have vague intent, between passive feeds and structured search.
User-centric evaluation of explainability methods for AI-based medical image diagnosis systems.
Investigation of CLIP embedding contributions to memorization in Stable Diffusion text-to-image models.
Analysis of how verification errors impact reinforcement learning with verifiable rewards for LLM reasoning tasks.
Vision-language model approach for video anomaly detection with interpretable reasoning and spatial localization.
Research showing guard models lose safety alignment through standard fine-tuning on benign data, affecting agentic AI pipelines.
Foundation model for cardiotocography analysis using self-supervised learning on clinical fetal monitoring data.
EvoJail uses evolutionary algorithms to generate diverse jailbreak prompts for LLMs, addressing adaptability and diversity in automated attacks.
Analyzes generalization bounds of Spiking Neural Networks through Rademacher complexity for theoretical understanding of neuromorphic computing.
DeRelayL proposes sustainable decentralized relay learning to democratize large-scale model training across resource-limited participants.
Proteo-R1 develops reasoning foundation models for de novo protein design that explicitly model functional residues and interactions.
Multi-agent framework for multimodal controversy detection in videos by modeling diverse audience perspectives without training data.
PrismAgent is a zero-shot multi-agent framework for detecting harmful content in memes through interpretable case-analysis reasoning.
Healthcare AI GYM provides comprehensive training environment for medical AI agents through multi-turn RL with clinical reasoning tasks.
Studies pass-rate rewards in critic-free RL for code generation with LLMs, addressing sparse reward signals on challenging problems.
RouteHijack demonstrates routing-aware attacks on Mixture-of-Experts LLMs by exploiting expert selection mechanisms to bypass safety alignment.
Proposes Kernel Affine Hull Machines to replace neural inference with lightweight analytical estimators for efficient query-side semantic encoding in retrieval.
Introduces techniques to detect LLM jailbreaks by tracing refusal dynamics through latent trajectories, enabling robust adversarial detection.
Reward Hacking Benchmark evaluates safety vulnerabilities in RL-trained LLM agents with tool access through multi-step tasks with shortcut opportunities.
AutoRAGTuner automates optimization of Retrieval-Augmented Generation pipelines through declarative configuration, eliminating manual tuning of architecture and hyperparameters.
Framework analyzing gradient transport in large language model pretraining across scales using five observables separating cascade dynamics and efficiency metrics.
Cross-lingual safety alignment framework for LLMs via self-distillation, addressing vulnerability in low-resource language jailbreak attacks.
ARIS: open-source autonomous research harness for multi-agent LLM collaboration with assurance mechanisms for long-horizon research workflows.
MechaRule: extract symbolic rules from LLM internals via contrastive hierarchical ablation, grounding decision logic in neuron circuits.
MedStruct-S benchmark for semi-structured information extraction from OCR clinical reports: key discovery, conditioned QA, and key-value pair extraction.
Transformer inference acceleration exploiting low-rank token activation manifold with gated subspace decomposition for language models.
ARISE: repository-level graph representation system enabling AI agents for fault localization and automated program repair with semantic precision across code dependencies.
PIIGuard: webpage-level defense mechanism against PII harvesting by browsing-enabled LLM assistants using adversarial sanitization techniques.
Choreographic programming language for designing protocols in multi-agent agentic systems handling self-interest, private information, and untrusted participants.
Analysis of community-developed LLM applications from 2025 hackathon for materials science and chemistry, categorizing emerging usage patterns across research workflows.
Economic analysis of how AI systems affect labor markets and human-provenance verification as infrastructure. Societal implications rather than technical.
Survey of confidential computing security for LLM-driven agents handling secrets, credentials, and sensitive context across tool-calling and multi-agent protocols.
Self-mined hardness approach for safety fine-tuning of LLMs, scoring prompts by model rollout harm rates to improve robustness.
MAGE framework protecting LLM agents from long-horizon attacks using shadow memory to detect multi-turn exploitation patterns.
Ortho-Hydra method using orthogonalized mixture-of-experts LoRA to prevent style bleed in multi-style diffusion transformer fine-tuning.
RLDX-1 technical report on vision-language-action models for robotic control with improved memory, motion awareness, and physical sensing.
SHIELD dataset and distilled language models for clinical text de-identification, balancing performance with enterprise deployment constraints.
Cryptographic provenance system for AI package registries to prevent dependency confusion attacks in software distribution.