REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations
Method for generating realistic adversarial prompts that elicit LLM hallucinations through constrained optimization of semantically coherent inputs.
Method for generating realistic adversarial prompts that elicit LLM hallucinations through constrained optimization of semantically coherent inputs.
Foresight Learning approach for training LLMs on longitudinal clinical notes from MIMIC-III for predicting future patient events.
Research testing whether LLM-based agent simulations generate mechanistically plausible phenomena in social simulations and game-theoretic scenarios.
Orthrus framework unifying autoregressive LLM fidelity with diffusion model parallel token generation for efficient high-throughput inference.
Deep learning approach for detecting image manipulation and AI-generated content using forensic routing and multi-path evidence fusion.
Research on model extraction attacks against Graph Neural Networks and defense mechanisms, addressing GNN theft vulnerability as cloud services.
Proposes Bayesian approach for merging task-specific expert models without retraining. Leverages anchor model bias for parameter estimation in model merging.
Hypothesizes emergent misalignment in fine-tuned LLMs involves persona-model collapse. Tests behavioral degradation using moral susceptibility and robustness metrics.
Multi-agent reinforcement learning system for RTL hardware code generation. Addresses vendor air-gap security and proprietary codebase training constraints in chip design.
Introduces language-based agent control (LBAC) programming model using static typing and runtime enforcement for agentic applications. Brings security concepts from programming languages to agent control.
Proposes survival analysis framework for evaluating LLM robustness to repeated jailbreak attacks. Captures temporal dynamics of adversarial pressure beyond binary success metrics.
Studies how LLM web-search agents' influence depends on trajectory through multiple documents. Extends generative engine optimization to agentic multi-step browsing scenarios.
Proposes RISED framework for pre-deployment evaluation of clinical AI systems across reliability, inclusivity, sensitivity, equity, and deployability dimensions.
Studies data difficulty's role in LLM fine-tuning from empirical and theoretical perspectives. Examines tradeoff between generalization and extrapolation in supervised fine-tuning.
Proposes multi-agent coordination through dialogue-based world model alignment in partially observable environments. Addresses communication as bridge for embodied agent collaboration.
Identifies 'lucky pass' problem in SWE-Agent evaluation where agents pass tests through trial-and-error rather than principled solutions. Analyzes 2,614 trajectories to show outcome-only metrics miss process quality.
Comparative analysis of probabilistic circuits and transformers for autoregressive language modeling, identifying expressivity gaps.
Statistical framework for determining optimal release timing in iterative LLM generate-verify workflows using always-valid inference.
Test-time multimodal agent for language-guided segmentation using iterative reasoning between MLLMs and foundation models.
Efficient long video understanding through adaptive relevance-diversity frame sampling with zero-cache memory overhead.
Algebraic method projecting LLM hidden states into Galois Field F2 to control logical consistency and verify ontological relations.
Contrastive perspective on reinforcement learning with verifiable rewards, reformulating GRPO as weighted score difference optimization.
Protocol-driven development framework for governing generated software through algebraic invariants and evidence collection.
Investigation of sycophancy in multi-agent LLM pipelines, showing alignment training alone insufficient to fix disagreement-induced errors.
Method for image inpainting with pretrained diffusion models using reusable offline-trained guidance modules.
Analysis and acceleration methods for training masked diffusion models to match autoregressive model training speed.
Technique to reduce feature drift and improve performance of merged expert models through feature calibration.
Semantic fuzzing framework identifies safety specification violations in LLM-powered agent skills where benign inputs cause guardrail breaches.
Framework to evaluate alignment between vision-language models and human perception using counterfactual semantic saliency analysis.
Method for adapting LLMs to downstream tasks via context optimization and active information seeking without weight updates.
Framework for cross-domain offline reinforcement learning that expands source data coverage while aligning to target distribution.
Study of proprioceptive encodings for robust robotic manipulation under distribution shift and unseen test conditions.
Differentiable approach for graph partitioning and parameter initialization in quantum approximate optimization algorithm.
Scaling few-shot spoken word classification to 1000 classes with five shots per class using generative meta-continual learning.
Framework for allocating responsibility in multi-agent systems using counterfactual reasoning in probabilistic games.
Analysis of Muon optimizer showing spectral flattening mechanism enables larger learning rates and faster convergence than other optimizers.
Meta-learning approach for multilingual few-shot spoken word classification using generative meta-continual learning.
Complexity-tiered benchmark for Indian language speech recognition identifying studio-bias in multilingual ASR models.
Analysis of watermarking as monitoring primitive for generative models, examining attribution and detection under adversarial conditions.
Classifier guidance method for chemical synthesis planning using margin-calibrated steering of retrosynthesis models.
Fully automated multi-agent framework for venture capital due diligence combining LLMs with real-time web retrieval and data extraction.
Audit framework for measuring gender bias in text-to-image generative models using risk-tiered use-case profiles.
Approach integrating dexterous grasping with language-guided semantics for reliable robotic manipulation.
Framework combining VLM agents for high-level reasoning with specialized VLA tools for long-horizon embodied robotic tasks.
Graph neural network method addressing over-squashing in multi-label graphs using information bottleneck principles.
Global premise retrieval system for Lean 4 theorem proving that identifies relevant library lemmas needed for complete proofs.
Data generation method using acquisition functions to synthesize targeted training samples and improve model quality.
Unsupervised 3D instance segmentation method bridging synthetic-to-real domain gap using evolving object-centric representations.
Framework for improving MLLM spatial understanding using 360-degree panoramic sensing for navigation and 3D scene understanding tasks.
Benchmark for studying multi-agent coordination paradigms in hierarchical, dynamically evolving industrial scheduling systems with coupled constraints.