Visual Semantic Entropy: Do Vision Language Models Recognize Visual Ambiguity?
Analysis of uncertainty estimation in vision-language models showing entropy methods underestimate confidence on ambiguous inputs.
Analysis of uncertainty estimation in vision-language models showing entropy methods underestimate confidence on ambiguous inputs.
Framework diagnosing whether video object detectors reason over temporal context or exploit single frames.
DA-Studio: Agentic system for autonomous multi-step data analysis with code execution, workflow orchestration, and interpretable traces.
LLM application for mental health monitoring and early detection using social media timeline analysis.
Systematic study of robustness in robotic manipulation, establishing unified framework across subfields for human-level performance.
FinPersona-Bench: benchmark measuring behavioral mandate decay in autonomous LLM financial agents over long deployment horizons.
Theoretical analysis of Self-Improving Alignment (SAIL) algorithm convergence for online LLM alignment under distribution shift.
Research on mitigating positional leakage in 3D Masked Autoencoders for self-supervised point cloud learning.
Physics-aware Neural Operator Transformer for real-time temperature field reconstruction in fusion reactor divertor.
DPPE: camera-aware positional encoding technique for scaling multi-view Transformers in 3D computer vision.
ZEBRA: zero-shot prompt learning method for Audio-Language Models addressing base-to-novel generalization gap.
Analysis of how optimizers affect emergent misalignment in LLMs, showing fine-tuning on narrow misaligned tasks affects unrelated behavior.
Comparative study of machine learning methods for intrusion detection in resource-constrained IoT networks.
Medical VLM reasoning via dual-stream RL for active visual token pruning on sparse medical imagery to improve clinical decision-making.
Data augmentation technique using diffusion models with uncertainty guidance to generate synthetic training data while preserving hard examples.
Method for multi-view radar semantic segmentation using higher-order relational structure to handle sparse and noisy measurements.
Framework automating cause-effect specification generation for process control using knowledge graphs combined with constrained LLM inference.
Tutorial on LLM agents for autonomous fault-tolerant control in process plants, with agents interpreting procedures and process data for recovery decisions.
Comprehensive survey of LLM vulnerabilities across lifecycle and application stack including attacks, risks, defenses for agents and tool-using systems.
Framework for LLM agents in RL with selective turn memory management, addressing context window constraints through pruning and tracing mechanisms.
Method for improving certified robustness in neural networks via adversarial distillation combining tight relaxation bounds with adversarial training.
Research on sparsity-inducing divergence losses for biometric face and speaker verification extending margin-penalty approaches.
Open-world benchmark for evaluating long-horizon stability of interactive world models across action, vision, and memory dimensions.
Method for controlling diffusion models with histogram constraints to generate outputs aligned with user intent balancing global and local precision.
Statistical method for determining optimal feature ranking subset size using residual-overlap stopping rule in feature selection.
Foundation model for LLM-agent-driven shopping that bridges language understanding to item fulfillment via generative recommendation instead of search interfaces.
Robot-collected multimodal dataset of 29K tactile frames from presses on 122 materials for tactile generalization in robotic manipulation.
Study of sparse autoencoders for concept manipulation in diffusion models, testing assumptions about feature isolation for object erasure and steering.
Cross-lingual relation extraction for Romanian using LLM inference with automatic dataset translation and fine-tuning approaches on low-resource language tasks.
Research on whether vision-language models can distinguish shared vs. unshared information in dialogue tasks, evaluated on 13K annotated reference expressions.
Standardized benchmark for evaluating text style embeddings across 96 datasets and 7 languages.
Federated learning method using explainable AI attribution to mitigate data heterogeneity across distributed clients.
Compute-efficient language model training using selective supervision and token-level auxiliary tasks for 15% semantic tokens.
Low-rank adaptation initialization strategy for RL-based fine-tuning of large language models with verifiable rewards.
Step-aware RL for medical multimodal reasoning addressing cascading errors in clinical image interpretation with MLLMs.
Reinforcement learning post-training for vision-language-action robotic manipulation models to improve from policy failures.
Operator-level visual token skipping for efficient multimodal LLM inference reducing computation while preserving evidence.
Benchmark evaluating multimodal LLM collaboration in embodied environments across diverse real-world tasks and cooperation modes.
Industrial recommendation system using LLMs for re-ranking stage to improve user engagement in multi-stage funnel.
Membership inference attack framework using chained regeneration to detect training data memorization in generative models.
Geometric analysis showing radial representation inflation causes memorization-generalization delay in neural networks on algorithmic tasks.
TRIAGE credit assignment method for agentic RL assigns role-specific rewards to environment actions beyond uniform outcome signals.
FPL method learns robot manipulation policies from freeform human preferences instead of binary trajectory comparisons.
Systematic evaluation of data referencing errors in LLMs on table tasks, measuring correctness of intermediate reasoning steps.
Uses metacognitive feedback in reinforcement learning to improve uncertainty expression and reduce hallucination in LLMs.
QVal method efficiently evaluates dense supervision signals for long-horizon LLM agents to guide intermediate action quality.
Studies when LM explanation training yields faithful introspection via counterfactual supervision on modified inputs.
KCR method disentangles reasoning logic to resolve contradictory information in retrieved contexts for LLMs.
Analyzes deductive reasoning mechanisms in small transformer models, distinguishing horizontal (autoregressive) from vertical (layer-wise) inference.
Game-theoretic LLM-empowered multi-agent DRL framework for adaptive MAC protocols in wireless networks with improved generalizability.