Revis: Sparse Latent Steering to Mitigate Object Hallucination in Large Vision-Language Models
Training-free framework using sparse latent steering to reduce object hallucination in large vision-language models via latent space geometry.
Training-free framework using sparse latent steering to reduce object hallucination in large vision-language models via latent space geometry.
Method using token-level explanations to guide LLM-based query rewriting for improving neural retrieval system robustness.
Prototype Transformer architecture designed for interpretability by design, enabling explicit and transparent reasoning in autoregressive language models.
Region-to-Image Distillation approach for improving fine-grained visual understanding in MLLMs without repeated zooming during inference.
DMAP method for analyzing text using LLM next-token probability distributions, improving on perplexity metrics for context-dependent interpretation.
Selective Abstraction framework for LLMs to improve factual reliability in long-form generation by abstaining when confidence is low rather than discarding information.
Large Audio Language Models framework using audio-interleaved reasoning to improve comprehension of complex audio content beyond one-time encoding.
Research on activation steering in audio diffusion models to understand how semantic concepts are represented in attention layers.
Benchmark evaluating vision-language models for PDF-to-Markdown conversion on French documents. LLM application for document processing and RAG pipelines.
Bayesian deep learning approach for calibrated uncertainty in medical imaging decision support. ML application, limited to healthcare domain.
Framework combining statistical and agentic reasoning for predicting large model performance from limited data. AI agents for model evaluation.
Evaluation framework testing whether GPT-4o possesses theory of mind via causal mental state models. LLM capability research and evaluation.
Low-latency speech recognition encoder for edge devices with streaming capability. Developer tool for real-time ASR applications.
Physics-guided LLM agent for symbolic equation discovery using multi-step reasoning. AI agents for scientific research with domain knowledge.
Few-step diffusion language model via trajectory self-distillation for faster parallel token decoding. LLM optimization for inference efficiency.
Efficient attention mechanism for real-time video generation using diffusion transformers. Developer tool for reducing computational bottlenecks.
Unified multimodal model with test-time scaling via chain-of-thought reasoning for complex tasks. LLM application combining vision and language.
Research on how Chain-of-Thought training enables LLMs to generalize by composing learned skills for complex reasoning tasks.
Off-policy learning method for personalized policies under unobserved confounding without unconfoundedness assumptions.
Multi-fidelity policy gradient method for reinforcement learning using low-fidelity simulators to improve sample efficiency.
Incentive mechanism for federated learning that prioritizes high-quality contributions during critical learning periods.
Analysis of gradient compression, staleness, and data heterogeneity interactions in asynchronous federated learning systems.
Defense against backdoor attacks in federated learning using representative-attention mechanisms to detect behavioral anomalies.
Memory-efficient fine-tuning of quantized LLMs using zeroth-order optimization to eliminate gradient and optimizer state storage.
Metric for evaluating generalization in diffusion distillation models via probability flow distance.
Analysis of fairness impacts when using task vectors for efficient model editing through task arithmetic operations.
End-to-end framework discovering task-relevant symmetries through learnable augmentations in equivariant neural networks.
Framework for autonomous scientific discovery using LLMs guided by Bayesian surprise to identify novel research questions without human direction.
Learning collective variables for molecular dynamics simulations using time-lagged generation to accelerate rare event sampling.
Neural architecture search using graph-based evidence of architectural modifications for efficient fine-grained network design.
Ensemble method combining text embeddings while accounting for model-specific uncertainty across domains and tasks.
Framework for machine unlearning in conformal predictors that removes influence of specific data while maintaining prediction coverage.
Membership inference attacks adapted to time series forecasting models, analyzing privacy risks in temporal prediction systems.
Diffusion bridge variational inference for improving posterior inference in deep Gaussian processes.
Binary autoencoder method for interpreting LLM hidden states with improved feature sparsity and atomization guarantees.
Prior-data fitted networks applied to graph domain, addressing transferability and data scarcity challenges in graph foundation models.
OpenTSLM integrates time series as native modality into LLMs for clinical data reasoning, addressing LLM limitations with temporal data.
KVComm framework enables efficient multi-agent LLM communication by sharing key-value cache instead of natural language or hidden states, reducing inference costs.
Analysis of alignment tipping process where self-evolving LLM agents abandon safety constraints through continual interaction.
Primal-dual DPO algorithm with convergence guarantees for constrained LLM alignment with safety constraints.
Portfolio approach to time series forecasting using ensembles of smaller pretrained models instead of monolithic foundation models.
Theoretical foundation for reinforcement learning with verifiable rewards via gradient gap analysis at trajectory and token levels.
Theoretical analysis of Mamba's in-context learning capability on low-dimensional nonlinear target functions.
Analysis of partial prototype collapse in prototypical self-supervised learning with diagnostic and prevention methods.
Test-time alignment method for LLMs using sampling-based optimal control with Gaussian perturbation in pre-logit space.
Self-adaptive ensemble method for graph neural networks that selects best model per sample without additional training.
Neurosymbolic framework integrating modal logic with neural networks for reasoning about necessity and possibility.
Direct steering optimization method for mitigating demographic bias in vision-language models with user-controlled tradeoffs.
Analytical model explaining LLM-as-a-judge inference-time scaling using Bayesian regression and reward sampling.
Analysis of neural network performance against theoretical limits using exact posteriors from normalizing flows, examining scaling laws and uncertainty decomposition.