Beyond Relevance: Utility-Centric Retrieval in the LLM Era
RAG systems should optimize for utility (task completion) rather than topical relevance when retrieving documents for LLMs.
RAG systems should optimize for utility (task completion) rather than topical relevance when retrieving documents for LLMs.
MuTSE: Human-in-the-loop evaluator tool for systematically comparing LLM text simplification outputs across different prompting strategies and architectures.
WOMBET: Framework for reinforcement learning that generates and transfers experience data between source and target robotic tasks for sample efficiency.
Aligned Agents, Biased Swarm: Empirical study measuring how multi-agent system topologies and feedback loops amplify bias in emergent behaviors.
Litmus ReAgent: Benchmark and agentic system for evaluating multilingual LLM performance prediction across 1,500 questions spanning six tasks and five evidence scenarios.
PerMix-RLVR: Training method for aligning LLM personas with reward models while preserving output diversity, avoiding inference-time computation overhead.
PinpointQA dataset and benchmark for evaluating small object localization and spatial reasoning in video MLLMs.
ASTRA: adaptive semantic tree reasoning architecture for LLM-based complex table question answering.
Regime-conditional retrieval with transferable router for two-hop question answering with theoretical foundations.
Noise-aware in-context learning approach to mitigate hallucinations in auditory large language models.
ImageProtector prevents multi-modal LLMs from analyzing images via visual prompt injection attacks.
Vision-language models for image geolocation with structured geographic reasoning and autonomous self-evolution.
CONDESION-BENCH evaluates LLM decision-making with compositional action spaces and conditional feasibility constraints.
Watt Counts: open-access energy consumption benchmark for LLM inference across 50 models and 10 GPU architectures.
PDYffusion combines diffusion models with physics-informed dynamics for long-horizon spatiotemporal prediction.
Vision-Language-Action models for autonomous driving combining perception, reasoning, and temporal dynamics modeling.
Method integrating graph-based embeddings into event sequence models for improved user prediction on digital platforms.
DeepGuard improves secure code generation by LLMs through multi-layer semantic aggregation to mitigate vulnerable patterns.
CLIP-Inspector detects backdoor attacks in prompt-tuned vision-language models through out-of-distribution trigger inversion.
Research on detecting covert misaligned AI behavior in real-world settings using open-source intelligence methods.
TensorHub introduces Reference-Oriented Storage for efficient weight transfer in LLM reinforcement learning across heterogeneous computational resources.
PS-TTS method for phonetic synchronization in automated dubbing, addressing duration and lip-sync challenges in AI-based video translation.
Interactive ASR system with human-like interaction and semantic coherence evaluation, replacing WER metric with agent-based correction mechanisms.
EquiformerV3: SE(3)-equivariant graph attention Transformer for 3D atomistic modeling, improving efficiency, expressivity, and physical consistency.
CORA framework for risk-controlled GUI automation agents using conformal prediction to provide formally verified, user-tunable safety guarantees for VLM-powered mobile automation.
LLM-based agents for scaffolding diagnostic reasoning in educational settings, combining scenario-based learning with learning analytics and personalized support.
Dataset for personality-shaped emotional responses to text events, addressing limitations of LLM role-playing and personality illusion in affective computing.
Theoretical analysis of generalization and scaling laws for Mixture-of-Experts Transformers, separating active capacity from routing combinatorics with covering-number bounds.
Symbolic-Neural Consistency Audit framework extracting and formalizing LLM self-stated safety policies.
GNN-based deep reinforcement learning scheduler for cloud workflow DAG assignment minimizing time and energy.
GRM gradient-ratio masking attack on audio LLMs balancing jailbreak success with utility preservation.
Mosaic multimodal jailbreak attack against closed-source VLMs via multi-view ensemble optimization.
SkillMOO multi-objective optimization framework automatically evolving agent skill bundles for coding tasks.
Visually-guided policy optimization improving visual faithfulness in vision-language models via reinforcement learning.
LLM-Rosetta hub-and-spoke intermediate representation for cross-provider LLM API translation and interoperability.
BadSkill: backdoor attack formulation exploiting model artifacts bundled in agent skills.
AI Codebase Maturity Model framework for systematic progression from assisted coding to self-sustaining systems.
Experimental evaluation framework for quantum-inspired 1024-D document embeddings in RAG and information retrieval applications.
Instruction Hierarchy in LLM Agents arXiv paper addressing multi-source conflicting instructions in LLM systems. Examines privilege levels for safe instruction following.
ECHO arXiv paper on one-step diffusion model for chest X-ray report generation. Compresses multi-step denoising to single parallel generation step.
SafeAdapt arXiv paper on provably safe policy updates in deep RL for non-stationary environments. Addresses safety preservation during policy changes.
Attack method demonstrating model poisoning vulnerabilities in federated learning without requiring collusion between adversarial clients.
Post-training approach enabling LLMs to effectively retrieve and use long-context information for improved reasoning capabilities.
BERT-based evaluation method for LLM outputs that addresses limitations of rigid lexical evaluation and formatting-dependent scoring.
Agentic system for visual retrieval-augmented generation with iterative search and multi-step reasoning across visually rich documents.
Theoretical framework showing how agents with different computational capacities can develop distinct semantic alphabets for communication.
Method for predicting future scene evolution by modeling uncertainty and simulating trajectories rather than dense pixel-level changes.
Technique for decoupled confidence calibration in large vision-language models to reduce hallucinations and improve reliability.
Approach using synthetic images to improve visual perception capabilities in vision-language models for spatial reasoning tasks.
Method for robust prompt learning in vision-language models that leverages visual content to handle label noise effectively.