Prompt Injection as Role Confusion
Analysis of prompt injection vulnerabilities traced to role confusion in LLMs, with novel detection probes and mitigation approaches.
Analysis of prompt injection vulnerabilities traced to role confusion in LLMs, with novel detection probes and mitigation approaches.
LLM-augmented graph learning approach for detecting miscitations in scholarly networks using context-aware analysis.
Training-free method for controlling LLM behavior via cross-layer activation steering with improved consistency.
Neural network architecture for learning causal reasoning with hierarchical primitives and dynamic composition.
Multi-agent agentic framework for video quality evaluation and iterative improvement across diverse generation tasks.
Framework applying non-equilibrium thermodynamics to formalize curriculum learning in reinforcement learning.
Reinforcement learning exploration method using maximum entropy without rollouts for efficient state space coverage.
Formally verified evaluation framework for AI-guided scientific candidate selection with budget constraints, accounting for LLM proposal reliability.
MLLM framework for pixel-level video grounding with improved spatial precision and temporal tracking consistency.
Test-time optimization strategies for agentic RAG systems to reduce inefficient retrieval and improve accuracy on complex multi-hop questions.
Research on connecting layers across Vision Foundation Models (CLIP, etc.) to measure representational compatibility via stitching.
NLP framework maps cyber incidents to MITRE ATT&CK techniques using automated classification for small enterprise risk management.
Topology-regularized benchmark for evaluating LLMs on multi-hop medical reasoning, addressing shortcut learning in knowledge graph navigation.
Industrial case study evaluating AI-assisted API design workflow comparing AI-generated vs human-authored API specifications with 16 experts.
One-Step Flow Policy framework distills generative flow models into single-step robotic policies for low-latency visuomotor control.
TRACE framework combines knowledge graphs, temporal rules, and LLM-guided reasoning for interpretable stock movement prediction.
Naïve PAINE improves text-to-image generation by evaluating and optimizing prompts for diffusion models to reduce generation cycles.
Q-DIG framework red-teams Vision-Language-Action models via quality diversity prompt generation to improve robotic policy robustness.
Study showing LLM judge correlation with global metrics can be misleading for best-of-n selection tasks, with implications for deployment.
LLM BiasScope is a web platform for real-time bias analysis comparing outputs from multiple LLM providers (Gemini, DeepSeek, Llama, etc).
Method for learning optimal stopping points in chain-of-thought reasoning to reduce compute overhead from overthinking in large reasoning models.
Mixture-of-experts architecture for referring image segmentation using spatio-semantic expert routing to match diverse reasoning requirements.
Framework training distributed RL policies under realistic network conditions including delays and packet loss for edge-cloud deployment.
RL training framework for diffusion language models using entropy-guided step selection to optimize non-autoregressive sequence generation.
Security study revealing unsafe recommendation drift in tool-augmented LLM agents when tools are corrupted, hidden by standard ranking metrics.
Method for personalized RLHF with user-specific preferences using swap-guided learning to address posterior collapse in variational preference learning.
AI agent pipeline for scalable diagram generation using knowledge-infused multi-modal systems to create high-quality visual designs.
Dataset and training method improving vision-language grounding models' ability to interpret negative semantics and complex expressions.
Framework for evaluating AI ethical reasoning using literary narratives and moral dilemmas to test genuine moral understanding beyond surface responses.
Novel approach combining speculative decoding with online learning to improve LLM inference speed by adapting draft models to better approximate target distribution.
Vision-language model framework for multimodal recommendation systems using semantic representation alignment between items and user preferences.
Research on budget-aware value tree search for LLM agents to optimize token and tool usage during execution.
Research on reducing memory overhead of Mixture-of-Experts LLMs through expert replacement technique.
Survey of continual learning methods for LLMs to adapt dynamically while preventing catastrophic forgetting.
Research paper on fusing textual information with time-series forecasting by bridging modality gaps.
RetroReasoner uses reasoning LLMs for strategic retrosynthesis prediction in organic chemistry with step-by-step planning.
MetaKE improves knowledge editing in LLMs via bi-level optimization to align semantic targets with feasible execution regions.
Empirical study of model collapse in ChatGPT through recursive training on synthetic data, demonstrating progressive self-convergence.
Federated hierarchical clustering method automatically determines optimal cluster numbers while preserving privacy across distributed data.
Eye2Eye framework enables human-AI collaboration through shared first-person perspective to reduce communication and understanding gaps.
Optimizes multimodal LLM inference cost by partitioning vision and language components across heterogeneous GPU tiers.
CognitionCapturerPro reconstructs visual stimuli from EEG using multi-modal information and asymmetric alignment.
Graph In-Context Operator Networks enable neural networks to infer solution operators from examples for spatiotemporal prediction.
MoKus performs knowledge-aware concept customization by transferring cross-modal knowledge to bind textual knowledge to visual concepts.
TaoBench evaluates generalization of automated theorem prover LLMs beyond MathLib using novel definitional frameworks.
Vision-Language Models enhance underwater image enhancement by improving semantic cue extraction for downstream vision tasks.
RIGID Framework integrates research-based knowledge into AI-mediated instructional design workflows for education.
arXiv paper presenting Cheers, unified multimodal model decoupling patch details from semantics for image comprehension and generation.
arXiv paper on Residual SODAP for continual learning in domain-incremental settings using prompt-based adaptation.
arXiv paper proposing UAV Scene Change Captioning task using multimodal learning to describe semantic changes in aerial imagery.