CARE: Multi-modal medical reasoning framework using visual grounding models to provide evidence-grounded agentic workflows improving clinical accountability and interpretability.
ToolRLA: Three-stage post-training pipeline (SFT→GRPO→DPO) with multiplicative reward decomposition for aligning tool-integrated agents in domain-specific deployment.
SEED-SET: Framework for scalable ethical testing of autonomous systems using automated benchmarking to evaluate alignment in high-stakes domains.
CDD: Method for detecting data contamination in small language models by measuring output distribution peakedness, with experiments on GSM8K, HumanEval, and MATH benchmarks.
UIS-Digger: Research agent system for unindexed information seeking, addressing limitations of LLM agents relying on search-engine-indexed knowledge for real-world information retrieval tasks.
RetroAgent: LLM-based agent trained with RL that uses retrospective dual intrinsic feedback to improve continuous adaptation and explicit knowledge retrieval beyond standard RL paradigms.
Analysis showing LLM activation steering requires non-linear interventions, challenging the linear representation hypothesis for controlling model behavior.
Research on continual learning using pre-trained models to mitigate catastrophic forgetting in systems learning under changing environments.
Survey of explainability and interpretability methods for deep learning models in NLP and information retrieval, covering transparency techniques for neural networks.
Application of LLMs to travel behavior prediction through natural language reasoning, comparing frameworks for transportation demand modeling.
Optimal transport method for aggregating distributed mixture-of-experts models trained independently on decentralized datasets.
Philosophical argument for large language models as scientific models of language, examining LLMs' role in understanding language as social entity.
Fine-tuning-free compression compensation technique for LLMs using eigenspace low-rank approximation to recover accuracy post-compression.
Fine-grained token-level data selection method for LLM supervised fine-tuning, removing redundant/uninformative tokens within samples.
Diffusion-based neural combinatorial optimization solver with inference-time adaptation for improved cross-problem generalization on NP-complete problems.
Multi-task learning framework for resource-constrained autonomous agents using spiking neural networks with adaptive task-switching policy.
Systematic review of psychometric evaluation methods for LLMs, measuring psychological constructs and establishing human-centered evaluation approaches.
Benchmark REI-Bench evaluates LLM-based robot task planners on handling vague human instructions, addressing real-world instruction ambiguity.
Improves instruction following in LLMs by training with pseudo-code instead of natural language, particularly for compositional tasks.
Data-driven survey of 14,648 papers on LLM limitations from 2022-2025, using keyword filtering and LLM-based classification to systematize research.
Benchmark comparison of ML models for retail sales forecasting, evaluating statistical baselines, tree-based ensembles, and deep learning architectures.
Self-improving visual robotic planning agents using video generative models, addressing generalization to unseen tasks through self-supervised learning loops.
Comprehensive survey of differential privacy techniques in machine learning, from symbolic AI to modern LLMs, covering foundational definitions and applications.
Study of psychological risks and feedback loops between AI chatbots and users experiencing mental illness, examining concerning edge cases.
Survey of ethical issues in code generation models, covering licensing, privacy, fairness, and environmental impact of AI-assisted software development.
Privacy analysis reveals KV-cache vulnerabilities in LLM inference enable reconstruction of sensitive training data and demonstrates mitigation strategies.
Theoretical analysis of sigmoid contrastive loss in SigLIP models explains temperature and bias synchronization for representation alignment in pretraining.
MonitorVLM vision-language framework detects unsafe worker behaviors in mining operations, automating safety monitoring in high-risk industrial environments.
Theoretical framework predicts kernel regression learning curves from empirical data covariance and polynomial decomposition without full training.
KVTC lightweight transform coder compresses KV caches for efficient storage and reuse across LLM inference turns, reducing GPU memory overhead.
DeepEyesV2 agentic multimodal model integrates text, images, code execution, and web search through reinforcement learning for structured reasoning and tool invocation.
Dataset-agnostic and gradient-guided augmentation approach improves out-of-domain robustness in computer vision across background, style, and instrument shifts.
REMSA agent automates foundation model selection for remote sensing tasks using constraint-aware reasoning across diverse unimodal and multimodal architectures.
Hierarchical dual-strategy unlearning framework for LLMs removes specialized knowledge while preserving medical competencies and protecting privacy-sensitive patient data.
Calibrated uncertainty estimation in controllable video generation models reduces hallucinations in instruction-guided editing and robotic world modeling.
GTR-Turbo enables efficient multi-turn RL training for agentic vision-language models by leveraging merged checkpoints as free teachers without costly privileged models.
Pretrained battery transformer foundation model for predicting battery cycle life using transfer learning to address data scarcity across battery chemistries.
First-order analysis of how cross-entropy training shapes attention mechanisms in transformers to enable Bayesian probabilistic reasoning.
Investigation of geometric scaling properties in Bayesian inference across production-grade language models including Llama, Mistral, and Pythia families.
Analysis of over-searching problem in retrieval-augmented LLMs, where models invoke search unnecessarily, causing inefficiency and hallucinations from irrelevant context.
Secure multi-tenant architecture with burn-after-use mechanism for preventing data leakage in enterprise LLM deployments across departments.
Multi-turn DoS attack exploiting tool-calling chains in LLM agents. Demonstrates stealthy resource amplification and cost attacks in agent-tool interaction loops.
Theoretical analysis of LLM hallucinations using rate-distortion theory. Connects memorization to sparse facts and establishes optimal memory efficiency tradeoffs.
Benchmark for evaluating long-term conversational memory in multi-party LLM applications with realistic interaction patterns across groups and channels.
Vision-language model that automatically edits HTML to fix WCAG2 accessibility violations while preserving original design.
RL algorithm for compressing verbose chain-of-thought reasoning in LLMs using fine-grained group policy optimization.
Unified discrete tokenizer with 2^128 codebook size for multimodal LLMs supporting high-fidelity reconstruction and semantic extraction.
Aperture-guided agent using RL to stabilize fine-grained visual reasoning in MLLMs by sequentially acquiring evidence from regions of interest.
Analysis of conformal prediction tradeoffs beyond coverage, examining deployment metrics like commitment frequency and deferral rates.
Zero-shot video anomaly detection method using MLLMs to address dataset scarcity and context-dependent anomaly semantics.