FedRot-LoRA: Mitigating Rotational Misalignment in Federated LoRA
FedRot-LoRA addresses rotational misalignment in federated fine-tuning of LLMs on decentralized data to reduce aggregation error.
FedRot-LoRA addresses rotational misalignment in federated fine-tuning of LLMs on decentralized data to reduce aggregation error.
AudioCapBench: benchmark for evaluating audio captioning of multimodal LLMs across sound, music, speech with 1,000 samples and LLM-as-Judge evaluation.
MedMAP pre-training framework for vision-language models on 3D MRI data with modality-specific alignment for multi-organ abnormality detection.
ProtoDCS framework for test-time adaptation of vision-language models under distribution shift in open-set scenarios.
TRIZ-RAGNER applies retrieval-augmented LLMs for named entity recognition in patent contradiction mining for systematic innovation.
Analysis of transformer training geometry showing parameter updates organize into dominant drift direction with oscillatory transverse dynamics.
SAGE-LLM architecture combining LLMs with formal safety verification (Fuzzy-CBF) and graph-structured knowledge for safe UAV autonomous decision-making.
System design for distributed LLM inference across device, RAN-edge, and cloud tiers with latency constraints for 5G embodied AI applications.
Agent-centric benchmarking paradigm where autonomous agents dynamically generate, validate, and solve problems to evaluate LLM reasoning capabilities.
BDGxRL uses Diffusion Schrödinger Bridge to enable cross-domain reinforcement learning policy transfer when target domain interaction is unavailable.
UPath proposes learning-based heuristics for A* pathfinding on grid maps using deep neural networks across heterogeneous topologies.
MPU framework for privacy-preserving machine unlearning in LLMs without sharing server parameters or client forget sets.
Research on applying causal discovery algorithms to real-world longitudinal data with institutional workflow constraints.
Sea² framework using vision-language models and autonomous agents for unsupervised cross-domain visual adaptation without fine-tuning on target data.
Theoretical research on offline reinforcement learning with general function approximation and parametric policies, extending beyond finite action spaces.
Q-learning approach for learning safe policies from expert demonstrations with unknown constraints and non-observable costs.
Federated learning optimization achieving consistency between local and global model flatness under data heterogeneity.
Continual fine-tuning method for LLM-based vulnerability detection addressing catastrophic forgetting under temporal distribution shift.
Multi-layer intrusion detection framework with incremental learning for Industrial IoT networks handling novel cyber threats.
Adversarial benchmark testing MLLM visual reasoning and grounding capabilities in referring expression comprehension tasks.
Multi-agent cascaded framework for breast ultrasound screening reducing unnecessary biopsy referrals through selective decision-making.
Transferability estimation method for selecting optimal medical foundation models for segmentation without retraining.
Multimodal benchmark evaluating MLLMs on explicit 3D geometric reasoning with point clouds, exposing geometric hallucinations.
Hierarchical concept embedding models improving neural network interpretability through human-readable concept representations.
Benchmark for agentic search systems balancing quality and efficiency, addressing underspecified user preferences in LLM-powered retrieval.
Controlled experimental studies identifying what triggers and prevents sycophancy in LLMs, an alignment failure in advisory contexts.
Vision for foundation world models enabling autonomous agents to learn, verify, and adapt reliably in open, non-static environments.
Multi-agent workflow system that translates jailbreak papers into executable modules for unified benchmarking of LLM robustness techniques.
Model-agnostic interpretable debiasing method for vision-language models to mitigate unintended social bias in black-box reasoning processes.
RewardUQ: Framework for uncertainty quantification in reward models used to align LLMs with human preferences, reducing annotation costs.
Data-driven optimization pipeline for scheduling and caching hundreds of LLM adapters in distributed serving to maximize GPU throughput.
Quant Experts: Post-training quantization method for vision-language models using mixture of experts for token-aware adaptive error reconstruction.
Empirical study evaluating whether reasoning capabilities universally improve LLM performance across sentiment analysis tasks of varying complexity.
ACWI framework for dynamically balancing intrinsic and extrinsic rewards in sparse reward reinforcement learning through adaptive scaling.
Preference packing technique for efficient batch training of LLMs during preference optimization (RLHF), improving resource utilization.
ARGUS framework studying how narrative features in argumentative texts influence persuasion using corpus analysis and modeling.
Formal verification approach for vision-language models drafting radiology reports to ensure logical consistency in clinical reasoning.
Systematic evaluation of LLM machine translation for Ancient Greek technical prose, showing terminology rarity predicts translation failures.
CoME: Mobile agent architecture with four expert modules for hybrid reasoning including screen understanding, planning, and action execution.
ArgLLM-App: Interactive web system implementing argumentative reasoning agents with LLMs for explainable binary decision-making tasks.
TASC framework for accelerating small language models through task-adaptive sequence compression and vocabulary enrichment during fine-tuning.
Method for training reasoning models to follow instructions in reasoning traces to prevent unintended leakage of private information in AI agents processing sensitive user data.
SafeGen-LLM framework enhances safety in robotic task planning by combining LLMs with safety constraints, addressing generalization challenges in classical and RL-based planning methods.
TREC 2025 DRAGUN track resources for evaluating RAG systems that help readers assess news trustworthiness with attributed reports.
Exploration of recurrent architectures with growing memory as subquadratic alternatives to Transformers for sequence modeling.
LoRA-Pre method reducing memory overhead in optimizers like Adam via low-rank approximation of momentum states.
CUDA Agent system using large-scale agentic RL to generate optimized GPU kernels, bridging gap between LLMs and compiler-based systems.
Study comparing standard multi-turn prompting with user-turn-only prompting to determine if LLMs benefit from their own prior responses.
Research on offline-to-online multi-agent reinforcement learning with offline value function memory and sequential exploration strategies.
CowPilot framework enabling autonomous and human-agent collaborative web navigation with preference modeling and human oversight.