Task-agnostic Low-rank Residual Adaptation for Efficient Federated Continual Fine-Tuning
Federated continual fine-tuning with low-rank residual adaptation, enabling efficient parameter-efficient learning across new classes in federated settings.
Federated continual fine-tuning with low-rank residual adaptation, enabling efficient parameter-efficient learning across new classes in federated settings.
Proxy model framework for efficient post-hoc interpretability of LLMs, reducing computational costs of model-agnostic explanations.
Theoretical analysis of OPTQ/GPTQ post-training quantization for LLMs, providing rigorous quantitative guarantees for PTQ algorithms.
Configuration-aware LoRA adaptation for quantized LLMs enabling efficient edge device deployment with heterogeneous capabilities.
RECAP: RL method for safety alignment in large reasoning models, teaching critical evaluation of flawed premises via counter-aligned prefilling.
LLM-based flight delay prediction integrating textual aeronautical data and aircraft trajectories for air traffic management.
Graph neural network architecture using selective state space modeling to address over-smoothing in deep GNNs via node-specific representation evolution.
Optimization of continuous attractor neural networks for brain-inspired path integration, reducing computational redundancy in navigation systems.
Vision-Language-Action model with active visual attention for robotic manipulation, extending from Markov to partially observable decision processes.
Multi-agent RL framework for adaptive traffic signal control, replacing static controllers with learning-based optimization for complex traffic dynamics.
Multi-agent RL for graph-based coordination with bandwidth constraints, addressing what information agents should transmit under communication limits.
Analysis of self-reflection emergence in LLMs through RL post-training, using gradient attribution to explain distinct solution generation and revision capabilities.
Imitation learning framework for combinatorial optimization problems, examining how expert demonstrations affect policy learning in sequential decision problems.
FP8 low-precision quantization for LLM reinforcement learning, addressing memory and compute bottlenecks in rollout generation with engineering and algorithmic solutions.
Demonstrates layer pruning limitations for LLM reasoning tasks, showing pruned models lose algorithmic capabilities despite compression on classification tasks.
dnaHNet foundation model for genomic sequence learning with tokenizer-free design preserving biological motifs while handling long contexts efficiently.
Reinforcement-aware knowledge distillation method for distilling RL-trained reasoning LLMs into smaller models while preserving chain-of-thought capability.
Distributed prompt caching technique for accelerating local LLM inference on resource-constrained edge devices via inter-device state sharing.
Analyzes implicit regularization of Deep LDA objective for scale-invariant discriminative metric learning.
Theoretical analysis explaining Adam's empirical advantage over SGD through second-moment normalization using stopping-time/martingale analysis.
Enables exact gradient computation for spiking neural networks via differentiable ODE solving in JAX, supporting arbitrary neuron models.
Proposes prototypical exemplar condensation for memory-efficient continual learning, reducing stored samples per class from 20+ to single digits.
ALMAB-DC framework combines active learning, multi-armed bandits, and distributed computing for expensive black-box optimization.
Investigates mechanisms of introspective awareness in LLMs, where models detect injected steering vectors with minimal false positives.
Analyzes distributional reinforcement learning with applications to healthcare, moving beyond expectation-based objectives for uncertain domains.
Proposes hierarchical SVG tokenization approach for improved scalable vector graphics modeling with LLMs via geometric-aware token design.
ALTO system for adaptive hyperparameter tuning and orchestration of LoRA fine-tuning jobs across heterogeneous multi-tenant environments.
Proposes CMRM, a framework for improving classification under label noise without privileged knowledge, using quantile-calibrated regularization.
Combines LLMs with Graph Neural Networks to enhance fMRI brain network analysis by leveraging LLM representations.
Method for constraining sequential editing of LLMs to prevent knowledge degradation using editing anchor compression.
Agentic system for generating and validating synthetic image data to address data scarcity and label noise in vision tasks.
Evaluates LLM reasoning capabilities in social deduction game Avalon using Bayesian inference with graph-informed models.
arXiv paper on Bayesian ego-graph inference for decentralized multi-agent reinforcement learning with constrained communication.
arXiv paper on interactive program synthesis for collaborative physical task modeling from narrated demonstrations.
RESample: Data augmentation framework for Vision-Language-Action models in robotic manipulation, addressing limited distribution in demonstration datasets.
Generative View Stitching: Method enabling camera-guided video generation with bidirectional conditioning to prevent collision with generated scenes.
BRIXEL: Approach to reduce computational cost of dense feature maps from vision foundation models like DINOv3 while maintaining performance.
Fed-Sparse-BNSL: Federated method for learning Bayesian network structures with differential privacy, addressing decentralized data challenges.
AV-SpeakerBench: Benchmark evaluating multimodal LLMs on fine-grained audiovisual speech understanding with 3,212 multiple-choice questions.
DRAM: Framework combining mechanism design and online learning for sequential multi-agent settings to ensure truthful reporting with cost-optimality.
Measurement-Consistent Langevin Corrector: Method stabilizing latent diffusion models for inverse problems by reducing discrepancy with learned reverse diffusion.
ConvoLearn: Dataset of 2,134 tutor-student dialogues for fine-tuning LLM-based AI tutors, grounded in dialogic learning theory and Earth Science curriculum.
Tiled Prompts: Method addressing prompt misguidance in text-conditioned diffusion models for image and video super-resolution by handling localized details.
PACED: LLM distillation method that weights training problems by student competence using gradient signal-to-noise ratio to improve distillation efficiency.
Framework addressing causal confusion in end-to-end autonomous driving models through causal intervention during training to improve reliability and safety.
Methodology for detecting prompt injection across multi-agent LLM pipelines. Stage-level kill-chain tracking for attack resilience evaluation.
Detection and mitigation of object hallucinations in vision-language models. Bayesian approach analyzing attention weights and token confounders.
3D Gaussian splatting for weather prediction downscaling. Proposes scale-aware vision transformer for arbitrary-resolution atmospheric forecasting.
Training-free semantic segmentation using vision-language models. Global context-aware framework for dense prediction without additional training.
Experiment using Claude to autonomously build a website designed to generate traffic, exploring AI agent capabilities and decision-making in open-ended tasks.