Adaptive Computation Depth via Learned Token Routing in Transformers
Token-Selective Attention mechanism enabling adaptive computation depth via learned per-token routing in transformers.
Token-Selective Attention mechanism enabling adaptive computation depth via learned per-token routing in transformers.
Theoretical analysis of compositional steering in Sparse Autoencoders, examining non-linear interference in feature activation.
Unlearnable examples via semantic perturbations for privacy protection across training paradigms including pretraining-finetuning.
Load balancing technique for multimodal MoE LLMs addressing information heterogeneity and stragglers in expert parallelism inference.
Framework converting outcome-level supervision into process-level signals for reasoning tasks via reinforcement learning credit assignment.
Online data reweighting during LLM training outperforms offline curation methods, improving generalization without preprocessing overhead.
Evolutionary algorithms for fine-tuning quantized convolutional models for IoT and edge device deployment.
Attribution-based method to mitigate catastrophic forgetting in LLMs during continual learning by selectively updating parameters.
Graph Normalization: differentiable dynamical system for approximating NP-hard Maximum Weight Independent Set with convergence guarantees.
Analysis of feature starvation in sparse autoencoders for LLM interpretability, proposing geometric solutions to dead neuron problems.
Method addressing interference between sequentially trained early-exit classifiers in neural networks via stability-plasticity tradeoffs.
Technique for verifying GNN model ownership and detecting unauthorized mimicry across different architectures and embeddings.
ML method for drug discovery that learns from sparse ligand-protein binding data to accelerate candidate selection.
Research on interpretable hidden state structures in recurrent RL policies for partial observability using Pontryagin methods.
Zero-shot conditional sampling with pretrained diffusion models for linear inverse problems, providing Langevin mixing guarantees with measurement consistency.
Neural routing method for vehicle routing problems on multigraphs using Node-Edge Policy Factorization to handle parallel edges with varying trade-offs at scale.
Gradient-based parameter learning method for differential-algebraic equations with state-dependent events, handling implicit algebraic variables and discontinuous resets.
Information-theoretic adversarial training methods for improving LLM robustness against adversarial prompting with improved computational efficiency over existing approaches.
Active learning framework for conditional generative compressed sensing combining prompt-conditioned generative models with adaptive sampling strategy selection.
Semantic loss fine-tuning approach preventing catastrophic model collapse when training Transformers on causal reasoning tasks like transitivity and d-separation.
Study of Graph Self-Supervised Learning robustness to real-world noise from automatic text-extracted biomedical knowledge graphs, identifying failure modes.
Benchmark framework evaluating knowledge graph construction methods and GNN robustness to noise/fragmentation/semantic inconsistencies in automatically extracted graphs.
GRALIS unified mathematical framework via Riesz representation theory establishing formal connections between major XAI attribution methods (GradCAM, SHAP, LIME, Integrated Gradients).
Approximate Next Policy Sampling method solving conservative policy update problem in deep RL by estimating state-visitation distribution of improved policy.
Flux Neural Operator augmented with ViT-based context injection via hypernetworks for robust inference of conservation law dynamics across initial conditions.
MEMOA framework for scaling federated learning of massive agent populations using mean-field decentralized Nash equilibria with minimal communication overhead.
Study of how Transformers learn shortcut solutions that impair compositional reasoning in continual learning, examining failure modes in systematic generalization.
Online Localized Conformal Prediction framework providing valid uncertainty quantification for time-series via localized calibration under covariate heterogeneity.
Non-Myopic Pathwise Policy Gradients (NM-PPG) method for active feature acquisition formulated as POMDP with sequential decision-making perspective.
MOSAIC method for discovering interpretable causal modules in scientific time series via sparse additive identifiable causal learning with post-hoc semantics.
Benchmark framework for evaluating adversarial attacks and defenses on Graph Neural Networks, addressing inconsistent experimental settings in GNN robustness research.
One-step generative modeling for efficient autoregressive dynamical system forecasting with long-trajectory structure preservation.
Adaptive Q-chunking approach for offline-to-online reinforcement learning with state-dependent action chunk sizes.
Federated knowledge distillation with energy-based gating for robustness in heterogeneous environments.
Optimization method accelerating linear minimization oracle-based optimizers via implicit gradient transport.
Joint-Embedding Predictive Architecture for scalable 3D aerodynamic surrogate modeling with semantic latent representations.
Carbon footprint modeling framework for LLM inference on solar-powered LEO satellites.
Out-of-distribution detection using frozen pretrained model representations without fine-tuning.
Theoretical analysis of expressive power in diagonal plus low-rank neural networks including LoRA structures.
Distributionally robust multi-objective optimization formulation accounting for distributional shifts in data.
Temporal Functional Circuits framework for interpreting Kolmogorov-Arnold Networks in time-series forecasting with mechanistic explanations.
Proposes Budgeted Attention Allocation with head-gating mechanism to create multiple cost-quality operating points in transformer inference.
Theoretical analysis proving pre-training is essential for weak-to-strong generalization in high-dimensional models using spiked Gaussian framework.
Framework for federated cooperative inference across models using unsupervised consensus embeddings while preserving data and parameter privacy.
Introduces CRAFT framework for continual LLM adaptation using low-rank interventions on hidden representations to avoid catastrophic forgetting.
Proposes CoMemNet combining non-topological space modeling with temporal learning for continual traffic prediction on dynamic streaming networks.
Behavioral evaluation framework for agentic LLM stock prediction systems using LLM judges and RL feedback to assess interdependent decision quality.
Proposes RVPO for risk-sensitive multi-objective alignment in LLMs by penalizing reward variance to prevent constraint neglect in critic-less RLHF.
Proposes adaptive LoRA component selection in federated learning under differential privacy to reduce aggregation error and improve fine-tuning stability.
Analyzes gradient noise imbalance across LLM modules and proposes SNR-based Adam calibration to improve training convergence and stability.