WriteSAE: Sparse Autoencoders for Recurrent State
WriteSAE: First sparse autoencoder decomposing and editing state-space and hybrid recurrent LLM cache writes for interpretability.
WriteSAE: First sparse autoencoder decomposing and editing state-space and hybrid recurrent LLM cache writes for interpretability.
Graph neural network approach for financial fraud detection with calibrated risk scoring and structural regularization.
ToolMol: LLM-based agentic framework using evolutionary algorithms for de novo drug discovery with improved molecular generation validity.
Port-Hamiltonian neural networks for identifying nonlinear string dynamics, combining physical knowledge with neural networks.
Systematic review of neural tangent generalization attacks, a clean-label data poisoning technique for protecting unauthorized training data.
Study of emergent misalignment in fine-tuned LLMs, analyzing how harmful datasets cause behavioral spillover through data-mediated transfer.
Benchmark for evaluating agentic AI systems on data reuse and integration tasks across fragmented neuroscience datasets with diverse formats.
Method for token-level attribution of LLM outputs to training data using influence functions in orthogonal latent spaces for healthcare applications.
AGOP-Weighted attribution method explaining image classifier predictions using Average Gradient Outer Product from feature learning theory.
Extension of Foresight Learning to clinical prediction using MIMIC-III longitudinal notes with natural language questions about future events.
Machine learning framework for coarse-grained molecular dynamics augmenting force matching with Hessian matching for improved potential accuracy.
Orthrus framework combining autoregressive LLM fidelity with parallel diffusion model generation for efficient token generation.
Method for quantifying missing observations in inverse reinforcement learning datasets, addressing data quality in behavioral analysis.
Discrete Stochastic Localization framework for non-autoregressive sequence generation using continuous diffusion with unit-sphere embeddings.
Bayesian approach to merging multiple task-specific expert models without retraining, using strong anchor models as inductive bias.
SMA framework using submodular optimization for efficient multimodal learning with limited paired data across modalities.
Analysis of descriptive collision problem in sparse autoencoders where single explanations describe multiple features in language model interpretability.
Randomized smoothing framework providing robustness certificates for multimodal models against heterogeneous perturbations across modalities.
ASAP method for efficient doubly-stochastic attention in Transformers via sliced dual projection, reducing computational cost in inference.
Pre-deployment evaluation framework (RISED) for clinical AI systems covering reliability, equity, sensitivity, and deployability dimensions.
VIP-COP framework for optimizing context in tabular foundation models, enabling in-context learning on structured data with long sequences.
Systematic empirical and theoretical analysis of how data difficulty affects generalization and extrapolation in LLM fine-tuning.
Study of DAgger algorithm for training long-horizon LLM agents, addressing covariate shift in multi-turn interactions with teacher supervision.
Study of efficiency gap in byte-level language modeling comparing byte-level and masked diffusion approaches to traditional subword tokenization.
Comparison of probabilistic circuits and transformers in autoregressive language modeling reveals expressivity gaps and bottlenecks.
MANGO framework for multi-agent LLM systems reducing error propagation through reinforced collaboration in workflow networks.
Fixed-pool data recipe search for SFT formulates supervised fine-tuning as optimization of operator combinations rather than instance ranking.
Minimal model study separating shortcut feature reliance from OOD failure in binary classification.
Algebraic Ontology Projection projects LLM hidden states into Galois Field F2 to verify and control logical relations.
Contrastive perspective on RLVR for LLM reasoning, reformulating GRPO as weighted positive-negative score difference.
Study showing multi-agent LLM sycophancy under disagreement stems from base model behavior, not RLHF, with activation patching analysis.
DP-Muon formulates differentially private training using matrix-valued optimizer with momentum and orthogonalization.
F-GRPO algorithm optimizing LLMs for unified candidate generation and ranking in retrieval pipelines using group-relative policy optimization.
DRIFT benchmark for task-free continual graph learning addressing catastrophic forgetting with continuous distribution shifts.
Joint embedding diffusion world model for online model-based reinforcement learning balancing computational efficiency and performance.
Theoretical framework for learning Nash equilibria in offline two-player zero-sum Markov games using KL regularization instead of explicit pessimism.
Study analyzing why masked diffusion models train slower than autoregressive models and proposing methods to accelerate MDM training while maintaining performance.
FeatCal addresses performance gaps in merged models by analyzing feature drift between merged and expert models, proposing methods to calibrate features during model merging.
Tide method for graph OOD detection via tri-component information decomposition to identify spurious signals in features and structure.
Evaluation showing LLMs lack temporal awareness in medical knowledge because benchmarks are atemporal while medical knowledge continuously evolves.
Target-aligned Coverage Expansion framework for cross-domain offline reinforcement learning that leverages source data while reducing distribution mismatch.
Amortized learned surrogate (Deceptron) for solving nonlinear inverse problems by encoding curvature information into reusable reverse operator.
Analysis of Muon optimizer showing spectral flattening mechanism enables larger learning rates and faster convergence through momentum orthogonalization.
Graph Information Bottleneck approach for multi-label graphs reducing over-squashing in deep message passing of graph neural networks.
Entropy regularization extension to multi-agent PPO addressing non-stationary observations in multi-dimensional cooperative reinforcement learning environments.
Game-theoretic analysis of collaborative multi-agent bandit learning where strategic agents balance information sharing against free-riding incentives.
Decision Pattern Shift framework analyzing how deep neural network internal decision mechanisms evolve from training to test data for understanding generalization failures.
Parameter-efficient fine-tuning method for LLMs using program memory to balance rapid adaptation and knowledge retention in continual learning settings.
Adversarial attack methods against multi-agent reinforcement learning systems using Jacobian-based gradient information to identify vulnerable communications.
arXiv paper: Hybrid Tucker-LSTM tensor network for battery state-of-charge prediction in electric vehicles.