PreFT: Prefill-only finetuning for efficient inference
Parameter-efficient finetuning method optimizing prefill-only updates to improve inference throughput for personalized LLMs.
Parameter-efficient finetuning method optimizing prefill-only updates to improve inference throughput for personalized LLMs.
Identifies and diagnoses training-inference mismatches in LLM reinforcement learning from implementation differences in token probabilities.
Empirical evaluation of quantum entanglement benefits in decentralized multi-agent reinforcement learning with variational quantum policies.
Active learning approach to improve pairwise ranking prompting reranking from LLMs with noisy and intransitive judgments.
Evaluation of AI-generated text detection methods' resilience to paraphrasing attacks across fine-tuned and classifier approaches.
Router for LLM agents selecting among functionally equivalent tool providers based on latency and quality tradeoffs.
Framework for predicting and optimizing energy consumption during multi-GPU LLM inference without expensive profiling.
Auditing protocol to verify that deep neural network heatmap explanations faithfully reflect actual model decision drivers in industrial inspection.
Analysis of how computation propagates through transformer layers by studying residual stream dynamics and spectral geometry in LLMs.
MetaMoE unifies independently trained domain-specialized experts into privacy-preserving MoE using public proxy data for federated settings.
KV-cache compression study with diversity-penalty survivor method for efficient LLM inference on long-form reasoning tasks.
Mixed gradient policy optimization for hybrid discrete-continuous action spaces in robotics and control problems.
Domain adaptation framework leveraging expert textual descriptions as language-induced priors to prevent negative transfer in cold-start scenarios.
Matrix-Space RL reuses local transition geometry through positive semidefinite matrix descriptors for compositional generalization.
Continuous semantic alignment framework for GUI critic models in test-time scaling for generalist AI agents beyond binary classification.
Dynamic Latent Routing composes optimal sub-policies temporally for MDPs and proposes LLM post-training method based on General Dijkstra Search.
AIM-DDI predicts drug-drug interactions using model-agnostic multimodal integration with focus on unseen-drug generalization.
Exemplar Partitioning constructs interpretable feature dictionaries from LLM activations using Voronoi partitioning with 1000x fewer tokens than sparse autoencoders.
Distributionally robust multi-task reinforcement learning addresses imbalanced learning across tasks via adaptive task sampling.
RQ-MoE combines residual quantization with mixture-of-experts for efficient input-dependent vector compression of embeddings.
MoRe proposes modular representations for continual learning on sequential data with minimal catastrophic forgetting.
LoMETab extends rank-1 ensembles to rank-r multiplicative factorizations for improved tabular deep learning performance.
Coherent Coordinate Descent optimizes zeroth-order optimization for memory-constrained on-device learning without backpropagation.
Optimal Pattern Detection Tree introduces interpretable rule-based classification model for pattern discovery in healthcare and maintenance domains.
Multi-agent exploration strategy for imperfect-information games like StarCraft using data-augmented game starts to accelerate policy gradient learning.
NodeSynth generates socially aligned synthetic data for AI model evaluation using a fine-tuned taxonomy generator anchored in real-world evidence.
Novel OOD detection method using class-wise Mahalanobis distance variance for neural networks in safety-critical applications.
Counterfactual time series forecasting method incorporating textual conditions for future events influence.
Federated actor-critic framework for collaborative policy training with shared representations and personalized local policies.
FrontierSmith: Method for synthesizing open-ended coding problems at scale to train stronger LLM coders.
QAOD: Single-pass white-box framework for hallucination detection in LLMs using question-answer orthogonal decomposition.
LiSA: lifelong safety adaptation framework for AI agents executing workflows, maintaining guardrails against contextual safety failures in tool use and data access.
Method for positive-unlabeled learning from highly imbalanced datasets, applicable to disease identification, fraud detection, and recommender systems.
EvoLib: test-time learning framework enabling LLMs to accumulate and evolve knowledge across problem instances via extracted modular skills without parameter updates.
Offline-to-online reinforcement learning method using bi-level optimization for adaptive data mixing between datasets.
Autonomous agentic workflow for end-to-end machine learning interatomic potential development with adaptive active learning.
Mechanistic interpretability study using activation patching to examine how LLMs process relative geographic space.
Benchmark for evaluating LLM medication recommendations at prescription level with per-timepoint predictions and information-rich inputs.
Weight space analysis of neural PDE operators identifying reusable physical structure across multiple regimes in pretrained models.
Multi-objective prompt optimization technique using pure-exploration bandits for efficient LLM prompt selection.
Token-level credit assignment method for agentic RL using energy-based perspectives to improve PPO and GRPO training signals.
Studies impact of plasticity interventions on backdoor attack vulnerabilities in deep reinforcement learning agents.
Identifies silent collapse phenomenon in recursive learning systems where models trained on self-generated data degrade undetected by standard metrics.
Statistical results for entropy-regularized inverse reinforcement learning with linear reward classes in finite-horizon MDPs.
Analysis of energy efficiency comparing neural combinatorial optimization solvers to CPU metaheuristics accounting for training costs.
Deep reinforcement learning framework combining HMMs with neural networks for forecasting multivariate hidden Markov processes.
Study of grokking phenomenon showing Transformers achieve fastest validation accuracy at intermediate dataset sizes, not largest ones.
Studies Goldstone-like modes in equivariant deep neural networks, demonstrating how symmetry breaking enables coherent information propagation across layers.
Proposes ReMIA, efficient membership inference attack against synthetic data generators that avoids shadow model overhead for privacy evaluation.
Characterizes tradeoff between reconstruction accuracy, feature efficiency, and interpretability in sparse autoencoders for mechanistic interpretability.