arXiv paper benchmarking deep learning models for malaria diagnosis, evaluating efficiency, robustness, and explainability alongside accuracy.
OrcaRouter is production-oriented LLM router using contextual bandits with hybrid offline-online learning to route requests to optimal models based on capabilities and cost.
arXiv paper presenting FLAG, a maximum entropy RL approach using latent augmented guidance with flow policies for high-dimensional action spaces.
arXiv paper proposing DUAL framework using diffusion models with uncertainty awareness for offline-to-online reinforcement learning.
arXiv paper on AbstainGNN enabling graph neural networks to abstain from predictions under uncertainty for safer graph classification.
arXiv paper showing smaller LLM models exhibit higher policy-level diversity in GRPO, improving rollout diversity without token-level noise.
arXiv paper proposing reliability score metric for evaluating conditional generative models that accounts for inherent uncertainty in outputs.
Learns permutation-invariant macroscopic dynamics of high-dimensional particle systems without assuming fixed ordering of microscopic degrees of freedom.
Proposes constrained optimization framework for unlearning in diffusion models using KL divergence and likelihood constraints to remove concepts while preserving utility.
Unifies SVD-based LLM compression methods but reveals weight reconstruction improvements fail to translate to downstream task performance gains.
Introduces CoMem framework decoupling memory management from agent workflows to reduce latency overhead in long-context agentic models.
Lecture notes covering inverse reinforcement learning foundations and connections to dynamic discrete choice models from econometrics.
Develops ForecastCompass for agentic forecasting with adaptive factor memory mechanism to transfer knowledge from resolved forecasts to future predictions.
Proposes DARTS to accelerate LLM reinforcement learning by identifying and addressing intra-prompt distribution inefficiencies in response length rollouts.
Adapts Variational Preference Learning to federated settings for personalizing LLM alignment to individual user preferences while preserving privacy.
Optimizes bandwidth allocation and device partitioning for federated learning over Industrial IoT wireless networks to improve convergence efficiency.
Proposes DensityFlow for generating robust counterfactual explanations on tabular data by adhering to high-confidence data manifolds.
Extends inverse reinforcement learning to handle multiple imperfect demonstrators with varying suboptimality levels using feasible reward set framework.
Studies reinforcement learning from verifiable rewards to improve LLM generation of formally verified programs and proofs in verification-aware languages.
Models AI benchmark aggregation as principal-agent game, showing uniform item averaging is suboptimal and proposes welfare-aware aggregation methods.
Proposes de-attribution method for LLM unlearning to address over-forgetting and model utility degradation when removing inappropriate training data.
Extends adjoint-based trajectory optimization to discrete domains for unsupervised diffusion-based combinatorial optimization without requiring near-optimal solution datasets.
Mathematical study of gradient-based learning for overparameterized Gaussian mixture models, analyzing convergence rates under overparameterization.
Unified framework reinterpreting zeroth-order Hessian approximation through policy optimization for derivative-free bilevel optimization.
Parallel tempering initialization improves sequential Monte Carlo sampling for inference-time reward alignment in generative models.
Training-free expert router for sparse mixture-of-experts LLMs using eigenvector decomposition to prevent expert collapse.
Differentiable multi-agent coordination framework using sheaf-ADMM where agents solve convex subproblems and coordinate via neural encoders.
Deep Q-learning framework for cost-aware multi-omics classification that selectively acquires expensive data modalities in clinical settings.
Efficient graph condensation method compressing large graphs into compact synthetic versions for resource-constrained GNN deployment.
Local learning algorithm for training deep networks via energy minimization with layer-local constraint tracking instead of backpropagation.
Analysis of why group-based policy optimization methods like GRPO perform well without explicit uncertainty tracking in bandit settings.
Unified representation learning approach combining RTL code and control data flow graphs for hardware design acceleration.
Analysis of challenges deploying reinforcement learning for controlling real-world industrial thermal energy systems beyond simulation.
CHECKMATE tool generates optimization algorithms via code evolution to solve combinatorial problems without expert-designed heuristics.
Bayesian optimization framework using best-arm identification for trust region selection on multimodal and high-dimensional functions.
Self-supervised contrastive method for learning interpretable embeddings of progressive time series data on manifolds.
JEPA architecture that decomposes latent world models into progression and content subspaces for improved representation learning.
SiGMA architecture unifying multi-scale time-series forecasting methods with continuous scaling operator to improve temporal dynamics modeling.
Time-series out-of-distribution detection using hyperspherical embedding representations to handle distributional shifts in temporal data.
Causal discovery foundation models using pretraining across causal environments for tabular data to recover directed causal relations from observational/interventional data.
Convergence analysis of two-timescale stochastic approximations with applications to temporal difference learning and actor-critic RL methods.
Method for selecting diverse retriever portfolios from large pools to adaptively handle heterogeneous RAG queries with principled query distribution coverage.
Overview of concept drift detection evaluation metrics for data streams; identifies limitations of classification accuracy as quality measure.
Rule-based generalized additive modeling framework with human-readable sparse bases for transparent tabular prediction in high-stakes domains.
Empirical study on ResNet teacher-student pairs showing student capacity moderates knowledge distillation effectiveness on CIFAR-10.
Geometry-based multimodal fusion robust to noisy/incomplete data and conflicting inputs using intrinsic data quality assessment.
Tensor Separation Learning for interpretable regression modeling feature interactions beyond additive decompositions used in GAMs/SHAP.
Fixed-point masked generative models enable efficient parallel decoding with dynamic denoiser computation allocation per refinement step.
Multivariate distributional RL using sliced divergences to model full return distributions beyond single-dimensional settings.
Survey of 70 on-device learning works characterizing post-deployment distribution shift and adapting ML models on microcontroller devices.