Multi-Agent Reasoning with Consistency Verification Improves Uncertainty Calibration in Medical MCQA
Multi-agent framework with verification for improving calibration and accuracy in medical multiple-choice question answering.
Multi-agent framework with verification for improving calibration and accuracy in medical multiple-choice question answering.
Study evaluating RAG systems on AI policy analysis showing retrieval improvements don't guarantee better answers on complex regulatory documents.
Inverse-forward differentiation method to reduce memory requirements for backpropagation by avoiding activation storage.
Learning-theoretic framework for coded computing in distributed systems to handle slow, faulty, or compromised servers.
Visualization technique for understanding RNN internal dynamics during training using multislice PHATE algorithm.
Physics-informed neural networks using wavelet decomposition to improve training on differential equations with rapid oscillations and steep gradients.
arXiv paper on Symmetry-Guided Memory Augmentation (SGMA) improving efficiency of RL-based legged locomotion training.
arXiv paper on machine learning techniques to detect and localize power/radiation leakage of cryptographic keys from hardware implementations.
arXiv paper on multi-agent reinforcement learning for adaptive traffic signal control in heterogeneous urban networks.
arXiv paper: GraphOmni benchmark framework evaluating LLM reasoning on graph-theoretic tasks with diverse formats and serializations.
arXiv paper introducing Distance Explainer method for post-hoc interpretability of embedded vector spaces in ML models.
arXiv paper on Bottlenecked Transformers: KV cache consolidation technique for scaling inference-time reasoning in LLMs.
arXiv paper interpreting neural networks as dynamical systems on latent manifolds, analyzing autoencoder vector fields.
arXiv paper on scalable longitudinal patient pathway modeling from multimodal EHR data using neural networks for condition forecasting.
Research paper demonstrating LLMs perform in-context reinforcement learning during inference. ICRL prompting framework enables inference-time self-improvement.
TimeRecipe benchmarks module-level effectiveness of components in time-series forecasting architectures.
DART adds server-side robustness to federated learning for edge devices without expensive client-side computation.
Theoretical analysis of federated distillation with weighted aggregation of client predictions under class mismatch.
PromptLoop refines prompts for diffusion models using sequential reinforcement learning feedback during sampling.
Generative method for synthetic financial time series data to address data shortage in ML models for trading and investment.
Proposes future summary pretraining for LLMs as alternative to next-token prediction, addressing limitations in long-horizon reasoning and planning tasks.
Addresses distribution shift in time-series forecasting by identifying concept drift and temporal shift, proposing mitigation strategies for generalization.
OffSim proposes model-based offline inverse RL framework to learn environmental dynamics and reward functions from offline data without manual definition.
Applies deep RL to dynamic origin-destination matrix estimation in traffic simulations, addressing credit assignment across temporal vehicle dynamics.
Proposes curiosity-driven quantized Mixture-of-Experts framework using Bayesian uncertainty for deploying neural networks on resource-constrained devices.
ContagionRL is a Gymnasium-compatible RL platform for reward engineering in spatial epidemic simulations, enabling systematic study of learned behavioral strategies.
Develops Hessian-free actor-critic algorithm for bi-level RL optimization with applications to LLM fine-tuning, addressing second-order information requirements in policy optimization.
Introduces continual learning task for GUI agents that must adapt to shifting domains and resolutions over time, identifying failure modes in existing agent methods.
Study of variance in agentic system evaluations using 60,000 trajectories on SWE-Bench-Verified, showing pass@1 estimates vary significantly across runs, questioning single-run reliability assumptions.
AceGRPO proposes adaptive curriculum learning with group relative policy optimization for autonomous ML engineering agents, addressing behavioral stagnation in LLM-based agents through RL with efficient data selection.
Framework for learning inspectable alignment through inverse RL without direct policy modification, improving reusability and transparency.
Soft advantage policy optimization using smooth gate functions instead of hard clipping for stable LLM training and reasoning.
Comprehensive benchmark comparing state space models, transformers, and recurrent networks for US power grid electricity demand forecasting.
Continual learning architecture for LLMs preventing catastrophic forgetting during sequential updates using thalamically routed cortical columns.
Offline reinforcement learning with parametric policies under general function approximation beyond state-wise mirror descent.
Federated learning algorithm addressing statistical heterogeneity and non-IID data with proximal-balanced scaling for privacy-preserving training.
Sample-efficient hypergradient estimation for decentralized bi-level reinforcement learning in strategic decision-making and environment design.
Masked discrete diffusion model with self-aware Markov transition kernels enabling adaptive reasoning and error correction in discrete tasks.
Stable end-to-end joint embedding predictive architecture learning world models from raw pixels without representation collapse.
Multi-scale convolutional architectures for time series classification using diverse input representations and multi-representation learning.
Theoretical framework for population-based neural network training combining fast within-model optimization with slower population-level adaptation.
Multi-task supervised fine-tuning algorithm addressing heterogeneous overfitting across dataset mixtures with overfitting-aware data allocation.
Precipitation nowcasting model combining radar observations with weather foundation model priors to improve long-lead forecasting accuracy.
Analysis of systematic biases in Chinchilla scaling law fitting method applied to LLM training, showing parameter allocation errors in compute-optimal estimates.
Cloud-edge collaborative system for photovoltaic power forecasting using large models with latency constraints and robustness to weather distribution shifts.
Method for routing prompts to optimal LLMs/generative models using diversity-aware adaptive selection beyond fidelity scores.
Survey on enterprise financial risk prediction using big data and LLMs, covering AI/computer science approaches to finance and management risk analysis.
Theoretical study of feature learning in Leaky ResNets via Hamiltonian mechanics. Analyzes representation geodesics and bottleneck structures in infinite-depth limits.
Set2Seq Transformer for temporal multiple-instance learning with permutation-invariant set representations. Models internal structure and temporal relationships across timesteps.
Coded computing schemes for distributed systems with probabilistic stragglers. Extends exact computation frameworks to handle approximate recovery scenarios.