Multi-AUV Cooperative Target Tracking Based on Supervised Diffusion-Aided Multi-Agent Reinforcement Learning
Multi-agent RL approach for cooperative AUV target tracking using diffusion models to address non-stationarity and coordination challenges.
Multi-agent RL approach for cooperative AUV target tracking using diffusion models to address non-stationarity and coordination challenges.
Framework for high-quality dataset generation from closed-loop automotive data collection for ML model development.
Novel neural architecture (Metriplector) based on metriplectic field dynamics enabling gradient-free computation.
Framework (PRoSFI) for generating verifiable step-by-step reasoning in LLMs using structured formal intermediaries and process rewards.
Benchmarks language models on child-scale datasets to understand data efficiency and linguistic knowledge emergence.
Agentic system for automated medical coding from clinical text using scalable, explainable approach that adapts to new codes.
Statistical learning approach for unbounded density ratio estimation and covariate shift adaptation without assuming bounded ratios.
mlr3mbo: modular R toolbox for Bayesian optimization supporting single/multi-objective, parallelization, and custom algorithm construction.
Reasoning-driven approach for generating synthetic multi-modal training data without manual prompts, addressing scarcity of specialized AI training datasets.
DIAL proposes decoupling intent and action in Vision-Language-Action models via latent world modeling to improve decision-making and training stability in end-to-end robotic control.
Hybrid machine learning framework for graduate admission prediction and university-program recommendation using 13,000 GradCafe records.
Method using epistemic uncertainty to identify unreliable explanations in post-hoc XAI methods, reducing explanation generation costs.
Proposes adaptive reasoning allocation during code generation for LLMs, addressing limitations of upfront thinking approaches in handling code complexity.
Early exiting predictive coding neural networks optimized for edge AI devices with resource constraints and privacy requirements.
GenOL framework for online learning with only concept names (name-only setup) enabling real-time adaptation to data distribution shifts in continual learning scenarios.
Introduces WEATHER-5K dataset and benchmarks physics-informed time-series forecasting models for global weather prediction.
Control-theoretic approach to reinforcement learning with convergence guarantees, new gradient theorem, and gradient ascent algorithm.
Information-theoretic analysis of transformer in-context learning on variable-order Markov chains with finite-sample accuracy bounds.
Diffusion sampler using value functions with invariant symmetries for sampling from unnormalized target densities.
Critical evaluation of model inversion attack assessment frameworks, identifying flaws in standard evaluation methodology.
Neural Graduated Assignment method for solving Maximum Common Edge Subgraph problem with improved scalability.
Training-free framework for compiling sparse Mixture-of-Experts variants with predicted expert utility metric for deployment optimization.
Framework for characterizing epistemic errors in uncertainty-aware multitask learners under distribution shift.
Probabilistic inference speedup for Hidden Markov Models by filtering low-probability states in temporal sequences.
Research on dynamic reward weighting for multi-objective RL alignment in LLMs, addressing non-convex Pareto fronts in preference learning.
Theoretical convergence analysis of Muon optimizer for matrix-structured parameters in neural network training.
Out-of-distribution detection for regression tasks in scientific AI using score-based diffusion models on joint likelihood estimation.
Transformer-based inter-atomic potential model for molecular simulations without explicit equivariance constraints.
Analysis of implicit models with infinite-depth weight-tied networks that match explicit models while reducing memory consumption.
Empirical study comparing message passing neural networks and graph transformers for atomistic property prediction.
Theoretical analysis of how attention head count influences transformer approximation properties and expressive power.
Adaptive rollout and routing method for data-driven weather forecasting with improved spatiotemporal modeling.
Learning-to-optimize Transformer framework for scalable beamforming in multi-user wireless systems.
Automated algorithm design using machine learning to optimize hyperparameter auto-tuning for high-performance applications.
Continual Transformers architecture for real-time low-latency inference on streaming data with reduced redundant computation.
Survey of deep unfolding techniques combining classical optimization algorithms with neural networks for signal processing.
Research on normalization-free transformer architectures using Dynamic Tanh as alternative to standard normalization layers.
Learning-theoretic approach to extracting interpretable features from superposition in complex ML models.
Split learning system using hybrid-order optimization to reduce memory overhead for collaborative LLM training on edge devices.
Architecture separating energy-based world models from language generation in LLMs to improve understanding vs. fluency tradeoff.
Addresses machine unlearning for sparse LLMs to remove memorized sensitive information while maintaining model sparsification benefits for efficient deployment.
ECHO-2 is distributed RL framework for LLM post-training via reinforcement learning, optimizing cost-efficiency of rollout generation across distributed resources.
VJE introduces reconstruction-free latent-variable framework for self-supervised learning using symmetric conditional ELBO on paired embeddings.
LP-FNO uses Fourier Neural Operators as surrogate model for laser welding simulations, enabling faster parametric solution learning for industrial process optimization.
DGPO: RL-guided graph diffusion model for neural architecture search using reinforcement learning steering.
Study using finetuned LLMs for topic-conditional sentiment extraction to forecast aluminum commodity prices.
Survey of privacy-preserving ML techniques for IoT including federated learning and differential privacy approaches.
Decentralized bi-level RL algorithm for environment design with sample-efficient hypergradient estimation.
MDM-Prime-v2: Improvements to masked diffusion language models through binary encoding and index shuffling.
FIPO: RL algorithm improving token-level credit assignment for reasoning in LLMs beyond outcome-based rewards.