Strategically Robust Multi-Agent Reinforcement Learning with Linear Function Approximation
Game-theoretic approach to multi-agent RL using risk-sensitive equilibrium for robust and efficient agent coordination.
Game-theoretic approach to multi-agent RL using risk-sensitive equilibrium for robust and efficient agent coordination.
Optimal control formulation for reasoning in language models to enable planning and goal-directed action selection.
Methods for efficient reasoning in Transformers at fixed test-time cost using attention priors and training techniques.
Spiking neural networks for spatiotemporal event-based data classification with improved energy efficiency and temporal decoding.
Theoretical analysis connecting training dynamics in Gaussian mixture models to surrogate systems using Gordon comparison theorem.
Reward-Zero: implicit reward mechanism using language embeddings to derive progress signals for reinforcement learning without explicit reward functions.
Graph neural network model for detecting anomalies across multiple domains with adaptive testing-time mechanisms to handle domain shift.
Dataset condensation method for clinical ML models using synthetic data and differential privacy to democratize healthcare AI while protecting patient privacy.
Contrastive learning approach for attributed hypergraph clustering with direct clustering supervision integration.
SPAARS: offline-to-online RL safety method for robotics using abstract exploration and refined action space exploitation.
Analysis of MDP design choices impact on sim-to-real transfer in reinforcement learning for industrial process control.
Nonparametric off-policy evaluation method for contextual bandits addressing limitations of inverse probability weighting.
Temporal-conditioned normalizing flows framework for multivariate time series anomaly detection with uncertainty modeling.
XLA-compatible state space model inference implementation enabling O(1) autoregressive caching without NVIDIA hardware dependency.
Constraint-based structure learning for Markov and Bayesian networks with unreliable conditional independence oracles.
Optimal control-theoretic framework for transformer training with structured constraints and McKean-Vlasov dynamics.
Routing mechanism for online continual learning in transformers without forgetting, addressing non-stationary streaming data.
Theoretical analysis of memorization capacity in deep ReLU networks characterized by width and depth parameters.
MM-algorithms for non-negative matrix factorization with Tweedie and Negative Binomial cost functions for unsupervised learning and feature extraction.
FreqCycle framework for time series forecasting using multi-scale time-frequency analysis to capture mid to high frequency patterns.
Research on how label and selection bias impact ML classification model evaluation, performance, and mitigation strategies.
Open-source framework evaluating graph neural networks for time series anomaly detection with critical benchmarking.
Empirical study of catastrophic forgetting in LoRA and parameter-efficient fine-tuning methods during sequential learning.
Active learning pipeline for efficiently generating preference data annotations for RLHF-based LLM alignment.
Federated knowledge distillation approach for AI-native radio access networks in multi-access edge computing systems.
Adaptive channel pruning scheme for split learning to reduce communication overhead in distributed training.
Bayesian optimization algorithm for optimizing probability distributions and mixtures on the probability simplex.
In-context reinforcement learning approach that uses high-quality reasoning traces as better demonstrations for improving LLM reasoning.
Lightweight pseudo-projector modification for transformer-based language models to reduce noise sensitivity in hidden representations.
GAST combines gradient-aligned sparse tuning with data-layer selection for parameter-efficient fine-tuning of large language models.
MSSR is a memory-aware replay strategy for continual LLM fine-tuning that reduces catastrophic forgetting during sequential task learning.
OptEMA optimizer improves exponential moving average with adaptive stepsizes, achieving zero-noise optimality for stochastic optimization.
Study of learning rate sensitivity in PPO actor-critic methods, analyzing early structural signals to predict training stability.
Neural debugger for Python that trains LLMs on execution traces to predict line-by-line program execution for debugging workflows.
Analysis of neural network optimizers (AdamW, Muon) as steepest descent under matrix norms, addressing width scaling stability.
Layer-wise representational analysis comparing diffusion language models and autoregressive LLMs, examining layer-skipping capabilities.
Novel techniques (OAS, MBS) improve MXFP4 quantization accuracy for efficient LLM inference, addressing gaps versus NVIDIA's NVFP4.
KernelCraft benchmarks agentic LLM systems for generating low-level kernels for novel AI accelerator instruction set architectures.
ALADIN framework for design-space analysis of mixed-precision quantized neural networks on resource-constrained embedded AI accelerators.
Review of ultra-low-power edge AI processors including SoCs, neural accelerators, and in-sensor architectures for embedded inference.
Research on dataflow-based CNN accelerators on FPGAs addressing data-rate inefficiencies in layers with reduced output dimensions.
Auralink SDC deploys autonomous edge AI agents for electric vehicle charging infrastructure management, achieving autonomous operation with edge computing latency requirements.
Sensitivity-guided compression framework for reservoir computing enabling design-space exploration of quantization, pruning, and hardware efficiency trade-offs.
AetherFloat family proposes block-scale-free quad-radix floating-point architectures reducing silicon area and power overhead in AI accelerators.
Permutation-equivariant 2D state space models for multivariate time series, formalizing permutation symmetry principle for exchangeable variables.
Formal analysis proving that no verification procedure can simultaneously satisfy soundness, completeness, and decidability for AI alignment certification.
MASEval extends multi-agent evaluation beyond model-centric benchmarks to evaluate LLM-based agentic system components including topology, orchestration, and error handling.
APPLV automates parameter tuning for autonomous navigation by learning from vision-language-action models, balancing safety assurances with learning flexibility.
FedLECC proposes cluster and loss-guided client selection for federated learning under non-IID data, improving convergence in distributed AI systems.
Vision-language models encode clinical guidelines for interpretable medical reasoning in Concept Bottleneck Models, enabling transparent AI in medical imaging.