Stability Annealing Selects the Implicit Bias of Smoothed Sign Descent: A Rate-Indexed Barrier Path on Separable Data
arXiv paper on stability annealing in smoothed sign descent; theoretical ML research on convergence and implicit bias.
arXiv paper on stability annealing in smoothed sign descent; theoretical ML research on convergence and implicit bias.
arXiv paper on learning resource allocation between automated chatbots and human agents in sequential task processing systems.
arXiv paper on SplineNet, an isogeometric deep learning method for designing complex shell structures using B-spline representations.
arXiv paper on online self-supervised learning for echo state networks with perturbation-based methods for autonomous adaptation.
arXiv paper on anomaly detection in cyber-physical systems using joint latent clustering to model normal behavior rather than faults.
Training-free method for accelerating diffusion and flow matching model sampling using x-prediction without retraining or distillation.
Novel optimizer combining extragradient methods with sharpness-aware minimization to improve generalization in deep learning by finding flatter minima.
Token-level explainability method for electronic health record foundation models using transformer attention analysis for clinical interpretability.
Theoretical analysis of infinite-width Gaussian-process limits for random neural networks using tensor programs with quantitative convergence bounds.
Physics-informed neural network framework for learning embeddings of PDE solution families using multihead architecture.
TILDE: Novel machine learning method for concept unlearning in text-to-image diffusion models while preserving model quality.
EntroPath method for manifold learning using maximum entropy diffusion path ensembles. ArXiv research on geometric recovery from graphs.
Spectral analysis of attention mechanisms for graph denoising showing linear attention alignment with denoising objectives.
Product-of-experts bridge for parallel decoding in diffusion language models improving generation quality over standard importance sampling.
Dynamical mean field theory explaining generalization in overparameterized neural networks via broken ergodicity and fluctuation-dissipation violation.
Performance optimization and comparative analysis of generative AI models across heterogeneous accelerators for deployment efficiency.
Life cycle assessment of Lucie 7B LLM pre-training covering operational, embodied emissions, and water consumption on HPC infrastructure.
Analysis of naturally occurring statistical patterns in ImageNet that function like backdoor triggers without malicious insertion.
Challenges retention-centered continual learning paradigm, proposing adaptation-focused approach for non-stationary environments.
REVIVE multi-modal framework for detecting and recovering autonomous vehicle cameras from vandalism-induced occlusion attacks.
Foundation model enhanced with GNSS-derived atmospheric water vapor data for improved precipitation forecasting.
ML approach using physics-based regularization for IMU-based vehicle localization without GNSS.
EVC-Mamba method for GNSS-denied vehicle localization using evidential deep learning to correct inertial drift.
Controlled study comparing BPE and Unigram-LM tokenizers for chemistry SMILES representations in chemical language models.
Association Restoration Test diagnostic for evaluating whether unlearned label-attribute shortcuts remain functionally usable by classifiers.
Coupled digital-twin framework for autonomous microscopy combining predictive models of sample response and instrument detection.
Onnes: multi-agent LLM simulator with physics-grounded digital twin for cryogenic fault diagnosis in quantum computing infrastructure.
Mixed-mode advantage regularization technique to mitigate factual hallucinations in large reasoning models during question answering tasks.
Few-medoids coreset selection method for efficient few-shot knowledge distillation using simple representative subset identification.
Dynamic Voltage Frequency Scaling methods for energy-efficient fine-tuning of small language models on embedded GPU devices.
Techniques using retrieval-augmented generation and constrained decoding to improve LLM accuracy in generating correct web API invocation code.
Case study evaluating domain adaptation benefits for sentiment analysis with frozen pre-trained language model backbones of varying sizes.
Multi-channel spread-spectrum watermarking scheme for LLM-generated code to enable attribution and provenance tracking with high payload capacity.
Systematic study of reinforcement learning reward function design for improving LLM-generated BPMN process models.
Automatic audio annotation pipeline for converting unlabeled domestic audio into labeled training data for classification.
Training-free acceleration technique for Vision-Language-Action models using action caching for faster robotic inference.
Theoretical analysis of neural networks outperforming neural tangent kernel on compositional tasks, quantifying performance gaps.
ELSA3D unified 3D foundation model with elastic semantic anchoring for 3D generation and understanding.
Mean-field theory analysis of three-layer neural network training dynamics and convergence properties.
DIRA-SS self-supervised adaptation for DNNs under domain shift in autonomous systems without retraining.
Trust-free personalized decentralized learning framework for federated settings without centralized coordinators.
Theoretical study of benign overfitting in over-parametrized binary linear classification models.
Analysis of o3 reasoning efficiency: improved LLM performance from better reasoning rather than longer chains.
Supervised reward inference approach handling diverse human behaviors for goal inference.
Theoretical analysis of selective SSMs like Mamba through passivity and stability properties with token-dependent gating.
Distance Explainer method for post-hoc interpretability of embedded vector spaces in ML models.
TV-INRs probabilistic framework combining implicit neural representations with latent variables for time series.
Learning minimum action distance as state representation metric from trajectories without rewards or actions.
Variational framework for learning disentangled representations that separate shared and condition-specific factors.
Adaptive approach for continuous conditional GANs addressing imbalanced label distributions in generative modeling.