Ghosts of Softmax: Complex Singularities That Limit Safe Step Sizes in Cross-Entropy
Analysis of complex singularities in cross-entropy loss that limit safe optimization step sizes during LLM training.
Analysis of complex singularities in cross-entropy loss that limit safe optimization step sizes during LLM training.
Semantic-aware feature extraction framework combining geometric and semantic information for improved 3D reconstruction and feature matching.
Empirical study on how scientists reuse pre-trained deep learning models to reduce computational costs and training complexity.
Physics-informed CNN for volumetric radar motion estimation in precipitation nowcasting with altitude-dependent motion fields.
NCCL EP: unified expert parallel communication API for Mixture-of-Experts LLM training using GPU-initiated RDMA operations.
Vision-language model approach for image geo-localization using locatability-guided reasoning to reduce hallucinations and improve accuracy.
Study demonstrating gender and pronoun bias in LLM moral judgments using controlled sentence-level counterfactual analysis.
Benchmark evaluating LLM performance on bibliographic reference extraction in multilingual SSH texts with complex citation formats.
Unsupervised domain adaptation framework for 3D lesion detection in medical imaging with label shift compensation between PET tracers.
SHAMISA: self-supervised framework for no-reference image quality assessment using structured relational supervision on unlabeled data.
SyMPLER: explainable piecewise linear regression model for nonstationary time series forecasting with VC-theoretical generalization bounds.
CAP-TTA test-time adaptation framework using context-aware LoRA updates to address out-of-distribution bias in narrative generation when bias risk is detected.
τ-Voice benchmark for evaluating full-duplex voice agents on real-world tasks requiring multi-turn conversations, policy adherence, and environment interaction.
QuarkMedBench realistic medical benchmark for evaluating LLMs on unstructured, ambiguous real-world healthcare queries beyond standardized exam questions.
Graph spectral decomposition approach balancing channel-independent and channel-dependent strategies for improved time series forecasting generalization.
REFINE-DP method fine-tuning diffusion policies via reinforcement learning for humanoid loco-manipulation with improved tracking and long-horizon task execution.
R3-REC framework using retrieval-augmented LLMs with multi-level user intent reasoning for sequential recommendation with sparse and noisy data.
Implicit maximum likelihood estimation approach accelerating diffusion-based trajectory planning for real-time model predictive control applications.
Curriculum learning approach using dynamic difficulty estimation based on gradient information to improve training efficiency and task performance.
Knowledge distillation framework compressing Qwen 3B to 0.5B using chain-of-thought reinforcement learning across English, Spanish, and code datasets.
Retrieval-feedback-driven distillation framework transferring query expansion behavior from teacher to student LLM for efficient deployment in retrieval systems.
Generate-then-correct approach for aspect sentiment quad prediction extracting aspect terms, categories, opinions, and sentiment polarity from text.
AD-Copilot system enhancing multimodal LLMs for industrial anomaly detection through visual in-context comparison and domain-specific adaptation.
IGU-LoRA method for adaptive rank allocation in LoRA fine-tuning using integrated gradients and uncertainty scoring to improve parameter-efficient LLM adaptation.
Federated unlearning framework reducing computational and communication costs for removing participant data from federated learning models while maintaining privacy.
ML research: memory-efficient continual learning via prototypical exemplar condensation, reducing samples-per-class requirements below 20.
ML research: proposes Temporal Aggregated Convolution to efficiently execute spiking neural networks on SIMD architectures.
ML research: evaluates semantic robustness of text-to-audio generation under prompt variations; reveals fragility to linguistic changes.
ML research: parallel framework combining imitation and reinforcement learning for autonomous driving, addressing policy drift in sequential fine-tuning.
SWhisper: Framework for inaudible near-ultrasonic prompt injection attacks against speech-driven LLMs using commodity hardware.
APEX-Searcher: Multi-round retrieval system combining agentic planning and LLM execution for complex information seeking tasks.
Step-CoT dataset and method for medical visual question answering with expert-curated structured reasoning supervision.
Multi-modal method for Chinese text character localization and extraction addressing complexity of scene text recognition.
Study showing LLMs reproduce racial stereotypes in text annotation across 19 models and 4M+ judgments, demonstrating bias issues.
UVLM: Unified framework for loading and benchmarking multiple vision-language models across architectures in Google Colab.
Self-training framework with label correction for training neural networks on noisy datasets using bilevel optimization.
Visual state representation for robotic agents encoding what-is-where composition from video for sequential decision making.
FedPBS federated learning algorithm addressing statistical heterogeneity and uneven client participation in distributed training.
SmoothVLA: Training method for vision-language-action robotic models balancing stability and exploration via smoothness constraints.
Discriminative flow matching framework reformulating classification and object detection as conditional transport problems.
Recommendation system using LLMs for semantic reasoning from individual to group interests for generative recommendations.
Speech enhancement framework combining audio-visual processing with reinforcement learning and LLM-based interpretable reward models.
Chunk-Guided Q-Learning algorithm for offline reinforcement learning addressing bootstrapping error and policy class restrictions.
FLUX: Data preprocessing pipeline for LLM training that balances scale and quality without sacrificing token efficiency or data integrity.
Behavioral benchmark comparing self-supervised vision transformer object grouping with human perception using same/different judgments on naturalistic scenes.
Benchmark for multi-party negotiation games with configurable generator and value-function approximations for evaluating agent decision-making.
Dynamic honeypot deployment system using AI agents to adaptively select assets based on evolving attacker tactics and threat intelligence.
Reinforcement learning fine-tuning for diffusion models using centered reward distillation to improve prompt fidelity and compositional correctness.
Analysis of training-inference mismatch in logic gate networks using hard vs soft selection and Gumbel noise.
Citation-enforced RAG system for tax compliance enabling explainable and cited knowledge retrieval from fiscal documents.