Low-Rank Online Dynamic Assortment with Dual Contextual Information
ML research on dynamic assortment optimization with dual contextual information for e-commerce recommendation systems.
ML research on dynamic assortment optimization with dual contextual information for e-commerce recommendation systems.
ML research on continuous-time Q-learning for jump-diffusion models using Tsallis entropy regularization instead of Shannon entropy.
SaVe-TAG uses LLM-based interpolation to address long-tailed class imbalance in text-attributed graphs for GNN generalization.
Theoretical analysis of Transformers as measure-to-measure maps implemented as interacting particle systems on the unit sphere.
Post-hoc method adding probabilistic uncertainty to vision-language models like CLIP to better handle domain shifts in downstream tasks.
Hierarchical retrieval method offering interpretability and efficiency improvements over embedding-based similarity search for large-scale systems.
PoliCon benchmark evaluates LLMs on achieving political consensus objectives using deliberation records from European Parliament.
Statistical diagnosis and training methods for LLM-based web agents addressing multi-step interactions and reducing post-training compute costs.
Technical analysis of load-balancing designs for AI training workloads, comparing approaches and establishing optimality bounds for distributed training.
Highlight & Summarize method for RAG systems to prevent jailbreaking and model hijacking of LLMs through prompt injection defense.
ToolACE-MT enables non-autoregressive generation for multi-turn LLM agent interactions with complex function calls, reducing data generation costs.
Bayesian method for decentralized multi-agent reinforcement learning over networked graphs with dynamic neighborhoods and constrained communication.
HEART framework uses emotional cues during test-time scaling to improve LLM problem-solving by preventing repetitive thought patterns through alternating critical and encouraging tones.
Research on physical embodiment constraints for AI agents in eldercare and disaster response scenarios, addressing generalization and care provision under uncertainty.
VoiceAgentBench evaluates speech language models on agentic multi-turn tasks and adversarial robustness beyond isolated capabilities.
Knowledge graph completion using attention mechanisms and diffusion models for few-shot learning on long-tailed relations.
Standardized framework for tuberculosis detection from cough audio combining machine learning models with clinical variables and uncertainty quantification.
Framework and tools for building and standardizing provenance-based intrusion detection systems with consistent evaluation metrics.
Width scaling approach for multi-agent systems addressing broad information seeking through organizational capability rather than single-agent depth.
Policy-gradient framework optimizing internal attention distributions in multimodal LLMs for improved reasoning without verbose rationales.
Method for identifying visual concepts in large multimodal models to audit medical decision-making and uncover shortcut behaviors in skin lesion classification.
Domain-specialized financial language model for Indian digital payment systems, developed by NPCI using multi-stage training on 68B tokens.
Space-time regularization using learned finite element methods for inverse electrocardiographic imaging of cardiac electrical activity.
Deep temporal neural hierarchical model for predicting open source software sustainability using contribution patterns and community metrics.
GPU-Fuzz fuzzer for finding memory errors in deep learning frameworks (PyTorch, TensorFlow, PaddlePaddle) by modeling operator constraints.
Statistical provability theory explaining why agentic theorem provers combining reasoning models with library retrieval, planning, and proof verification achieve strong mathematical reasoning performance.
Theoretical characterization of trainability for instantaneous quantum polynomial circuit born machines as quantum generative models.
Privacy-preserving algorithm for computing top singular vectors using adaptive power iteration with differential privacy guarantees.
Training-free inversion stabilization for rectified-flow generative models, improving reconstruction and editing tasks without additional training.
Selective Abstraction framework for reducing factual errors in LLM long-form generation by enabling partial uncertainty-based abstention instead of binary all-or-nothing approach.
Theoretical analysis of logit regularization in linear classifiers, examining implicit bias mechanisms of label smoothing and related convex penalties.
Trajectory self-distillation method for improving few-step decoding in diffusion language models, enabling faster parallel token generation with maintained quality.
arXiv paper on lightweight, parameter-efficient fine-tuning framework for disaster tweet classification using LLMs in resource-constrained settings.
arXiv case study examining how demographic-based persona assignments affect LLM agent robustness and task performance, highlighting operational risks beyond text generation bias.
arXiv paper proposing retrieval-augmented self-taught reasoning model with adaptive chain-of-thought for correcting named entity errors in automatic speech recognition using LLMs.
arXiv paper on multimodal LLMs covering fundamentals, emblematic models, and practical techniques for preprocessing, prompt engineering, and building pipelines with LangChain and LangGraph.
propella family of small multilingual LLMs annotating documents across 18 properties for flexible LLM pretraining data curation.
RankLLM framework quantifying question difficulty in benchmarks to better differentiate LLM capabilities across model comparisons.
RBCorr method to correct response bias in language models' performance on fixed-response evaluation questions.
Evaluation of HiFloat low-bit formats for efficient LLM inference on Ascend NPUs with precision-efficiency comparisons.
CLASE hybrid method for evaluating stylistic quality of legal text generated by LLMs without manual expert annotation.
Reinterprets partition function in GFlowNet-based RL for LLMs as difficulty scheduler to improve reasoning diversity while maintaining performance.
Novel reward modeling paradigm combining generative and discriminative approaches for probabilistic preference-based LLM alignment.
Experiential Knowledge Distillation framework for compressing LLMs by learning from teacher model's original training environment.
MedXIAOHE medical multimodal foundation model with entity-aware pretraining for clinical image and text understanding tasks.
ReFilter method improving robustness of retrieval-augmented generation systems through gated filtering mechanisms for LLM integration.
Lamer-SSL framework using Layer-Aware Mixture of LoRA Experts for multilingual expansion of speech models without catastrophic forgetting.
Error analysis methodology for evaluating and improving sequence labeling systems with diagnostic and predictive insights.
BERT-MoE framework for aspect-based sentiment analysis on Persian tourism reviews addressing low-resource language challenges.
RAT-Bench benchmark for evaluating text anonymization tools' effectiveness at removing PII from data used with LLMs.