Efficient Domain Adaptation for Text Line Recognition via Decoupled Language Models
Modular transformer approach for efficient domain adaptation in optical character recognition with reduced computational requirements.
Modular transformer approach for efficient domain adaptation in optical character recognition with reduced computational requirements.
Research on how LLMs perform scientific reasoning tasks and how prompting affects their internal reasoning processes.
Deep reinforcement learning framework using PPO to train virtual agents for guiding fish school collective motion.
Multi-agent pipeline for literature analysis using Deleuzian ontology to identify non-linear patterns in research landscapes.
Framework for evolutionary GPU kernel optimization using evaluation-driven agent and evolutionary techniques for operator generation.
Analysis of prompt framing artifacts in vision-language model evaluation on clinical neuroimaging tasks.
Survey of network performance modeling approaches comparing traditional simulation with deep learning methods.
Diffusion distillation approach using reinforcement learning to improve student model performance beyond teacher anchoring.
Medical AI Scientist system that autonomously generates hypotheses, conducts experiments, and writes manuscripts in clinical medicine.
Genetic programming pipeline for automatically evolving interpretable composite features for music tagging tasks.
Method for improving diversity in text-to-image diffusion transformers through contextual space repulsion to address typicality bias.
Study on few-shot learning and RNNs applying asymptotic equipartition property from information theory to machine learning.
Theoretical analysis of inexact Langevin algorithm convergence for score-based generative models with KL divergence guarantees.
Survey on continual graph learning covering incremental learning from streaming graph data with experience and generative replay approaches.
Mathematical analysis of auto-differentiation reliability in neural-ODE training with high-order numerical methods.
Novel prior learning method for neural networks using structured posteriors to improve generalization and uncertainty estimation.
Proves asymptotic optimality of new restless bandit policies with O(1/√N) gap under unichain and aperiodicity conditions.
Theoretical analysis of sample complexity for model-based Q-learning, establishing finite-time convergence bounds for model-learning algorithms.
Paper proposing Explaining-Away Variational Autoencoders to improve uncertainty representations in deep generative models for visual inference tasks.
Survey of multimodal continual learning methods that enable models to learn from new data across multiple modalities while retaining previous knowledge without catastrophic forgetting.
Transformers learn variable-order Markov chains in-context with finite-sample accuracy analysis using context-tree weighting.
Wavelet subspace compression for optimizer states reduces memory during LLM training, improving upon low-rank approaches.
Steering vectors applied to LLM activations for bias mitigation across social dimensions like age, gender, and race.
Survey of LLM integration with Computer-Aided Design tools, covering applications in 3D modeling and design workflows.
Foundation models for time-series prediction often use simple parroting strategies rather than learning physics, revealing shared failure modes.
Structured Agent Distillation compresses large LLM-based ReAct agents into smaller models while preserving reasoning and action consistency.
CoDec kernel optimizes LLM decoding by sharing prefix computation across multiple prompts to reduce memory-intensive KV cache access.
MicroMix: mixed-precision quantization method using microscaling formats for efficient LLM inference on NVIDIA Blackwell hardware.
Novel fine-tuning mechanism for LLMs that addresses data quality/volume issues through controlled forgetting to improve domain adaptation.
PENGUIN: Transformer variant with periodic-nested group attention mechanism for improved long-term time series forecasting.
Empirical study of initialization schemes for Kolmogorov-Arnold Networks, proposing theory-driven approaches to improve training of spline-based KANs.
Training-free framework for deferring predictions to multiple experts using conformal prediction without retraining.
ReTrack enables data unlearning in diffusion models via importance sampling to remove memorized training data influence.
Algorithms for distributed RL with policy gradients under asynchronous parallel computation and communication.
Uses LLMs to programmatically synthesize anomaly detectors for tabular data without direct processing of raw data for privacy.
ACE framework evolves context for self-improving LLM agents, addressing brevity bias and context collapse in iterative refinement.
Mitigates premature exploitation in particle filtering for inference-time scaling of language models using process reward models.
TabPFN-Wide extends prior-data fitted networks for tabular data with extreme feature counts in biomedicine applications.
Constraints-of-Thought framework enables LLMs to perform constrained multi-step reasoning while satisfying symbolic constraints and user intent.
PANTHER applies generative pretraining to model user behavior sequences beyond language, using multi-dimensional action attributes.
Bandit algorithm for high-stakes sequential decision-making that learns when to abstain from actions with irreparable consequences.
RL algorithm for learning policies that maximize return while inducing dispersed state distributions across multiple reward sources.
Transformer-based symbolic regression method for discovering interpretable mathematical expressions from observed data.
Two-stage entropy approach for noise-tolerant multimodal LLM training using reinforcement learning with verifiable rewards.
Object-centric world models for reinforcement learning using decomposed representations to improve sample efficiency in multi-object environments.
UniGame addresses inconsistency in unified multimodal models between understanding and generation through adversarial framework.
SAFLe framework enabling scalable non-linear federated learning in a single round with heterogeneous data distribution invariance.
Domain adaptive retrieval using prototype-based semantic consistency alignment to transfer knowledge from labeled to unlabeled domains.
Research on measuring noise in LLM evaluations using statistical methods to separate signal from noise in prediction, data, and combined noise.
Day-ahead electricity price forecasting combining linear models, neural networks and online learning for volatile market prediction.