Switching Successor Measures for Hierarchical Zero-shot Reinforcement Learning
arXiv paper: Hierarchical reinforcement learning with switching successor measures for zero-shot RL without fixed temporal abstractions.
arXiv paper: Hierarchical reinforcement learning with switching successor measures for zero-shot RL without fixed temporal abstractions.
arXiv paper: Bilingual pre-training outperforms hyperparameter tuning in data-constrained low-resource language settings.
arXiv paper: Teacher-Guided Policy Optimization for LLM distillation using Reverse KL to improve student-teacher convergence.
arXiv paper: EMO framework for progressive training of Mixture-of-Experts models addressing memory and communication efficiency bottlenecks.
arXiv paper: Generalization bounds analysis for Physics-Informed Neural Networks (PINNs) and variational variants.
arXiv paper: Chem-GMNet geometric transformer for molecular property prediction using domain-native architecture over generic SMILES models.
LightSplit privacy-preserving split learning using orthogonal projections to reduce communication overhead and prevent reconstruction attacks.
Exploration algorithm balancing uncertainty resolution with action budget using expected improvement and surprisal gating.
Contextual bandits algorithm optimized for resource-constrained devices using probabilistic learning.
Real-time AI agents using asynchronous I/O and speculative tool calling for sub-1-second latency in interactive applications.
Phasor Memory Networks architecture enabling stable backpropagation for explicit memory in language models via unitary dynamics.
Analysis of signal propagation in GNNs addressing information loss through oversmoothing and oversquashing phenomena.
Teaching framework for machine learning that accounts for deductive errors in learners like LLMs during few-shot learning.
Adversarial training approach addressing long-tail data imbalance through adaptive perturbations and theoretical analysis.
Novel encoder architecture replacing VAE encoders with diffusion models for improved latent representation learning.
Data augmentation method for offline reinforcement learning using trajectory-based techniques to train models from limited suboptimal data.
Research on warmstarting techniques for scaling language models, analyzing initialization constraints and growth strategies for training efficiency.
Data-free second-order preconditioning for differentially private deep learning without privacy budget consumption.
Pipeline combining pretrained LLM table extraction with fine-tuned small model error repair for low data requirements.
Asynchronous SGD with gradient rescaling for distributed optimization under data and system heterogeneity.
Flow-based policy for stable and expressive reinforcement learning without backpropagation through solvers.
Bijective representation learning framework for robust inversion of continuous forward processes.
Improved delta rule with online preconditioning for linear attention in state-space models.
Method for discovering hidden miscalibration in model confidence across different input types.
Analysis of how tokenization and representation choices affect transformer context window effectiveness and information exposure.
RL-based LLM approach for high-level synthesis code generation using comparative rewards for quality optimization.
Inference-time alignment technique using temperature adjustment to mitigate reward hacking in LLM outputs.
Foundation model for dynamic graphs across multiple domains using decoupled prompts for multi-domain pretraining.
Self-supervised contrastive reinforcement learning algorithm for discrete action spaces without hand-crafted rewards.
Neural Low-Degree Filtering (Neural LoFi) provides spectral theory of hierarchical feature learning in gradient-based deep learning.
Proves first Õ(ε⁻²) sample complexity for single-loop off-policy actor-critic methods under minimal assumptions.
Reward-Decorrelated Policy Optimization (RDPO) for multi-task and mixed-reward reinforcement learning with heterogeneous rewards.
Geometric and spectral study of low-rank pre-training for LLMs, comparing generalization to full-rank training beyond perplexity.
Three-stage learning approach for long-term time series forecasting using simple linear models and MLPs without complex architectures.
Proposes sampling method for flow language models using marginal-conditioned bridges to preserve posterior marginal structure.
Establishes scale-sensitive generalization of PAC learning fundamental theorem using fat-shattering dimension at optimal scales.
Analyzes hierarchical synthetic languages with exact k-gram ansatz to derive explicit scaling laws and benefits of reasoning in transformers.
Analyzes regret in online learning over combinatorial actions using convex relaxations, governed by polyhedral instability.
MILM extends large language models to handle multimodal irregular time series with asynchronous observations and textual channels.
Derives tight sample complexity bounds for best-policy identification in risk-sensitive reinforcement learning with entropic risk measures.
Learns POMDP world models from observation-action trajectories using language model priors for agents in partially observable environments.
Quantized matrix multiplication technique for weight-only post-training quantization of LLMs using available covariance information.
Managed infrastructure system for LoRA post-training and serving of millions of LLM variants using shared base model deployments.
Stateful transformer architecture for efficient streaming inference with persistent KV cache reducing prefill cost to O(|q|) complexity.
Multi-level annotator modeling framework to improve reproducibility of LLM evaluation by accounting for human rater biases and subjectivity.
Theoretical analysis of vector quantization using randomized Hadamard transform for similarity search, federated learning, and KV cache compression.
Entropy-guided approach for efficient test-time scaling of agentic systems on software engineering tasks like code generation and bug fixing.
Review of foundation models for Earth science applications integrating multimodal data for perception and scientific discovery tasks.
Lightweight CNN for brain tumor classification in MRI images. Medical imaging application.
Swin Transformer with federated learning for privacy-preserving UAV image transmission in low-altitude networks.