RELOAD: A Robust and Efficient Learned Query Optimizer for Database Systems
RELOAD: reinforcement learning-based query optimizer for database systems with robust per-query performance.
RELOAD: reinforcement learning-based query optimizer for database systems with robust per-query performance.
World-Value-Action model for vision-language-action embodied agents with implicit planning capabilities.
Systematic classification and analysis of compression techniques exploiting correlations in federated learning.
Nautilus tensor compiler with automated scheduling for efficient GPU kernel generation from high-level specifications.
Bandit best-arm identification algorithm robust to both stochastic and adversarial reward distributions.
Theoretical analysis of regret tail behavior in multi-armed bandit algorithms with stochastic rewards.
arXiv paper analyzing reasoning dynamics and visual-textual information integration in 18 vision-language models.
arXiv paper evaluating multilingual text embedding models for hate speech detection across Lithuanian, Russian, and English.
arXiv paper proposing mixture-of-experts flow matching for faster language model inference while maintaining generation quality.
arXiv paper on Route to Rome Attack, demonstrating black-box adversarial suffix attacks on cost-aware LLM routers.
arXiv paper on Atropos, optimizing cost-performance trade-offs for LLM-based agents using small models with early termination and model hotswapping.
Feature selection method based on modified Shapley values for non-linear models with dependent features.
Uncertainty quantification framework for long-form LLM generation addressing factuality and coherence in open-ended text.
Machine unlearning method targeting class removal by identifying and removing forget-specific representational directions in neural networks.
Exposes vulnerability in LLM-as-judge systems where contextual framing about downstream consequences influences evaluation independent of content.
Symbolic superoptimizer for tensor programs using hierarchical symbolic graphs to represent and optimize families of implementations.
Diagnostic framework using conformal prediction and transitivity analysis to measure reliability of LLM-as-judge systems for NLG evaluation.
Controlled study examining whether LLMs can generalize systematically using shortest-path planning as a testbed to isolate training, architecture, and inference factors.
Online incremental learning method using optimal transport to manage multimodal class distributions in latent space with continuous data streams.
Survey on generative models applied to connected autonomous vehicles for predictive modeling, simulation, and decision-making.
Bilevel DPO approach for hierarchical RL addressing non-stationarity and infeasible subgoals through preference optimization.
DiffGap framework for molecule generation integrating adaptive sampling and pseudo-molecule estimation to address exposure bias in diffusion models.
Applies GNNs with human mobility data for COVID-19 forecasting, analyzing when spatio-temporal architectures outperform simpler baselines.
IMPACTX leverages XAI techniques as automated attention mechanism to improve model performance without external knowledge or manual intervention.
AutoRAN framework automating hijacking of safety reasoning in large reasoning models using weaker model simulation and iterative refinement.
Logo-LLM adapts LLMs for time series forecasting by combining local and global modeling to capture both short-term and long-range dependencies.
First unsupervised learning model for Maximum Independent Set in dynamic graphs using GNNs with learned distributed update mechanisms.
Method for estimating optimal loss value in diffusion models to distinguish between large optimal loss and insufficient model capacity.
Time-RA reformulates time series anomaly detection as reasoning task using LLM feedback, introducing RATs40K dataset for fine-grained categorization.
SPaCe applies curriculum learning to LLM fine-tuning with RL, reducing data/compute requirements by sampling examples by difficulty and learning value.
EEGDM uses latent diffusion models for self-supervised EEG representation learning, capturing global dynamics beyond masked reconstruction.
DPQuant combines quantization scheduling with differentially-private SGD/Adam to reduce training time and energy while protecting privacy.
Compares two strategies for integrating safety filters in RL: safeguarding environment vs embedding in policy through differentiable optimization.
Studies reinforcement learning under random sensor delays in POMDPs where observations arrive out-of-sequence, addressing real-world RL challenges.
Research on reduced-order modeling using deep learning to compute linear subspaces for parametric systems with offline/online stages.
Presents PreScope, a prediction-driven scheduling system for efficient MoE inference on commodity hardware with CPU offloading.
Proposes layered prefill scheduling for MoE LLM inference to optimize time-to-first-token and throughput while managing compute/memory constraints.
PatMD approach for detecting harmful memes by learning from misjudgment patterns in multimodal content with implicit rhetorical devices.
Graph-topological active learning using Balanced Forman Curvature for coreset construction under label budget constraints.
Reinforcement learning approach for language model reasoning that learns from trial-and-error to overcome exploration stagnation in RLVR.
Interlat enables LLM-based agents to communicate in latent space instead of natural language, improving information transfer depth.
AccelOpt is a self-improving LLM agent that autonomously optimizes kernels for AI accelerators using iterative generation and optimization memory.
Active learning framework for PDE surrogate modeling with selective time-step acquisition to reduce training data generation costs.
Function-word De-Attention method improves robustness of vision-language models against cross-modal adversarial attacks.
Cornfigurator automates deployment planning for any-to-any multimodal models with heterogeneous computation paths and component scaling.
Hierarchical approach combining reinforcement learning with MPC planning for sample-efficient decision making in structured planning problems.
Federated learning approach for spectral clustering in decentralized environments, capturing latent correlations across tasks.
NNGPT framework uses LLMs for neural architecture synthesis through iterative supervised fine-tuning cycles generating validated PyTorch networks.
ORBIT system for controlling reasoning budget in Large Reasoning Models via on-policy exploration-exploitation to reduce computational cost.
Theoretical analysis of differential privacy limitations in DP-SGD using f-differential privacy framework with shuffled sampling.