Irminsul: MLA-Native Position-Independent Caching for Agentic LLM Serving
Position-independent caching system (Irminsul) for agentic LLM serving that prevents cache invalidation from token shifts during multi-turn interactions.
Position-independent caching system (Irminsul) for agentic LLM serving that prevents cache invalidation from token shifts during multi-turn interactions.
Active learning method for optimizing communication structures in LLM multi-agent systems to reduce token usage and improve performance.
Framework combining LLMs with reinforcement learning for unified 3D scene generation and interactive user interaction in multimedia systems.
Research on linear decodability versus correction of medical LLM failure modes, showing limitations of fixed residual-stream steering.
Research on Fourier feature methods for nonlinear causal discovery in mixed data combining scoring and constraint-based approaches.
Research on machine learning interatomic potentials with polarizable atomic multipoles for modeling long-range electrostatics.
Theoretical work proving transformers can implement in-context reinforcement learning with policy improvement via explicit constructions.
Research on supremum-norm generalization error and uniform inference bounds for kernel gradient flow methods.
Survey of ratio-based loss functions for supervised and unsupervised learning algorithms.
Research on anytime-valid statistical inference for controlling error in LLM self-consistency aggregation methods.
Research on causal fairness methods for reducing bias in AI systems while respecting mediating variables.
Training-free diffusion sampler for image-to-image translation reducing function evaluations via exponential integrators.
Research using multi-task learning with mixture-of-experts for medical image anomaly detection from self-supervised and pseudo-labeling tasks.
Research on flow-based activation steering for inference-time control of language model behavior without parameter updates.
Research on personalized review summarization systems that adapt to individual user preferences in e-commerce using online learning.
Research on integrating quantum computing with large language models using Cayley unitary adapter parameters on quantum hardware.
Research on quantum neural network trainability showing architecture shape affects parameter efficiency and gradient flow.
Research on prototype alignment methods for heterogeneous federated learning with different data distributions and model architectures.
SIREN protocol addresses selection bias in adaptive LLM evaluation benchmarking through selection-aware repeated-split reporting.
TabCF method uses tabular foundation models for control function estimation in causal inference with unmeasured confounding.
Optimization technique for MoE inference communication bottlenecks using pooled HBM relay buffer approach on Ascend hardware.
Pipeline for end-to-end acceleration of power-of-two quantized DNNs on resource-constrained edge devices.
Prologue approach prepends learnable tokens to visual sequences to bridge reconstruction-generation gap in autoregressive image generation.
Method for jointly training discrete image tokenizers and autoregressive priors using Wasserstein gradient flow.
Proposes graphlet-based structural vocabulary for knowledge graph foundation models to enable discrete symbolic representations.
PAGE method optimizes LoRA adapter placement in low-rank adaptation for efficient fine-tuning of pretrained models.
Evaluation framework testing 14 LLM families on 500 C verification tasks, analyzing program semantics learning via symbolic execution traces.
Multimodal framework combining CNNs and LLMs to generate clinically interpretable diagnostic narratives for brain tumor classification.
Research on variable codebook size quantization for discrete visual tokenizers to overcome information-theoretic limits in autoregressive image generation.
TIDE revisits single-injection token embedding in LLMs; proposes per-layer token lookups to address rare token under-training and improve efficiency.
LatentRAG improves agentic RAG by having LLM generate latent reasoning and subqueries in lower-dimensional space for efficient multi-step retrieval.
Multimodal deep generative model for semi-supervised learning handling class imbalance via addressing bias in pseudo-labels.
Demonstrates log-likelihood signals for detecting machine-generated text are non-uniform across hidden space, revealing Simpson's paradox in token-level analysis.
Identifiable and consistent learning of recurrent switching dynamical systems using non-variational estimators for sequential data with regime changes.
Framework to certify safety audits against strategic platform manipulation via semantically equivalent content routing under online safety regulations.
Paired-prompt protocol measures evaluation-context divergence in open-weight LLMs, revealing alignment-pipeline-specific behavioral differences.
TinyBayes applies closed-form Bayesian inference with Jacobi priors for real-time image classification on edge devices, tested on plant disease detection.
MANTRA synthesizes SMT-validated compliance benchmarks for tool-using LLM agents from natural language procedural manuals to test agent behavior.
Benchmark for strategic gaming in continuous compliance audits under regulations like EU AI Act, formalizing delayed reporting and metric cherry-picking.
Analyzes how class imbalance and data heterogeneity affect training dynamics and memorization in diffusion models.
eX2L framework uses visual explanations as regularization to address distribution shifts and spurious correlations with improved interpretability.
Finite-sample analysis of Deep Q-Learning under τ-mixing to handle temporal dependencies in replayed data instead of independence assumption.
Nash equilibrium learning in partially observable Markov games with decoupled dynamics for multi-agent reinforcement learning without centralization.
Compares latent space choices (VAE vs semantic) for robotic world models using action-conditioned video diffusion for policy evaluation.
Reinforcement learning method for fine-tuning diffusion models to balance multiple reward criteria simultaneously.
Study of pricing agents in revenue management showing how standard agents can achieve high returns while failing at market-like behavior.
Evaluation framework for cooperative multi-agent RL emphasizing coordination diagnostics beyond aggregate outcome metrics.
Vision-language pretraining method combining DINO distillation with ranking-consistency loss for improved CLIP performance.
Bilevel optimization framework for robot motion retargeting using reinforcement learning to ensure physical feasibility.
Multi-agent reinforcement learning approach for cross-modal embodied navigation using modality-specialized lightweight agents.