Formalizes optimal control over natural images as reinforcement learning task, derives conditions for sufficient visual information for policy implementation.
Survey of positional encoding techniques in transformer-based time series models for forecasting, anomaly detection, and classification tasks.
Graph-conditional diffusion model for generating relational databases with support for multi-table dependencies.
Graph Tsetlin Machine extending interpretable logical learning to graph-structured inputs for pattern recognition.
Adaptive ensemble aggregation algorithm for actor-critic reinforcement learning that dynamically balances bias and variance.
Theoretical framework reinterpreting transformer attention computation through Pavlovian conditioning with linear attention analysis.
Pipeline for LLMs to discover and prove novel mathematical theorems in Lean 4 using in-context proof learning.
Signal-aware dynamic wavelet positional encoding for transformers to better capture non-stationary temporal dynamics in time series.
Method using empirical neural tangent kernel eigenanalysis to identify interpretable feature directions in trained networks.
Spiking neural network for online continual learning on Intel Loihi 2 neuromorphic hardware with minimal power consumption.
System for accelerating long-context LLM inference through context reuse for retrieval-augmented generation and agent memory.
Comprehensive multimodal fusion benchmark across specialized domains to address evaluation gaps in current fusion methods.
Theoretical analysis showing row-stochastic matrices outperform doubly stochastic matrices in decentralized learning with heterogeneous node weights.
DVPO algorithm for stable LLM post-training via distributional value modeling, handling noisy supervision in RL settings.
RL-trained Conductor model that orchestrates multiple LLM agents by discovering coordination strategies and optimizing communication topologies.
Variance-reduced domain adaptation method via stratified sampling to address domain shift in real-world ML deployment.
Transformer-based autoregressive model for 3D molecular generation with permutation invariance properties.
Method for dynamic hyperparameter importance analysis in multi-objective optimization to improve model selection efficiency.
Analysis of interaction between supervised fine-tuning and reinforcement learning during LLM post-training, examining their non-decoupling effects.
Multi-task learning approach for virtual sensors using time series foundation models to predict signals from available measurements.
Framework integrating foundation models with disentangled autoencoders to mitigate spurious correlations and improve robustness in vision tasks.
Identifies norm-feedback loops causing sequential model edits to fail, proposes norm anchors to stabilize edit composition.
Augmented Lagrangian-guided diffusion for safe offline reinforcement learning with multimodal action distributions.
Quant VideoGen: 2-bit KV-cache quantization enabling autoregressive video generation on memory-constrained hardware.
Prompt-efficient RLVR via rare-event amplification and bidirectional pairing for improved optimization and transfer in reasoning tasks.
DFPO: distributional flow method for robust LLM post-training via distributional RL, improving OOD generalization with fine-grained value modeling.
UniComp: unified evaluation framework comparing LLM compression techniques (pruning, quantization, distillation) across performance, reliability, efficiency.
Addresses class-incremental learning with pre-trained models through vision-language calibration to balance adaptation and stability.
Theoretical analysis of training dynamics in reinforcement learning with verifiable rewards for reasoning models, showing natural curriculum emergence.
Addresses intransitive preferences in LLM preference fine-tuning by modeling multi-objective trade-offs with principled scalarization.
Framework for mapping unsafe regions in LLMs using quality-diversity search to understand failure modes and improve AI safety.
Empirical analysis of LoRA as parametric knowledge memory for continuous LLM updating, comparing with ICL and RAG approaches.
RDB-PFN: first relational foundation model trained on synthetic data to overcome data scarcity in relational database pre-training.
Introduces compute-efficient pipeline for LLM data mixture optimization using capacity-aware scaling laws to improve downstream performance.
Analyzes attention sinks and massive activations in Transformers through backpropagation perspective, explaining gradient regulation mechanisms.
ARTA: Adversarially robust time-series anomaly detection via sparsity-constrained perturbation training for improved detector robustness.
GEGCN: Graph convolutional networks enhanced with Ricci flow for dynamic geometric representation learning.
Massively parallel O(N) algorithm for Hawkes process inference using sparse matrices, exploiting GPU parallelization.
Malliavin Calculus for Adaptive Inverse Reinforcement Learning: Novel passive Langevin-based algorithm for recovering loss functions from gradients.
Master Key Hypothesis: Post-trained capabilities transfer across models via linear subspace alignment without retraining, enabling cross-model transfer.
CausalGaze: Counterfactual graph intervention to detect and unveil LLM hallucinations by exposing underlying causal mechanisms.
Solver-Sampler Mismatch in multi-agent LLM negotiation: Reasoning models hurt behavioral simulation; stronger reasoning ≠ better sampling.
AutoOR: Synthetic data and RL pipeline training LLMs to autoformalize operations research problems for industrial optimization.
CAP: Controllable Alignment Prompting for unlearning sensitive information from LLMs without parameter access, enabling regulatory compliance.
Sub-Token Routing in LoRA: Fine-grained compression via sub-token routing combined with KV compression for efficient transformer adaptation.
Supernodes and Halos: Study showing loss sensitivity concentrates in ~1% of channels in transformer FFNs, enabling targeted compression.
Knowledge Distillation Must Account for What It Loses: Position paper arguing student models should preserve teacher reliability beyond task metrics.
LLMs vulnerable to semantically-invariant row/column permutations in tabular data, demonstrating fragility in Table Question Answering tasks.
Near-optimal first-order algorithm for multi-task learning with shared linear representations, addressing non-convex matrix factorization challenges.
AI Alignment via Incentives and Correction: Framework applying law-and-economics deterrence models to prevent misconduct in agentic AI systems.