Revisiting Adam for Streaming Reinforcement Learning
Adam optimizer revisited for streaming reinforcement learning without replay buffers, enabling online policy updates from continuous interactions.
Adam optimizer revisited for streaming reinforcement learning without replay buffers, enabling online policy updates from continuous interactions.
Calibration method for Process Reward Models using conditional optimal transport to improve inference-time scaling in LLM reasoning.
Error attribution framework for multi-agent LLM systems using conformal prediction with distribution-free coverage guarantees.
Monge Inception Distance metric for evaluating generative models using sliced Wasserstein distance, addressing FID limitations.
Model-to-Data framework improving transparency and explainability of Graph Neural Networks by shifting complexity from models to data representations.
PAC learning theory for autoregressive chain-of-thought reasoning in LLMs, extending online learning analysis to token generation mechanisms.
Self-evolving trading agent using interpretable rubric policy with bounded prompt optimization for financial markets.
Unified measure-theoretic framework showing diffusion, score-based, and flow matching as instances of vector field learning.
Shadow Mask Distillation technique compressing KV cache during RL post-training of LLMs for memory efficiency.
Analysis of benchmark-utility gap in generative AI across 28 real-world deployments, identifying evaluation failures.
Dataset watermarking technique for closed LLMs enabling provable detection of proprietary or benchmark data usage.
Method to adapt autoregressive LMs to diffusion LMs via representation alignment without retraining.
TraXion proposes new pre-training framework for mobility data reflecting structural properties of human trajectories rather than importing language modeling objectives.
Direction-informed adaptive test-time compute for LLM agents uses directional signal interpretation to decide when additional computation improves performance.
Studies internal inconsistencies in LLM probabilistic reasoning, analyzing whether LLMs update beliefs consistently with Bayesian principles as evidence changes.
Tyche efficient probabilistic weather forecasting using one-step flow models, reducing inference cost compared to diffusion-based ensemble methods.
Target-aware data augmentation for SAT prediction reduces labeling costs by improving learning-based solvers for NP-hard Boolean satisfiability problems.
MAGIQ proposes multi-agentic AI governance system with post-quantum cryptography for secure agent communication and accountability.
Learned Lyapunov shielding augments adaptive control with learned quadratic Lyapunov functions and physics-informed neural networks for safety filtering.
Reproducible calibration workflow for prompt-based LLMs in evidence synthesis tasks, separating task rules from mutable prompt harness with explicit metrics.
Analyzes systematic bias in LLM-as-a-Judge evaluation, proposing estimators to correct bias and improve calibration stability for model comparisons.
C3PO network applies causal-aware foundation models to bilevel optimization for dynamic pricing and assortment selection in discrete choice settings.
ProtoSSL enables interpretable time-series prediction through prototype learning from unlabeled data, providing case-based explanations.
RL-based approach for generating physically stable brick structures without external simulators, using learned constraint satisfaction during generation.
Kurtosis-guided denoising score matching method for detecting anomalies in tabular data by learning score functions from noise-corrupted samples.
Unified theoretical analysis of f-divergence regularization in RLHF for LLM post-training, exploring alternatives to reverse KL divergence.
PLOT advances causal abstraction for neural network interpretability using optimal transport for efficient localization of relevant neural sites.
FastOmniTMAE proposes parallel clause learning for efficient Tsetlin Machine embeddings in NLP, offering interpretable alternative to BERT and Word2Vec.
Framework using response time alongside binary choice data to align LLMs with heterogeneous user preferences without pooling feedback.
Analysis of why safety measures in multi-task AI agents fail to generalize despite successful task execution generalization.
Echo: KV-cache-free method using spectral Koopman operators for associative recall in long-context reasoning and tool-calling.
GRU-gated Graph Attention Network for identifying vulnerable transmission lines and predicting cascading failures in power grids.
Dual-agent framework using implicit adversarial preference optimization to improve AI health coaches based on motivational interviewing.
Delulu: Verified multi-lingual benchmark of 1,951 code hallucination samples in fill-in-the-middle tasks across 7 languages.
Mechanistic study of how RL-based adversarial attacks successfully jailbreak LLMs through multi-step sequential optimization.
Port-Hamiltonian approach to risk-aware navigation policies that adapt evasive maneuvers based on local scene context.
PACEevolve++: Reinforcement learning framework enabling test-time policy adaptation for LLM-driven evolutionary search agents.
Theoretical framework for differentially private reinforcement learning with general function approximation beyond tabular/linear settings.
Method for discovering minimal Markovian states from causal DAGs in reinforcement learning without assuming states are pre-provided.
Dr. Post-Training framework reconceptualizes data selection in LLM fine-tuning as regularization to prevent overfitting on scarce target data.
Framework for selecting best pretrained model for new tasks from thousands of open-source models using transferability estimation without expensive evaluation.
Approach for discovering and composing learned concepts from diffusion model score functions at test time for compositional generation.
Actor-critic algorithm optimizing behavior policy via importance sampling to reduce variance in policy gradient estimation.
Method for efficient model evaluation using cached responses from previously-evaluated models to reduce queries needed for benchmark assessment.
Theoretical analysis of fundamental limits on reward improvement for LLM alignment via RL and best-of-N selection methods.
Graph neural network approach for solving max-cut combinatorial optimization within branch-and-bound using learned semidefinite relaxations.
Optimal rollout allocation strategy for group-based RLVR improving LLM reasoning by dynamically distributing compute based on prompt saturation.
Error analysis and applications of neural solvers for Hamilton-Jacobi-Bellman equations in continuous-time model-based reinforcement learning.
Theoretical analysis of in-context reinforcement learning with chain-of-thought, explaining convergence and emergence of adaptation capabilities at inference time.
Empirical study revealing LLMs fail at retrieving last items in short lists despite strong few-shot performance, characterizing the 'Position Curse' failure mode.