Reward Shaping and Action Masking for Compositional Tasks using Behavior Trees and LLMs
Combines behavior trees with LLMs for automated reward shaping and action masking in compositional RL tasks, improving subtask learning efficiency.
Combines behavior trees with LLMs for automated reward shaping and action masking in compositional RL tasks, improving subtask learning efficiency.
Introduces Selective Rollout to reduce training costs in GRPO by early-terminating rollouts with zero reward variance in multi-turn dialogue agent training.
Proposes INTRA framework enabling attention-based models to perform retrieval-augmented generation using only internal representations without external systems.
Theoretical analysis of constant collapse in variational autoencoders using teacher-guided training, establishing exact thresholds for this failure mode.
Introduces HCInfer, an inference system using error compensation to enable efficient LLM deployment on memory-constrained devices without significant accuracy loss.
Proposes MDN for parallelizing momentum-based optimization in linear attention models to improve convergence and information retention in long-sequence LLM scaling.
Studies how LLMs generate and update hypotheses from partial information using the number game task, analyzing inference optimality and mechanisms.
Proposes Gradient-Momentum Coupling metric to measure learning progress in curiosity-driven RL exploration.
Proposes SOPE to stabilize off-policy evaluation in online reinforcement learning with prior data without manual tuning.
Addresses machine unlearning via retain-neutral surrogates in min-max formulation to remove training data influence.
Proposes VARS-FL for client selection in federated learning on non-IID data using validation-aligned scoring.
Presents VisMMoE system for efficient offloading of visual-language mixture-of-experts models on memory-constrained platforms.
Proposes Near-Policy Distillation to accelerate knowledge distillation for autoregressive models via asynchronous generation and selective packing.
Studies failure mode where LLMs suppress factual corrections when embedded in task-oriented requests; benchmarks 8 models with suppression rates 19-90%.
Proposes Hyperspherical Confidence Mapping for uncertainty estimation in neural networks without sampling or distributional assumptions.
Theoretical analysis improving regret bounds for kernelized bandit optimization under model misspecification.
Research on training transformers to be more compressible for KV cache compression, addressing long-context language modeling bottlenecks.
Paper on compressing high-fidelity flow-matching models for dynamical systems into efficient student networks for real-time inference.
Research on training steering vectors for LLM behavior control without sacrificing generation quality or requiring per-vector tuning.
DiBA proposes diagonal and binary matrix factorization for efficient neural network weight compression in linear layers and attention mechanisms.
Theoretical analysis of uniform convergence for halfspace learning, improving on standard VC dimension bounds.
Paper proving the efficiency of randomized Hadamard transforms for quantization in gradient compression, KV-cache compression, and vector search.
Research paper presenting matrix-decoupled concentration bounds for autoregressive LLM sequences with improved dimension-free guarantees.
Principled approach analyzing multi-agent LLM decision-making using Blackwell informativeness theory with formal guarantees.
Large-scale empirical study on synthetic data augmentation for time series forecasting across architectures and datasets.
Framework using optimal transport theory to train robust reward models for RLHF despite noisy preference data.
Studies trade-offs between batch size and prefix homogeneity in LLM inference to optimize memory-bound token generation.
Input-space adapter for tabular foundation models enabling efficient task-specific adaptation without full fine-tuning.
Cross-site fMRI graph learning method addressing out-of-distribution generalization across brain imaging sites via transient neurodynamics.
Generation-efficient uncertainty estimation method for LLMs reducing inference cost of hallucination detection without full autoregressive generation.
CoExVQA provides self-explainable Document Visual QA through chain-of-explanation predictions grounding answers to document evidence.
MTG-Causal-RL benchmark for causal reinforcement learning on Magic: The Gathering with partial observation, masked action space, and structural causality.
nGPT architecture with unit hypersphere constraints enables stable 4-bit LLM training without random transforms or per-tensor scaling.
PRISM applies iterative cross-modal posterior refinement to dynamic text-attributed graph representation learning using multimodal fusion.
Position paper arguing diffusion model generalization requires fundamentally new theoretical frameworks beyond classical learning theory and benign overfitting.
Fast Gauss-Newton decomposes multiclass softmax cross-entropy curvature into true-vs-rest and within-competitor terms for scaled optimization.
Extended Decision Transformer architecture conditioning method that injects Return-to-Go outside sequential modeling for improved offline RL efficiency.
BoostLLM applies boosting paradigm to LLM fine-tuning for few-shot tabular classification, improving performance in low-data regimes.
Listwise Policy Optimization reveals group-based RLVR methods for LLM post-training share common geometric structure as target-projection on response simplex.
SymDrift enables one-shot generative modeling of physical systems via equivariant diffusion models that preserve global symmetries like rotations.
Mathematical proof extending equivalence between augmented Lagrangian and optimistic primal-dual methods to matrix-valued corrections for constrained optimization.
Theoretical unification of goal-conditioned RL and unsupervised skill learning through control-maximization framework explaining why mutual information skill learning supports downstream goal-reaching.
AdaGamma proposes state-dependent discount factor method for deep reinforcement learning to improve planning horizon and bootstrapping stability in actor-critic algorithms.
Proposes unified scoring mechanism for joint parameter and data selection in LLM fine-tuning to reduce computational overhead.
In-context black-box optimization handling unreliable side information from multiple feedback sources. Generalizes across feedback types.
Budget-constrained contextual bandit algorithm for adversarial contexts with cost constraints and realizability assumption.
Bandit learning framework for open multi-agent systems with dynamic agent arrivals/departures and endogenous non-stationarity.
Federation of Experts architecture reduces communication bottlenecks in distributed MoE LLM inference by clustering experts per KV head.
Extension of identification-in-limit learning to contrastive settings with partially labeled data. Theoretical learning framework.
Framework for analyzing neural networks as continuous piecewise affine functions. Improves interpretability and expressivity analysis.