Adaptive negative reinforcement method for improving LLM reasoning by dynamically balancing correction of errors and diversity in sampled trajectories.
Neurosymbolic imitation learning combining neural networks with symbolic reasoning using privileged human guidance information.
Parallel multimodal search agent using reinforcement learning to dispatch concurrent queries with efficiency awareness.
Post-training method for creating multiple nested reasoning LLM variants efficiently with computational budget control.
One-step generative model for discrete sequences using coupling between discrete structures and Gaussian latent variables.
Foundation model for zero-shot causal discovery on tabular data using transformer architecture with structured graph inference.
Framework for forecasting academic research impact using LLMs, evaluating frontier models on prospective manuscript evaluation.
Sample complexity analysis for stochastic optimization with integer variables, establishing when integer optimization requires more or fewer samples than continuous counterparts.
Mutual Reinforcement Learning framework enables concurrent RL post-training of heterogeneous LLMs with shared experience exchange and tokenizer alignment across incompatible vocabularies.
Counterfactual routing analysis evaluates mixture-of-experts routing decisions in MoE language models by comparing standard routes against equal-compute alternatives on token prediction.
PerCaM-Health discovers personalized dynamic causal graphs for individual patient healthcare reasoning from short, noisy, irregular temporal trajectories.
Practical implementation of G-bispectrum group invariants for machine learning on signals, images, and spherical data with applications to classification and pooling layers.
Bifurcation models use weight-tied dynamics to learn set-valued solution maps where different initializations converge to different equilibria for multi-solution scientific problems.
Mask2Cause end-to-end framework discovers causal graphs in time series via constrained causal attention during forecasting, avoiding post-hoc graph extraction and spurious correlations.
Convergence gap diagnostic measures per-layer next-token distribution distance showing instruction-tuned models commit to predictions later than pretrained counterparts.
Analysis of finetuning pretrained models showing optimization occurs in low-dimensional subspace; investigates why certain parameter directions remain unexplored and contain task-relevant structure.
SparseRL-Sync enables ~100x communication reduction for policy weight synchronization in large-scale RL systems via sparse weight transfer in decoupled trainer-rollout architectures.
Research on importance sampling design for LLM policy optimization in reinforcement learning, addressing bias-variance tradeoff in token-level IS ratios for PPO and GRPO.
Theoretical analysis of in-context reinforcement learning in softmax transformers without linear attention simplification, first such analysis of ICRL with standard attention mechanisms.
CellScientist framework uses LLM-assisted workflows for virtual cell modeling with closed-loop refinement, addressing routing of prediction discrepancies to relevant model components.
Multi-axis evaluation framework (Mage) for LLM-generated game code beyond compile-pass rate, testing 4 LLMs on 858 generation attempts with compile, runtime, structural, and mechanism fidelity metrics.
MISA optimization for sparse attention inference reduces indexer computational cost via mixture-of-indexers design for long-context LLM decoding.
FlightSense MLOps platform with rotation-chain propagation features and agentic AI for real-time flight delay prediction in aviation networks.
Training-free zero-shot metric using sample-wise neural activation patterns for NAS and neural network evaluation without computational overhead.
Large-scale empirical study of multi-LLM routing across benchmarks examining evaluation artifacts and unsolvability ceiling in cost-quality tradeoffs.
ROPD framework enables on-policy distillation using semantic rubrics instead of teacher logits for scalable black-box model alignment.
SR²-LoRA addresses catastrophic forgetting in class-incremental learning by analyzing inter-layer relation drift in parameter-efficient fine-tuning.
GameGen-Verifier uses parallel keypoint-based verification and runtime state injection to validate LLM-generated game correctness beyond syntax.
Causal discovery method leveraging physics simulations as interventions for molecular design and materials science under latent confounders.
Machine unlearning method for LLMs removing memorized content via self-distillation without requiring retain-set examples.
RL framework for adaptive chain-of-thought compression in large reasoning models using experience-guided difficulty-aware penalties.
Investigates transferability of neural scaling laws across domains and tasks to reduce compute requirements for fitting new model-task pairs.
Combines latent-space prediction with masked language modeling for protein sequence encoders at 35-150M parameters.
Novel RL method for large reasoning models using policy model's internal states for baseline estimation, reducing computational cost of variance reduction in reinforcement learning.
Causal Energy Minimization framework recasts Transformer layers as optimization steps on conditional energy functions to analyze parameterization choices.
CIKA framework uses LLMs as interventional simulators via causal discovery to identify concepts causally contributing to correct mathematical reasoning answers.
Research on learning modular addition functions in neural networks by controlling training difficulty through increased zeros in sequences.
STMD accelerates diffusion model inference without a teacher network while preserving probabilistic generation quality through stochastic transition-map distillation.
Improves molecule generative models through better geometric representation learning in multi-stage generation pipeline.
Addresses grammar-constrained LLM generation with speculative decoding by analyzing projected distribution gaps.
Extends LoRA with Bayesian fine-tuning in projected subspaces for uncertainty quantification in large model adaptation.
Proposes hybrid CPU-GPU sparse attention mechanism for efficient long-context LLM inference with disaggregated systems.
Theoretical study showing synthetic data collapse in recursive generative retraining can be mitigated with pluralistic reward preferences.
Benchmarks MLOps retraining strategies on semiconductor manufacturing quality prediction using five years of real data.
Presents gradient-based bilevel optimization method to learn composite loss weights online during pretraining.
Studies TabPFN tabular foundation models as summary networks for simulation-based Bayesian inference without training.
Analyzes mean-field dynamics of Transformers showing how training reshapes token clustering behavior across layers.
Proposes POETS framework combining uncertainty quantification and policy optimization for black-box LLM optimization with compute efficiency.
Studies uncertainty dynamics in LM reasoning through Chain-of-Thought by analyzing intermediate token sequences as evolving model states.
Post-hoc method for improving classification by perturbing Hessian spikes to rebalance class learning.