Prediction-Powered Conditional Inference
Develops prediction-powered conditional inference method combining localization and prediction-based variance reduction for scarce labeled data.
Develops prediction-powered conditional inference method combining localization and prediction-based variance reduction for scarce labeled data.
Presents RACAS, an agentic system for controlling diverse robotic platforms with unified API and autonomous behavior pipeline.
Introduces interpolated FID metric to improve correlation between VAE reconstruction and diffusion model generation quality.
Proposes error enumeration as reward signal for RL post-training in reference-free settings without ideal answer references.
Comprehensive analysis of parallelization strategies for dense LLM deployment, evaluating tradeoffs between tensor/pipeline/sequence parallelism.
Mechanistic analysis of LLM safety mechanisms revealing decoupling between harmfulness recognition and refusal via disentangled geometry.
LLMs trained via RL to self-reflect and correct generated code without external oracles, improving complex algorithmic task performance.
Improved one-shot LLM pruning using optimal weight reordering instead of predefined order, advancing SparseGPT methodology.
Addresses ecological fallacy in language models by incorporating author context through specialized LM pretraining tasks.
Framework for fine-tuning small language models with stylized personas using structured style-rewriting to improve character consistency.
Interpretable models using LLMs to predict mental health and well-being from longitudinal social media data by integrating psychological traits.
Reinforcement learning approach for off-road autonomous driving handling unmapped terrain, variable dynamics, and long-horizon planning.
Diffusion Language Models adapt generation length dynamically, reducing computational waste on short responses in reasoning tasks.
Method for efficient vector search that generalizes across multiple K values in top-K retrieval without retraining, improving serving performance.
Adaptive sampling method using Gaussian Mixture Models to improve Physics-Informed Neural Networks training on stiff PDEs.
Framework for monitoring and preventing alignment drift in recursive self-improvement systems using multi-signal detection and constraint preservation.
Serverless deployment system for efficient serving of Mixture-of-Experts LLMs by optimizing sparse activation patterns.
Diffusion Transformer variant with dynamic token chunking that adapts compute allocation based on image content detail and denoising stages.
Kinetic-based regularization extension for learning spatial derivatives from noisy data with provable accuracy for PDE applications.
Framework enabling LLMs to execute scientific workflows with schema-gated constraints ensuring determinism, provenance, and governance.
RL-based method for retrieving diverse, property-aligned result sets using diffusion models for set-valued retrieval objectives.
Reference architecture framework analyzing 18 RL implementations to establish common patterns and standardization for RL frameworks.
Continual learning strategy for online adaptation of interactive segmentation models in medical imaging with low-parameter updates.
Method for computing certified bounds on function space norms of deep neural networks applied to PDE solutions.
Optimization technique using semantic-aware caching to improve instance retrieval efficiency in concept learning on knowledge bases.
Two-stage hybrid framework combining logical options with deep reinforcement learning to improve agent alignment and prevent over-exploitation of early reward signals.
SCOPE incremental few-shot 3D point cloud segmentation addressing catastrophic forgetting by leveraging unlabeled background scenes in sparse supervision settings.
BEVLM distills semantic knowledge from LLMs into bird's-eye view representations for autonomous driving with improved spatial consistency and reduced computation.
Tutorial and survey on predictive coding networks based on neuroscientific framework viewing brain as hierarchical Bayesian inference model minimizing prediction errors.
PACE combines parameter-efficient fine-tuning with consistency regularization to improve generalization by reducing gradient norms during transformer adaptation to downstream tasks.
FragFM hierarchical framework using fragment-level discrete flow matching and coarse-to-fine autoencoders for efficient scalable molecular graph generation.
System-level DPO method for aligning compound AI systems with multiple interacting components including LLMs, foundation models, and external tools to human preferences.
Context-Aware Priority Sampling using VQ-VAEs to improve data efficiency and handle imbalanced datasets in imitation learning for autonomous driving systems.
Controlled study examining how LLM tokenizer bias and backbone capability affect time series forecasting performance using pre-trained language models as backbone.
Federated Learning survey covering privacy-preserving distributed machine learning enabling multiple clients to collaboratively train models without centralizing sensitive data.
FourierSpecNet hybrid framework combining Fourier spectral methods with neural networks to approximate collision operators for solving the Boltzmann equation efficiently.
State Space Neural Operator for learning solution operators of time-dependent PDEs using structured state space models with adaptive damping and learnable frequency modulation.
arXiv paper providing theoretical analysis of GRPO (Group Relative Policy Optimization) for LLM fine-tuning from human feedback.
arXiv paper analyzing EM algorithm behavior under model misspecification in mixture models with excess components.
arXiv paper analyzing GNN-based SAT solvers through graph Ricci curvature geometric perspective to explain performance degradation.
arXiv paper on efficient world models for heterogeneous multi-task planning, addressing gradient conflicts and plasticity loss.
arXiv paper on assessing performance of language model applications in healthcare, addressing evaluation methodology.
arXiv paper introducing Answer-Then-Check safety alignment method to defend LLMs against jailbreak attacks using reasoning.
arXiv paper on prompt-based federated continual learning addressing class-wise and temporal forgetting across distributed clients.
arXiv paper on training diffusion language models with planner-aware path learning to optimize generation strategies.
arXiv paper formulating diffusion model alignment as variational EM to reduce reward over-optimization and mode collapse.
arXiv paper on adapting decoder-only LLMs to partial differential equations via cross-modal learning for scientific machine learning.
arXiv paper introducing KLASS sampling method for masked diffusion models using token-level KL divergence to accelerate inference.
SQDF applies soft Q-function RL to fine-tune diffusion models with KL regularization, mitigating reward over-optimization.
Analyzes diversity loss in RL-trained LLMs caused by mode-seeking reverse KL; proposes forward KL filtering for reasoning tasks.