Framework for evaluating tabular foundation models using proper scoring rules for probabilistic predictions beyond point estimates.
Agentic framework for autonomous multimodal query processing with dynamic supervisor delegating tasks to specialized tools across modalities.
Benchmark for evaluating vision-language models on unsafe action detection for household robots in embodied AI scenarios.
Introduces Proof-Carrying Materials framework with formal safety certificates for machine-learned interatomic potentials, demonstrating 93% recall improvement over standard filters.
Introduces yat-product kernel operator with self-regularization properties, proposing Neural Matter Networks replacing conventional activation-normalization blocks with single operation.
Proposes MOGP-MMF multi-objective genetic programming framework for protein secondary structure prediction through automated feature selection and fusion.
Reviews synthetic data generation for brain-computer interfaces using deep learning, benchmarking approaches to mitigate limited physiologically-plausible neural recording data.
Proposes GER-steer, training-free activation steering framework for LLMs using cross-layer consistency to improve control precision and reduce noise sensitivity.
Proposes optimization framework integrating Minimum Description Length principle as active driving force in neural network training guided by geometrically-grounded cognitive manifold.
Introduces Hierarchical Causal Primitive Dynamic Composition Network enabling self-improving causal reasoning including interventions, counterfactuals, and mechanism understanding.
Applies non-equilibrium thermodynamics to formalize curriculum learning in reinforcement learning, proposing geometric framework interpreting reward parameters as task coordinates.
Develops efficient exploration methods for reinforcement learning maximizing entropy of steady-state visitation distribution without requiring rollouts during training.
Proposes TreeKD knowledge distillation method transferring tree-based specialist model knowledge into generalist LLMs for molecular property prediction in drug discovery.
Introduces Budget-Sensitive Discovery Score, a formally verified metric for evaluating AI-guided scientific candidate selection with 20 theorems proving correctness.
Proposes PDE-aware selective state-space model with nested memory for spatiotemporal traffic forecasting in cellular networks addressing scalability and heterogeneity.
Establishes theoretical connection between drifting generative dynamics and gradient flows induced by Sinkhorn divergence with cross-minus-self decomposition.
Proposes NeuroLoRA, a context-aware neuromodulation approach for parameter-efficient multi-task LLM adaptation improving upon static LoRA routing mechanisms.
Investigates performance gap between multimodal and unimodal models in context-aided forecasting, proposing solutions for improving context quality in datasets.
Studies length generalization degradation in Mamba sequence models through controlled image reconstruction tasks, analyzing performance when inference lengths exceed training lengths.
Two-phase curriculum sampling improves flow matching training efficiency by balancing convergence speed and asymptotic quality.
Shows LLM judge scores correlate poorly with best-of-N selection performance, highlighting evaluation metric limitations.
TERMINATOR learns optimal early stopping points for chain-of-thought reasoning in large reasoning models to reduce overthinking.
Shows transformer depth dynamics can be approximated by low-order linear surrogates, improving interpretability of language models.
CALF framework trains distributed RL policies under realistic network conditions including latency, jitter, and packet loss.
RL post-training for diffusion language models using entropy-guided step selection and stepwise advantages for sequence generation.
Empirical scaling laws for single-layer physics-informed neural networks on nonlinear PDEs identifying optimization pathologies.
Lyapunov stable graph neural flow uses control theory to defend GNNs against adversarial perturbations without adversarial training.
CA-HFP enables device-specific pruning in federated learning using curvature-aware significance scores and model reconstruction.
Swap-guided preference learning addresses posterior collapse in personalized RLHF for aligning diverse human preferences.
Consensus aggregation method for policy optimization in PPO using Fisher information geometry to reduce noise drift.
Feynman agent generates knowledge-infused diagrams at scale using multi-modal AI for visual design applications.
Speculative decoding for LLM inference acceleration combined with online learning to improve draft model quality and token acceptance rates.
Human-AI collaborative system using Bayesian optimization and proxy modeling for autonomous experimental design.
Budget-aware tree search method for LLM agents optimizing token/tool usage during inference with value estimation.
Surrogate modeling approach using diffusion models for uncertainty quantification in high-dimensional nonlinear dynamical systems.
Method for reducing Mixture-of-Experts redundancy through expert replacement to lower memory requirements in large language models.
Reasoning LLM for retrosynthesis prediction in organic chemistry using strategic bond disconnection analysis.
Neural surrogate model for parameterized PDEs using disentangled latent dynamics for parameter generalization and temporal extrapolation.
Method for enzyme function annotation using active learning and protein language models to predict biochemical reactions.
Optimization technique for multimodal LLM inference using heterogeneous GPU tiers to reduce costs by partitioning vision and language tasks.
Benchmark and methods for evaluating LLMs on scientific inverse design problems across multiple domains.
In-context operator learning method for spatiotemporal prediction enabling generalization across different problem instances.
Benchmark evaluating automated theorem prover LLMs on novel mathematical frameworks beyond standard libraries.
Theoretical analysis of learning coefficients for three-layer neural networks in singular learning theory.
Large reasoning model for molecular science integrating scientific logic with deep learning for knowledge-guided predictions.
Prompt-based continual learning method addressing catastrophic forgetting in domain-incremental settings with structural knowledge preservation.
Research investigating whether the MNIST handwritten digit dataset is linearly separable in high-dimensional space.
Research on test-time reinforcement learning for LLM alignment, revealing that benchmark performance may reflect task familiarity rather than genuine capability. Proposes train-before-test approach.
LLM-based framework for drug-drug interaction prediction integrating adaptive drug knowledge to handle imbalance and improve generalization.
DirPA addresses prior shift in few-shot crop classification under class imbalance using distribution shift correction.