Universal Hypernetworks for Arbitrary Models
ArXiv paper introducing Universal Hypernetworks that generate weights for arbitrary model architectures using descriptors.
ArXiv paper introducing Universal Hypernetworks that generate weights for arbitrary model architectures using descriptors.
ArXiv paper applying diffusion denoising objectives to causal structure learning from observational data.
ArXiv paper on model-based reinforcement learning for control systems with time-varying dynamics.
ArXiv paper on in-context agentic reinforcement learning enabling LLM agents to internalize skills at inference time.
ArXiv paper on lightweight diffusion transformer for crystal structure generation using subatomic tokenization.
ArXiv paper unifying group-relative and self-distillation policy optimization for LLM post-training with improved credit assignment.
ArXiv paper proposing Head-Calibrated Clipped-Linear Softmax as efficient surrogate for attention softmax in edge inference.
ArXiv paper on exact parameterization of doubly stochastic matrices for learned mixing in neural networks.
Single-stage training paradigm for efficient LLM reasoning that reduces token consumption in chain-of-thought without degrading quality.
Learning-based cooperative coevolution framework addressing heterogeneous large-scale global optimization via adaptive low-dimensional optimizers.
Neural-symbolic framework for discovering constitutive closures in nonlinear PDEs from spatiotemporal data while avoiding spurious physical recovery.
Research on regularizing attention scores in vision transformers using bootstrapping to improve interpretability and reduce noisy attention maps.
Analysis of safety, security, and cognitive risks in world models used for autonomous decision-making in robotics, autonomous vehicles, and agentic AI systems.
Study of reliability gaps in AI-assisted medication systems, highlighting risks in healthcare decision support.
Framework combining LLMs with infeasibility detection for NP-hard combinatorial optimization problems.
Equivariant transformer architecture for modeling agent behaviors in autonomous driving with SE(2) symmetry.
New optimizer deriving design principles from Muon, improving LLM training efficiency through surrogate model analysis.
Method for efficiently adapting closed-box LLM APIs to target tasks by priming followed by local optimization.
Benchmark dataset for evaluating AI coding agents based on production workloads, addressing language distribution and codebase structure gaps.
Study of interactions between normalization methods and optimizers in LLM training at 1B parameters.
LiteInception: lightweight interpretable deep learning framework for fault diagnosis on edge devices.
LiveMathematicianBench: benchmark for evaluating LLM mathematical reasoning capabilities with proof sketches.
Study on language pre-training bias improving performance on general vision tasks through cross-modality transfer.
Analysis of permutation-invariant discrete representation learning for spatially aligned images using vector quantization.
Woosh: open-source sound effects foundation model from Sony AI with architecture, training details, and benchmarks.
Theoretical analysis of multi-head self-attention transformers using particle systems and homogenization limits.
Curia-2 foundation model using self-supervised learning for medical imaging analysis on CT and MRI data.
Comparison of centralized and decentralized RL controllers for traffic signal control in urban corridor networks.
Mining instance-centric vision-language contexts for human-object interaction detection, leveraging VLMs to improve semantic understanding and contextual reasoning.
LatentUM unified model for interleaved cross-modal reasoning combining visual understanding, generation, and world dynamics in latent space.
Hybrid framework combining LSTM workload prediction with game-theoretic heuristics for cloud cost optimization during dynamic workload changes.
AstroConcepts corpus of 21,702 astrophysics abstracts for multi-label classification research addressing extreme class imbalance with specialized terminology.
Shared task participation comparing lexical and contextual approaches for cross-document software mention coreference resolution under mention noise.
Analysis of Mixture-of-Experts LLM interpretability at expert level, comparing sparsity properties to dense feed-forward networks using probing methods.
Characterization of exact Pareto fronts in average-cost multi-objective MDPs, extending prior work on discounted settings.
Method for verifying operational design domain coverage for safety-critical AI systems in aviation, addressing EASA certification requirements.
ASK framework combines smaller language models with RL policies to enhance out-of-distribution generalization by gating LM assistance based on agent uncertainty.
Extended publication on learning state machines from data streams with PAC-bounds analysis and heuristic improvements.
Bayesian vertical federated learning approach for multimodal survival prediction with privacy preservation and uncertainty quantification.
De Jure is an automated pipeline for extracting structured regulatory rules from legal documents using LLM self-refinement, requiring no human annotation or domain-specific data.
Analysis of token initialization strategies for new vocabulary in language models used for generative recommendation systems.
Theoretical analysis showing transformers can solve non-linear non-Markovian stochastic filtering for conditionally Gaussian signals.
Moonwalk: inverse-forward differentiation technique addressing backpropagation's memory requirements for training deeper networks.
Comprehensive benchmark of 17 graph pooling methods across 28 datasets evaluating effectiveness, robustness and generalization.
Study of in-context learning in LLMs with spurious correlations, examining transformer robustness to spurious features in classification.
Automated neural architecture selection for time series forecasting comparing LSTMs, GRUs, Transformers, and State-Space Models.
ML research on classification metrics accounting for confidence in incorrect predictions for safety-critical applications.
Scientific ML research on training neural differential-algebraic equation systems extending neural ODE methods.
GradPower: lightweight gradient transformation technique for accelerating LLM pre-training via sign-power transformation, requiring minimal code changes.
ML research introducing Reliable Policy Iteration variants that maintain theoretical guarantees under function approximation in reinforcement learning.