Estimating Tail Risks in Language Model Output Distributions
Framework for estimating tail risks and rare harmful outputs in LLM distributions at population scale deployment.
Framework for estimating tail risks and rare harmful outputs in LLM distributions at population scale deployment.
Dynamic routing technique for offline reinforcement learning that balances Q-value improvement with dataset support constraints.
Research on how LLMs detect and correct their own errors using second-order confidence signals from decision neuroscience.
FETS Benchmark: Foundation models outperform dataset-specific approaches for generalizable energy time series forecasting.
Japanese medical foundation model balancing scaling laws and task-specific efficiency for clinical risk prediction on longitudinal data.
SOC-ICNN: Input convex neural network architecture generalizing from linear programming to second-order cone programming.
Extension of neural activation coverage technique for uncertainty estimation in regression tasks of pre-trained neural networks.
Hidden failure modes of gradient modification under Adam optimizer in continual learning with adaptive decoupled moment routing solution.
Analysis of distance-misaligned training failure modes in graph transformers with adaptive control mechanism.
HubRouter: Sub-quadratic routing module replacing O(n²) attention with O(nM) hub-mediated routing for hybrid sequence models.
FeatEHR-LLM: Framework using LLMs for automated feature engineering in electronic health records with irregular temporal patterns.
SOLAR-RL: Semi-online reinforcement learning framework for training multimodal LLM-based GUI agents on complex navigation tasks.
Data-free contribution estimation in federated learning using gradient von Neumann entropy to identify client importance without privacy leakage.
SpikingBrain2.0: 5B brain-inspired foundation model using spiking neural networks for efficient long-context inference with reduced training overhead.
Proposes adaptive head budgeting mechanism for multi-head attention in Transformers to improve efficiency by selectively activating attention heads based on task requirements.
Self-supervised learning approach using action-conditioned world models for cardiac disease detection, shifting from invariance-based to dynamics-aware objectives.
Research evaluating Shapley value variants for explainable AI in high-stakes applications, comparing theoretical formulations and human utility alignment.
WG-SRC white-box probe diagnosing graph neural network mechanisms and dataset feature requirements for node classification.
Transformer model extracting historical lexical structure and cognates from modern Bantu language morphological data.
Active experiment selection method for budget-aware scaling law fitting, reducing costs of planning large-scale ML training runs.
Evaluation framework testing whether LLMs perform genuine mathematical reasoning vs pattern matching on novel abstract mathematical problems.
Code generation pipeline study showing execution feedback outweighs topology complexity in 1-3B model composition on HumanEval.
MambaCSP: Hybrid state-space model combining attention and Mamba for hardware-efficient channel prediction with subquadratic scaling.
RE-CONFIRM framework evaluating robustness of brain biomarkers extracted by foundation models from dynamic functional connectivity data.
LLM prompt sensitivity investigation comparing instruction-based vs example-based prompting through shared lexical task representations.
Lightweight RAG and LLM approach for clinical patient-trial matching over heterogeneous EHR data with improved scalability.
LoRA adapter placement study in hybrid language models combining attention and recurrent components (Qwen, Falcon architectures).
Sovereign Agentic Loops: Control-plane architecture decoupling LLM reasoning from execution to improve safety in API-calling agents by emitting structured intents instead of direct outputs.
ArXiv paper on unsupervised anomaly detection for retinal OCT imaging without labeled data. Medical imaging ML research.
ArXiv paper on multi-armed bandits optimizing statistical utility functionals via influence-function gradients. Theoretical ML research.
ArXiv paper on adaptive control for constrained linear quadratic regulator achieving optimal regret bounds. Control theory and learning research.
ArXiv paper using multimodal diffusion models to enhance polarized light and EBSD microscopy data. Domain-specific ML application.
ArXiv paper on generative AI for marine propeller design using physics-based data generation. Applied ML research for engineering design.
ArXiv paper analyzing benchmark hacking in ML contests—gaming evaluation metrics without genuine improvement. Research on evaluation methodologies.
ArXiv paper on learning-augmented robotic control for manufacturing with real-world production validation. Applied ML research for industrial automation.
ArXiv paper on algorithmic feature highlighting for human-AI decision-making systems. Research on interpretable AI agent design.
ArXiv paper on LUT-aware neural network training for FPGA inference with ultra-low latency. Hardware-efficient ML architecture research.
ArXiv paper on pliable rejection sampling with kernel estimators for sampling difficult distributions. Machine learning research with new approach.
ArXiv paper on adaptive dictionary learning for kernel ridge regression to reduce space complexity. Machine learning research addressing scalability.
ArXiv paper on Conformalized Super Learner ensemble method with interval prediction uncertainty quantification. Machine learning research with theoretical contributions.
Research formalizing hidden randomness in LLMs via background temperature concept, analyzing nondeterminism from batch variation, kernel issues, and floating-point arithmetic.
Empirical evaluation of collective intelligence in large-scale LLM agent societies using MoltBook platform with 2M+ agents, testing emergence of intelligence at scale.
SnapLog approach extracts event data from video streams using image embeddings for process mining and business process management applications.
Research on neuron labeling for interpretability using contrastive examples to produce faithful textual descriptions of deep network internal units.
Research paper on federated learning frameworks for SPDnet models operating on symmetric positive definite matrices with geometry-preserving aggregation strategies.
Studies nonrobust predictive features in deep medical imaging models. Shows networks learn adversarially vulnerable patterns useful in-distribution.
Variational autoencoder for long-term customer revenue forecasting from sparse transaction data. Combines probabilistic and ML approaches.
Framework aligning dense retrievers with LLM utility via distillation for improved RAG performance. Balances precision and computational cost.
Training method for neural network surrogate models that improves their embedding in optimization problems. Enhances tractability of MILP formulations.
Studies privacy side-channels in ML models via output label space and proposes differentially private continual learning defenses.