QuasiMoTTo: Quasi-Monte Carlo Test-Time Scaling
Quasi-Monte Carlo test-time scaling method for language models reducing redundancy in parallel sampling while improving inference efficiency.
Quasi-Monte Carlo test-time scaling method for language models reducing redundancy in parallel sampling while improving inference efficiency.
RLVR framework combining verifiable rewards and human demonstrations for LM training, addressing diversity collapse from objective-only optimization.
Neural Certificate Pricing applies unsupervised learning to combinatorial optimization by leveraging asymmetry between search and verification complexity.
Empirical comparison of quantum machine learning models versus classical machine learning approaches across benchmarks.
TiRex-2 extends univariate time series foundation model to multivariate forecasting using recurrent xLSTM with streaming capability.
Language-critique framework for imitation learning from suboptimal demonstrations using natural language feedback instead of scalar signals.
Study showing single transformer layer training matches full-parameter RL fine-tuning for LLMs, revealing unequal layer-wise contribution during RL post-training.
Identifies vocabulary gap in modern encoders for sparse retrieval and proposes approach to bridge gap between dense and sparse retrieval.
Framework using steering vectors and latent space analysis to control and calibrate language model behavior for trustworthy deployment.
Black-box attack recovering private vision-tokenizer configurations of vision-language models through side-channel analysis.
SLIM-RL proposes risk-budgeted random-masking reinforcement learning for diffusion LLMs, eliminating trajectory slicing overhead from prior TraceRL method.
Analysis of sample complexity for estimating watermark proportions in documents under Gumbel-max LLM watermarking mechanism.
Computer vision neural networks for radioisotope identification from gamma-ray spectrograms in urban environments.
Rosetta: Composable multimodal pretraining approach addressing gradient conflicts when integrating new modalities without catastrophic forgetting.
DiscoLoop: Method for internalizing multi-hop reasoning in LLMs within single forward pass using discrete embeddings and continuous states.
Systematization of knowledge on attack and defense landscape for mobile on-device AI systems.
Structured evaluation showing text-to-image diffusion safety alignment methods create illusion of high utility through coarse metrics.
Mechanistic investigation of authority bias in LLMs showing how models prioritize source credibility over factual consistency.
Information-regularized attention mechanism to improve visual grounding and reduce hallucination in vision-language models.
StochasT: Visual instruction tuning method addressing visual attention decay in multi-turn vision-language model conversations.
Deep neural networks and ensemble methods to predict mortality outcomes and identify biomarkers in acute myocardial infarction.
Causal auditing framework to detect whether deleted facts persist in limited memory language models through parametric memory or retrieval artifacts.
RL framework enabling interactive real-time control of agent behavior during gameplay through coachability mechanisms instead of learning single optimal policy.
Lightweight backbone-agnostic segmentation benchmark adapter enabling fair comparison of transformer backbones independent of decoder and pretraining.
Recursive Vision Transformer approach using soft mixture-of-recursions to build deeper models with better parameter efficiency and performance.
Method for explaining financial LLM decisions using Shapley values combined with domain expertise, addressing regulatory explainability requirements.
Graph-native RL system for materials discovery generating scientifically valid hypotheses through multi-step reasoning with traceable intermediate steps.
Multitask learning framework handling mixed-type outcomes with shared sparsity by unifying task-specific losses through transformation.
Method to identify attention heads in LLMs that synthesize answers from context meaning rather than literal copying, improving long-context model interpretability.
Research on message passing between LLM threads for efficient parallel reasoning, reducing computational cost of long chains-of-thought.
GRINCO uses group-invariant coresets for active learning that respect data symmetries and transformation groups.
FAR enables robots to learn from failures at test time, adapt behavior, and improve policy without human intervention.
Cartridge distillation method exposes hidden biases in LLMs that favor specific entities or viewpoints.
Compares PPO and SAC reinforcement learning algorithms for fault tolerance in autonomous machines.
Invariance Pair Guidance improves robustness to spurious correlations through corrective gradients without dense labels.
Studies inherent many-to-many multiplicity in multimodal learning relationships beyond deterministic alignment.
FusionFactory fuses capabilities of multiple LLMs using multi-LLM log data for improved performance.
FLAT reveals hidden backdoor failures in federated learning through latent-conditioned reliability stress testing.
FedIA improves federated graph learning through importance-aware aggregation on distributed social media networks.
rBridge predicts reasoning performance of large LLMs using small proxy models under 1B parameters.
K-Merge enables online merging of LoRA adapters for efficient on-device LLM deployment with limited storage.
Studies computable PAC learning and derives analogs of the Fundamental Theorem of Statistical Learning in the computable setting.
FlowPath: invertible flow-based method for learning manifolds from irregularly-sampled time series, improving neural controlled differential equations robustness.
Analysis showing 8-bit quantization unexpectedly improves continual learning in LLMs compared to FP16, reducing catastrophic forgetting with replay buffers.
Study using AlphaEarth foundation model embeddings from satellite imagery to improve hydrological river flow prediction in data-sparse regions.
OpFML pipeline for operationalizing ML-based climate and Earth science models with data acquisition, preprocessing, and failure handling infrastructure.
KAGE-Bench: JAX-native benchmark for systematically evaluating RL agent visual generalization by independently controlling observation distribution shifts.
PaAno: lightweight patch-based representation learning for time-series anomaly detection, outperforming large transformer models on constrained hardware.
Attentive kernel smoothing approach for efficient Neural Controlled Differential Equations, reducing function evaluations via smoother path construction.
Unified framework for geometry-preserving neural architectures on manifolds with boundary, organizing constraint enforcement strategies.