SiameseNorm: Breaking the Barrier to Reconciling Pre/Post-Norm
SiameseNorm architecture reconciles pre-norm and post-norm trade-offs in Transformers for improved stability and capacity.
SiameseNorm architecture reconciles pre-norm and post-norm trade-offs in Transformers for improved stability and capacity.
CoFEH: LLM-driven automated feature engineering with Bayesian hyperparameter optimization for flexible ML pipelines.
Analysis of robustness and consistency issues in RL-finetuned vision-language models under visual perturbations and hallucinations.
Theseus: training-free method to transfer task-specific parameter updates across models with different architectures and widths.
Method to discover and interpret implicit alignment objectives in LLMs without pre-defined rubrics, addressing reward hacking risks.
Discrete Stochastic Localization framework for non-autoregressive sequence generation using continuous-state diffusion.
MapTab benchmark evaluates multimodal LLMs on multi-criteria route planning reasoning tasks in heterogeneous graphs.
arXiv paper on interpreting and steering state-space models (Mamba) using activation subspace bottlenecks and mechanistic interpretability.
arXiv paper on InnerQ: hardware-aware quantization of KV cache for efficient LLM text generation with reduced memory footprint.
arXiv paper on Heterogeneous Agent Collaborative RL where agents share verified rollouts during training for mutual improvement.
arXiv paper on adaptive subgraph denoising for zero-shot graph learning using LLMs as predictors, improving cross-modal alignment.
arXiv reproduction paper of FairDICE algorithm for fair multi-objective offline reinforcement learning from demonstrations.
arXiv paper on MDM-Prime-v2, scaling masked diffusion language models via binary encoding and index shuffling for improved generalization.
arXiv paper on MemReward: graph-based experience memory for LLM reward prediction with limited labels in reinforcement learning scenarios.
Claude Opus 4.6 with Rocq proof assistant tools via Model Context Protocol autonomously solved 10/12 Putnam 2025 competition problems.
arXiv paper on Bayesian framework for compliance monitoring in rule-governed domains without labeled outcomes, handling missing observations.
arXiv paper on uncertainty-aware generative models for scientific imaging using stochastic flow matching with out-of-distribution detection.
arXiv paper using reinforcement learning to train biological constraint-respecting generative models for virtual cell simulation.
arXiv paper extending k-means clustering with Minkowski distance and feature weighting for adaptive feature extraction.
arXiv paper on target-weighted cross-validation for spatial prediction models addressing biased performance estimation.
arXiv paper introducing Robust Reasoning Benchmark testing 8 LLMs against textual perturbations on AIME math problems.
Analysis of how chain-of-thought reasoning decomposes complex tasks by breaking classification into smaller subtasks to reduce error.
Geometric probing framework for automated algorithm selection in continuous black-box optimization using landscape sampling.
Meta-learning approach for offline black-box optimization from small datasets, targeting molecular and materials discovery.
Study on multi-timescale credit assignment in PPO reinforcement learning, addressing surrogate hacking through representation-based approaches.
Research on token importance in on-policy knowledge distillation for LLMs, identifying which token positions provide useful learning signals during training.
Proposes Prototype-Grounded Concept Models grounding interpretable concepts in learned visual prototypes for verifiable concept alignment in deep learning.
Introduces SceneSelect for selective trajectory prediction across heterogeneous scenes via expert scheduling and scene classification.
Philosophical examination of true target existence assumptions in ML paradigms proposing framework for evaluation under democratic supervision.
Proposes uncertainty-aware predictive safety filters using probabilistic neural networks for enforcing constraints during deep reinforcement learning.
Addresses fairness in dataset distillation via cross-group barycenter alignment to preserve predictive signals for demographic subgroups.
Establishes mathematical correspondence between decision trees and diffusion models via Global Trajectory Score Matching optimization principle.
Applies permutation invariant Bayesian optimization to carbon capture problems, improving Gaussian Process surrogate models for unordered set inputs.
Presents EdgeRazor framework compressing LLMs via mixed-precision quantization-aware distillation for sub-4-bit deployment on resource-constrained devices.
Proposes Jordan-RoPE extending relative positional encodings using complex Jordan blocks and nilpotent operators for improved transformer attention mechanisms.
Proposes bi-objective decision tree learning for generating actionable recourse summaries enabling global auditing and bias detection across population subgroups.
Introduces Particle MCTS enabling principled parallel Monte Carlo tree search for neural network inference scaling with preserved sample efficiency.
Develops online multicalibration algorithm adapting between benign and worst-case sequences via dyadic grid refinement with optimal convergence rates.
Proposes Metis framework reformulating LLM jailbreaking as inference-time policy optimization using metacognitive POMDP approach for automated red teaming.
Addresses offline black-box optimization using diffusion estimation and proximity constraints to handle out-of-distribution extrapolation in design discovery.
Introduces Holder Policy Optimisation improving on GRPO by adaptively aggregating token-level probabilities for trajectory-level advantages in LLM training.
Proposes Discrete Stochastic Localization for non-autoregressive generation using continuous diffusion with unit-sphere embeddings, improving over masked discrete diffusion models.
Investigates task-aware layer pruning effects on model generalization. Finds pruning improves out-of-distribution accuracy while harming in-distribution performance across polynomial tasks and LLMs.
Analyzes rank-1 activation steering for controlling LLMs. Shows effectiveness variability reflects search difficulty rather than concept limits, formalizing steering as budget-constrained optimization.
Introduces Identifiable Token Correspondence to improve temporal consistency in token-based transformer world models for long-horizon visual RL.
Develops speech-to-text system optimized for medical terminology, contextual ambiguity, and specialized clinical notation in healthcare applications.
Analyzes weight drift and activation sparsity in neural network training dynamics caused by interactions between standard losses and activation functions.
Proposes General Preference Reinforcement Learning to unify online RL for verifiable tasks with preference optimization for open-ended generation in LLM alignment.
Position paper critiquing graph condensation methods and proposing solutions beyond full-dataset training.
Biologically plausible neural networks for blind source separation using local plasticity and dendritic computation.