Reasoning Models Don't Just Think Longer, They Move Differently
Study of hidden-state trajectories in reasoning-trained LLMs showing they follow different internal paths during chain-of-thought.
Study of hidden-state trajectories in reasoning-trained LLMs showing they follow different internal paths during chain-of-thought.
Analysis of when sparse Mixture-of-Experts routing benefits vision models, identifying compute-leverage patterns.
X-SYNTH system for AI agents to synthesize enterprise context from human attention patterns rather than retrieval.
Theoretical analysis proving RoPE positional embeddings lose effectiveness in long-context LLMs.
Distributional process reward model predicting step-level success probability and reliability for reasoning tasks.
Non-autoregressive text generation model using draft-conditioned latent refinement with flow networks.
Multi-agent orchestration framework dynamically switching between parallel and sequential LLM agent collaboration.
Calibration method for LLMs using semantic-level rewards to improve uncertainty estimation in high-stakes tasks.
Study evaluating LLM code generation transfer to unseen programming languages through fine-tuning analysis.
Prompting technique for vision-language models addressing base-new class trade-off in transfer learning.
Offline policy learning framework for contextual bandits optimizing under general risk criteria.
Few-shot LLM approach for triaging online patient inquiries into clinical action categories with minimal labeled data.
Statistical framework improving Concept Activation Vectors for interpretability in deep learning models.
Adaptive autoscaler for container orchestration that learns cold-start duration using EWMA estimation.
Preference optimization framework for flow models addressing intra-group variance decay in RL alignment of generative models.
ML framework for HPC performance prediction using merged execution traces to overcome hardware counter limitations.
Multi-layer cloud intrusion detection system combining LLMs with Q-learning for improved performance in real deployments.
Tool for detecting and interpreting domain shifts in high-dimensional data by identifying anomalies in feature subspaces.
Causal analysis of format inconsistencies in LLM-as-judge scoring using PEAP to investigate internal mechanisms.
RecMem: memory consolidation system for long-running LLM agents that reduces token consumption through lazy consolidation.
Uses property-guided synthesis to reduce LLM inference costs in program synthesis for planning, guiding generation with formal properties.
Asteria: runtime system for scalable LLM training using second-order optimization, decoupling preconditioner state from GPU path.
Combines formal methods with ML for auditing and monitoring AI systems across development lifecycle, enabling compliance checks.
Controlled study of compound LLM agent design trade-offs in adversarial environments, analyzing context, reasoning, and task decomposition.
Technique to characterize language model organization by lesioning parameters, inspired by neuroscience aphasia studies.
arXiv: Framework combining generative AI, smart metering, quantum optimization for energy infrastructure and billing.
arXiv: FORGE—population-based protocol for self-evolving LLM agent memory without gradient updates. ReAct agents with Reflexion.
arXiv: Studies how LLM-mediated communication influences collective opinion formation on social platforms.
arXiv: Data reconstruction attacks against federated learning systems. Privacy and security analysis of FL protocols.
arXiv: Convergence analysis of federated Q-learning in heterogeneous multi-agent environments. Theory for distributed RL.
arXiv: Explores Kolmogorov Superposition Theorem alternatives for neural network design beyond Kolmogorov-Arnold Networks.
arXiv: Tube Loss function for prediction interval estimation in regression. Novel loss function with theoretical guarantees.
arXiv: Semi-supervised learning approach for sparse reward shaping in reinforcement learning. Uses SSL and data augmentation.
GPU scheduling algorithm for LLM inference that manages KV cache memory constraints and minimizes latency under $700k daily inference costs.
Comparison of LLM routing strategies showing simple k-NN outperforms complex learned routers for selecting specialized models.
Comprehensive benchmark for learning surrogate models of stochastic PDEs with complex spatio-temporal dynamics.
Active learning approach for LLM alignment that selectively samples preference annotations, reducing cost while maintaining alignment quality.
Proposes gradient-free neural network training via projection operators and feasibility-seeking, alternative to conventional loss minimization.
Extends double Q-learning to deep RL, improving target bootstrap decoupling in value function estimation.
Combines supervised and reinforcement fine-tuning for LLMs using prefix sampling, addressing trade-offs between behavior cloning and performance gains.
Trains LMs via RL to reason about uncertainty rather than just correctness, improving performance on question answering by penalizing low-confidence outputs.
Applies policy gradient with decision transformers for reliable path planning in stochastic transportation networks with uncertain travel times.
Uses LLM reasoning to improve decision tree induction for tabular data, balancing interpretability with performance over black-box foundation models.
Examines data leakage issues in ML models for bearing fault diagnosis, highlighting methodological flaws preventing real-world generalization in condition monitoring.
Method to train small open-weight advisor models that generate dynamic prompts to improve black-box LLM performance. Demonstrates 27.4% improvement on GPT-5.2 tax tasks.
Graph-based gradient boosting approach for insurance fraud detection using inductive inference on networked claim data.
Strategy for efficiently pre-training large LLMs by orthogonally expanding Mixture-of-Experts parameters from existing checkpoints.
Post-training method enabling LLMs to improve reasoning without external rewards through self-generation and self-training.
LLM-based semantic optimization approach for expensive black-box problems incorporating domain knowledge and heuristics.
Regularization technique for neural network pruning that improves robustness under aggressive sparsity levels.