Predict-then-Diffuse: Adaptive Response Length for Compute-Budgeted Inference in Diffusion LLMs
Predict-then-Diffuse method for diffusion LLMs enabling adaptive response length under compute budget constraints with parallel generation.
Predict-then-Diffuse method for diffusion LLMs enabling adaptive response length under compute budget constraints with parallel generation.
DASE stopping heuristic for LLM ensembles enabling early commitment on consensus and adaptive deliberation with calibrated signals.
Universal Semi-Supervised Learning framework addressing scarce labeled data and unknown unlabeled distributions using structural inference.
PIQL framework integrating privileged information to accelerate learning and improve generalization in tabular foundation models.
Non-monotonic latency behavior in Apple MPS transformer decoding with KV cache interactions, identifying 21x latency spikes.
Multi-task bilevel optimization extending bilevel learning beyond single-task settings with equality constraints for complex ML problems.
RubricRefine improves tool-use agent reliability through training-free pre-execution refinement using rubric-based feedback for code generation.
Full-pipeline FP4 quantization training for large language models on native FP4 hardware, studying MXFP4 in transformer pretraining.
Tree-of-Thought reasoning acceleration for LLMs via speculative exploration, addressing reward dependency bottleneck in complex task solving.
Thompson sampling algorithm for offline-to-online learning addressing distribution shift between offline data and online environments.
Theoretical framework of entropy mechanics in LLM reinforcement learning with verifiable rewards analyzing token-level policy updates.
GEAR: granularity-adaptive advantage reweighting for LLM agents using self-distillation for fine-grained credit assignment.
Supervised fine-tuning analysis for procedural-skill learning across Qwen3.5 model scales (0.8B-4B).
Sparse-to-dense reward principle for LLM post-training combining GRPO sparse rewards with dense token-level distillation.
Continual learning method combining parameter updates and in-context learning for LLMs to adapt without catastrophic forgetting.
WriteSAE: sparse autoencoder decomposing recurrent language model cache writes for interpretability.
ToolMol: evolutionary agentic framework using LLMs with molecular tools for multi-objective drug discovery.
Benchmarking agentic AI for neuroscience data reuse and format standardization across fragmented experimental datasets.
TIDE: graph neural network out-of-distribution detection via information decomposition for robust node classification.
MLGIB: Graph Neural Network method addressing over-squashing in multi-label graphs via information bottleneck.
EMO: progressive training method for Mixture-of-Experts models addressing efficiency paradox in sparse MoE scaling.
Speculative Interaction Agents: framework for real-time LLM agents using asynchronous I/O and speculative tool calling under 1-second latency.
Novel projected gradient methods for nonconvex smooth optimization with improved iteration complexity and auto-conditioned stepsizes.
OMAC framework automatically optimizes multi-agent LLM systems through cost-effectiveness analysis and collaborative patterns.
ActivePusher combines active learning with learned residual dynamics models for nonprehensile robotic manipulation.
Scalable subset selection method for linear mixed models with thousands of candidate predictors.
VER combines multiple vision foundation models through distillation and dynamic routing for flexible robotic learning tasks.
TRIM method uses token-wise attention saliency to identify important samples for efficient LLM instruction tuning with reduced data requirements.
Uses generative models as acquisition functions for batch Bayesian optimization, enabling large-scale optimization of non-continuous and high-dimensional design spaces.
Interactive physical reasoning agent learning human-like causal understanding from game interaction with visual domain gaps.
OPT-ENGINE benchmark evaluating LLM capabilities in optimization modeling across LP to MIP with controlled complexity scaling.
Agent-designed agentic workflows via reinforced canvas editing with graph-level feedback and in-loop error repair for complex multi-step tasks.
Conformal prediction framework for adaptive reasoning in LLMs, controlling risk-accuracy tradeoff when allocating compute budget for inference.
Proxy compression training scheme for language models preserving efficiency of tokenization while enabling raw-byte inference interface.
Top-W: geometry-aware decoding for LLMs using Wasserstein distance over token embeddings to balance diversity and coherence.
Framework extracting distribution maps from LLM next-token probabilities for better statistical text analysis beyond perplexity metrics.
Combines Mamba state-space models with LLMs for dynamic fMRI graph learning in autism diagnosis using multimodal reasoning.
Multi-agent LLM framework for robotic manipulation with closed-loop visual feedback, replacing specialized models with general reasoning.
Scalable reward modeling framework for robotics using trajectory comparisons to handle failed and suboptimal trajectories.
Speculative decoding optimization for LLM inference in high-concurrency serving using elastic sparse gating and dynamic trees.
Scalable framework for training and evaluating agents in claw-style environments with file systems, tools, and persistent workspace.
Evaluation study of frontier LLM metacognitive failures under adversarial pressure, examining cognitive collapse in high-stakes scenarios.
Large-scale benchmark for evaluating AI agents on workspace tasks with file dependencies, testing real-world file system operations.
Decentralized framework organizing coding agents in co-evolving populations for algorithmic discovery and skill evolution.
Memory-efficient continual learning framework for malware detection that avoids catastrophic forgetting while adapting to new threats.
Diagnostic framework for predicting multi-agent LLM system behavior across different communication topologies using successor representation.
Training-free method to improve Vision-Language-Action models' handling of temporal dynamics in non-stationary scenarios.
Analysis of watermarking as monitoring primitive for generative models, examining internal attribution and safety monitoring.
Test-time self-training method for LLMs that updates parameters at inference to adapt to specific queries and correct misconceptions.
Proposes ledger-based system extending Git to coordinate humans, AI agents, and automation in software repositories.