ProactBench: Beyond What The User Asked For
ProactBench evaluates LLM conversational proactivity: ability to infer and act on implied user needs beyond explicit requests.
ProactBench evaluates LLM conversational proactivity: ability to infer and act on implied user needs beyond explicit requests.
Theoretical analysis of variance reduction in MeanFlow one-step generative modeling, addressing training instability and gradient variance.
Study of LLM multi-turn dialogue understanding, analyzing failures in context switching and topic detection across conversation turns.
Theorem-SFT approach for supervised fine-tuning that improves LLM generalization on mathematical reasoning by targeting principles over instances.
Representational Effective Theory framework for understanding LLM computation through learned macrostates from hidden-state trajectories.
Neural optimization method using differentiable optimal transport for vehicle routing problems, replacing autoregressive decoding.
Graph Neural Network framework with attention mechanism for interpretability and computational efficiency on heterogeneous graphs.
Theoretical framework embedding information causality into representation learning via query-separated computation, tangential to core interests.
arXiv paper investigating spurious correlations in agentic memory systems where retrieved context propagates erroneous reasoning in LLM decision-making.
arXiv paper proposing latent chain-of-thought reasoning compression using rule-based priors to replace multi-step reasoning with efficient one-step inference.
Benchmark dataset for multimodal knowledge graph question answering integrating LLMs with knowledge graphs to reduce hallucinations and leverage external knowledge.
arXiv paper on optimizing reusable skills for agentic LLMs through reinforcement learning with two-level credit assignment, moving beyond prompt engineering.
Framework for verifying LLM-generated PDE simulation code by detecting comprehension-generation gaps where code executes but encodes wrong physics.
Operational analysis of 504-GPU production cluster documenting hardware failures, recovery patterns, and distributed training reliability over 55 days.
EduStory framework for pedagogically-consistent multi-shot STEM video generation using pedagogical state tracking and script guidance.
LiteMedCoT-VL enables parameter-efficient medical VQA reasoning on compact models through knowledge distillation of reasoning steps.
Kinetic-optimal scheduler and moment correction method for discrete flow matching in zero-shot text-to-speech generation.
Study of how FFN sparsity patterns reshape attention computations in small Transformers across arithmetic and counting tasks.
RePO-VLA framework improves vision-language-action models for robotic manipulation by learning from recovery trajectories.
AtteConDA improves conditional diffusion models through attention-based conflict suppression for data augmentation applications.
Contrastive pruning technique for vision-language models preserves essential low-attention tokens for compositional reasoning tasks.
SWIFT enables efficient long video generation with prompt-adaptive memory management for continuous semantic transitions.
Training-free acceleration of identity-preserved image generation by transferring frozen adapters to distilled diffusion backbones.
APCD method improves LLM generation reliability by adaptively branching decoding paths to mitigate hallucinations from error accumulation.
LASSA architecture applies LLMs to autonomous fault-tolerant control of underwater vehicles in communication-constrained environments.
Spectral Transformer Neural Processes extend TNPs with frequency-aware Spectral Aggregator for time series and spatial data with strong periodicity patterns.
Swarm-attack: open-source framework where multiple LLM agents coordinate to discover safety bypasses and software vulnerabilities at minimal cost, demonstrating systemic risks.
Empirical comparison of RAG versus fine-tuning for industrial LLM-based question-answering systems, evaluating cost-accuracy trade-offs in domain-specific enterprise scenarios.
Design science framework for governing AI-assisted security operations, addressing accountability, privacy, cost, and auditability in high-risk operational decision support.
TAD framework addresses accuracy-parallelism trade-off in diffusion LLMs through temporal-aware trajectory self-distillation for faster, accurate parallel text generation.
KANMultiSign uses Kolmogorov-Arnold Networks with multi-scale supervision to generate 2D pose sequences from HamNoSys sign language notation.
CLR-voyance framework evaluates clinical LLM reasoning as sequential decision-making under partial observability using outcome-aware rubrics instead of closed-form retrieval.
Method for efficient ensemble selection from multiple AI systems using minimal model calls and human evaluation, formulated as distributional multiwinner voting.
SmartEval benchmark for evaluating LLM-generated Solidity smart contracts against specifications, with 9,000 contracts and five-dimensional rubric for code quality assessment.
DiffKT3D uses diffusion models with knowledge transfer from vision domains for voxel-wise dose prediction in radiotherapy planning, demonstrating cross-domain generalization.
Framework for dynamic DNN partitioning and offloading across edge-cloud devices, tested on real hardware rather than simulation for resource-constrained AI deployment.
MedMeta benchmark evaluates LLMs on higher-order reasoning by synthesizing medical meta-analysis conclusions from study abstracts, addressing a gap in medical LLM evaluation.
Data selection framework using multi-indicator weights for efficient instruction tuning of LLMs with task-model adaptation.
Hierarchical benchmark for medical VLMs and tool-augmented agents on 3D CT tumor analysis with stage-wise evaluation.
Red-teaming methodology for coding-agent monitors exposing attack generation gaps with attack taxonomy for coverage.
Fine-tuned Qwen VLM demonstrates spatial reasoning and mental imagery on visual puzzles like tangram and 3D rotation.
Benchmark for evolutionary LLM-based kernel search on Apple Silicon Metal compute across scientific computing tasks.
Knowledge distillation framework transferring 3D spatial reasoning from 7B to 2.29B vision-language model via CoT.
Systematic security analysis of tool-enabled AI agents in privileged cloud environments identifying attack vectors and risks.
Transformer architecture for in-context RL enabling cross-domain generalization without parameter updates via in-context learning.
Studies tool-use learning in LLMs by fine-tuning Llama 3.1 with trajectory supervision across sequential API domains.
KV-cache management optimization for static-graph LLM decoders to reduce memory overallocation and latency.
Addresses zero-shot LLM classification bias by aggregating semantic neighborhoods to recover discarded probability mass.
TIDES: state space model improvements for irregular time series by preserving physical meaning of time discretization steps.
Novel decoding strategy for LLMs that adaptively branches based on entropy to improve output quality while reducing computation.