FlashPrefill: Instantaneous Pattern Discovery and Thresholding for Ultra-Fast Long-Context Prefilling
FlashPrefill: ultra-fast long-context prefilling framework using instantaneous pattern discovery and thresholding for sparse attention.
FlashPrefill: ultra-fast long-context prefilling framework using instantaneous pattern discovery and thresholding for sparse attention.
CoE: training-free multimodal summarization via chain-of-events without domain-specific supervision for video, transcript, and image fusion.
GazeMoE: Mixture-of-Experts architecture for gaze target perception combining multi-modal cues from vision foundation models.
Learning-based two-stage optimization method DeCoST for solving orienteering problem with time windows and variable profits.
HiPP-Prune: hierarchical preference-conditioned structured pruning framework for efficient vision-language model deployment.
Agentic retrieval-augmented reasoning pipelines for clinical radiology QA, analyzing reliability under model variability.
Neural implementation of fuzzy cognitive maps using Langevin dynamics for learning causality patterns.
STEM: sparse attention mechanism rethinking causal information flow to reduce quadratic complexity of self-attention in long-context LLMs.
Physics-informed neural networks using Gaussian Mixture Model adaptive sampling to improve modeling of stiff PDEs.
DEX-AR framework for explaining decision-making in autoregressive vision-language models through dynamic explainability methods.
Method to post-train LLMs to efficiently infer calibrated uncertainty estimates using entropy-based approaches in three-stage pipeline.
Field study comparing contextual bandits versus LLM architectures for personalized health behavior intervention message selection.
K-MaT prompt-learning framework transferring vision-language models from high-end to low-end medical imaging modalities.
MoEless system for efficient serving of Mixture-of-Experts LLMs using serverless computing to address sparse activation overhead.
Dynamic Chunking Diffusion Transformer allocating variable compute to image regions based on information density and denoising timestep.
ESAA architecture for event-sourced security audits of AI-generated code using agent-assisted verification with immutable audit trails.
Prompt sensitivity reduction approach for SAM3 medical image segmentation using group-wise consistency training.
Method for video generation that incorporates physical simulation constraints to prevent violations of gravity, inertia, and collision.
Reference architecture analysis of 18 reinforcement learning frameworks, proposing standardized architectural patterns for comparison and integration.
Study comparing LLM performance on abductive reasoning tasks involving syllogistic forms, examining shared biases between humans and models.
Prosodic boundary-aware post-training strategy for streaming TTS with streaming text input.
Analysis showing vision-language models encode geometry in frozen features better than text outputs.
PONTE system for personalized natural language explanations from LLMs addressing faithfulness and hallucinations.
NOBLE architecture augmentation adding nonlinear low-rank branches to transformer layers for efficient pretraining.
COLD-Steer framework for training-free steering of LLM activations using one-step learning dynamics.
Vision-language model benchmark for surgical reasoning and interpretation from surgical video.
Method for distilling semantic knowledge from LLMs into bird's-eye view representations for autonomous driving.
Philosophical analysis of LLMs' ontological status and characterization as agents.
Position paper critiquing anthropomorphization of intermediate token generation as reasoning traces in LLMs.
Activation steering technique to reduce reasoning biases in LLMs caused by content plausibility.
VisioMath benchmark for evaluating multimodal models' figure-based mathematical reasoning and fine-grained visual comparison.
Benchmark assessing moral competence in LLMs across multiple dimensions beyond verdict prediction.
ContextBench benchmark for generating inputs that activate specific latent features and behaviors in language models.
Method for improving safety compliance in frozen LLMs using adaptive system prompts without costly fine-tuning.
Multi-agent system using multimodal LLMs for automated chemical information extraction from scientific literature.
Post-training pipeline using knowledge graphs as implicit reward models for compositional multi-hop reasoning in specialized domains.
L-ICL method for localizing and correcting constraint violations in LLM-generated plans using iterative in-context learning demonstrations.
Survey of uncertainty quantification in LLM agents with framework for safety guardrails in interactive multi-step deployments.
Framework for explainability in agentic AI systems focusing on multi-step trajectories and decision sequences rather than single predictions.
MERIT feedback framework and AgoraBench for improving LLM bargaining agents with utility feedback across nine negotiation scenarios.
Literature review examining subjectivity in data annotation and flaws in ground truth paradigm across ML research venues.
Analysis of 43 AI agent benchmarks and 72,342 tasks measuring alignment between agent development efforts and real-world labor distribution.
Multimodal mixture-of-experts with retrieval augmentation for protein active site identification at residue level.
MOOSEnger tool-enabled AI agent for MOOSE multiphysics simulation using RAG and natural language to generate HIT input files.
SEA-TS framework using self-evolving AI agents to autonomously generate, validate, and optimize time series forecasting code.
RAG-Driver framework using retrieval-augmented in-context learning in multimodal LLMs for explainable autonomous driving decisions.
Method for detecting visual hallucinations in VLM outputs on cartoon images using pose information.
Study of LLM-based pricing agents reaching supracompetitive outcomes in oligopoly settings with analysis of prompt influence on behavior.
SpecFuse method for ensembling multiple LLMs via next-segment prediction to overcome individual model limitations and improve long-range semantic collaboration.
Survey of LLM applications across scientific discovery, literature search, idea generation, experimentation, content creation, and multimodal artifact generation.