OrbitFlow: SLO-Aware Long-Context LLM Serving with Fine-Grained KV Cache Reconfiguration
OrbitFlow system dynamically reconfigures KV cache across CPU/GPU for long-context LLM serving with adaptive runtime memory management.
OrbitFlow system dynamically reconfigures KV cache across CPU/GPU for long-context LLM serving with adaptive runtime memory management.
MAS-Orchestra framework for orchestrating multi-agent systems with holistic reasoning and controlled benchmarks for evaluating agent coordination.
MSP-LLM unified framework for material synthesis planning using LLMs, addressing precursor identification and operation sequencing for materials discovery.
SpotAgent grounds vision-language models in geo-localization through agentic reasoning with grounded tool use for sparse and ambiguous visual cues.
Global Hypothesis Space method improves Chain-of-Thought reasoning by enabling holistic exploration of solution space rather than greedy step-by-step generation.
Studies whether Large Reasoning Models transfer step-by-step inference benefits to Theory of Mind tasks assessing belief and intention inference.
Multi-dimensional alignment framework for medical LLMs combining RLHF and RLVR paradigms for high-stakes question answering accuracy.
REMem introduces episodic memory and spatiotemporal reasoning to language agents, enabling recollection and reasoning over interaction histories.
Arbor framework decomposes decision tree navigation for LLMs in high-stakes workflows like healthcare triage, preventing instruction-following degradation.
CoreCraft RL environment simulates enterprise customer support organization for training generalizable agentic AI with 2500+ entities and 23 tools.
Phase-Aware Mixture of Experts architecture for RL-trained LLM agents, addressing simplicity bias in policy networks for complex task solving.
Interprets LLM softmax layer as Energy-Based Model to track 'energy spills' correlating with factual errors and biases during decoding.
MagicAgent framework enables generalized agent planning across heterogeneous tasks using LLMs, addressing data scarcity and task conflicts.
ProactiveMobile benchmark evaluates proactive intelligence in mobile AI agents that anticipate user needs and autonomously initiate actions.
SideQuest framework optimizes KV cache management for long-horizon agentic reasoning tasks requiring multi-hop reasoning over retrieved documents.
OmniGAIA benchmark for evaluating multi-modal AI agents combining vision, audio, and language with reasoning and tool usage capabilities.
Agent architecture with recurrent persistence loop and affect proxy for machine consciousness research. Theoretical work on inspectable AI agents.
Formal analysis demonstrating optimization-based systems and LLMs trained via RLHF cannot be norm-responsive agents without architectural changes.
First Turing test for speech-to-speech conversational systems, evaluating 9 state-of-the-art systems against human participants.
Study on modifying training data distribution to reduce simplicity bias and improve in-distribution generalization for neural networks.
GLEE framework and benchmark for evaluating LLM behavior in economic and strategic interactions requiring natural language communication.
Survey of deep reinforcement learning approaches for network intrusion detection systems reviewing DRL frameworks and recent applications.
FSW-GNN graph neural network architecture achieving Weisfeiler-Leman equivalence with improved separation quality using bi-Lipschitz constraints.
Neuro-symbolic architecture for discovering high-level action symbols from low-level skill demonstrations using neural networks and symbolic planning.
SimpleToM benchmark exposing gap between LLMs' explicit theory-of-mind reasoning and implicit application in predicting human behavior.
Decision Transformer approach for offline reinforcement learning using return-conditioned supervised learning across different data distributions.
Interaction2Code benchmark for evaluating MLLMs on generating interactive webpage code from prototypes with dynamic user interactions.
Multi-PA benchmark assessing privacy risks and preservation capabilities of large vision-language models across multiple dimensions.
Study of polynomial, trigonometric, and tropical activation functions as alternatives to standard neural network activations.
WorldSense benchmark for evaluating multimodal LLMs on video understanding tasks requiring synchronized visual, audio, and text inputs.
Sparse shift autoencoders for interpretability of LLM activations, enabling unsupervised discovery and identification of internal concepts.
GradientStabilizer method addressing training instability from gradient-norm spikes in deep learning without requiring threshold tuning.
Framework analyzing LLM impact on Wikipedia content and page views with simulations exploring potential risks from automated contributions.
Meta-analysis of 92 open-source pretrained LLMs examining how architectural decisions and data curation impact model performance beyond just scale.
Research on LLaVE multimodal embedding models using hardness-weighted contrastive learning. Improves distinguishing hard negatives in vision-language tasks.
Token-efficient item representation via images for LLM-based recommender systems, balancing efficiency and effectiveness in item encoding.
Vision-R1 applies reinforcement learning to enhance reasoning capabilities in multimodal large language models, building on DeepSeek-R1-Zero.
SemHiTok proposes a unified image tokenizer using semantic-guided hierarchical codebook for multimodal understanding and generation tasks.
Open-Sora 2.0 demonstrates training a commercial-grade video generation model for $200k, addressing cost-effectiveness in large-scale model training.
ROMA: specialized hardware accelerator for QLoRA-based on-device LLM inference with hybrid read-only-memory architecture for privacy-preserving deployment.
GateLens: LLM agent framework for automotive software analytics combining reasoning with structured tabular data handling and ambiguity resolution capabilities.
AdaRank: adaptive rank selection method for SVD-based model merging that reduces cross-task interference and improves multi-task learning performance.
Analysis of compression effects (quantization, distillation, pruning) on large reasoning model capabilities with mechanistic interpretation and performance benchmarking.
Spiking neural network approach for multi-task reinforcement learning in resource-constrained autonomous agents with adaptive task-switching to reduce interference.
SwallowCode and MathShallow datasets with systematically rewritten pre-training data significantly improve LLM performance in code and math reasoning under open licenses.
Explainability framework for Wasserstein distances analyzing dataset shifts and transport phenomena by decomposing distance contributions.
Theoretical analysis connecting initialization bias and trainability in deep neural networks through mean field theory and gradient behavior characterization.
Defense mechanism against harmful fine-tuning attacks on LLMs by reducing model trainability on harmful data while maintaining safety and normal performance.
FreeKV: efficient KV cache retrieval method for long-context LLM inference that improves compression methods and reduces accuracy loss during context extension.
Analysis of expert offloading in Mixture-of-Experts LLMs on memory-constrained devices, examining local routing consistency and practical deployment challenges.