Language Models as Messengers: Enhancing Message Passing in Heterophilic Graph Learning
Method using language models to improve message passing in heterophilic graph neural networks by leveraging semantic node text.
Method using language models to improve message passing in heterophilic graph neural networks by leveraging semantic node text.
MLE-Live framework for evaluating LLM agents in ML engineering that engage with research communities through knowledge sharing and communication.
Position paper examining opportunities and limitations of integrating LLMs into agent-based social simulations from computational social science perspective.
Multi-agent system for clinical diagnosis that accumulates self-learned clinical knowledge across agent interactions for improved LLM performance.
Analysis of failure modes in multi-agent workflows built on low-code orchestration platforms, examining propagation across heterogeneous nodes.
RE-PO framework for robust LLM alignment that handles noisy preference data and unreliable annotations in RLHF-style training.
MITS algorithm using pointwise mutual information to improve tree search reasoning in LLMs with better step quality assessment.
Method for reducing hallucinations in multimodal LLMs by reallocating attention across layers to balance perception and reasoning.
AutoSpec framework for automatically refining logical specifications in reinforcement learning through exploration-guided search strategies.
Agentic framework orchestrating specialized tools for automated radiology reporting, combining vision-language models with multi-step reasoning.
Research on whether LLMs can mediate online conflicts by fostering empathy and constructive dialogue beyond content moderation.
Evaluation of visual UI design factors influencing web agent decision-making and task performance.
Real-time alignment technique for RLHF reward models to prevent overoptimization and maintain human intent capture.
Study on whether Large Reasoning Models know when to stop thinking, addressing redundancy in long chains-of-thought.
Training method for Large Reasoning Models using adaptive reflection and length penalties to reduce unnecessary token consumption.
ForesightSafety Bench evaluates frontier risks in autonomous AI with unpredictable and difficult-to-control behaviors.
IntentCUA framework for computer-use agents with intent-aligned planning and multi-agent coordination over long horizons.
Layered execution structures for tool orchestration in agentic systems with reflective error correction mechanisms.
Aletheia AI agent solved 6/10 FirstProof mathematics challenges autonomously using Gemini 3 Deep Think reasoning.
Framework for causal embeddings enabling multiple detailed models to map into sub-systems of coarser causal models.
Examines AI agents with persistent state, tool access, and skills for autonomous execution of social science research pipelines.
ConstraintBench evaluates whether LLMs can directly solve constrained optimization problems without solver access.
Study on LLM vulnerability to jailbreak attacks using classical Chinese prompts to bypass safety constraints.
Dispatcher/Executor principle for multi-task reinforcement learning using abstraction to improve generalization across tasks.
R2GenCSR uses LLMs with visual feature extraction from X-ray images for automated radiology report generation.
Research framework for sparse counterfactual explanations using optimal transport and Shapley values for model interpretability.
Research on robust watermarking techniques for distinguishing generated from real content in generative models.
Research paper on grounding LLMs with real-time financial data for knowledge-aware financial agent applications.
Semantic parallelism technique for efficient MoE LLM inference via model-data co-scheduling reducing communication bottlenecks.
Optimization perspective on reward model quality in RLHF showing accuracy alone doesn't capture effective teacher properties.
Domain decomposition approach for neural operators to improve geometry generalization and transferability in PDE solving.
LLM-empowered hierarchical RIC controller for O-RAN addressing cooperation, computational demands, and domain-specific adaptation.
FineScope framework using SAE-guided data selection for domain-specific LLM pruning and finetuning with maintained performance.
Feature selection method using permutation-invariant embeddings and policy-guided search for complex feature interactions.
Agentic Predictor using multi-view encoders for performance prediction in LLM-based agentic workflows without exhaustive evaluation.
Method bridging target-free and target-based deep reinforcement learning to reduce memory requirements and improve update propagation.
Framework converting generative multimodal LLMs into zero-shot discriminative embedding models without extensive pre-training.
OM2P offline multi-agent reinforcement learning using flow-based generative models with improved sampling efficiency.
Framework for improving LLM context-aided forecasting with diagnostic tools and reduced computational costs for practical deployment.
AC3 reinforcement learning framework for long-horizon robotic manipulation using continuous action chunking with sparse rewards.
LumiMAS framework for real-time monitoring and observability of multi-agent systems with LLMs, addressing system-wide failure detection.
SAT reduction approach for automating input/output logics, a family of deontic logics for reasoning over norms and obligations.
Latent Self-Consistency: method for reliable majority voting in LLM outputs handling both short and long-form reasoning tasks consistently.
Once4All: LLM-synthesized test generator for SMT solver fuzzing using skeleton guidance to uncover bugs in evolving solver versions.
Veritas: pattern-aware deepfake detection system with HydraFake dataset bridging gap between academic benchmarks and industrial deployment.
Draw-In-Mind: multimodal model rebalancing designer and painter roles to improve precision in image editing tasks.
LLaDA diffusion-based large language model applied to automatic speech recognition with deliberation-based post-processing for Whisper transcripts.
E-CIT: plug-and-play ensemble framework for conditional independence testing to reduce computational bottlenecks in constraint-based causal discovery.
Investigation of in-context learning emergence in world models for environmental dynamics prediction beyond static zero-shot performance.
Study of activation function design's role in preventing plasticity loss during continual learning, beyond catastrophic forgetting.