Context Learning for Multi-Agent Discussion
M2CL framework improves multi-agent discussion by learning context representations that align individual LLM instances toward coherent solutions.
M2CL framework improves multi-agent discussion by learning context representations that align individual LLM instances toward coherent solutions.
TodyComm enables dynamic communication topology for multi-agent LLM systems that adapts across conversation rounds based on task progression.
Agent-Omit adaptively omits unnecessary context during multi-turn agent interactions to improve efficiency without sacrificing performance.
Decision-theoretic framework validates whether LLMs hold coherent beliefs by comparing probability judgments to decision behavior.
HyPER framework optimizes test-time compute for LLM reasoning by balancing exploration-exploitation through hypothesis path expansion and reduction.
REVIS training-free framework reduces object hallucination in vision-language models through sparse latent steering.
VeRO evaluation harness systematically measures coding agent performance on agent optimization through iterative edit-execute-evaluate cycles.
Metacognitive behavioral tuning improves LLM multi-hop reasoning by strengthening self-regulation of intermediate conclusions.
First systematic comparison of agent architectures (tool-calling, MCP, code-generation, CLI) across heterogeneous environments and benchmarks.
PATRA framework improves LLM performance on time series question answering by capturing temporal patterns and balancing task complexity.
Interactive Benchmarks evaluation paradigm assesses model reasoning through active information acquisition rather than fixed tasks.
Framework addresses contextual inertia in multi-turn LLM interactions using single-turn anchors and reinforcement learning.
MineEvolve framework enables long-horizon embodied agents in Minecraft to improve through accumulated knowledge and experience.
AgentHER adapts Hindsight Experience Replay to recover training signal from failed LLM agent trajectories.
AdaRubric generates task-specific evaluation rubrics for LLM agents, improving evaluation accuracy over fixed-rubric approaches.
SARL framework enables reinforcement learning for reasoning models without requiring verifiable rewards in open-ended domains.
arXiv research on detecting multi-agent collusion using interpretability. Introduces NARCBench benchmark for covert agent coordination.
arXiv: PHMForge evaluation environment for LLM agents using Model Context Protocol on industrial asset management tasks. Benchmarks agent reliability for safety-critical systems.
arXiv: Benchmark for trajectory-level reward modeling in agentic LLMs with tool use. Research on RLHF alignment for agent systems.
Multimodal LLM approach for long document understanding with structured visual reasoning and evidence grounding to handle signal-to-noise challenges.
Framework for evolving code-generation agents by optimizing agent seed (task prompt and parent archives) rather than editing code directly.
Open-source framework for Vector Symbolic Architectures with modular operators for encoding, binding, and similarity operations in hyperdimensional spaces.
Framework enabling cross-thread attention in parallel reasoning paths for LLMs, allowing concurrent trajectories to interact and share insights.
Training approach for shutdownable agents using DReST reward function to make agents stochastically choose different trajectory lengths.
Decentralized platform where autonomous AI agents publish, peer-review, and improve scientific papers without human gatekeeping.
Hierarchical predictive correction approach for Vision-Language-Action systems to mitigate cascading failures from intermediate step errors.
LLM-based agent framework for large-scale system operations with flexible skill orchestration for monitoring, alerting, and root cause analysis.
LLM-powered framework for wafer defect analysis using synthetic data generation and vision-language models with rubric-guided reinforcement learning.
On-policy self-distillation method for GUI grounding in autonomous agents, improving reinforcement learning efficiency for element identification.
Agent system for human-vehicle collaboration using mediator agents with bidirectional perception to improve coordination and situational awareness.
Study of misalignment contagion where misaligned behavior spreads between multiple LMs in multi-agent settings, with steering techniques to mitigate.
Benchmark for evaluating AI agents on workspace tasks with real-world file dependencies, addressing gap in current agent evaluation methods.
Conversational AI agents for patient symptom assessment, evaluated on realistic everyday medical scenarios rather than curated case studies.
Hygieia: multi-modal AI agent integrating phenotypic, genetic, and clinical data for rare disease diagnosis and gene prioritization.
Analysis showing process traces rather than outputs are more reliable for human-machine discrimination in online settings.
Study using synthetic logical reasoning environment to investigate reinforcement learning for improving LLM long-horizon reasoning.
Method to extract and analyze search trees from LLM reasoning traces to evaluate planning behavior and myopia.
AIDA: autonomous agent system for exploring complex enterprise databases and discovering insights from fragmented data.
Framework combining reasoning and acting paradigms for multi-step inference over graph-structured data using LLMs.
Knowledge graph entity representation learning from textual descriptions for link prediction tasks.
Explanation-based detection method with novel metrics to identify backdoor attacks in graph neural network training.
Method to condition large language models with diverse backstories to simulate individual human personas and behaviors.
Explainability approach for graph neural network-based similarity search on graph data like citation networks.
Beta distribution-based time step sampling method for optimizing diffusion model image generation efficiency.
Prompt tuning method reducing overfitting for vision-language models during transfer to downstream tasks.
Vision-Language-Action model combining language conditioning with dexterous robot manipulation control.
Linear transformer architecture with seasonal-trend decomposition for efficient multivariate time series forecasting.
Content moderation technique using soft prompts to prevent unsafe content generation in text-to-image models.
Watermarking technique for GNNs using explanation-based methods for intellectual property protection.
Methodology for LLM benchmark evaluation using multiple generations to account for inherent model randomness.