GATS: Planning framework combining tree search with layered world models to reduce LLM inference calls during agent planning, improving efficiency and reducing stochasticity.
Long-Horizon-Terminal-Bench: Benchmark with 46 long-horizon terminal tasks and dense reward grading to evaluate agent capabilities beyond simple well-specified problems.
AI-assisted formalization of mathematical proofs in Lean 4 proof assistant, framed as a game where AI directs formalization of research results.
ARCANA: Multi-agent framework for ARC-AGI-2 task solving using perception agents, program synthesis, symbolic execution, and reflective refinement under computational constraints.
Neuro-agentic control framework combines LLMs with neural networks for industrial IoT security monitoring while mitigating hallucination risks in closed-loop control.
L-MAD framework systematically evaluates multi-agent debate structures for legal reasoning tasks, testing different agent personas and aggregation methods.
MedRealMM is a large-scale benchmark for evaluating LLMs in Chinese online medical consultation using real multimodal data and clinical quality metrics.
Efficient process reward modeling via KV-cache transfer for scaling multi-agent systems, reducing quadratic complexity in long trajectory scoring.
Verification mechanism for agentic context evolution in long-horizon LLM deployments with persistent system instructions under distribution shift.
Protocol for making LLM agents auditable in scientific discovery by explicitly tracking hypothesis evolution, tests, and belief updates.
Open-source system for LLM-driven automated theorem proving with Lean 4 verification using Planner-Worker-Verifier architecture.
Benchmark for evaluating LLM-based medical agents on long-horizon clinical decision-making using real EHR data across repeated visits and evolving treatments.
Coordination framework for heterogeneous LLM embodied agent teams with communication-efficient mechanisms for physical AI deployments.
Multi-agent LLM collaboration system for fictional worldbuilding using hierarchical context compression and iterative review for consistency.
LLM agent with author-critic architecture for solving open mathematical problems through agentic workflow inspired by mathematical practice.
Study of reward hacking in multimodal LLM reinforcement learning across VQA and safety tasks with varying reward designs and model scales.
Memory mechanism for multi-turn agentic LLM systems that selectively persists configuration, domain constraints, and tool-use patterns across sessions.
Self-evolving agent for multimodal survival prediction that actively reasons about which diagnostic modalities to acquire given cost constraints.
Examines structural limitations in AI systems for reasoning and code generation when representational frames are fixed rather than open-ended.
Combines knowledge graphs and explainable AI techniques for pre-demolition assessment in urban mining decision support.
Risk classification framework for governing agentic AI systems across enterprise and public sector applications with seven system types.
Framework for LLM agents to orchestrate expert models and tools via auction-based task allocation considering performance and cost efficiency.
Benchmark for evaluating LLM reverse engineering capabilities on decompiled binary function naming tasks, addressing measurement gaps in security applications.
Unified framework analyzing knowledge distillation mechanisms in LLMs through interaction decomposition to understand why KD methods work.
Interpretable LLM-guided mixture-of-experts model for Alzheimer's disease survival prediction combining neuroimaging with natural language reasoning.
Quantization technique for few-bit integer representation addressing asymmetric clipping issues in signed symmetric quantizers.
Training method for Mixture-of-Experts models using differentiable routing consistency loss to reduce weight swapping during edge device inference.
Technique using optimal transport coupling in flow matching to embed controllable structure for molecular property generation.
Distributed serving system for Mixture-of-Experts models using online proactive expert placement to optimize communication and computation latencies.
Standardized benchmark library for federated continual learning with consistent evaluation methodology across heterogeneous client settings.
GPU kernel optimization for sparse matrix multiplication enabling efficient LLM inference acceleration with moderately unstructured pruned weight matrices.
Uses LLMs as mutation and crossover operators to evolve multi-objective Bayesian optimization algorithms automatically with hyperparameter tuning.
Framework decoupling patient dynamics learning from treatment optimization for sepsis using generative EHR model digital twins and inference-time control.
Diffusion-based synthetic image generation pipeline for sand boil defect detection on earthen levees using ControlNet and DreamBooth fine-tuning.
Unified corpus for biology domain combining heterogeneous biological databases and resources for pretraining specialized large language models.
Method using LLMs and VLAs as exploration guidance in reinforcement learning to escape weak policies through natural language prompts.
Production-deployed agentic LLM system using graph-guided multi-agent framework for reliable Standard Operating Procedure execution in warehouse operations.
Framework analyzing specification ambiguity in natural language supervision for LLMs, introducing NL-PAC to measure minimax risk bounds in label generation.
Benchmark dataset for evaluating vision-language models' ability to integrate multi-view observations into coherent 3D scene understanding.
Research on adapting pretrained vision-language models to vision-language-action models for robotics with minimal architectural changes to preserve VLM contributions.
Identifies structural incoherence in LLM-generated code where locally valid patches fail globally due to missing configs, imports, or authentication guards.
SCATE framework teaches coding agents to generate better tests by addressing lazy generation problem that causes premature task termination and low code coverage.
Analyzes AlphaZero performance gap in sparsely rewarded games like Connect Four and Chomp, studying auxiliary supervision effectiveness.
Model-agnostic graph prompt learning reduces parameters and domain expertise requirements for crystal property prediction with Graph Neural Networks.
Studies contextual bandits with correlated arms and surrogate rewards for LLM routing, handling noisy auxiliary reward information.
Reviews evolutionary computation as basis for autonomous scientific discovery systems that integrate experimental feedback and human guidance in open-ended exploration.
Examines emerging AI agent skill repositories and marketplaces, analyzing what software engineering activities become reusable agent skills.
OmniMapBench introduces benchmark with 2,096 manually annotated examples for evaluating visual-centric reasoning in LVLMs on map documents.
Integrates LLMs with Graph Convolutional Networks to improve semi-supervised image classification by addressing graph construction challenges.
RAG-based system augments fundamental company analysis by combining LLMs with SEC filings and macroeconomic data via API calls to GPT-4o.