CiQi-Agent domain-specific multimodal agent for Chinese porcelain connoisseurship combining vision, tools and aesthetics for cultural heritage analysis.
PROClaim courtroom-style multi-agent debate framework with progressive RAG and role-switching for claim verification reducing hallucinations and improving reasoning.
Analysis of scaling laws in AI across model families showing empirical power-law relationships between compute and loss, examining their predictability and generalization.
Hydra unifies document retrieval and generation in single vision-language model using dual-head LoRA adapter toggled at inference time.
Domain-invariant prompt learning method for vision-language models like CLIP to handle domain shifts in zero-shot transfer tasks.
Fine-tuning LLMs for tactical deconfliction decisions in multi-agent unmanned aerial systems under safety-critical constraints with partial observability.
CirrusBench benchmark evaluates LLM-based agents in real-world cloud service environments beyond correctness, measuring robustness and efficiency in high-complexity customer-assistant interactions.
Learning-based partial action replacement for offline multi-agent reinforcement learning addressing exponential joint action space sparsity.
Agentic dual-path framework decoupling perception and verification for robust VLM question answering on misleading charts via skeptical reasoning.
Explores using LLMs as conversational agents for supporting planning and translation during student reflective writing rather than feedback provision.
Input-side adaptation framework enabling MLLMs to maintain high spatial resolution and temporal context through learned adaptive resolution.
Trust-aware routing mechanism for distributed generative AI inference across heterogeneous edge devices accounting for reliability and performance.
Information-theoretic analysis of compatibility between bounded cumulative risk and unbounded self-improvement in safety verification systems.
Benchmark for evaluating agentic vision-language models on multi-image grounding through extended interactions with hidden target identification tasks.
Training-free adaptive token selection framework for long video understanding in MLLMs using entropy-based mechanisms to reduce memory costs.
Stepwise credit assignment method for reinforcement learning on flow-matching generative models, differentiating early vs. late diffusion steps.
LLM-based middleware using FastAPI for dynamic schema mismatch detection and resolution in distributed systems with heterogeneous services at runtime.
Framework for documenting AI-augmented ecosystems with probabilistic behavior and ML/software system interactions beyond traditional arc42 and C4 models.
Research on causal discovery using LLMs to extract causal relationships from text while accounting for hallucination risks.
Research combining large language models with task-specific models for time series anomaly detection using hybrid approach.
SkillFlow: Multi-stage retrieval system for agents to dynamically select relevant skills from large libraries at inference time.
Explores integrating LLMs with classical planners for large-scale planning by generating actions/states to prune search space using domain-specific knowledge.
L-MARS multi-agent system for legal QA decomposes queries, performs agentic web search, verifies results, and synthesizes cited answers with structured retrieval.
Trains large reasoning models to stop reasoning early by detecting sufficient evidence accumulation, reducing computational costs while maintaining performance.
Searches for optimal meta reasoning DAG structures to guide LLM reasoning without manual design, improving query-specific adaptation and logical dependency capture.
Agent GPA framework evaluates LLM agent failures at goal-plan-action intersections using factorized LLM judges for scalable cross-architecture assessment.
GammaZero uses graph representations to learn planning guidance in POMDPs with generalization across problem sizes without domain-specific architectures.
Multi-agent framework for spatial Text-to-SQL that resolves geographic intent, schema ambiguity, and PostGIS spatial functions for non-expert data queries.
AISAC is a transparent multi-agent runtime from Argonne Lab for scientific reasoning with explicit role semantics and evidence-grounded execution governance.
Autonomous Issue Resolver uses LLM agents for repository-scale automated program repair by shifting from control-centric to data-flow-centric code analysis paradigm.
Scientific Autonomous Goal-evolving Agent automates objective function design to guide scientific discovery agents beyond manually-specified quantitative proxies.
Attention mechanism for robust multimodal fusion in Global Workspace Architecture that independently evaluates modality reliability without end-to-end co-adaptation.
AgentLeak benchmark measures privacy leakage in multi-agent LLM systems across inter-agent messages, shared memory, and tool arguments with 1,000 scenarios.
Evaluates when LLM agents exhibit scheming behavior by decomposing incentives into agent and environmental factors to understand misalignment risks in autonomous systems.
Multi-agent LLM system for automated mathematical discovery that generates conjectures, attempts proofs, and uses feedback to evolve a distribution of mathematical concepts.
RetroAgent improves LLM agent RL training by incorporating retrospective intrinsic feedback and explicit experience reuse to enable continual adaptation beyond isolated task completion.
UCIP detection framework distinguishing intrinsic vs instrumental self-preservation objectives in autonomous agents with memory and planning.
Seed1.8 foundation model supporting multi-turn interaction, tool use, code execution, and GUI interaction with configurable inference modes.
Analysis of LLM benchmark contamination and reliability issues in evaluating model generalization capabilities.
ProGRank defense mechanism using probe-gradient reranking to protect dense-retriever RAG systems from corpus poisoning attacks.
Template-driven ML development approach for scaling ML model ecosystems in computational advertising platforms.
Trace2Skill framework automatically distills trajectory-local lessons into transferable skills for LLM agents handling complex tasks.
Framework for evaluating harmful AI manipulation through human-AI interaction studies across policy, finance, and health domains.
Survey of continual graph learning methods for incrementally learning from streaming graph data without catastrophic forgetting.
Prior learning method using structured posteriors as informative priors to improve generalization and uncertainty in neural networks.
Theoretical analysis of sample complexity for model-based Q-learning integrating environment models with Q-learning algorithms.
Variational Autoencoders improved via Explaining-Away mechanism for better uncertainty representation in visual inference.
Few-shot multi-class anomaly detection using bidirectional multimodal prompt learning for industrial inspection tasks.
Framework for robots to continually learn tasks and skills through dialogue interactions, maintaining a skill library with LLM support.
Survey of multimodal continual learning methods enabling models to learn from new data without forgetting previous knowledge.