NEURON: A Neuro-symbolic System for Grounded Clinical Explainability
Neuro-symbolic system combining SNOMED CT ontology with machine learning for interpretable clinical AI predictions.
Neuro-symbolic system combining SNOMED CT ontology with machine learning for interpretable clinical AI predictions.
Benchmark for evaluating process reward models across diverse reasoning tasks beyond mathematics, enabling detection of intermediate reasoning errors.
Framework improving faithfulness in vision-language GUI agents by grounding actions in screen evidence and user instructions via guided advantage estimation.
Position paper proposing agentic systems be designed as token allocation economies with specialized layers for routing, planning, and action selection.
Interactive simulation environment for training multimodal agents to perform Earth observation analysis with tool use and uncertainty resolution.
Neuro-symbolic framework for inducing executable skills from agent interactions, combining LLM reasoning with programmatic logic for long-horizon planning in dynamic environments.
Policy optimization method aligning RL credit assignment with natural reasoning steps in multi-modal tasks at segment granularity rather than token or sequence level.
Study of in-group favoritism biases in persona agents facing contradicting information and methods to mitigate adverse effects on factual accuracy.
Multimodal dataset and recognition framework for non-standard system-level chip design diagrams to improve MLLM understanding of architectural specifications.
Evaluation of cognitive plausibility for computational models of analogy and metaphor including SME, CogSketch, and LLMs using the Minimal Cognitive Grid framework.
Systems framework analyzing AI safety through irreversibility control and deployment friction reduction in autonomous decision-making systems.
Hierarchical tokenization framework for time-series generation enabling user control over temporal granularity from sketches or scratch.
Formal theory of Artificial Jagged Intelligence modeling uneven optimization pressure across capability domains during training as finite-budget gradient allocation.
Method for auditing and composing LoRA adapter libraries with residual merging and reliability assessment for task-level reuse and instance-level selection.
Formal framework for explaining entailments in description logic knowledge bases with user-centered contrast-based approach beyond traditional justifications.
Few-step generative model for offline multi-agent reinforcement learning enabling coordinated inference without sacrificing inter-agent coordination.
Framework grounding multi-hop fact verification in structural causal models using group relative policy optimization to improve LLM reasoning and reduce hallucinations.
Multi-turn legal consultation agent using coverage-driven retrieval control to determine sufficient evidence and relevant legal issues.
Deep research agents for automated scientific discovery using post-training on information-seeking tasks and iterative problem-solving capabilities.
Agent system for human-vehicle collaboration using bidirectional perception and alignment to improve driver-automation coordination and situational awareness.
Analysis of inference scaling strategies including self-consistency, self-refinement, and multi-agent debate for compute-efficient LLM performance improvement.
Evaluation framework for production agentic AI systems addressing compounding errors, tool failures, output drift, and long-horizon task evaluation beyond lab-scale benchmarks.
Multi-agent system using LLMs to translate natural language into constraint programming models with synthesized validation checkers to reduce semantic errors.
Framework for designing latent state representations in world models for agents, categorizing methods by functional purpose rather than implementation approach.
Research on routing mechanisms in AI systems and their impact on trust, cost, quality, and accountability of responses across different service tiers and endpoints.
Study of genre bias in LLM credibility assessment, showing models misclassify entertainment news more than hard news.
Defense mechanism against infectious jailbreak attacks in multi-agent systems through foresight-guided strategies.
Momentum: game with runtime procedural content generation evaluated by autonomous agents for balance and playability.
DataEvolver: closed-loop visual data generation system using goal-driven agents for iterative dataset creation and improvement.
Decision-propagation method integrating Answer Set Programming with neural networks for scalable neuro-symbolic reasoning.
NeuroState-Bench: human-calibrated benchmark evaluating commitment integrity in multi-turn LLM agent tasks via side-query probes.
Categorical sheaf-theoretic framework for resilient multi-agent autonomous systems planning in stochastic environments.
AI-driven cybersecurity system for financial institutions using LLMs to enhance SOC reasoning capacity and alert investigation.
Adversarial self-play defense against persona-based jailbreak attacks in LLMs through intent-role disentanglement.
Formal language specification for encoding LLM agent context composition, standardizing context engineering practices.
Hierarchical reinforcement learning with language for pair trading, addressing credit assignment in long-horizon semantic tasks.
Multi-agent LLM benchmark using 12 AI agents with film personas to evaluate deliberation and reasoning capabilities.
Self-supervised learning framework for brain tumor classification using SSL methods (SimCLR, BYOL, DINO, Moco v3) on MRI data.
Personalized digital health modeling framework using adaptive weighting for heterogeneous user data with limited annotations.
Position paper on externalizing implicit knowledge in AI systems for reliability through human-AI collaboration infrastructure.
Model Spec Midtraining improves LLM alignment generalization by training on specification-aligned behavior before final fine-tuning.
NORA: multi-agent autonomous research system specialized for spatial data science workflows, automating scientific research end-to-end.
Proposes DGMM architecture addressing memory, temporal grounding, and interpretability limitations in LLMs through gist-based memory mechanisms.
arXiv paper proposing multi-agent framework with planner, actor, and memory manager roles for long-horizon LM-based task automation.
arXiv paper evaluating 1M-token context window retrieval and multi-hop reasoning in frontier LLMs on classical Chinese texts.
arXiv paper proposing intervention complexity as canonical reward measure for general intelligence in computable environments.
arXiv paper on uncertainty-guided exploration for stabilizing multi-turn reinforcement learning in agentic LLMs.
arXiv paper introducing MEMAUDIT protocol for evaluating long-term memory writing in budgeted LLM agents independent of retrieval and reasoning.
arXiv paper on clean-label backdoor attacks against vision-language models using diffusion model-generated triggers.
arXiv paper formalizing efficient benchmark selection for LLM evaluation as submodular maximization problem.