Can we automatize scientific discovery in the cognitive sciences?
Proposes automating scientific discovery in cognitive sciences using agentic AI to explore computational models faster than manual researcher-driven cycles.
Proposes automating scientific discovery in cognitive sciences using agentic AI to explore computational models faster than manual researcher-driven cycles.
Formalizes intelligent disobedience in shared autonomy using Stackelberg games and MDPs, modeling when automated assistants should override human instructions for safety.
Framework for integrating LLMs into Pepper robot for low-latency multimodal interaction, addressing cascaded pipeline latency and improving agentic control.
Proposes Knowledge Boundary Discovery, an RL framework to map LLM knowledge boundaries by generating questions the model can and cannot confidently answer.
Presents KLDrive, using knowledge graphs for fine-grained 3D scene reasoning in autonomous driving with LLMs, reducing hallucinations and improving reasoning transparency.
Open-source 560B MoE model advancing formal reasoning in Lean4 via agentic tool-integrated reinforcement learning, decomposing tasks into auto-formalization, sketching, and proving.
Proposes ORACLE, a method for generating high-quality synthetic reasoning data to train LLMs by validating intermediate reasoning steps beyond final answer correctness.
Studies adversarial attacks on text-attributed graphs combining GNNs and language models, exploring vulnerabilities in graph learning systems.
Addresses scaling failures in AlphaZero-style tree search for LLM reasoning, proposes ReSC using Gumbel sampling and sequential halving for budget-scalable inference.
Introduces ConsRoute, a semantic-aware adaptive routing framework for distributing LLM queries across cloud-edge-device tiers to reduce latency and inference costs.
Proposes Graph of States framework for LLMs to perform abductive reasoning through structured state representation and explicit state control, addressing gaps in logical reasoning.
Formalizes transformer context windows as I/O pages, proving tool-augmented agents with indexed external memory achieve exponential retrieval cost improvements over sequential scanning.
Approach improving LLM agent coherence for system optimization by addressing evolutionary bias and coherence ceiling.
Physics-constrained composable world model architecture ARYA designed for deterministic planning and AI safety.
RoboAlign approach improving embodied reasoning in vision-language-action models through test-time reasoning optimization.
Framework for AI scientific research using swarms of virtual lab agents enabling decentralized collaborative exploration.
AgentHER adapts hindsight experience replay to recover training signal from failed LLM agent trajectories.
AdaRubric generates task-specific evaluation rubrics on-the-fly for LLM agent assessment with step-by-step feedback.
PivotRL framework improving agentic post-training efficiency by combining supervised fine-tuning with reinforcement learning.
Method for measuring and steering LLM strategic behavior in games using activation vectors for personality traits.
Theoretical framework extending Myhill-Nerode theorem to bounded agents in finite POMDPs for canonical abstractions.
Study of error detectability in instruction-tuned LLM agents, proposing governability metric for agent security.
Agent architecture combining knowledge graphs and case-based reasoning for domain-specific code generation with LLMs.
Formalizes vendor value alignment constraints in AI decision support systems as behavioral feasible sets.
Task-oriented dialogue system approach treating safety certification as computational primitive for answer reuse.
Framework for automatically generating domain-specific agent nodes to improve multi-agent systems in specialized fields.
Neuro-symbolic approach stabilizing iterative self-training in LLMs by preventing recursive drift through verified reasoning.
Framework for solving credit assignment in collaborative multi-agent LLMs via counterfactual policy optimization to reduce free-riding.
Research on robust credit assignment for multi-agent LLM collaboration using reinforcement learning with heavy-tailed reward handling.
Study of multimodal LLM spatial reasoning capabilities, showing limitations in mental navigation and long-range spatiotemporal planning despite embodied agent deployment.
Multi-agent AI team (Cerebra) coordinating specialized agents for clinical decision support on heterogeneous patient data across EHR, notes, and medical imaging.
Method for improving uncertainty quantification in RAG systems by addressing entropy-based estimation failures caused by internal tug-of-war between induction and grounding.
Full-stack platform for enterprise AI agent development unifying tool integration, data generation, and training in closed-loop framework with privacy preservation.
Critical analysis of LLM benchmark contamination and overfitting, arguing benchmark scores conflate test-oriented competence with genuine capability generalization.
Research revealing gaps in multimodal AI systems' visual understanding mechanisms, showing frontier models generate plausible but unsupported reasoning traces.
Analysis of AI tokens as commodities and design of derivative contracts for inference compute, treating token consumption as infrastructure raw materials.
Framework for analyzing reasoning behavior and decision-making patterns in autonomous AI agents through structured behavioral analytics and operational tooling.
Bayesian method for detecting hallucinations in medical VQA using confidence-evidence gains without multiple stochastic generations.
MIND: Multi-agent framework using Theory of Mind for realistic consensus-building in travel planning with heterogeneous preferences.
Self-evolving coding agent blueprint for discovering surrogate ML pipelines for vehicle aerodynamic drag prediction.
Method leveraging LLM language guidance to address tail class scarcity in long-tail class incremental learning scenarios.
CurvZO: Efficient LLM fine-tuning via adaptive curvature-guided zeroth-order optimization with reduced memory overhead.
EvoIdeator: RL approach for autonomous scientific idea generation using checklist-grounded refinement instead of scalar rewards.
Framework identifying four structural properties of representational systems demanded by different types of reasoning.
Philosophical analysis of how LLMs differ from prior cognitive systems in representation genesis without clear transitions.
Framework for agentic personas that provide adaptive scientific explanations using knowledge graphs based on expert goals.
Empirical analysis testing whether LLMs exhibit genuine moral reasoning or produce superficially similar outputs through alignment training.
GenAI SECI Model: Framework for using generative AI to manage tacit knowledge in organizations.
Oph-Guid-RAG: Multimodal RAG system for ophthalmology clinical decision support retrieving guideline pages as evidence units.
Braid theory approach to multi-agent trajectory prediction for autonomous vehicles without extensive computation or heuristics.