Bounding Probabilities of Causation with Partial Causal Diagrams
Proposes methods to bound probabilities of causation with partial causal diagrams for individual-level explanation and decision-making.
Proposes methods to bound probabilities of causation with partial causal diagrams for individual-level explanation and decision-making.
COOL-MC framework formally verifies and explains sepsis treatment RL policies through model checking with explainability for healthcare decision-making.
Examines knowledge conflicts in multimodal LLMs during long chain-of-thought reasoning, using internal representation probing to identify linear separability of conflict types.
Mechanistic analysis of LLM failures distinguishing knowledge existence from behavior expression in entity-based factual queries.
MATEO benchmark evaluates multimodal LLMs on temporal reasoning and planning for complex goal achievement with ordered steps and precondition dependencies.
Tabular foundation models applied to association rule mining for knowledge discovery in tabular data, addressing scalability and low-data regime performance limitations.
Arbor framework decomposes decision tree navigation for LLMs to maintain structured workflows in high-stakes domains like healthcare, addressing instruction-following degradation in long prompts.
Multi-plan dataset generation approach to remove planner bias in goal recognition for autonomous agents.
Evolutionary system prompt learning method for joint improvement of LLM contexts and weights via reinforcement learning.
Large-scale open-web simulator trained on 1M+ interactions for web agent training supporting reasoning and multi-format data.
Evaluation of frontier models' strategic reasoning, theory of mind, and metacognitive abilities in simulated nuclear crises.
Framework for building complete knowledge graph datasets with schema-level information for algorithm evaluation.
LLM-based StarCraft II agents enhanced with learnable action-conditioned world models for policy refinement.
Framework embedding web agents into customized UIs with lightweight frontend hooks for improved robustness and action expressiveness.
Training data attribution method using concept-based interpretability to identify which training examples drive LLM behaviors.
Joint learning and reasoning approach for probabilistic inference in first-order relational domains without constructing explicit models.
In-depth analysis of chain-of-thought prompting trace dynamics to understand driving forces behind LLM reasoning performance.
Position paper on linguistic self-reflection and conversational environments as foundation for robust AI reasoning inspired by developmental psychology.
Framework for dynamic workflow construction in agentic AI using standardized, modular workflow segments and dual knowledge architecture.
Multi-agent collaboration system for antimicrobial peptide design balancing activity, toxicity, and novelty objectives.
AI agents for drug discovery that scout non-English biotech sources globally, addressing >85% of patents originating outside the U.S.
Study revealing dense retrievers exhibit bias toward LLM-generated content, investigating relationship between perplexity and source bias in retrieval systems.
Simulation study of AI-assisted channel adaptation in UAV cellular networks with ground base stations and aerial repeaters.
Traffic simulation model for UAV ad-hoc networks with generative AI-based channel adaptation across packet sizes and transmission parameters.
Closed-loop framework combining causal LLMs, knowledge graphs, and digital twins for proactive telecom failure mitigation and simulation.
Safety-constrained reinforcement learning framework preventing unsafe emergent behaviors in wireless systems, UAVs, and IoT applications.
Framework integrating LLMs with reinforcement learning for wireless network optimization, addressing high-dimensional state spaces and computational demands.
Hierarchical multi-agent reinforcement learning for overlay multicast routing with network situational awareness decoupling multi-objective problems.
Quest Graph framework analyzing computational capabilities of agentic systems with finite context, establishing equivalence to Turing machines and pushdown automata.
Agentic AI control plane architecture for 6G network slice orchestration using intent-driven and economically programmable approaches.
Real-world deployment of generative AI-powered 9-1-1 training system addressing staffing crisis and reducing training time requirements.
Multi-LLM evaluation system for K-12 science instructional materials with human validation, designed to inform future GenAI-based material design agents.
Framework for responsible AI implementation in business organizations, structured around four focal areas for SMEs.
Global audit of Llama-3 8B model evaluating geographic and socioeconomic biases across 213 countries and eight technical metrics.
Boltz foundation model for atom-level representation learning in molecular systems, bridging protein and small-molecule modeling approaches.
Study measuring implicit bias against transgender populations in LLMs through word association tests and scenario evaluation.
Feedback control optimizer for training spiking neural networks with sparse activity and local learning rules, addressing energy efficiency in neuromorphic computing.
Novel uncertainty quantification framework for generative models using directional concentration approach, more flexible than existing heuristic methods.
MergePipe system for efficient parameter management in LLM merging reducing disk I/O and enabling scalable multi-expert model integration.
LLM-enhanced rumor detection framework using virtual node edge prediction to capture textual coherence across rumor propagation paths.
Graph-based spatio-temporal network for cellular traffic prediction capturing complex temporal dynamics and spatial correlations.
Large-scale empirical study of AI-only social platform with 27,269 agents examining emergent behavior, governance, safety, and inter-agent dynamics.
Explanatory interactive machine learning approach to mitigate bias and spurious correlations in visual gender classification models.
Study of how post-training quantization compression affects reliability and accuracy in multimodal LLMs for VQA tasks on edge devices.
Agentic architecture for autonomous decision-making in beyond 5G network management balancing efficiency, user satisfaction, and energy optimization.
Multi-agent AI simulation framework for coordinated space exploration and settlement addressing communication delays, resource scarcity, and safety constraints.
VisPhyWorld framework evaluating physical reasoning in multimodal LLMs through code-driven video reconstruction instead of VQA-style benchmarks.
Controlled comparative study of VGG, ResNet, and GoogLeNet architectures analyzing how convolutional depth affects classification performance and efficiency.
WildfireVLM framework combining satellite imagery and vision-language models for early wildfire detection and risk assessment.
Fine-tuned vision-language model for automated artistic creativity assessment of paintings with explainable feedback.