Multimodal learning framework for generalized traveling salesman problem in robot task planning. Combines vision and language for mobile robot applications.
LLM-based generalized planning via strategy refinement and reflection. Framework generating Python programs from natural language for PDDL domains.
EvA-RL framework optimizing policy evaluation accuracy during training for safer RL deployment. Addresses variance and bias in policy evaluation.
Benchmark evaluating unified multimodal models for vision understanding and generation. Comprehensive evaluation of architectural synergies.
Score-based diffusion models for inverse problems with decoupled posterior annealing. Novel Bayesian perspective on diffusion inference.
Neuro-symbolic reasoning for image region localization combining LLMs and vision models with verification. Original research on compositional reasoning.
Study of sample-efficient generalized planning using learned transition models, comparing symbolic approaches with Transformer-based planners like PlanGPT.
Unified framework for assessing cultural intelligence and competence of generative AI systems across different cultural contexts.
Framework using generative AI to assist participatory modeling in socio-environmental planning, translating natural language stakeholder descriptions into quantitative models.
Benchmark with 2,700+ conflict stimuli to evaluate whether audio multimodal LLMs process acoustic signals or rely on text-based inference.
Manifesto establishing conceptual foundations for Agentic Business Process Management (APM), extending BPM with autonomous agent governance for organizational processes.
GAN-based simulation framework measuring racial bias propagation in predictive policing systems across multiple cities.
Uses Stochastic Gumbel AlphaZero to evaluate difficulty in Tetris Block Puzzle variants.
Comprehensive survey on vector database technologies, storage methods, and retrieval techniques integrated with LLMs for modern AI systems.
AI models for analyzing police-public interactions from bodycam footage to improve government transparency and accountability in law enforcement.
Framework for evaluating information security awareness in LLMs, covering security knowledge, attitudes, and behaviors to help models understand security context and reject unsafe requests.
Deep reinforcement learning-based adversarial attacks against XSS detection models, exploring mutation strategies to evade deep learning classifiers.
Analysis and optimization of multi-stage LLM inference pipelines including RAG, KV cache retrieval, dynamic routing, and multi-step reasoning with diverse computational demands.
Pseudo-simulation evaluation paradigm for autonomous vehicles combining efficiency of open-loop evaluation with realism of closed-loop simulation.
Hierarchical reinforcement learning approach for scalable adaptive traffic signal control in smart cities with thousands of interconnected sensing nodes.
Study on LLM-based chatbot design for Alzheimer's and dementia caregivers, evaluating mental health support capabilities and identifying gaps in caregiver needs.
Method leveraging superclasses to mitigate spurious correlations and improve group robustness.
Semantic-driven topic modeling framework for analyzing creativity and insights in virtual brainstorming.
Diffusion world models for refining robotic manipulation policies via reinforcement learning.
Benchmark and evaluation framework for knowledge graph extraction from financial SEC 10-K filings using LLMs.
Lightweight module for vision-language models to adaptively select image resolution based on query context.
Scalable robot benchmarking via real-to-sim translation for evaluating diverse robotic agents.
Framework for decoding original text from single LLM token representation to understand model internals.
Audit of Google's AI Overviews and Featured Snippets quality on baby care and pregnancy queries.
Efficient RL training method for reasoning LLMs using adaptive drafter to handle long-tail response distribution.
LLM-generated dataset of 23,100 emails for benchmarking phishing, spam, and valid email classification.
Framework analyzing interaction between temperature settings and text perturbations in RAG system robustness.
LLM with reinforcement learning for dementia prognosis from clinical notes with longitudinal reasoning.
Benchmark and moderation model for evaluating LLM safety, adversarial robustness, and bias detection.
Theoretical analysis using sheaf theory and topology to understand feature distribution and attention mechanisms in graph neural network architectures.
RL framework using GRPO to generate adversarial paraphrases that evade AI-text detectors, stress-testing their robustness against evasion attacks.
Framework for curating and measuring ambiguity in long-horizon AI agents to improve their ability to seek clarification during extended task execution.
Browser-based interactive platform for teaching federated learning concepts. Developer tool and educational resource.
Investigates efficient chain-of-thought reasoning in LLMs through reward shaping and RL. Optimizing LLM inference cost and accuracy.
MADQA benchmark evaluates whether multimodal agents use strategic reasoning or stochastic search on document collections. Agent reasoning research.
Research on prompt injection attacks via role confusion. Analysis of how LLMs infer authority from text structure. Security for LLM applications.
ClawWorm demonstrates self-propagating attacks across LLM agent ecosystems via OpenClaw platform. Security research on multi-agent systems.
FEAT foundation model for structured data with linear complexity. Scalable model for tabular/structured data in healthcare and finance.
RAG-based system using LLMs for automated cybersecurity incident analysis from multiple log sources. LLM application for security workflows.
PlanTwin enables privacy-preserving planning for cloud-hosted LLM agents by abstracting private environment details. Research on agent security and privacy.
SO(3)-equivariant neural potential for molecular systems handling long-range electrostatic interactions. Machine learning research for materials modeling.
Formal specification for cryptographic admission control layer governing autonomous agent actions in enterprise systems.
Expert prefetching technique to accelerate MoE model inference by predicting expert access patterns during decoding.
Visualization framework for comparing regression model performance across multiple metrics.
Self-improvement framework for LLM personalization maximizing mutual information between contexts and responses without additional training data.