Agentic AI-based Coverage Closure for Formal Verification
LLM-enabled agentic workflow automating coverage analysis and gap identification for IC formal verification.
LLM-enabled agentic workflow automating coverage analysis and gap identification for IC formal verification.
Extends RAG paradigm to time-series foundation models for predictive maintenance with covariate dynamics.
Adopts goal recognition heuristics for classical planning problems to improve plan search prioritization.
Multi-agent routing system aware of cascading failures in tree versus cyclic graph topologies with geometry-switching.
Knowledge graph-driven multi-agent LLM framework for semantic geospatial data discovery with improved retrieval.
Cerebra: multi-agent AI system with specialized agents for EHR, clinical notes, and multimodal data in dementia assessment.
Evaluates ChatGPT (GPT-3.5/4) effectiveness on extracting research challenges from HCI literature at scale using two-step approach.
Method for reliable out-of-distribution virtual screening in drug discovery using extrapolatory pseudo-label matching.
HFLDD: hybrid federated learning framework using dataset distillation to handle non-IID data distribution skew.
LOGSAFE: logic-guided defense mechanism for federated learning in time-series cyber-physical systems against poisoning attacks.
Attention calibration method to reduce object hallucinations in vision-language models through vision token reordering.
BalanceKV: streaming algorithm using discrepancy theory to approximate attention for efficient long-context LLM token generation.
Studies use of deliberation-enhancing chatbot to help groups detect deepfake text through human-AI collaboration.
Agentic system autonomously generates, evaluates, and refines quantum feature maps for quantum machine learning using LLMs.
Information-theoretic framework to characterize and quantify information leakage in concept-based models for interpretability.
Meta-optimization framework for LLMs to generate generalizable heuristics for combinatorial optimization without manually predefined evolutionary operators.
CyberGym large-scale benchmark with 1,507 real-world vulnerabilities for evaluating AI agents' dynamic cybersecurity capabilities.
Learns minimum action distance metric from state trajectories alone to capture environment structure for MDPs without rewards or action labels.
RedTopic framework for topic-diverse red teaming of LLMs to identify vulnerabilities across broad range of harmful topics adaptively.
Reasoning-guided LLM function completion using context when docstrings are absent in real-world code repositories.
MARS proposes efficient multi-agent collaboration framework for LLM reasoning, reducing computational overhead of Multi-Agent Debate while maintaining reasoning capabilities.
VL-KnG constructs spatiotemporal knowledge graphs from egocentric video using vision-language models for persistent scene understanding without 3D reconstruction.
Self-correction Loop with Structured Output framework enhances GPT-based VLMs for generating reliable dental radiological findings in medical image interpretation.
Information Gain-based Policy Optimization uses RL to train LLM agents for multi-turn search with tool use, addressing reward sparsity in exploration-based tasks.
MCP Security Bench systematically evaluates attacks against Model Context Protocol in LLM agents, measuring resistance of tool-calling systems to adversarial inputs.
GUIrilla is a scalable framework for automated desktop UI exploration generating large-scale training data for LLM-based GUI understanding and automation.
MeasureBench benchmarks vision-language models on visual measurement reading tasks with real-world and synthesized instrument images.
Xmera framework evaluates adversarial man-in-the-middle attacks on LLM factual recall through prompt injection, measuring vulnerability of question-answering systems.
MOON2.0 addresses multimodal imbalance in MLLMs for e-commerce product understanding through dynamic modality-balanced representation learning.
HumorChain: Theory-guided multi-stage reasoning framework for interpretable multimodal humor generation using LLMs.
Study on spatial reasoning in LLMs for 3D scene understanding, examining attention masking mechanisms for order-agnostic objects.
ThinkDeeper: Framework for autonomous vehicle grounding using world models for 3D spatial reasoning and scene prediction.
Research on metaphor-based jailbreak attacks against text-to-image models' safety defense mechanisms.
Zero-shot object navigation for robots using ensemble prediction of future states in unseen, cluttered environments.
Empirical study examining reproducibility gaps in code generated by LLM coding agents and missing dependency specifications.
VLM-CAD: Collaborative agent design workflow for analog circuit sizing using vision-language models with spatial reasoning.
Information-theoretic analysis of trade-offs between fairness, privacy, and accuracy in machine learning using Chernoff Information.
HAVEN: Framework for long-video understanding using agentic search and audiovisual entity cohesion to maintain global coherence.
Analysis of representational homomorphism in transformers to predict and improve compositional generalization in language models.
Vision-DeepResearch: Framework augmenting multimodal LLMs with tool-calling capabilities for visual and textual search.
1S-DAug: One-shot data augmentation method for improved few-shot learning generalization using generative synthesis.
Residual Decoding: Training method to reduce hallucinations in vision-language models using history-aware residual guidance.
FlyPrompt: Brain-inspired routing method for continual learning from non-stationary data streams without task boundaries.
Study evaluating behavioral consistency of LLM agents in stock market simulations against real market participant behavior.
Energy-aware reinforcement learning for robotic manipulation of articulated objects in infrastructure maintenance and smart cities.
KDFlow: Framework for efficient knowledge distillation of large language models into smaller models with heterogeneous training backends.
MA-RAG: Multi-round agentic RAG system for medical reasoning with LLMs, addressing hallucinations and outdated knowledge through iterative refinement.
Augmenting Proximal Policy Optimization with temporal sequence models for robust reinforcement learning under sensor drift and partial observability.
NCCL EP, unified communication API for mixture-of-experts architectures in large language models built on NCCL with GPU-initiated RDMA.
Training-free fine-grained visual recognition using large vision-language models with sample-wise adaptive reasoning for subordinate-level category disambiguation.