Production-validated framework for clinical conversational AI agents grounded in signals from 115M+ patient interactions optimizing for real-world healthcare scenarios.
Vision-language model-based reranking system addressing modality gap in hybrid text-image retrieval pipelines using domain-specific adaptation.
Hardware accelerator design optimizing systolic arrays for skewed matrix multiplication workloads common in LLMs and modern AI/ML inference.
End-to-end deep learning framework combining segmentation and dual-mode compression for efficient wind turbine inspection image transfer.
Demonstrates LLM capability to semi-autonomously compute Bethe Ansatz solutions for integrable spin chain models with minimal human intervention.
Study of AI risk perception and responsible adoption among 139 students in CS/Data Science programs through explicit ratings and scenario assessments.
Research on using LLM-derived abstractions to improve analogical reasoning over narrative structures, addressing prompt sensitivity and entity extraction challenges.
Pythonic framework for aggregate computing with machine learning integration for distributed systems.
PanDA workflow system for CERN detector design optimization using distributed computing and AI/ML pipelines.
Position paper on system-level defenses against indirect prompt injection attacks on LLM-powered AI agents.
Hybrid framework combining RL for control and LLMs for planning in robotic manipulation tasks.
Theoretical generalization of attention mechanisms including GQA and MLA, advances LLM architecture understanding.
Transformer-based approach to identify parallelizable loops in source code, combines ML with developer tools.
MindCube benchmark evaluates Vision-Language Models on spatial reasoning from limited views, revealing gaps in 3D scene understanding.
Multi-agent framework applying teamwork theory for medical reasoning with LLMs, enabling efficient clinical decision-making without frontier models.
CLAUSE agentic neuro-symbolic framework for knowledge graph reasoning using dynamic learnable context engineering to balance accuracy, latency, and provenance.
Introduction to symbol grounding in neuro-symbolic AI, covering how neural networks and symbolic reasoning combine for trustworthy and constrained predictions.
Survey examining adaptive reasoning in LLMs, showing models apply uniform strategies regardless of task complexity and need dynamic reasoning allocation.
Analysis of 25K+ chain-of-thought trajectories revealing domain-specific phase transitions in LLM reasoning across law, science, code, and math.
Graph of Concept Predictors distills LLM reasoning into compact student models using active distillation to reduce inference costs and latency.
Generative data transformation approach to address domain gaps when combining multi-domain data for improved recommendation model training.
MultiGen introduces persistent external memory for diffusion-based game engines enabling user control and shared multiplayer world interactions.
Method for attributing multi-agent system outputs when execution traces unavailable, using implicit execution tracing for accountability and debugging.
Empirical comparison of communication protocols for multi-agent orchestration, evaluating tool integration versus inter-agent delegation approaches.
CoMaTrack uses competitive multi-agent reinforcement learning with game theory for embodied visual tracking, improving generalization beyond imitation learning.
UniAI-GraphRAG enhances GraphRAG with ontology-guided extraction and multi-dimensional clustering for improved multi-hop reasoning and domain-specific QA.
Trace2Skill framework automatically distills transferable skills for LLM agents from trajectories, enabling scalable domain-specific capability development.
GUIDE framework addresses domain bias in GUI agents through real-time web video retrieval and annotation, improving interface understanding and task execution.
Language-conditioned multi-game level generation using shared representations for procedural content generation across multiple game domains.
Neuro-symbolic approach for predictive process monitoring that incorporates domain-specific constraints and knowledge into machine learning models for process prediction.
Closed-loop ranking optimization system using agent-based influence exchange for recommendation systems.
Multi-agent pipeline using rhizomatic process-relational approach for non-linear systematic literature analysis.
Early exiting mechanism for predictive coding neural networks to enable efficient inference on edge devices.
Name-only online learning with generative example synthesis for continual learning under distribution shifts.
Evaluation of LLMs on diagnostic reasoning from unstructured clinical narratives in epilepsy using multiple models.
LLM-driven conversational recommender system for leisure events with user-centric evaluation in SME context.
Neuro-symbolic feedback approach to improve text-to-video generation consistency for complex multi-object prompts.
German-language LLM pre-training dataset combining heuristic filtering, model-based curation, and synthetic data generation.
Meta-learning framework enabling LLMs to design selection operators for evolutionary symbolic regression algorithms.
Benchmark for evaluating vision foundation models on atomic visual abilities, addressing evaluation gaps in VFM+LLM pipelines.
SlowFast Sampling strategy to accelerate diffusion-based language models with parallel token generation and flexible decoding.
Defense mechanism against GUI attacks on multimodal LLM-based agents using layer-wise scaling to prevent pop-up injection attacks.
Research on Question Augmentation strategy improving LLM reasoning capacity via reinforcement learning on harder problems.
Research on applying LLM embeddings to generate variant-level representations across 8.9 billion genetic variations.
Research benchmark SecureVibeBench evaluating code generation security of LLM-powered agents on realistic vulnerability scenarios.
Research paper proposing semantic voting method for LLM self-improvement on open-ended tasks without self-evaluation. Addresses pseudo-label generation for unverifiable outputs like translation.
Multi-stream generative policy framework for sample-efficient robotic manipulation using inference-time composition.
Theoretical analysis of implicit models' expressive power, showing infinite-depth networks with test-time scaling capability.
Study on applying generative AI to automate bug tracking, reporting, classification and resolution in software development.
Low-attention transformer optimization technique reducing computational overhead by exploiting architectural redundancies.