Learning to Select Like Humans: Explainable Active Learning for Medical Imaging
Active learning framework for medical imaging that selects informative samples while ensuring models learn clinically meaningful features for expert annotation.
Active learning framework for medical imaging that selects informative samples while ensuring models learn clinically meaningful features for expert annotation.
Adaptive value decomposition for multi-agent reinforcement learning coordinating varying numbers of agents in urban systems with asynchronous actions.
Visual Para-Thinker: parallel reasoning strategy for visual comprehension tasks extending LLM test-time scaling to multimodal domain.
PeroMAS: multi-agent system orchestrating perovskite solar cell discovery through closed-loop workflows integrating literature, data, and experiments.
Agentic approach for spatio-temporal video grounding using collaborative multi-agent reasoning instead of frame-wise localization.
Sim2Radar: framework synthesizing radar training data from RGB images using VLMs for radar perception without manual annotation.
IDPruner: visual token pruning method balancing importance and diversity to accelerate multimodal LLM inference.
Zero-shot framework cascading object detection with lightweight vision-language models for edge robotics with limited training data.
HiST-VLA: hierarchical vision-language-action model for autonomous driving with improved spatial reasoning and trajectory prediction.
MedScope: multimodal LLM system with coarse-to-fine tool calling for clinical reasoning using long-form surgical videos.
CellMaster: AI agent leveraging LLM knowledge for zero-shot cell-type annotation in single-cell RNA sequencing analysis.
FOREST: diffusion-based world model for robotic warehouse stow operations using latent diffusion transformers to predict post-action bin configurations.
Production pipeline for deploying text-to-image models with automated and human-in-the-loop quality control for brand-safe marketing content.
Deep learning approach for generating image captions in Hindi using computer vision and NLP techniques.
Research on calibrating predictive distributions in probabilistic regression to improve uncertainty estimates in safety-critical applications.
Assessment of security risks from LLM coding agents generating malicious code, specifically spear-phishing websites.
Graph-grounded communication protocol for multi-agent LLM systems using structured graph operations instead of natural language to reduce hallucinations.
Reference-free evaluation framework for assessing vision-language models generating code from flowchart images in production settings.
Benchmark and defense methods for multi-turn safety risks in tool-using LLM-based agents, addressing gaps between capability and safety.
Neuro-symbolic approach using LLMs for retrosynthesis with fine-grained control to avoid chemically sensitive sites on molecules.
Security research on backdoor attacks that induce bias in LLMs, exploring white-box threat models where model builders are adversaries.
Evaluation of LLM and machine translation performance in crisis scenarios, focusing on preserving urgency in multilingual communication.
Research on social behavior and emergent dynamics in large-scale AI agent communities using MoltBook, a social platform designed for AI agents.
Study on how LLM embeddings store information in hidden layers, finding they contain relatively little input information compared to autoencoders trained for regeneration.
Diary study examining how multimodal LLMs support visual information access for blind and low vision users through conversational assistance.
Tutorial on PyCM library for comprehensive evaluation and comparison of multi-class classifier performance with multiple metrics.
Mechanistic interpretability method identifying prompt-specific circuits in language models revealing task-specific mechanisms vary by prompt.
Analysis of rank collapse phenomenon in federated low-rank adaptation with heterogeneous client resources and data distributions.
Review of edge AI adoption for autonomous biodiversity monitoring applications and barriers to practical implementation.
Optimizer improvement using trust-region adaptive scaling to address magnitude sensitivity in Newton-Schulz orthogonalized momentum methods.
Fine-tuned BERT classifier for detecting AI-generated content in Turkish news media with empirical prevalence analysis.
LLM-powered SQL agents enhanced with database-specific tribal knowledge to improve natural language to SQL translation on real-world databases.
Mechanistic interpretability study showing why and when singular vectors of attention heads align with learned features in language models.
Research on LLM calibration distinguishing between response-level and capability-level confidence estimation for reliable deployment.
Lightweight defense mechanism against jailbreak attacks that activates latent safety behaviors in LLMs without fine-tuning.
Approach to balance safety and utility in LLM alignment by adapting safe context learning without explicit safety rules in training data.
Training-free method to improve retrieval-augmented generation systems by using LLM confidence scores for reranking retrieved documents.
Co-evolutionary framework for LLM alignment using multi-agent competition and adaptive opponent pools to replace static reward models.
Research on vulnerability in LLM evaluation pipelines where edits to natural language rubrics can cause systematic preference shifts despite passing benchmarks.
Benchmark evaluating LLMs on software design and implementation tasks bridging requirements to code generation.
PT-RAG: retrieval-augmented generation system preserving academic paper hierarchical structure for improved evidence allocation in QA.
KorMedMCQA-V: multimodal benchmark of 1,534 Korean medical exam questions with images for evaluating vision-language models.
LeafNet: multimodal dataset and benchmark for vision-language models applied to plant disease detection in agriculture.
Multi-agent LLM system framework enabling dynamic test-time adaptation using retrieval-augmented generation and feedback loops.
Reinforcement learning framework for LLM-based audio agents to intelligently select and use tools for acoustic reasoning tasks.
Two-step generative policy combining flow matching and diffusion for faster robotic manipulation with better precision.
Large-scale multimodal dataset of scientific images with annotations for improving MLLM interpretation of domain-specific content.
Cross-embodiment robotic transfer learning using action motifs for few-shot adaptation to different robot morphologies.
Framework combining mechanistic consensus and language models for predicting transcriptional responses to unseen genetic perturbations.
Mean Velocity Policy for one-step generative action sampling in reinforcement learning with improved efficiency.