Reliability Gated Multi-Teacher Distillation for Low Resource Abstractive Summarization
arXiv paper on multi-teacher knowledge distillation for low-resource abstractive summarization using inter-teacher agreement for supervision routing.
arXiv paper on multi-teacher knowledge distillation for low-resource abstractive summarization using inter-teacher agreement for supervision routing.
arXiv paper introducing PR3DICTR, open-access PyTorch/MONAI framework for 3D medical image classification and outcome prediction.
arXiv paper on server learning with client filtering to improve federated learning robustness against malicious attacks.
arXiv paper on WiseMind, multi-agent LLM framework inspired by Dialectical Behavior Therapy for reliable and empathetic psychiatric diagnosis.
arXiv paper on AutoCO, LLM-based method coupling OR principles with bidirectional coevolution for complex constraint optimization problems.
arXiv paper on Glia, multi-agent LLM architecture for autonomous computer systems design using specialized agents with empirical feedback loops.
arXiv paper introducing CostBench benchmark for evaluating LLM tool-use agents on cost-optimal planning and adaptation in dynamic environments.
arXiv paper on code-in-the-loop agentic tool use for image forgery detection, unifying low-level artifacts with semantic knowledge from MLLMs.
arXiv paper on ClinicalReTrial, multi-agent system using LLMs to redesign failing clinical trial protocols with actionable recommendations.
arXiv paper on AgenticRed, automated pipeline using in-context learning to evolve red-teaming systems without human-designed workflows.
arXiv paper analyzing gap between LLM math benchmark performance and real-world application through contextual reasoning benchmark ContextMATH.
arXiv paper on embedding authorization mechanisms directly into LLM reasoning to prevent data leakage and unauthorized command execution.
arXiv paper introducing framework for evaluating harmful AI manipulation through human-AI interaction studies across policy, finance, and health domains.
arXiv paper proposing PAPO, integrating process-level evaluation into policy optimization to improve reasoning quality beyond final-answer correctness.
arXiv paper on multi-agent RAG with adaptive orchestration and evolving agent prompts to handle complex multi-hop reasoning tasks.
Analysis showing LLM reasoning models encode decisions before generating chain-of-thought explanations via linear probes.
Study evaluating reliability and risk of AI systems in medication decision-making and healthcare workflows.
OSCAR framework for mitigating hallucinations in diffusion language models using self-verification during generation.
Research probing whether LLMs encode awareness of conversation continuity by generating user turns after assistant responses.
Novel framework using LLMs for causal graph discovery via breadth-first search, reducing query complexity from quadratic to linear.
Improves emotion intensity and speaker consistency in zero-shot LLM-based text-to-speech through expressive prompt design methods.
Multimodal LLM fine-tuned for interpretable image forgery detection and localization providing semantic understanding beyond low-level artifacts.
Proposes scale transformation method for transferable targeted adversarial attacks requiring minimal data without surrogate model feedback.
Zero-shot concept bottleneck models enabling interpretable predictions without target task training by leveraging zero-shot learning.
Improves text-to-video generation semantic and temporal consistency using neuro-symbolic feedback without retraining the model.
LMask framework uses dynamic masking with learning to solve constrained routing problems as combinatorial optimization tasks.
StructEval benchmark systematically evaluates LLM capabilities in generating structured outputs across JSON, HTML, React, SVG and other formats.
Formalizes mission-aligned learning-informed control framework for autonomous physical agents integrating learning with task objectives.
Proposes modular vision-language alignment architecture improving CLIP's handling of multi-object images and caption misalignment.
Introduces ReDef, high-confidence software defect prediction dataset from 22 C/C++ projects, evaluating code language model understanding of changes.
Compares psychometric questionnaire profiles with actual LLM generation behavior across eight open-source models to assess assessment validity.
Generates synthetic robot poses for RGB-D bimanual manipulation data augmentation to improve imitation learning policy training.
Analyzes political bias in LLM training data composition across pre and post-training stages to understand sources of model bias.
Proposes learning progress monitoring to improve exploration efficiency in reinforcement learning agents when encountering unlearnable noise sources.
Introduces attribution gradients technique to improve citation informativeness and evidence transparency in AI answer engines.
Forecasts expert selection patterns in Mixture of Experts LLMs to optimize data movement overhead in multi-unit serving systems.
Extends Forward-Forward algorithm to reinforcement learning using action-conditioned Q-functions and layer activity statistics as learning signals.
CQA-Eval evaluation framework for multi-paragraph clinical question answering systems with physician annotations and recommendations for resource-constrained settings.
f-INE hypothesis testing framework estimates sample influence on model performance while accounting for training randomness, addressing instability in existing influence estimation methods.
MusicRFM framework adapts Recursive Feature Machines to enable fine-grained control over frozen pre-trained music generation models via internal activation steering.
Deep learning approach fixing systematic S-wave detection failures in seismic phase picking via shape-aware loss functions.
SAGA framework for source attribution of AI-generated videos. Identifies specific generative model used instead of binary real/fake detection.
Research on contrastive fusion for higher-order multimodal alignment in joint representation learning across multiple modalities.
IMAgent: open-source visual agent trained with end-to-end RL for multi-image reasoning tasks, addressing limitations of single-image VLM agents.
FedVideoMAE: federated learning framework for privacy-preserving video moderation using self-supervised representations and differential privacy.
Open-source image generation model with improved reasoning for logic-intensive instruction following, closing gap to closed-source systems.
Multi-agent framework automating full computational catalysis research lifecycle from conception to publication.
Equilibrium propagation method for optimizing compound AI systems with multiple modules in long-horizon agentic workflows.
Framework using influence functions to craft training data perturbations inducing targeted model behavior changes.
Research on uncertainty quantification for ML interatomic potentials using evidential deep learning.