MultiEmo-Bench: Multi-label Visual Emotion Analysis for Multi-modal Large Language Models
Multi-label visual emotion analysis benchmark for evaluating multimodal large language models on image emotion prediction tasks.
Multi-label visual emotion analysis benchmark for evaluating multimodal large language models on image emotion prediction tasks.
AI agent framework (Automat) for automated design of material descriptors through iterative proposal, implementation, and evaluation for materials science.
Comparing neural machine translation and glossary-augmented LLM approaches for translating specialized rock art terminology in cultural heritage documents.
Benchmarking foundation models on EEG and brain-computer interface tasks with standardized datasets and metrics for clinical relevance evaluation.
arXiv paper on SceneFunRI: vision-language model benchmark for reasoning about occluded objects using spatial reasoning and context inference.
arXiv paper on IntentVLA: vision-language model for robot manipulation handling multimodal demonstrations and intent disambiguation.
arXiv paper investigating task-aware layer pruning for LLMs; shows pruning improves out-of-distribution performance while maintaining in-distribution accuracy.
arXiv paper on governance failures in LLM-based financial systems; proposes metrics for auditable decision-making compliance beyond task accuracy.
arXiv paper on Video2GUI: automated framework synthesizing large-scale GUI interaction data for pretraining generalized GUI agents using video.
arXiv paper extending intervention methods for LLMs beyond linear approaches to capture non-linear feature representations in model internals.
arXiv paper on EVA: defense against LLM jailbreaks through editing techniques to improve model safety alignment without computational overhead.
Knowledge distillation approach for identifying student misconceptions in educational settings with noisy labels.
Theoretical analysis of compositional sparsity as inductive bias enabling deep networks to overcome curse of dimensionality.
SpeechLLM system for streaming speech-to-text translation in real-time without waiting for complete utterances.
Data selection scheduling method that dynamically adjusts training data volume throughout model training.
Security analysis showing LLM browser agents can be fingerprinted through UI interaction traces and timing patterns.
LLM-based research idea generation using citation evolution graphs as structural supervision signal.
Autonomous AI agents for scientific discovery in cosmology using LLM-guided code evolution and multi-agent research labs.
Parameter-efficient fine-tuning method improving upon LoRA through isometric global parameter partitioning.
Dynamic weight quantization method for efficient LLM inference with adaptive codebook sizing and quality targets.
Multi-agent framework integrating operational plan generation and verification for complex battlefield planning.
Evaluation of whether coding agents understand least-privilege authorization principles for safe deployment.
Multi-agent AI system with recursion-of-thought for root cause localization in microservice systems.
Multi-step reasoning approach for text-to-image generation with closed-loop verification to handle complex semantics.
Analysis of noise in CLIP vision-language model embeddings using spectral decomposition of covariance matrices.
Distillation method for converting deep RL policies to interpretable surrogate models using Voronoi quantization.
MHSA: lightweight framework for mitigating hallucinations in vision-language models via steered attention mechanisms.
Viverra: text-to-code system generating verifiable, correct code with guarantees, reducing developer review burden.
MicroscopyMatching: framework for automated microscopy image analysis across diverse biological and imaging conditions.
Second-order actor-critic methods for discounted MDPs using policy Hessian decomposition for accelerated convergence.
Study quantifying premature closure in frontier LLMs—inappropriate commitment under uncertainty—with mitigation strategies.
Method for improving sample efficiency in RLVR by using randomly selected few-shot guidance for LLM chain-of-thought tasks.
COTCAgent: LLM-based clinical decision support using probabilistic chain-of-thought reasoning on longitudinal EHR data.
Generalized Priority-Aware Shapley Value: extension of Shapley value for valuation on arbitrary weighted priority graphs.
SemaTune: LLM-based system for online OS tuning that models cross-knob policy structure for long-running service optimization.
WARD: defense framework against prompt injection attacks on web agents, improving robustness across unseen domains and attack patterns.
Study of linguistic adaptation in multi-agent LLM systems under social observation, examining LLM behavior as communicative actors.
SpeakerLLM: audio-specialized LLM for speaker understanding, verification, and reasoning in audio-first agents and conversational robots.
TFGN: architectural method for continual pre-training of LLMs without replay buffers or task labels, addressing catastrophic forgetting at scale.
Survey of training algorithms for spiking neural networks with taxonomy and benchmarking framework.
Research identifying cultural anachronism in Vision-Language Models interpreting historical artifacts with temporally inappropriate concepts.
AsyncFC: Execution framework enabling concurrent function calling for LLM agents by decoupling decoding from function execution, reducing latency.
ML-Embed: Suite of inclusive multilingual text embedding models addressing computational costs, linguistic coverage, and transparency barriers in open-source framework.
DBS-Adam optimizer for deep learning on imbalanced/sequential datasets, applied to vehicular accident injury severity prediction.
Method improving LLM multi-turn dialogue consistency using self-recall thinking to track non-adjacent turn dependencies and reduce context bottlenecks.
Research on designing logging policies to minimize off-policy evaluation error for treatment policies like recommender systems.
CLOVER: Closed-loop value estimation and ranking for end-to-end autonomous driving planning with training-evaluation mismatch resolution.
Study on international students using conversational AI (ChatGPT, Gemini) for cross-cultural adaptation support.
Security research on adversarial attacks exploiting LLM quantization via outlier injection, showing quantized models can exhibit malicious behavior.
Pelican-Unified 1.0: Embodied foundation model using single VLM for unified understanding, reasoning, imagination, and action in shared semantic space.