Research on how LLM internal representations correlate with brain activity, showing left-right hemisphere asymmetry during language processing.
MentalBench evaluates psychiatric diagnostic capability in LLMs using MentalKG knowledge graph encoding DSM-5 criteria for 23 psychiatric disorders.
BaziQA-Benchmark from 200 professional problems evaluating symbolic and temporally compositional reasoning in LLMs using structured inference.
Vietnamese Medical Code-Switching Speech Dataset addressing ASR challenges when English medical terms appear in Vietnamese sentences.
Large-scale Bengali idioms dataset with 10,361 idioms annotated across 19 semantic, syntactic, cultural fields for evaluating figurative language understanding in LLMs.
Curriculum learning and pseudo-labeling techniques improve multi-label Arabic Dialect Identification models trained on single-label datasets.
ProbeLLM automated framework for principled diagnosis of LLM failures with benchmark-agnostic test generation and principled exploration of model weaknesses.
SciAgentGym benchmark for evaluating multi-step scientific tool-use in LLM agents, featuring 1,780 domain-specific tools across four natural science disciplines.
Analysis of homogeneity in keyphrase prediction models, evaluating generative models' ability to predict absent keyphrases not explicitly in documents.
Meta-cognitive framework for reliable knowledge augmentation in LLMs addressing knowledge-confidence gaps that cause overconfident errors or uncertain truths.
Study investigating bias in AI models detecting healthy multilingual speakers among cognitively impaired cohorts using real conversational speech.
TraceBack multi-agent framework for fine-grained cell-level attribution in table QA systems, providing verifiable grounding for answer transparency.
LLM-based competency modeling process for HR that automates analysis of interview transcripts, improving reproducibility and reducing manual effort.
NLP approach to classify Estonian learner proficiency levels (A2-C1) using interpretable feature selection for explainable language assessment models.
SCOPE framework for selective pairwise LLM judging with finite-sample statistical guarantees, addressing miscalibration and systematic biases in LLM evaluators.
OpenLID-v3 extends language identification classifier with additional training data to improve precision on closely related languages and low-resource scenarios.
Statistical model of natural language multi-scale structure relating to entropy rate, providing insights into redundancy and patterns that LLMs capture.
Study on fusing DNA foundation models with LLMs via early-stage integration rather than late-stage embedding alignment for improved DNA-language reasoning.
GT-HarmBench benchmark evaluates AI safety risks in multi-agent game-theoretic scenarios with 2,009 high-stakes situations, addressing coordination failures and conflicts.
Proposes Entity State Tuning for temporal knowledge graph forecasting using stateful entity representations to preserve long-term dependencies.
Proposes Context-Conditioned Delta Steering using sparse autoencoders for jailbreak mitigation through inference-time feature steering.
Proposes constraint-rectified training to optimize chain-of-thought reasoning paths, reducing inference costs while maintaining reasoning quality.
Proposes DiffuRank using diffusion language models for document reranking, improving efficiency over autoregressive LLM-based rerankers.
Decoder-only Conformer for ASR using modality-aware sparse MoE with separate expert pools for speech and text processing.
Proposes attention-driven self-compression for reducing vision tokens in MLLMs through learned token pruning compatible with efficient attention mechanisms.
Proposes adaptive cognitive depth framework for LLM agents enabling variable reasoning intensity based on task demands in multi-turn decision-making.
Proposes VimRAG, a multimodal memory graph framework for RAG systems handling long-context visual reasoning in agentic scenarios.
Introduces RADAR, an evaluation framework for diagnosing performance bottlenecks in MLLM pre-training by analyzing asymmetric ability development.
Proposes RGAlign-Rec for intent prediction in e-commerce chatbots using LLMs to align discrete user features with semantic intents in recommendation systems.
Proposes fine-grained MLLM-based evaluation framework for image editing models with improved interpretability over traditional metrics.
Proposes hierarchical RL method for learning adaptive temperature policies from LLM internal states to balance exploration-exploitation in RL-based training.
Introduces MeSP for memory-efficient on-device LLM fine-tuning through manually-derived backward passes, balancing gradient accuracy and memory constraints.
Proposes LCSB, a layer-cyclic selective backpropagation method enabling memory-efficient LLM fine-tuning on mobile devices by computing gradients for subset of layers per step.
Addresses LLM unlearning robustness under quantization, proposing low-rank adaptation methods to preserve unlearning updates in 4-bit quantized models.
Proposes CoPE-VideoLM using codec primitives to reduce computational overhead in video language models while maintaining temporal coverage beyond sparse keyframe sampling.
arXiv paper: CATP token pruning for multimodal model inference using cross-attention, achieving 12.1X accuracy improvement.
arXiv memoir: NLP foundations covering morpheme-based annotation for Korean and system evaluation methods.
arXiv paper: RAISE method for reinforced adaptive instruction selection during LLM fine-tuning.
arXiv paper: PReSS framework evaluating political stance stability in LLMs under argumentative pressure.
arXiv paper: LLM-powered embodied agents with personalization via memory utilization for object-rearrangement tasks.
arXiv paper: preventing jailbreaking and model hijacking in LLMs using RAG without compromising security.
PsyCrisis-Bench: Reference-free evaluation benchmark for safety alignment of LLMs in Chinese mental health crisis dialogues.
MLLM-CTBench: Benchmark for continual instruction tuning of multimodal LLMs with reasoning process diagnosis across six domains.
ToolACE-MT: Non-autoregressive method for generating multi-turn agentic LLM interactions with function calls and tool use.
First parallel corpus of five Romansh language varieties extracted from comparable schoolbooks using automatic alignment.
TASO: Task-aligned sparse optimization method reducing parameter redundancy in LoRA-based parameter-efficient fine-tuning.
HEART framework using emotional cues to guide test-time scaling in LLMs, alternating critical and encouraging tones to improve reasoning.
Continual pre-training with low-rank adaptation for dialect-specific LLM adaptation in low-resource French dialects.
Framework for generating multiple interpretation-answer pairs for ambiguous requests in LLMs using RL with customized reward functions.
SGM: White-box neuron-level safety intervention for multimodal LLMs using detoxification to remove toxic and NSFW signals.