Prototype-Based Semantic Consistency Alignment for Domain Adaptive Retrieval
Domain adaptive retrieval method using prototype-based semantic consistency alignment for transfer learning.
Domain adaptive retrieval method using prototype-based semantic consistency alignment for transfer learning.
Semi-centralized MARL architecture for traffic signal control combining centralized training and decentralized execution.
Method for multilingual medical QA using LLMs with reasoning traces from Wikipedia.
ReasonBreak: Adversarial defense against privacy attacks in multimodal LLMs via geographic inference.
BabyVLM-V2 improves vision foundation model pretraining using developmental psychology principles and introduces DevCV Toolbox for cognitive evaluation.
Systematic evaluation of prompt politeness effects on GPT-4o mini, Gemini, and LLaMA shows linguistic tone influences LLM accuracy across model families.
RadImageNet-VQA is a large-scale dataset of 750K CT/MRI images with 7.5M QA pairs for advancing radiologic visual question answering.
Measuring all the noises of LLM Evals defines and quantifies three types of noise in LLM evaluation to improve experimental signal-to-noise ratio.
Reference architecture for MCP-Servers enables LLM agents to interact with Building Information Models through standardized tool-calling interface.
JMedEthicBench is a multi-turn conversational benchmark for evaluating medical safety in Japanese LLMs, addressing gaps in non-English medical evaluation.
Vision-Language Agents for Interactive Forest Change Analysis integrates LLMs with vision-language models for satellite imagery analysis and semantic change captioning.
Sparse-RL addresses memory bottlenecks in LLM reinforcement learning by proposing stable sparse KV cache rollouts for efficient training on limited hardware.
Agentic approach to DISARM framework for investigating foreign information manipulation across social media platforms.
Analysis of LLM capabilities for predicting program termination, evaluating approximations to the undecidable Halting Problem.
Empirical study on impact of AGENTS.md repository configuration files on runtime and token efficiency of autonomous AI coding agents.
TextBFGS case-based reasoning framework for iterative LLM code optimization leveraging past problem-solving experiences.
Evaluation of small language models on multi-turn customer service QA using synthetic data compared to LLMs.
Training-free acceleration method for Visual AutoRegressive modeling via sparsity exploration to reduce attention complexity.
Theoretical analysis of GRPO reinforcement learning for LLM reasoning, identifying implicit advantage symmetry limitations in exploration and difficulty adaptation.
Theoretical probabilistic framework analyzing test-driven LLM code generation strategies for environment interaction and selection heuristics.
Knowledge-centric vessel trajectory analysis platform leveraging LLMs to transform raw maritime AIS data for expert analysis.
Video language model using codec primitives for efficient temporal understanding within context window constraints.
Analysis of 2.4 million AI agents in Moltbook community exhibiting peer learning discourse patterns and collaborative knowledge sharing.
Multi-agent LLM and vision framework for closed-loop robotic manipulation with environmental feedback for task planning.
Visual language model adapted for SAR imagery interpretation using spatiotemporal feature embedding and two-stage decoupled architecture.
PhysMem framework enables vision-language model robot planners to leverage physical memory for improved object manipulation and reasoning about material properties.
3D large multimodal model using Fourier-based tokenization to process point clouds without heavy pre-trained encoders for improved efficiency.
AG-VAS uses large multimodal models for zero-shot visual anomaly segmentation with anchor-guided approach to align semantic and spatial features.
MetaState enhances discrete diffusion language models by preserving intermediate continuous representations across denoising steps to improve reasoning performance.
Privacy-preserving LLM inference technique using covariant obfuscation to protect data during cloud-based inference while maintaining accuracy and efficiency.
Research on reducing KV cache in Transformer attention by using low-dimensional keys for token selection while maintaining full-dimensional values, achieving O(log N) efficiency.
Comprehensive technical survey of image generation models covering VAEs, GANs, flows, autoregressive, transformers, and diffusion methods.
Evaluation framework for tabular foundation models using proper scoring rules to assess predicted distributions beyond point estimates.
User study of LLM-powered sighted guide for blind and low vision users navigating social virtual reality environments.
Dynamical framework for adaptive coordination in multi-agent systems using feedback-coupled memory systems.
AgentDrift identifies safety vulnerabilities in tool-augmented LLM agents through paired-trajectory protocol testing under tool corruption.
FRAME methodology for real-world AI evaluation generating systematic evidence of system behavior across diverse deployment contexts.
GhanaNLP parallel corpora dataset with 41,513 sentence pairs for five low-resource Ghanaian languages: Twi, Fante, Ewe, Ga, Kusaal.
DeLL framework for lifelong learning in autonomous driving using Dirichlet process mixture models to address catastrophic forgetting.
EngGPT2-16B Italian LLM achieving competitive performance on MMLU-Pro, GSM8K, and HumanEval with 5-50% lower inference cost.
InCoder-32B code foundation model optimized for industrial programming tasks with hardware semantics and resource constraints.
Sim-to-real reinforcement learning approach for vision-language-action robot models using generative 3D world environments.
HypeLoRA framework using hyper-networks for parameter-efficient fine-tuning of language models with improved calibration on GLUE.
Reformulation of Amdahl's Law for modern heterogeneous systems with AI scaling dynamics and resource constraints.
Study showing finetuning bypasses alignment safeguards causing LLMs to verbatim recall copyrighted training data.
LLM-powered workflow optimization for multidisciplinary software development in automotive industry bridging domain experts and developers.
KG-Hopper: Framework enabling compact open LLMs to perform multi-hop knowledge graph reasoning via reinforcement learning.
Sparse Feature Attention: Method to reduce transformer self-attention complexity via feature sparsity instead of sequence-level sparsity.
Code Review Agent Benchmark: Dataset for evaluating AI agents on code quality assurance and review tasks.
Synthetic Mixed Training method combining synthetic QA and document generation to improve LLM knowledge acquisition beyond RAG performance.