DiPS: Dialogue Policy Selection for High-Stakes Persuasion Agents
DiPS: Q-learning framework for dialogue policy selection in high-stakes persuasion agents, dynamically adapting strategies based on personality.
DiPS: Q-learning framework for dialogue policy selection in high-stakes persuasion agents, dynamically adapting strategies based on personality.
ADVENT: LLM-driven predicate invention mechanism for Inductive Logic Programming combining LLM abduction with Prolog deductive verification.
VLAFlow: Unified flow-matching framework for training vision-language-action models enabling controlled comparison of VLA objectives with robot manipulation datasets.
Comprehensive benchmark for LLM-based data agents automating data science workflows. Evaluates agent capabilities on heterogeneous datasets.
Probabilistic framework for merging task-specific models into multitask solutions. Scores statistical utility of parameter updates across tasks.
Diagnostic framework for predicting closed-loop performance of world models in model-based RL without explicit validation metrics.
Bayesian reinforcement learning method addressing data scarcity through prior knowledge and belief updates in sequential decision-making.
Open-vocabulary object detection calibration using frozen VLMs. Vision-language model application with limited novelty.
Adaptive expert pruning and growing for efficient MoE fine-tuning using LoRA. Parameter-efficient training for large models.
Expander sparse autoencoders for mechanistic interpretability with reduced parameters. ML research on interpretability and efficient dictionaries.
Causal study of AI coding agent adoption on open-source projects, analyzing impact on newcomer participation. Developer tools and OSS ecosystem.
Continuously evolving multimodal benchmark using multi-agent pipeline for VLM evaluation. Developer tools and evaluation frameworks.
Memory-efficient training stack combining parallelism techniques for MoE models. ML research on scalable training infrastructure.
Evaluation of chunking strategies for RAG systems on academic texts using RAGAS framework. Directly relevant to LLM applications and retrieval techniques.
Empirical study of LLM-generated code and comments in real repositories, analyzing prevalence and quality concerns. Developer tools and LLM applications.
Binarization technique for vision-language models reducing memory/latency for deployment. Model optimization for efficient inference.
Reward-free reinforcement learning from video using VLM as progress scorer and GRPO objective. AI agents and novel training approach.
Brain disease diagnosis framework integrating LLM semantics with brain connectivity analysis via hypergraphs. LLM application with healthcare focus.
Pipeline for adapting Qwen 27B model to perform reasoning in Turkish rather than English, addressing multilingual LLM reasoning.
Dataset of 1,639 K-12 science explanations with human and LLM-generated alternatives for training risk assessment auditors.
CausalSTeward agentic divide-conquer-combine system for causal discovery integrating prior knowledge to identify causal models from high-dimensional data.
PhysMani framework combines physics-principled 3D Gaussian world model with action policy for dynamic object manipulation in embodied AI.
Conditional co-ablation technique reveals transformer self-repair mechanisms where dormant backups activate after primary component ablation.
Analyzes representational geometry showing LLMs become robust to science skepticism through problematic mechanisms rather than genuine understanding.
Object Aligner provides configurable JSON schema similarity scoring for measuring LLM output structure alignment in tool calling and agentic systems.
Evaluates Vision-Language Model reliability for medical image quality assessment under corruption and bias conditions.
MolSight vision-language model combines molecular LLMs with graph-aware visual understanding for molecular structure and drug discovery tasks.
Controlled study comparing nine lightweight CNN architectures across multiple datasets and hardware to assess efficiency claims.
Proposes load-aware prefill deflection technique to improve disaggregated LLM serving efficiency by balancing prefill and decode GPU pools.
OpenSafeIntent benchmark evaluates whether LLMs calibrate assistance appropriately across benign, dual-use, and malicious intent variants.
SPLIT benchmark evaluates LLM cross-lingual empathy and cultural grounding in emotional-support contexts across English and Ukrainian.
Demonstrates performance evaluation failures in spatiotemporally correlated domains due to data leakage from non-i.i.d. splits.
Proposes prompt coverage adequacy testing framework to guide LLM and autonomous agent testing when prompts replace traditional code.
kNNGuard presents training-free guardrail for LLMs using activation space of off-the-shelf models to detect unsafe/adversarial prompts with minimal labeled data.
Introduces emotional self-correction mechanism for vision-language models to improve reasoning reliability without post-training or engineered feedback.
Proposes test-time guidance framework for vision-language-action policies using learned critic to guide flow-matching inference without retraining base models.
Presents vLLM-based inference pipeline for unified audio understanding and generation in speech language models with multi-token prediction support.
Develops behavioral monitoring techniques to detect and analyze guardrail activations in LLMs, enabling black-box security testing of production AI systems.
Reviews AI risk assessment and management methodologies under EU AI Act and other regulatory frameworks, covering identification, analysis, and mitigation approaches.
Systematically categorizes 53 human-AI team studies into five clusters using psychological teaming taxonomies to understand collaboration patterns and diversity.
Presents open infrastructure for operationalizing AI audits by moving beyond taxonomies to actionable tests. Reviews 74 existing risk taxonomies and proposes executable evaluation framework.
Proposes CoFL-S framework for vision-language navigation using language-conditioned flow fields for low-level robot action generation and trajectory planning.
Analyzes limitations of LLM-as-a-Judge evaluation paradigm for multilingual and low-resource language tasks, showing proficiency gaps and validation challenges.
Proposes HERMES, a hierarchical multi-granularity labeling system for pre-training data mixtures to improve flexibility in corpus partitioning and semantic organization.
Introduces specialized benchmark for evaluating vision-language models on rare concepts and complex spatio-temporal video grounding beyond general datasets.
Research on security vulnerabilities in LLM-based agent skill marketplaces where benign skills can interact unexpectedly. Addresses fuzzing skill composition to discover implicit intent attacks.
Systematic study of AI agents patching compiler missed optimizations showing generalization beyond single cases is the key challenge.
Offline-first Android assistant application for people with visual impairment using multimodal models with personalized object retrieval.
ACID method uses inverse dynamics to verify trajectory realizability during decision-time planning with world models for embodied control.
Neuron-aware active learning method for LLMs identifies valuable unlabeled samples for annotation reducing human labeling costs in few-shot adaptation.