Token-Importance Guided Direct Preference Optimization
Token-Importance Guided DPO: Enhanced direct preference optimization for LLM alignment using token-level importance weighting beyond standard DPO.
Token-Importance Guided DPO: Enhanced direct preference optimization for LLM alignment using token-level importance weighting beyond standard DPO.
Adaptive location hierarchy learning method for long-tailed mobility prediction addressing biased predictions in human location forecasting.
CityLens: Benchmark for evaluating large vision-language models' ability to predict socioeconomic indicators from satellite and street view imagery.
VPI-Bench: Security research on visual prompt injection attacks against computer-use agents with full system access.
FAuNO: Federated asynchronous reinforcement learning framework for task offloading in decentralized edge computing systems.
Control Tax: Framework for measuring operational and financial costs of integrating AI control mechanisms into agentic AI pipelines for high-stakes applications.
Novel preference learning framework for aligning aggregate opinions proportionally with population distribution rather than majority preference.
Study of spiking neural networks' accuracy-efficiency trade-offs using Lempel-Ziv complexity to compare unsupervised, supervised, and hybrid learning paradigms.
Meta-Adaptive Prompt Distillation for few-shot visual question answering in large multimodal models using optimized in-context learning.
Study of generative agents for modeling consumer behavior in energy operations, examining their role in operational decision-making and uncertainty handling.
Dual-level framework for controlling behavioral diversity in multi-agent reinforcement learning systems accounting for agent composition structures.
SPIRAL: Self-play reinforcement learning framework where language models play zero-sum games against themselves to develop reasoning without human-curated data.
SpiroLLM: Fine-tuned LLM for analyzing spirogram time series to predict COPD respiratory disease with clinical validation.
Collab-REC: Multi-agent LLM framework with three specialized agents negotiating tourism recommendations to balance personalization, popularity, and sustainability.
Re4: LLM-based agent framework for scientific computing using rewriting-resolution-review-revision chain for complex mathematical and scientific reasoning tasks.
EigenBench: Black-box method for benchmarking language models' value alignment using constitutional scoring without access to model internals.
DeepMedix-R1 foundation model for chest X-ray interpretation generates diagnoses with step-by-step reasoning process for clinical explainability.
AIssistant: open-source agentic framework for human-AI collaborative generation of scientific review and perspective papers, combining autonomous workflows with human guidance.
Framework teaches multimodal agents to reliably interact with GUIs by identifying toggle controls, addressing key bottleneck in automated interface manipulation.
SciTrek benchmark probes long-context numerical reasoning in LLMs using full-text scientific articles, testing counting, sorting, aggregating, and comparing across documents.
Research shows bilinear representation in neural networks mitigates the reversal curse and enables consistent model editing, revealing knowledge encoding artifacts rather than fundamental limitations.
EHR-ChatQA benchmark evaluates LLM-powered agents for Electronic Health Record database access, addressing query ambiguity and terminology mismatch in clinical workflows.
ViTSP framework uses vision language models to solve large-scale Traveling Salesman Problems with improved generalization and scalability over traditional learning-based approaches.
Study analyzing answer attribution in large reasoning models, showing inconsistencies between chain-of-thought reasoning and final outputs.
G-reasoner foundation model enables unified reasoning over graph-structured knowledge, combining LLM capabilities with structured knowledge integration.
Framework for detecting and correcting harmful content in chain-of-thought reasoning of large reasoning models via corrective interventions.
TRACE method detects implicit reward hacking in reasoning models by measuring reasoning effort via truncated CoT analysis.
Analysis of training data conditions enabling test-time scaling in LLMs; examines when long chains-of-thought emerge during training.
FaithCoT-Bench benchmarks instance-level faithfulness of chain-of-thought reasoning to assess reliability in high-risk LLM applications.
Doctor-R1 agent uses experiential agentic reinforcement learning to master both clinical decision-making and strategic patient inquiry skills.
DRPO optimization technique reduces computational cost of reasoning models by mitigating overthinking and unnecessary lengthy reasoning chains.
Systematic investigation of bias in LLM-as-a-judge systems across 6 models; evaluates fairness issues in autonomous content quality assessment.
HardcoreLogic benchmark tests large reasoning models on non-canonical logic puzzle variants to evaluate flexible rule application and generalization.
ScholarEval: retrieval-augmented evaluation framework assessing research ideas on soundness and contribution using literature grounding.
DAG-Math models chain-of-thought reasoning as rule-based processes on directed acyclic graphs to analyze LLM mathematical problem-solving mechanisms.
Theoretical analysis of representation convergence in brains and neural networks through compression efficiency principle framework.
FM Agent: multi-agent framework combining LLM reasoning with large-scale evolutionary search for complex scientific and engineering discovery tasks.
Survey of language model capabilities and limitations for addressing cognitive science research challenges in integration, formalization, and conceptual clarity.
Hierarchical multi-agent system for medical triage addressing specialization, department heterogeneity, and real-time responsiveness in healthcare AI.
Analysis of adaptive reasoning in LLMs showing models apply uniform strategies regardless of task difficulty; proposes methods for adaptive compute allocation.
Knowledge graph-guided chain-of-thought framework improves disease prediction on EHR data by scaffolding reasoning with structured medical knowledge paths.
OVERTONBENCH framework measures viewpoint diversity in LLM outputs through set coverage metrics, validated against 1208-person human study across 8 models.
Multi-agent reinforcement learning framework handling lossy communication via generalized communication-constrained priors for scalable cooperative policy learning.
Introspection study on Llama-3.1 using activation steering to test whether LLMs can detect perturbations to internal states; finds artifact in binary detection paradigm.
AgentMath: Agent framework integrating LLM reasoning with code interpreters for accurate mathematical problem-solving, addressing computational limitations of reasoning models.
AWARE-US: Tool-calling agent system that resolves database query infeasibility by preference-aware repair, handling underspecified and infeasible queries.
Lossless Hierarchical Speculative Decoding: Improves LLM inference speed via hierarchical verification approach that overcomes joint intractability in token acceptance.
PsyAgent: Framework for building human-like AI agents using Big Five personality traits and social context conditioning to enable stable dispositions with adaptive behavior.
DRAGON combines LLM-guided decomposition and reconstruction agents for large-scale combinatorial optimization with improved scalability beyond 30 nodes.
Zero-shot LLM-guided framework for dynamic voltage/frequency scaling and core allocation on embedded systems using hierarchical multi-agent RL.