Causal Direct Preference Optimization for Distributionally Robust Generative Recommendation
Causal Direct Preference Optimization method for training LLMs to generate recommendations while mitigating spurious correlations.
Causal Direct Preference Optimization method for training LLMs to generate recommendations while mitigating spurious correlations.
Graph RAG framework combining labeled property graphs and RDF for retrieval-augmented generation over structured and semi-structured data.
T-MAP uses evolutionary search to red-team LLM agents by exploiting multi-step tool execution vulnerabilities in MCP ecosystems.
Analysis of feature importance bias in gradient boosting models under multicollinearity, affecting SHAP-based explanations.
WIST framework uses web-grounded iterative self-play with reinforcement learning to improve LLM reasoning in specific domains.
Study on using LLMs for algorithm synthesis with provable guarantees, combining mathematical reasoning with practical performance.
Research on improving conditional modeling in diffusion models, establishing equivalence between classifier-free guidance and alignment objectives.
Proposes Reasoner-Executor-Synthesizer architecture for LLM agents that maintains O(1) context window while avoiding hallucination and token cost scaling.
Research evaluating Vision-Language Models' ability to detect misleading data visualizations and deceptive captions in charts.
FAAR quantization method for NVFP4 ultra-low-bit format that adapts rounding to non-uniform numerical grid for efficient LLM edge deployment.
Study on multimodal fusion strategies for time series forecasting showing naive fusion fails and proposing constrained fusion approach.
MTEO method for few-step diffusion sampling by distilling layer-wise, step-wise time embeddings to accelerate inference.
AI Co-Scientist framework combining LLM agents with cloud computing to automate search ranking research from ideation through GPU training.
Cross-task evaluation study of LoRA adapters showing nominal instruction-tuning labels don't reliably predict realized instruction-following capabilities.
Symbolic Graph Network framework for discovering partial differential equations from noisy sparse data without numerical differentiation.
Adaptive temporal control system for autonomous agents that learns optimal action intervals using hyperbolic geometry predictive signals.
Open-source framework (CaP-X) for benchmarking and improving code-as-policy agents for robot manipulation tasks.
Token-level analysis of distributional shifts in RLVR fine-tuning of LLMs to understand mechanisms underlying reasoning improvements.
LLM-guided headline rewriting system that enhances reader engagement while maintaining editorial integrity and avoiding clickbait.
Ablation study analyzing specialization patterns in hybrid language models combining attention with state space models on sub-1B parameter models.
Framework for building language model general capabilities via automatic curriculum of cross-entropy game tasks for relevant skill discovery.
Inference-time scaling method using small latent verifiers instead of multimodal LLMs to score and select outputs while reducing computational cost.
Empirical study measuring semantic novelty of 13,847 IS papers (2020-2025) to assess whether LLM productivity gains translate to genuine intellectual advancement.
LLMON proposes a markup language for LLMs that preserves structure and semantics in prompts, distinguishing between instructions and data in input/output.
ChatP&ID is an agentic RAG framework enabling LLM interaction with engineering P&ID diagrams using knowledge graphs for cost-effective grounded reasoning.
Ego2Web benchmark for multimodal web agents grounded in egocentric video, evaluating agents performing real-world workflows with physical context awareness.
STRIATUM-CTF is an agentic framework using search-based reasoning for automated cybersecurity CTF challenge solving with multi-step stateful reasoning.
Study evaluating faithfulness of chain-of-thought reasoning in LLMs, finding models often produce misleading explanations despite correct outputs.
flexvec is a SQL vector retrieval kernel exposing embedding matrices and score arrays for programmatic manipulation by AI agents via Programmatic Embedding Modulation.
Method leveraging vision-language models to explain sparse autoencoder features in vision models through causal interventions instead of correlation-based approaches.
Theoretical work on causal discovery in chain-reaction systems using interventional data, proving identifiability under cascade-like structural assumptions.
Evaluation of medical vision-language models revealing a grounding-sycophancy tradeoff, analyzing hallucination and agreement behaviors across six VLMs.
Benchmark for evaluating attribution map faithfulness in semantic segmentation models, testing intervention-based faithfulness and perturbation robustness.
LGSE framework for adapting pretrained language models to low-resource languages using morphologically grounded subword embeddings instead of arbitrary segmentation.
Study examining whether humans can learn to recalibrate AI confidence signals through repeated interaction, testing four calibration conditions with 200 participants.
AwesomeLit proposes an agent-supported literature research system for hypothesis generation, designed for inexperienced researchers to identify gaps and propose feasible research directions.
Population-representative resume dataset for causal fairness auditing of LLM/VLM-based screening systems.
Quantitative model predicting when post-hoc fusion of independent LLM specialists outperforms individual models.
Persona-based data augmentation framework using LLMs for legal domain information retrieval in low-resource settings.
Bayesian visualization interface supporting multi-issue human-AI negotiation to manage cognitive load.
Study of neural network resilience to hardware bit-flip errors, comparing logic-based vs arithmetic architectures.
LLM fine-tuning framework addressing knowledge-action gap in personalized e-commerce search at Taobao.
Embodied AI agent integrating multimodal LLMs with chain-of-thought reasoning for robotic photography tasks.
PinPoint method for identifying instruction-relevant image regions in VLMs to reduce computational overhead.
Study evaluating whether frontier LLMs genuinely use reasoning steps or generate decorative narratives post-hoc.
Security analysis framework for LLM agent deployments covering model, tool code, credentials, and MCP configurations.
Dual-View Pheromone Pathway Network (DPPN) architecture for persistent structural memory in neural networks. Identifies coordinate system requirements.
Agent-Sentry: Security system for bounding LLM agents via execution provenance tracking. Addresses safety and security concerns in agentic systems.
Empirical study of sim-to-real transfer for robotic dexterous manipulation using vision-language-action models. Addresses synthetic-to-real gap.
Shows that confidence calibration fails when annotator disagreement exists. Proposes calibration against annotator distribution rather than majority labels.