ANSE method for selecting optimal initial noise in video diffusion models using Bayesian active selection and attention mechanisms to improve quality and prompt alignment.
JALMBench: adversarial benchmark and dataset for evaluating jailbreak vulnerabilities in audio language models with unified evaluation framework.
Training-free AI agent framework using DeepSeek LLM for 3D parametric CAD generation without fine-tuning, enabling flexible and efficient CAD design automation.
Research characterizing how LLMs rely on pattern-matching for compositional tasks and analyzing OOD generalization failures through behavioral studies with controlled task setups.
Study on augmenting LLM research ideation with data at multiple stages to improve feasibility and quality of generated ideas.
RefTool framework enabling LLMs to automatically create task-specific tools using reference data for knowledge-intensive reasoning.
Diffusion model with equivariance regularization for solving inverse problems in image restoration.
AReaL large-scale asynchronous RL system for training LLMs on reasoning tasks, enabling efficient distributed parallelization for online RL with language models.
Search techniques for imperfect-information games without common knowledge, achieving superhuman performance on Fog of War chess requiring information reasoning.
Framework for pure exploration using in-context learning, enabling efficient adaptive data collection for hypothesis testing and best-arm identification tasks.
Protap benchmark systematically comparing protein model architectures, pretraining strategies, and domain-specific designs across realistic applications.
Tru-POMDP planner combining LLMs for belief generation with open-ended POMDPs for home-robot task planning under ambiguous instructions and hidden object locations.
OmniSpatial benchmark evaluating vision-language models on comprehensive spatial reasoning tasks beyond basic relations like left/right and object counting.
RoboPARA is an LLM-driven framework for dual-arm robot task planning with parallel execution and task recomposition optimization.
Foundation model approach for RL using pre-trained occupancy flow models for intention-conditioned planning, enabling transfer learning in reinforcement learning.
InterSyn dataset of 1.8M high-quality multimodal samples for training models on interleaved image-text generation with comprehensive evaluation metrics.
Analysis of reward design principles for multi-agent RL teams, studying when heterogeneous specialist teams outperform homogeneous teams in cooperative task allocation.
AQUA framework for watermarking image knowledge in multimodal RAG systems to protect copyright of contributed data in service-oriented RAG platforms.
VITA uses test-time adaptation of vision-language models as zero-shot value functions for reinforcement learning, improving generalization and temporal reasoning.
Language agents with RL for clinical decision-making, enabling LLMs to support hypothesis-driven interactive diagnosis and treatment planning in dynamic clinical settings.
SPARE framework for efficient automated step-wise annotation of LLM reasoning processes using reference-guided evaluation, enabling process supervision and reward modeling.
Sparse attention mechanism for transformers to improve long-context generalization by preventing attention dispersion on non-informative tokens in extended sequences.
Framework combining partial supervision with RL for sequence generation tasks, reducing reliance on expert demonstrations while addressing sparse reward challenges in learning.
LongWriter-Zero uses reinforcement learning to improve ultra-long text generation in LLMs, addressing quality degradation at extended sequence lengths without relying on synthetic supervised fine-tuning.
Method to interpret value trade-offs in language models using cognitive science frameworks.
Study on classifier-free guidance scale optimization in diffusion models for image generation.
Research on fine-tuning diffusion models for reward-guided biomolecular design using iterative distillation.
Proposes Partial Model Collapse (PMC) method for machine unlearning in LLMs without requiring unlearning targets in objective.
Generates synthetic multi-table EHR time-series from latent space with minimal preprocessing while preserving temporal relationships.
SeC: Concept-driven video object segmentation using Large Vision-Language Models to construct semantic priors across frames.
Model Predictive Adversarial Imitation Learning unifying inverse reinforcement learning with planning for ambiguous demonstration data.
FMIP: Generative model using joint continuous-integer flow to learn ML-based heuristics for Mixed-Integer Linear Programming.
Causal representation learning method using delta embeddings to learn robust representations from interventional image pairs.
PiKV: Distributed KV cache management system reducing memory/communication overhead in Mixture-of-Experts LLM inference.
FROGENT: Multi-agent system for end-to-end drug discovery automating fragmented AI tools across web, desktop, and code interfaces.
Proposes biologically-grounded learning paradigm for spiking neural networks that optimizes synaptic weights and intrinsic neuronal parameters jointly.
MOON: Generative multimodal LLM for e-commerce product understanding using many-to-one alignment between images and text.
Proposes Next Visual Granularity (NVG) framework that generates images through structured sequences with varying token granularity levels.
Investigates how sparsity in Mixture-of-Experts models affects memorization vs reasoning capabilities, establishing new scaling laws for MoE architectures.
Novel training method (DQO) for post-training LLMs that enhances output diversity while maintaining performance on downstream tasks using reinforcement learning.
BinaryShield enables privacy-preserving threat intelligence sharing across LLM services to detect prompt injection attacks.
ERIS framework uses genetic algorithms to generate real-world interference patterns for adversarial testing of audio language models.
Flow matching for learning robotic policies with non-uniform time scheduling to improve multi-step inference performance.
Steering Vector Decoding adapts LLMs to downstream tasks via output-distribution alignment during decoding without weight updates.
GraphUniverse framework generates synthetic graph families to evaluate inductive generalization in graph learning models.
Rubric-based reward modeling for LLM post-training focusing on distinguishing high-quality responses to prevent reward hacking.
Group Tree Optimization improves speculative decoding for LLMs by aligning draft model training with tree-based verification during inference.
Analysis of N:M activation sparsity techniques for efficient LLM inference through post-training pruning methods.
HEAPr uses Hessian-based pruning for fine-grained expert pruning in MoE LLMs to reduce memory while maintaining accuracy.
Quantile Advantage Estimation stabilizes RL training for LLM reasoning by replacing mean baseline with quantile-based approach to prevent entropy collapse/explosion.