Confidence-Based Decoding is Provably Efficient for Diffusion Language Models
Analyzes efficient decoding strategies for diffusion language models as alternative to autoregressive models with parallel token generation.
Analyzes efficient decoding strategies for diffusion language models as alternative to autoregressive models with parallel token generation.
TiCo enables spoken dialogue models to generate responses with controllable duration for voice assistants and interactive agents.
3D-Layout-R1 uses scene-graph reasoning with LLMs/VLMs for text-conditioned spatial layout editing in images.
ThinkJEPA combines latent world models with vision-language models for improved long-horizon video forecasting and semantic understanding.
UniMotion: unified framework for simultaneous understanding and generation of human motion, language, and images with continuous latent representations.
UNITE: unified autoencoder enabling end-to-end training of tokenization and latent diffusion, eliminating staged training requirements for LDMs.
WorldCache: content-aware caching method accelerating video world model inference through feature reuse across denoising steps in Diffusion Transformers.
Privacy-preserving framework for clinical annotation extraction from EHRs using small-scale LLMs, balancing privacy regulations and computational efficiency.
Formula-R1: formula-driven RL framework improving LLM numerical reasoning over complex tabular data using spreadsheet formulas as executable reasoning interface.
SynPO: preference learning enhancement for vision-language models in fine-grained video captioning, addressing subtle dynamics and detailed content.
Multi-agent RL framework for reinsurance treaty bidding with autonomous learning-based agents improving efficiency over traditional broker-mediated placement.
CRAMF: conceptual retrieval-augmented LLM approach for automated theorem prover formalization, addressing hallucination and semantic gap challenges.
BuilderBench: benchmark for developing AI agents that learn through interaction and exploration beyond training data limits, addressing scalable learning mechanisms.
UniWM: unified memory-augmented world model for visual navigation integrating egocentric visual foresight and planning for embodied agent robustness.
DeepCompress: dual reward RL framework improving efficiency of Large Reasoning Models by balancing overthinking and underthinking while maintaining accuracy.
Study showing LLMs can accurately model correlational structure of human psychological traits from Big Five responses, predicting nine other psychological scales.
Self-report fine-tuning technique teaching LLMs to accurately report their hidden objectives, addressing model deception in safety testing.
RadHiera: semantic hierarchical RL framework for medical radiology report generation modeling dependencies between Findings and Impression sections.
LAMP framework integrating language into multi-agent reinforcement learning for economic decision-making under semantic ambiguity and contextual richness.
Comparative study of visual creativity between humans and Stable Diffusion image generation model using human and GPT-4o evaluation.
Evaluation framework testing LLM reasoning robustness on rule-based logic through four stress tests including rule deletion and contradiction injection.
MeG framework for large-scale knowledge editing in LLMs via dynamic weight generation, addressing reliability, generality, and locality challenges.
arXiv paper on probabilistic paradigm for AI benchmarking under uncertainty, addressing ground truth ambiguity in evaluation.
arXiv paper analyzing whether hierarchical reasoning models genuinely reason or pattern-match via mechanistic study of failure modes.
arXiv paper on embodied exploration benchmark with multimodal LLM-based RL framework leveraging long-term episodic memory.
arXiv paper on EvoOpt-LLM: using LLMs to evolve mixed-integer linear programming optimization models from natural language requirements.
arXiv paper on RE-MCDF: closed-loop multi-expert LLM reasoning for clinical diagnosis from noisy EMR data using multi-agent validation.
arXiv paper on architecting trust in epistemic agents: examining how LLMs function as knowledge curators and their reliability/calibration.
arXiv paper on S5-SHB agent: multi-model agentic blockchain framework for smart home governance within Society 5.0 vision.
arXiv paper on multimodal mathematical reasoning via unified perception-alignment-reasoning paradigm for visual+textual math problems.
arXiv paper on HECG framework for autonomous agents with LLM-based action generation using multi-dimensional error correction.
arXiv paper on FABRIC strategy for backward reachability analysis verification of neural feedback systems and dynamical systems.
arXiv paper on curveball steering: non-linear activation steering for LLM control, questioning linear representation hypothesis.
arXiv paper introducing UCIP framework to detect whether autonomous agents pursue self-preservation as objective vs instrumental strategy.
arXiv paper on using LLMs to construct representations for efficient supervised learning across multimodal data types via agentic pipeline.
Proposes agentic operating system using LLM agents with tool use and memory for clinical workflows, addressing safety and transparency challenges.
Evaluates state-of-the-art LLMs and prompting strategies for zero-shot vulnerability detection in Solidity smart contracts.
Studies representation alignment across independently trained large language models for cross-model applications in privacy-constrained settings.
Compares deep reinforcement learning algorithms for robust continuous control under partial observability in robot control tasks.
Presents policy optimization methods for reinforcement learning over general state and action spaces with convergence guarantees.
Introduces continual federated learning framework where distributed clients learn new tasks sequentially without storing historical data.
Reviews urban foundation models and their application to intelligent city services, covering machine learning and foundational model paradigm shifts.
Proposes HPE-CogVLM for head pose estimation using vision language models instead of CNN-based approaches for improved robustness in real-world scenarios.
Multi-scale Transformer foundation model for enhancing low-quality CT scan images with limited training data and improved generalization.
Training-free method to verify text provenance and detect LLM-generated content without model fine-tuning, addressing challenges from indistinguishable AI text.
Continual learning approach for vision-language models using general attribute descriptions to balance knowledge forgetting and new task adaptation.
Neural network pruning method using structured lasso with class-wise information for lightweight model development while preserving statistical information.
Simplified RLHF method using supervised fine-tuning approach to reduce complexity and computational cost of LLM alignment compared to PPO and GRPO.
Framework applying cognitive science methods and levels of analysis from neuroscience to understand and interpret large language models.
Multi-task learning approach using spiking neural networks with adaptive task-switching for resource-constrained autonomous agents in diverse environments.