Pareto-Conditioned Diffusion Models for Offline Multi-Objective Optimization
Pareto-Conditioned Diffusion framework for offline multi-objective optimization via conditional sampling from diffusion models on static datasets.
Pareto-Conditioned Diffusion framework for offline multi-objective optimization via conditional sampling from diffusion models on static datasets.
Analysis of optimization instability in deep neural networks caused by parametric singularities, proposing gradient Frobenius norm solutions.
Study explores how LLMs support healthcare professionals in analyzing patient-generated health data from wearables and smartphones for cardiac risk reduction.
AEGIS method for concept erasure from diffusion models balances robustness against concept reactivation with retention of model utility under adversarial conditions.
Framework establishes first scaling laws for LLMs in recommendation systems using principled synthetic data to replace noisy user interaction data in pre-training.
Learnable Chernoff Baselines enable efficient inference-time reward-guided alignment for generative models without architecture modifications or computational overhead.
Bielik Guard provides efficient safety classifiers for Polish-language LLM content moderation, offering 0.1B and 0.5B parameter model variants based on RoBERTa.
GISA benchmark evaluates information-seeking AI agents performing multi-turn web interactions, addressing limitations of existing benchmarks with naturally-constructed real-world tasks.
LLaDA2.1 improves text diffusion model decoding speed by combining Token-to-Token and Mask-to-Token editing schemes for 100B-parameter block-diffusion models.
Deep learning model with explainability for predicting open source software project sustainability from temporal contribution patterns.
Method for synthesizing diverse post-training data for LLMs using Feature Activation Coverage metric based on model representations.
Investigation of learning dynamics in transformers showing training trajectories collapse onto low-dimensional manifolds in modular arithmetic tasks.
Security analysis of mobile LLM agents identifying vulnerabilities in screen-interface paradigm and proposing intent-centric architecture.
Block floating-point data format for efficient LLM inference achieving 4.5 bits per value with hierarchical scaling.
Behavioral study examining user interaction with three LLM assistance modalities (Advisor, Coach, Delegate) in multi-party negotiation scenarios.
Lightweight 5B parameter multimodal model for image generation and editing competitive with much larger models.
Framework for synthesizing and optimizing CUDA kernels using LLMs combined with systematic exploration to achieve competitive hardware performance.
Method for identifying queries that cause LLM character specification violations using red-teaming approaches to detect deployment-level failures efficiently.
Research on adaptively merging multiple LoRA modules from open models to improve performance on downstream tasks.
Exploration method in reinforcement learning using ensemble error estimates to compute optimistic value bonuses for directed agent exploration.
Addresses LLM personalization by generating synthetic user-specific interaction data at scale to optimize prompts for individual user preferences and constraints.
Applies deep reinforcement learning to automate analog and mixed-signal circuit design optimization across diverse non-differentiable design spaces.
Studies soft contamination in LLM training data through semantic duplicates, showing typical decontamination filters fail to detect near-equivalent benchmark test data.
Demonstrates stable training of LLMs from scratch using exclusively low-rank weight factorization, matching dense model performance while reducing computational costs.
Safe reinforcement learning framework using recovery-based shielding with Gaussian process models for non-linear continuous control with provable safety guarantees.
Training-free guidance method enabling continuous diffusion language models to satisfy formal syntactic constraints like JSON schema matching via regular expressions.
Regularized meta-learning framework addressing redundancy and overfitting in deep ensemble methods through redundancy-aware projection and statistical weighting.
RNA sequence design reframed as conditional sequence generation task using language models instead of traditional optimization approaches.
Theoretical analysis of Mamba state space model training dynamics and generalization properties.
Analysis of robustness and reasoning consistency in vision language models fine-tuned with reinforcement learning.
Bench-MFG: standardized benchmark suite for mean field games and multi-agent reinforcement learning.
Multi-agent model-based RL with joint state-action learned embeddings for coordination in dynamic environments.
Constraint-rectified training method to reduce overhead and overthinking in chain-of-thought reasoning.
Flow-Factory: unified modular framework for reinforcement learning with flow-matching and diffusion models.
Adaptive steering method to balance modality preferences in multimodal large language models.
Concept-grounded transparent domain adaptation for clinical event prediction on electronic health records.
Fractional-order federated learning for battery electric vehicle energy consumption modeling.
Verifier-independent RL method for LLM reasoning without external verifiers using confidence-guided variance reduction.
Analysis of catastrophic forgetting in mixture-of-experts transformers with multi-head attention.
Causal ODE networks for explainable anomaly detection and root cause analysis in power grids.
RelBench v2 benchmark for relational deep learning on database-like data at scale.
Temporal graph neural networks optimized for continuous prediction on dynamic graphs.
Federated PCA with personalization and manifold optimization for anomaly detection in IoT networks.
Dual-granularity contrastive reward framework using generated episodic guidance for sample-efficient embodied reinforcement learning.
SLA2 improves sparse-linear attention for diffusion models with learnable routing and quantization-aware training.
Knowledge distillation via calibrated uncertainty preserves dark knowledge from teachers trained with uncertainty-aware losses.
Split-MoPE addresses sample misalignment in vertical federated learning using mixture of experts specialized for partial data.
ADEPT framework combines speech LLMs with evidence probing tools via RL-aligned agentic decoding for interpretable emotion reasoning.
Control-theoretic framework addressing stability issues in LLM-based time series forecasting through closed-loop feedback instead of naive autoregressive generation.
Placer applies message passing neural networks to network routing, improving explainability of ML-based telemetry-aware routing decisions.