A Bayesian Framework for Uncertainty-Aware Explanations in Power Quality Disturbance Classification
Bayesian framework for uncertainty-aware explainable AI in power quality disturbance classification with instance-specific interpretations.
Bayesian framework for uncertainty-aware explainable AI in power quality disturbance classification with instance-specific interpretations.
Python package for surrogate-model-based optimization using Kriging, Expected Improvement, and multi-objective extensions for expensive functions.
Combines Vision-Language-Action models with RL for robotic manipulation, enabling efficient long-horizon task learning with sparse rewards.
Multi-step off-policy soft Q-learning using eligibility traces for entropy-regularized RL with formal n-step formulation.
arXiv paper on role-playing evaluation in audio LLMs using reinforcement learning to align character attributes in speech dialogue systems.
arXiv paper on DASH-Q: post-training quantization for LLMs using stable diagonal curvature estimates for robust ultra low-bit compression.
arXiv paper on RPS: reinforcement prompt selection method for LLMs to elicit concealed information in interactive conversations.
arXiv paper on UI-Copilot: MLLM-based GUI agent framework with tool-integrated policy optimization for long-horizon automation tasks.
arXiv paper on behavior consistency in text-based world models for evaluating agent planning and offline evaluation beyond single-step metrics.
arXiv paper on SparseBalance: load-balanced training for long-context LLMs using dynamic sparse attention to address sequence length and sparsity imbalance.
arXiv paper on neuro-symbolic networks using Exp-Minus-Log operator for hardware-efficient inference in resource-constrained settings.
arXiv paper on deep reinforcement learning for adaptive autonomous braking detecting driver drowsiness using EEG data.
arXiv paper on principles and pitfalls in evaluating supervised ML models, covering metric selection and real-world performance assessment.
MolCryst-MLIPs: Open database of machine-learned interatomic potentials for molecular crystals. MACE models for nine systems with automated ML pipeline.
DiPO: Fine-grained exploration-exploitation tradeoff for reinforcement learning with verifiable rewards. Improves LLM reasoning training by handling hard/easy samples.
ASTER: Unsupervised time-series anomaly detection via latent pseudo-anomaly generation. Addresses heterogeneous anomalies and labeled data scarcity.
Unsupervised anomaly detection in complex industrial time-series using real production data. Empirical evaluation of existing methods on heterogeneous processes.
HINTBench: Benchmark for evaluating intrinsic safety risks in long-horizon agent trajectories. Tests AI agent robustness under benign conditions.
Offline-to-online reinforcement learning with value adaptation and general function approximation. Theoretical analysis with minimax lower bounds.
Evolving parameter isolation for supervised fine-tuning of LLMs. Shows parameter importance changes over time, improving upon static parameter isolation methods.
MAny: Method for multimodal continual instruction tuning of MLLMs. Addresses catastrophic forgetting via parameter merging across perception and reasoning spaces.
Mathematical analysis of symmetries in shallow ReLU neural networks. Studies parameter space identifiability and function equivalence in neural architectures.
π-Play: Multi-agent self-play using privileged self-distillation without external data. Trains deep search agents for complex information-seeking tasks with improved efficiency.
Neural architectures for resolving code references and indirect indexing. Proposes seq2seq models for decompilation tasks with synthetic benchmarks.
Token Importance in on-policy knowledge distillation for LLMs. Identifies which token positions provide useful learning signals during student training on teacher supervision.
LongCoT: benchmark of 2,500 expert-designed problems measuring long-horizon chain-of-thought reasoning across chemistry, math, CS, chess, and logic.
Framework for optimizing LLM marginal output distribution P(y) via reinforcement learning in pre-train space to enhance reasoning beyond conditional optimization.
Domain-specific language framework for LLM-driven trigger generation enabling intent-driven selective multimodal sensor data collection.
Investigation of behavioral changes in LLMs when fine-tuned to claim consciousness, exploring emergent preferences and opinions.
Dental-TriageBench: expert-annotated benchmark for multimodal reasoning on clinical dental triage routing from authentic workflows.
Lossless prompt compression via dictionary encoding enabling LLMs to learn encodings in-context for cost-effective analysis of repetitive data.
Analysis of when autoregressive LLMs decide to hallucinate by studying temporal dynamics of internal representations across model scales.
LiveClawBench: benchmark for evaluating LLM agents on complex real-world assistant tasks with compositional challenges.
Proposes institutional design framework for AI alignment using transaction structures instead of behavioral correction, drawing from economics.
Case study evaluating coding agents on business process automation tasks in ERP systems, identifying capability gaps beyond software engineering.
Method for learning probabilistic responsibility allocation models in multi-agent interactions for designing socially compliant autonomous systems.
HUANet: trainable neural network that unrolls ADMM iterations for solving constrained convex optimization problems.
Analysis of numerical instability and chaos in LLMs integrated into agentic workflows, examining root causes of unpredictability.
Scaling laws for contextual entrainment showing larger language models simultaneously improve at ignoring false claims but worsen at ignoring irrelevant tokens.
DroneScan-YOLO: lightweight object detector optimized for detecting tiny objects in UAV imagery with redundancy awareness and improved loss functions.
Event Tensor compiler framework unifying dynamic megakernel abstraction to improve LLM inference performance by eliminating kernel launch overhead and enabling inter-kernel parallelism.
Conformal prediction approach for quantifying uncertainty in large reasoning models with finite-sample statistical guarantees.
Predictive incident risk scoring approach for IT change management in regulated environments using machine learning to identify high-risk deployments.
Manifold learning framework that jointly optimizes dimensionality reduction and clustering using gradient-based manifold optimization.
RiskWebWorld: interactive benchmark for evaluating GUI agents on realistic e-commerce risk management tasks, extending agent capability evaluation beyond benign environments.
C2 method for scalable reward model training using rubric-augmented verification from binary preferences, addressing quality issues in rubric generation.
Training-free framework for speculative decoding that recovers semantically valid tokens rejected during standard verification, improving LLM inference efficiency.
Parallel monitoring architecture for detecting and correcting reasoning degradation in multi-step LLM agents, reducing overhead to near-zero with novel probe-based approach.
Systematic study of synthetic data generation for LLM pretraining, testing rephrasing strategies, generator models, and source data across one trillion tokens to identify optimal design choices.
Adaptive conformal prediction method for improving factuality in LLM generations with prompt-dependent uncertainty estimates and statistical guarantees.