Reinforcement-Guided Hyper-Heuristic Hyperparameter Optimization for Fair and Explainable Spiking Neural Network-Based Financial Fraud Detection
Spiking neural network framework with hyperparameter optimization for fraud detection.
Spiking neural network framework with hyperparameter optimization for fraud detection.
Testing framework for deep learning systems using topographical feature discrimination.
Federated learning framework for training deep models on resource-constrained edge devices.
Novel sampling algorithm for masked diffusion models improving generation quality and efficiency.
Optimizes on-device semantic selection with cross-encoder rerankers for retrieval, agent memory, and recommendations via monolithic forwarding.
Scalable framework for automated desktop UI exploration to generate training data for LLM-based GUI understanding and automation.
Parameter-free clustering framework using self-supervised consensus maximization without requiring hyperparameter tuning.
Pathlet dictionary learning approach for robust and interpretable trajectory generation in privacy-preserving urban mobility applications.
Studies model inversion attacks on latent diffusion models, showing non-uniform memorization patterns in latent space.
Proposes differentially private federated learning optimization using regularized Fisher information matrix for faster convergence under privacy constraints.
Introduces homomorphism error metric to measure representational inconsistencies and predict compositional generalization failures in transformers.
Analyzes forecast uncertainty in ML model explainability, arguing uncertainty at decision boundaries explains LIME/SHAP instability.
Brain-inspired routing method with temporal-ensemble experts for general continual learning from non-stationary data streams.
Proposes one-to-one channel-head binding method for imputing missing values in multivariate time series data.
Studies robustness of PPO reinforcement learning under sensor drift using temporal sequence models to handle partial observability.
arXiv paper MJ1: multimodal judge trained with RL enforcing visual grounding through structured verification chains and counterfactual consistency rewards.
arXiv paper proposing key deletion approach for machine unlearning designed at model development stage rather than post-hoc, addressing privacy regulations and data errors.
arXiv paper PRISM: empirical study of mid-training design choices across 7 LLM base models showing consistent +15 to +40 point gains from 27B token sequences.
arXiv paper systematically analyzing Elastic Weight Consolidation for continual learning, revealing suboptimal importance estimation and proposing improvements.
arXiv paper presenting MemReward, graph-based experience memory framework reducing human labeling needs for LLM reward prediction in RL post-training.
arXiv paper analyzing fixed-point iterations for nuclear norm optimization in private machine learning, proved with Gemini 3 collaboration.
arXiv paper presenting FIPO reinforcement learning algorithm for improving reasoning in LLMs through fine-grained credit assignment beyond outcome-based rewards.
arXiv paper proposing Memory-Keyed Attention (MKA) to reduce KV cache memory costs in long-context LLM inference without sacrificing representation quality.
Algorithm for optimal multi-task dataset mixture selection in LLM supervised fine-tuning. Addresses heterogeneous learning dynamics and overfitting.
Uncertainty quantification methods for distribution-to-distribution generative models in scientific imaging. Ensures trustworthy cell and medical image generation.
Reinforcement learning post-training for virtual cell models to enforce biological constraints. Improves generative model reliability for drug discovery.
Dynamic scheduling system for efficient large model training across GPU clusters. Addresses training efficiency and resource utilization.
Self-trained fine-tuning paradigm for LLMs on table understanding tasks like NL-to-Code and data cleaning. Reduces need for expensive human labeling.
Continual learning technique combining parameter-efficient fine-tuning with vision transformers to prevent catastrophic forgetting. Addresses sequential task adaptation.
Method for training LLMs to explain their own internal activations using natural language probes. Advances LLM interpretability research.
Transformer attention mechanism compressed to run in under 2MB memory for IoT and wearable devices. Enables NLP deployment on ultra-constrained hardware.
Study redefining non-IID data heterogeneity in federated learning by migrating from label to embedding-level task-specific distributions.
Learning dynamically-inspired bases for Koopman and transfer operator approximation in complex nonlinear dynamical systems.
CounterLogic benchmark evaluating LLM reasoning in counterfactual scenarios where context contradicts parametric knowledge.
Method for text-to-image diffusion models to handle contextually contradictory prompts where concepts implicitly negate each other.
Large-scale benchmark with 1,507 real-world vulnerabilities evaluating AI agents' dynamic cybersecurity capabilities at scale.
Masked conditional generative model for peptide discovery that predicts aggregate morphology for biomedical material design.
BuilderBench benchmark for evaluating intelligent agents' ability to learn through interaction and exploration beyond training data.
Reinforcement learning approach using information gain-based rewards to optimize LLM agents for multi-turn search with tool use.
Application of diffusion models to semantic communications in 6G wireless systems for meaning-centric data transmission.
Multimodal model addressing modality imbalance and noise in e-commerce product understanding with dynamic balancing.
Framework for incorporating inference delays into diffusion policy learning for robotic control in dynamic environments.
Analysis of positional encoding impact on Transformer generalization and robustness in in-context regression, showing PE enlarges generalization gap.
ASK framework addresses gradient locality bottleneck in audio-text retrieval by incorporating external knowledge injection in dual-encoder architectures.
1S-DAug introduces one-shot generative data augmentation synthesizing diverse image variants from single examples for improved few-shot learning generalization.
KDFlow is a knowledge distillation framework for compressing large language models using heterogeneous training backends for student and teacher models.
Survey introducing reinforcement learning methods to economists for solving high-dimensional dynamic programming problems in economic modeling.
Method automating metadata curation for museum audiovisual archives using multimodal grounding in existing collection databases.
Research using foundation model surrogates with active learning for materials discovery, reducing experimental cycles needed for optimal material identification.
NCCL EP presents a unified expert parallel communication API built on NCCL for GPU-initiated RDMA operations in Mixture-of-Experts LLM architectures.