Synaptic Activation and Dual Liquid Dynamics for Interpretable Bio-Inspired Models
Framework analyzing bio-inspired RNN models with chemical synapses and synaptic activation to improve interpretability of recurrent neural networks.
Framework analyzing bio-inspired RNN models with chemical synapses and synaptic activation to improve interpretability of recurrent neural networks.
FedHENet extends federated learning to image classification using fixed feature extractors, reducing privacy risks and computational costs in heterogeneous environments.
Curriculum-DPO++ applies curriculum learning to direct preference optimization for improved text-to-image generation model training.
AdaGrad variant using cumulative squared gradient differences instead of norms for adaptive stepsize learning.
Method leveraging low-rank structures and compressible parameter dynamics to reduce computational costs in overparameterized neural network learning.
LTSM-Bundle: toolbox and benchmark for using large language models as universal time series forecasting models on heterogeneous datasets.
Framework for incorporating physical priors into diffusion-based generative models to generate physically feasible dynamics.
Analysis of the modality gap in contrastive multimodal learning (CLIP-like models) and methods to explain and reduce representation misalignment between modalities.
B3C: minimalist offline multi-agent reinforcement learning approach using behavior cloning regularization to address overestimation in action selection.
MINJA attack: adversarial method for injecting malicious data into LLM agent memory banks via query-only interactions without direct memory access.
Theoretical analysis of identifiability and singularity in polynomial neural networks and their neuromanifolds.
Framework for systematic design and analysis of weight quantization formats for efficient deep learning model training and deployment.
Transformer model with dynamic structure learning for multivariate time series forecasting using causal relationships instead of all-to-all connections.
Research on learning verifiers for chain-of-thought reasoning in LLMs using formal verification to improve mathematical problem-solving.
N², a unified Python package for nearest neighbor-based matrix completion with theoretical guarantees and empirical benchmarks.
PeakWeather dataset of Swiss weather station measurements for training spatiotemporal deep learning models for weather forecasting.
Framework for continual learning that handles concept drift in real-world data streams while preventing catastrophic forgetting.
Research on improving AI vision robustness by training on human developmental visual patterns to reduce reliance on texture and increase shape recognition.
Foundation model for analyzing industrial signals from SCADA systems across multiple modalities for anomaly detection.
Instruction-based diffusion model for editing time series properties while preserving specified conditions, replacing rigid attribute vectors.
Federated learning approach using generative models to address data heterogeneity across distributed clients and improve model generalization.
Self-evolving LLM that autonomously generates training data and improves reasoning without relying on human-curated tasks or labels.
Semantic caching system for reducing LLM inference costs by retrieving cached responses based on query similarity rather than exact matches.
Hyperdimensional refinement method for LLM-generated reasoning graphs in video anomaly detection handling distribution-deficient structures.
Fast approximate softmax attention mechanism clustering queries and keys for 36% faster transformer pretraining on long sequences.
Online reinforcement learning framework using sparse Gaussian mixture model Q-functions with interpretable policy iteration.
Adaptive time series foundation model with parameter-efficient design handling temporal heterogeneity and varying sampling rates.
Vector diffusion wavelets integrated into geometric graph neural networks for point cloud and manifold data representation.
Statistical set-level inference framework to identify training data in large language models with controlled error rates for legal evidence.
Reinforcement learning method to fine-tune generative models avoiding decision boundary mode-seeking when paired with safety classifiers.
Weight decay regularization has greater impact than maximal update parameterization for transferring learning rates across model scales.
Active learning approach using task-driven representations to select relevant samples from uncurated, messy data pools.
Large language models used as in-context meta-learners to recommend model families and hyperparameters from dataset metadata.
Methods to detect if specific audio data was used in training generative audio models through membership and dataset inference attacks.
Graph neural network framework for handling heterophilic graphs without iterative message passing using multi-resolution community features.
Imitation learning framework for combinatorial optimization under uncertainty. Explores role of expert demonstrations in training policies for sequential decision problems.
Study of spectral ghost phenomenon in self-supervised learning. Analyzes representation learning using unlabeled data for downstream task transfer.
RL-based preference optimization for generative recommenders. Proposes SAGE to address symmetric conservatism failure in list-wise ranking with multi-objective feedback.
Gauss-Newton natural gradient descent method for shape learning addressing ill-conditioning in implicit neural surfaces.
Pareto-Conditioned Diffusion framework for offline multi-objective optimization using conditional sampling.
Method for erasing harmful concepts from diffusion models while maintaining robustness and generation quality.
tLoRA framework for efficient multi-LoRA training on frozen LLM backbones with elastic shared super-models.
Learnable Chernoff Baselines for efficient inference-time alignment of generative models using reward guidance.
Thermodynamic framework analyzing transformer attention through Lagrangian mechanics and entropy minimization.
LLaDA2.1 text diffusion improvement combining token-to-token and mask-to-token editing for faster generation.
Formalization of attention-constrained inference for screening and verifying candidates under limited review capacity.
Analysis showing transformer learning dynamics collapse onto low-dimensional manifolds despite high parameter dimensionality.
Multilingual embedding models using contrastive learning on diffusion-pretrained backbone for web-scale retrieval.
HiFloat4 block floating-point format for efficient LLM inference, achieving 4.5 bits per value with three-level scaling.
Online tensor inference method for real-time processing of sequentially arriving high-dimensional data with statistical capabilities.