FIRE: Frobenius-Isometry Reinitialization for Balancing the Stability-Plasticity Tradeoff
arXiv research on FIRE reinitialization method balancing stability and plasticity in neural networks trained on nonstationary data.
arXiv research on FIRE reinitialization method balancing stability and plasticity in neural networks trained on nonstationary data.
arXiv paper on Online Causal Kalman Filtering for stabilizing policy optimization in LLM reinforcement learning with high-variance importance sampling.
arXiv research on SWE-MiniSandbox: container-free reinforcement learning method for training software engineering agents at scale.
arXiv paper on MiniCPM-SALA: hybrid sparse and linear attention mechanism for 9B LLM enabling efficient ultra-long context modeling.
arXiv research on low-bit floating-point formats (HiFloat) for efficient LLM inference on Ascend NPUs with rigorous performance evaluation.
Framework for designing generative social robots in education, addressing hallucinations and safety concerns of LLM-powered tutoring systems.
Elo-Evolve framework for LLM alignment using co-evolutionary multi-agent competition instead of static reward functions, improving training stability.
UniWeTok: unified binary tokenizer for multimodal LLMs enabling high-fidelity visual reconstruction and semantic extraction with 2^128 codebook size.
Research on iterative refinement in Transformer architectures using inner loop inference to improve model capabilities without additional training.
Research paper on structural misalignment in Transformers between residual connections (tied to current token) and causal supervision (targets next token).
HIMM proposes non-parametric memory framework for embodied agents using MLLMs, disentangling episodic and semantic memory for long-horizon exploration.
Study evaluates LLMs as zero-shot annotators for Bangla hate speech detection, examining bias and reliability in low-resource identity-sensitive annotation tasks.
Graph Meta-Network learns on Kolmogorov-Arnold networks using weight-space models to predict neural network accuracy on new datasets.
Stable Asynchrony proposes variance-controlled off-policy RL for LLM post-training that handles stale rollouts and heavy-tailed importance weights in asynchronous training.
Agentic Unlearning removes sensitive information from LLM agent parameters and persistent memory in closed-loop systems, addressing parameter-memory backflow issues.
Framework measuring model propensities and behavioral tendencies alongside capabilities using Item Response Theory, relevant for safety evaluation.
Dynamic sample pruning method for spatio-temporal model training to reduce computational overhead on redundant large-scale datasets.
Survey of LLM integration with UAV systems for environmental understanding, swarm coordination, and task planning applications.
Architecture using stateful recurrent attention for full-piece symbolic music modeling with constrained computational resources.
Fine-tuning approach enhancing vision-language models by aligning structural/edge-based cues across modalities for improved cross-modal retrieval.
Dataset compression framework using color quantization to reduce image dataset storage demands while preserving model training effectiveness.
Foundation model for urban spatio-temporal forecasting addressing scenario-specificity in city-scale prediction tasks across multiple regions.
Pipeline for low-resource speech-to-text translation with punctuation restoration to mitigate ASR structural noise in Nepali-English systems.
Approach using LLM-generated textual relevance judgments to augment behavioral signals in large-scale app store search ranking systems.
Training paradigm using pseudo-contrastive learning to improve diagram comprehension and fine-grained visual understanding in multimodal models.
Analysis of transformer training dynamics under AdamW, identifying low-dimensional drift directions that capture 60-80% of parameter displacement.
System translating jailbreak papers into executable modules via multi-agent workflow for reproducible benchmarking and evaluation of LLM robustness.
Diffusion model approach for time series forecasting using learnable spectral trajectory schedules to improve noise handling and structure recovery.
Method for improving LLM-as-a-judge evaluation by accounting for correlated errors and latent confounders in ensemble aggregation mechanisms.
Research on 4-bit quantization-aware training for attention mechanisms to enable FP4 computation on emerging GPUs, addressing challenges with dynamic range and heavy-tailed activations.
Addresses extreme LLM compression by identifying spectral energy gain in binary quantization and fixing latent geometry misalignment issues.
Novel RL approach for control systems with probabilistic stability guarantees derived from Lyapunov methods using finite trajectory samples.
Property-driven evaluation framework for assessing GNN expressiveness using formal specification and systematic testing across large-scale datasets.
Proposes method to overcome factorization barrier in diffusion language models, enabling faster parallel token generation without sacrificing coherence.
REMIND addresses multi-modal learning with missing modalities in medical AI, treating missingness as a long-tailed distribution problem.
BiJEPA extends JEPA with bi-directional prediction for self-supervised learning, leveraging both forward and inverse relationships in representation learning.
Knowledge-guided generative surrogate modeling combines data-driven models with domain expert knowledge for design optimization under data scarcity.
Expert Divergence Learning addresses expert homogenization in MoE language models by encouraging functional specialization through auxiliary loss during pre-training.
M3-AD is a multimodal framework for industrial anomaly detection using MLLMs with reflection-aware self-correction mechanisms for fine-grained scenarios.
Representation-consistent gated recurrent framework for medical time-series classification with irregular sampling and missing values.
Certainty-Validity Framework: Diagnostic method decomposing model performance for discrete commitment systems beyond standard accuracy metrics.
SEval-NAS: Metric-evaluation framework for neural architecture search enabling hardware-aware objectives independent of search strategy.
LIDS: Method for evaluating LLM-generated summaries using BERT-based inference across model layers.
MAML-KT: Few-shot meta-learning approach for knowledge tracing addressing cold-start problem with new students.
LLM-augmented rebalancing system for shared micromobility using policy optimization and reinforcement learning for dynamic urban demand.
Engineering privacy-preserving generative AI applications for personalized healthcare using transformers while maintaining data privacy.
Task-driven subspace decomposition method for LoRA-based continual learning to enable knowledge sharing while preventing task interference.
Diagnostic methods for detecting instability in individual-level predictions from overparameterized ML models in healthcare.
Language model trained on 5.8M EHRs from 1.8M patients for automated medical coding to standardized ICD-10 codes.
CoPeP: Benchmark for continual pretraining of protein language models on dynamically updated biological databases.