Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning
Unified framework viewing LLM post-training methods (SFT, preference optimization, RL, distillation) through off-policy and on-policy learning perspectives.
Unified framework viewing LLM post-training methods (SFT, preference optimization, RL, distillation) through off-policy and on-policy learning perspectives.
Studies data mixing strategies for LLM training, questioning domain definitions, human-model alignment, and impact of domain weighting on generalization.
Discusses risks of LLM-generated peer reviews and automated editorial processes, proposing RAG-XAI detection framework for identifying machine-generated content.
DSCA method for lifelong vision-language model editing via dynamic subspace concept alignment, preventing degradation from sequential edits.
Decomposes long-context reasoning in LLMs into atomic skills, automatically identifying and improving fundamental capabilities for complex reasoning.
First comprehensive survey of abductive reasoning in LLMs, defining taxonomy and exploring inference of plausible explanations from observations.
PrivFedTalk privacy-aware federated framework for personalized talking-head generation using diffusion models with identity-stable adapters.
LINE uses LLMs iteratively to explain individual neuron concepts in vision models without predefined vocabularies, enabling interpretability of neural networks.
Training-free open-vocabulary semantic segmentation using global context awareness with pretrained vision and vision-language models without additional training.
Method for 2-bit LLM quantization via optimal codebook initialization, enabling extreme compression for edge deployment with O(1) lookup dequantization.
Tempo framework compresses long videos for multimodal LLMs by query-aware selection of frames, addressing context limits and lost-in-middle problems.
Diffusion model for virtual staining in histopathology that preserves cellular structures while translating immunohistochemistry images.
Study showing medical multimodal LLMs underperform traditional deep learning on image classification despite pretraining advantages.
RL primitive (Dataset Policy Gradient) optimizing synthetic data generators to produce targeted training examples for fine-tuning LLMs on differentiable metrics.
RL framework extending verifiable-reward training to general reasoning tasks in LLMs using natural instruction data for causal and temporal understanding.
Unified evaluation platform for prompt injection attacks and defenses, addressing benchmark gaps in comparing robustness across diverse tasks.
Study of language generation under differential privacy constraints, proving privacy can be achieved without qualitative cost for countable language collections.
Analysis of on-policy distillation failure modes in LLM training, identifying length inflation and truncation collapse as destabilizing factors.
Survey of tabular data generation comparing GANs, diffusion models, and LLMs across sample quality, privacy, and controllability dimensions.
Meta-learning approach with uncertainty quantification for limited-data task learning, addressing out-of-distribution scenarios in safety-critical settings.
Comprehensive survey of generative AI covering large language models, architectures, deployment protocols, and real-world applications as of early 2026.
Continuous online learning framework for activity recognition systems addressing model drift and domain shift in long-term deployments.
Privacy-preserving face recognition using information-theoretic Privacy Funnel model with end-to-end trainable representation learning.
Multi-agent LLM framework acting as generative adversarial network for synthesizing tabular data in low-data regimes, targeting healthcare applications.
Federated learning security method detecting backdoor attacks using data distribution inference to filter poisoned model updates.
Causal structure learning method for linear systems (LV-SEM-ME) addressing unobserved variables and measurement error simultaneously.
AdaProb: efficient machine unlearning method using adaptive probability to remove specific data with reduced computational overhead.
OpenGLT: comprehensive benchmark for evaluating graph neural networks on graph-level tasks across multiple domains.
Orthogonal representation learning method for estimating causal quantities from high-dimensional observational data with theoretical guarantees.
SALSA-RL: stability analysis framework for deep reinforcement learning agents to assess safe behavior in continuous action spaces.
Federated search mechanism for retrieval-augmented generation across distributed knowledge sources to reduce LLM hallucinations.
ShuffleGate: unified gating mechanism for feature selection, model compression, and importance estimation in recommender systems.
AEQ-RVAE-ST: recurrent variational autoencoder with progressive training for quasi-periodic time series generation.
Study of adversarial robustness in tabular foundation models like TabPFN and TabICL, examining test-time attacks and in-context defenses.
DiffGradCAM improves class activation maps for CNNs by addressing adversarial vulnerabilities in gradient-based explanation methods.
BTC-LLM: sub-1-bit quantization framework for LLMs using learnable transformations and binary codebooks for extreme compression.
Investigates Kolmogorov-Arnold Networks as interpretable alternatives to black-box models for clinical tabular data classification.
Neural diffusion method for estimating transfer entropy in time series addressing curse of dimensionality and convergence issues.
Applies XAI techniques (Grad-CAM, SHAP) to interpret PhaseNet deep neural network for microseismic event detection.
Generative model combining adversarial and flow-based families with native one-step/multi-step generation trained via adversarial objective.
Research on high-dimensional Bayesian optimization showing simple Bayesian linear regression outperforms complex BO methods after geometric transformation.
Case study applying LLMs to structured financial fraud detection data with focus on interpretability and feature analysis.
Tree-structured advantage redistribution method for group-based RL improving sample efficiency in LLM alignment on reasoning tasks.
Sample-efficient reinforcement learning algorithm for Value-at-Risk constrained optimization with safety guarantees during training.
Benchmark for evaluating multimodal LLMs on multi-criteria route planning reasoning tasks in heterogeneous graphs.
Extension of TabPFN foundation model to handle multimodal tabular data combining images, text and structured features.
Theoretical analysis explaining Adam optimizer's empirical advantages over SGD through second-moment normalization properties.
Theoretical analysis proving attention sinks are functionally necessary in softmax Transformers for trigger-conditional tasks.
vLLM Semantic Router architecture for optimizing LLM inference with routing mechanisms, semantic caching, and safety classification.
Computationally efficient classification algorithm with frequentist uncertainty bounds for safety-critical applications.