TTQ: Activation-Aware Test-Time Quantization to Accelerate LLM Inference On The Fly
TTQ: test-time quantization framework for LLM inference acceleration without retraining, addressing domain shift in calibration.
TTQ: test-time quantization framework for LLM inference acceleration without retraining, addressing domain shift in calibration.
CLaRE quantifies representational entanglement to predict unintended side effects from LLM editing techniques.
STEU enables parameter-efficient machine unlearning for clinical language models through sparse token embedding editing.
MemReward uses graph-based experience memory for LLM reward prediction with limited labels, addressing expensive human annotation costs.
LeWorldModel: stable end-to-end Joint Embedding Predictive Architecture for learning world models from raw pixels without auxiliary supervision.
MSNet and LS-Net: scalable multi-scale convolutional architectures for time series classification with multi-representation inputs.
Theoretical framework connecting neural network learning with Ternary Gamma Semirings to improve compositional generalization.
OXRL framework evaluates 51 post-training alignment algorithms across model scales, revealing scale-dependent ranking differences among DPO, SimPO, KTO, GRPO.
Proposes DAPA, a hardware-friendly activation function for on-device Transformer inference and training optimized for energy efficiency.
Research on replacing weighted summation with learnable nonlinear aggregation functions in neural networks to improve robustness against noisy inputs.
Empirical analysis of transformer layer heterogeneity in SmolLM2-135M using diagnostic metrics including weight predictability and ablation degradation.
Mathematical framework for information absorption and understanding, relating signal clarity to learner structural capacity.
Flow matching approach with warm-start initialization for faster inference in autoregressive LLMs and diffusion models.
Hierarchical reinforcement learning for allocating scarce public health resources across multiple disease outbreak clusters.
Framework for interpreting LLM token representations using geometric learning and regularizers inspired by harmonic analysis.
Deep learning approximation methods for nonlinear PDEs on Hilbert spaces using neural operators with universal approximation guarantees.
Convergence analysis of multiplicative updates in regularized nuclear norm optimization for private machine learning.
Off-policy correction technique for LLM reinforcement learning addressing policy staleness and training-inference distribution mismatch.
Diffusion-based method for recovering dense trajectories from sparse GPS data in urban mobility applications.
Neural network architectures with flexible equivariance to arbitrary symmetries for diverse data modalities.
In-context learning approach for tabular anomaly detection across multiple supervision regimes without retraining.
Graph filtering for decision-making on expanding networks with unknown attachment patterns and evolving topology.
Federated learning approach for training large foundation models on distributed supercomputer facilities with privacy and data sovereignty constraints.
Kernel learning framework for tensor data with mode-wise subspace comparison for large-scale structured multi-way data.
Unified geometric framework explaining adversarial vulnerability in vision models and hallucinations in LLMs via uncertainty principle.
Foundation models for wearable devices moving beyond static encoders toward temporal reasoning for health monitoring.
Defense mechanism against model poisoning attacks in continual federated learning for indoor localization systems.
Theoretical analysis of in-context learning in LLMs, examining demonstration selection, chain-of-thought prompting, and adaptation without parameter updates.
Federated learning with personalized constraints for heterogeneous agents in distributed optimization problems.
Deep reinforcement learning with policy regularizations for inventory management, improving training stability and hyperparameter sensitivity.
Continual learning framework for food image classification enabling incremental category addition without retraining from scratch.
Proximal sampling method using zeroth-order function queries and simulated diffusion dynamics, avoiding score estimation requirements.
Finite-time convergence analysis of stochastic approximation under heavy-tailed and long-range dependent noise for optimization and reinforcement learning.
Ensemble-based extension of Feature Guided Analysis for improving recall of DNN behavior explanations beyond precision limitations of existing FGA methods.
Proves KV cache in transformer inference is mathematically redundant; keys and values are deterministic residual stream projections with zero reconstruction error.
Theoretical analysis of reverse diffusion sampling under Wasserstein geometry, identifying scale-dependent radial contraction properties.
GoAgent generates task-specific communication topologies for LLM-based multi-agent systems to optimize coordination and group-based problem-solving.
Theoretical analysis of multi-armed bandits with competing players and time-varying availability, extending regret bounds to sleeping competing bandits framework.
Binary classification method using weak pairwise labels (similarity/dissimilarity) instead of explicit instance-level labels through probabilistic supervision framework.
FedRG addresses noisy label robustness in federated learning by leveraging representation geometry rather than scalar loss values for sample identification.
FedPDPO framework applies Direct Preference Optimization to federated learning for LLM alignment with decentralized, non-IID preference data while preserving privacy.
Novel attribution method (DPA) for understanding transformer LLM internals through efficient layer-wise target propagation, addressing computational costs of dense component attribution.
Coreset construction method for scalable multivariate conditional transformation models and density estimation.
Theoretical framework analyzing two-time-scale learning dynamics in population-based neural network training methods.
FIPO is an RL algorithm improving token-level credit assignment in LLM reasoning tasks beyond outcome-based reward methods.
NASimJax is a GPU-accelerated framework for training RL policies on penetration testing with large action spaces.
Test-time RL method for LLMs using selective-complementary reinforcement learning to improve reasoning under weak consensus conditions.
Meta-learning approach integrating meta-features with knowledge graph embeddings for pipeline performance prediction and dataset similarity.
Memori is an LLM-agnostic persistent memory layer for context-aware AI agents supporting multi-session interactions without vendor lock-in.
Model-driven learning for physical layer authentication in wireless IoT devices using hybrid approach.