Position: Zeroth-Order Optimization in Deep Learning Is Underexplored, Not Underpowered
Position paper argues zeroth-order optimization in deep learning is underexplored rather than fundamentally limited, reviewing variance and complexity issues.
Position paper argues zeroth-order optimization in deep learning is underexplored rather than fundamentally limited, reviewing variance and complexity issues.
IO-SVD compresses LLMs via input-output whitened SVD with adaptive rank selection for hardware-agnostic post-training reduction of model size and latency.
Perforated backpropagation applied to keyword spotting enables simultaneous improvements in edge model accuracy and size under strict memory constraints.
Proposes using off-the-shelf LM embeddings of PyTorch code as low-cost feature extractors for neural architecture search surrogate models.
Theoretical analysis of sample complexity for epsilon-best arm identification in linear bandits with adaptive vs. non-adaptive algorithms.
Rule2DRC benchmarks LLM agents for synthesizing design rule checking scripts from natural language with execution-guided test generation.
Interaction-aware influence functions extend training example attribution to groups, capturing redundancy and complementarity beyond summed individual influences.
FRWKV+ improves frequency-space time series forecasting with adaptive periodic-position branch interaction for long-term predictions.
SEED formulates data selection as weighted independent set problem to identify compact, high-quality, diverse subsets from training corpora.
Theoretical analysis of regret bounds for episodic reinforcement learning with context-dependent action sets using MVP algorithm extension.
CATS framework enables distributed transformer inference across multiple ultra-low-power IoT devices for collaborative model execution.
AGOP-IxG proposes a fast gradient-based feature attribution method for tabular data with controlled benchmark for evaluating explanation fidelity.
Differentiable Mixture-of-Agents framework enables adaptive multi-agent LLM systems with self-evolving communication topologies.
Unified perturbation framework analyzing robustness and manipulation vulnerability of LLM evaluation leaderboards using influence-based methods.
Continual learning method that learns domain-invariant representations to prevent catastrophic forgetting across domains.
Analysis of grokking in transformers as structural inference through Bayesian lottery tickets perspective on generalization delay.
Framework for learning context-conditioned Gaussian bounds for uncertainty quantification in neural network predictions.
AOT-POT pre-training method for neural operators handling multi-PDE datasets through adaptive operator transformation.
Martingale neural operators recover stochastic structure and variance in neural operator surrogates for stochastic PDEs.
Shapley Neuron Valuation framework quantifies neuron importance in continual learning using cooperative game theory.
Heterogeneous graph prompt learning method with structure-conditioned experts for cross-domain scenarios.
Framework using diffusion geometry to compare and analyze neural representations across layers and networks.
LoCO: parameter-efficient fine-tuning method for foundation models using low-rank compositional orthogonal updates while preserving geometric structure.
Research on adversarial training for physics-informed neural networks using neural tangent kernel theory to understand training mechanisms.
Framework unifying representation learning under competing constraints for temporal, multimodal, and partially observed systems.
CT-AGD: curvature-tuned accelerated gradient descent optimization for faster convergence in deep learning.
Variational autoregressive networks with probability priors for efficient Monte Carlo sampling with reduced critical slowing.
Looped SSMs: depth-recurrence architecture for time series classification matching performance of larger models.
Ada-Diffuser: latent-aware adaptive diffusion for decision-making modeling hidden environment dynamics.
SAFE-stabilized variational quantum classifier with amplitude encoding and learnable classical pre-encoding layer.
ITGPT: generative pretraining for irregular timeseries with missing values using transformer architectures.
MIND framework addresses model-induced label noise via latent manifold disentanglement for foundation model annotations.
MolCHG: multi-level self-supervised pretraining framework on hierarchical molecular graphs for property prediction.
Performance analysis comparing centralized and decentralized federated learning architectures for IoT and edge devices.
Federated learning approach for heterogeneous feature spaces where clients observe only partial feature overlaps.
Analysis of attention dispersion failure mode in dynamic graph transformers under temporal distribution shift with proposed fix.
Cascade refinement framework using multi-fidelity flow matching for parametric PDE solutions.
Neural architecture search package optimizing for FPGA-specific hardware costs including LUTs, DSPs, and latency.
Efficient RL training for vision-language-action policies using probabilistic chunk masking to reduce gradient computation cost.
Bayesian reinforcement learning method for non-stationary continuous control with adaptive robustness.
Imitation learning approach for clinical decision support in pediatric ECMO using trajectory-based action modeling.
Study of hidden bias in instruction-tuned LLMs showing fair outputs mask biased internal representations affecting high-stakes decisions.
Unified data mixing method for language models across pretraining, continual learning, and adaptation phases.
Authorization framework for autonomous AI agents using proof-derived policies to prevent unsafe actions in sovereign systems.
Kubeflow MLOps integration for adversarial robustness in AI models deployed on Kubernetes.
Transformer-based unified particle simulator for diverse physical phenomena without solver-specific redesign.
SMCEvolve applies Sequential Monte Carlo sampling to LLM-driven program search for automated scientific discovery.
Minerva-Ego benchmark for evaluating egocentric video reasoning with intermediate reasoning step annotations.
Belief Engine for auditable stance dynamics in multi-agent LLM deliberation with inspectable evidence-based belief updates.
Comprehensive analysis of neural activation patterns across six LLM architectures on cognitive tasks.