Switchable Activation Networks
Switchable Activation Networks that dynamically select activation functions for computational efficiency in LLMs and vision-action models during inference.
Switchable Activation Networks that dynamically select activation functions for computational efficiency in LLMs and vision-action models during inference.
Method to align LLM confidence scores with correctness using output token probabilities for reliable error detection and hallucination identification.
LegoNet compression technique for neural networks using block weight clustering to reduce memory footprint for embedded device deployment.
CapTrack benchmark for evaluating multi-faceted forgetting in LLM post-training beyond parametric knowledge loss, addressing domain adaptation challenges.
Research showing majority-voting and ensemble inference methods fail to improve LLM truthfulness without external verification, unlike in math/code domains.
OptiRoulette meta-optimizer that dynamically selects update rules during training via warmup locking and random sampling. Torch-compatible drop-in component achieving 5.3x faster convergence.
RACER system for efficient multi-model LLM routing formulated as risk-aware optimization problem. Extends base routers to minimize cost-performance trade-off.
Novel language model combining autoregressive and diffusion-based generation through latent trajectory modeling with evolving balance parameter.
Framework for token-efficient reinforcement learning in LLMs. Proposes NAT to reduce computational cost of backpropagation over long chain-of-thought trajectories during training.
Research on vulnerabilities in Process Reward Models used in LLM reasoning pipelines. Introduces diagnostic framework to quantify adversarial exploitability and fluency-logic dissociation.
Empirical comparison of ARIMA, LSTM, BiLSTM, and Transformer for short-term power load forecasting.
Survey of Group Relative Policy Optimization for aligning generative models with human preferences.
Grouter: decoupled routing method for accelerating Mixture-of-Experts training with structural priors.
Leakage-safe graph feature extraction for fraud detection in temporal transaction networks.
Graph property inference in small language models: effects of representation on structured reasoning.
SmartBench: evaluation benchmark for LLMs in smart home environments with anomalous device detection.
HEARTS: benchmark for evaluating LLM reasoning on diverse health time series tasks and modalities.
SR-TTT: test-time training for LLMs with infinite context via fast weights, improving long-context reasoning.
Trust-aware federated learning framework for bone healing classification in distributed medical environments.
Ensemble learning framework for financial risk detection in ERP systems with leakage-safe evaluation.
ATLAS: reinforcement finetuning framework for scaling agent capabilities with large toolspaces using small language models.
Synthetic EHR generation pipeline with clinical consistency validation for privacy-preserving health data sharing.
ProtAlign: contrastive learning framework for protein sequence-structure alignment using language models.
Hybrid approach combining time series foundation models with regression models for electricity price forecasting capturing temporal and cross-variate patterns.
Safe Transformer: Modular approach adding explicit safety bit as interpretable information bottleneck for controllable alignment in language models.
Orion: First open system enabling direct LLM training and inference on Apple Neural Engine (ANE) with compiler pipeline and on-device training support.
Reinforcement learning approach for safe neural navigation in variable-density crowds that generalizes beyond training conditions.
Super-resolution transformer optimization using FlashAttention with rank-factorized implicit neural bias to reduce computational burden.
Efficient decentralized framework for training diffusion models with heterogeneous objectives across isolated experts with reduced computational requirements.
Constrained generation framework for diffusion models handling complex feasible regions in robotics and autonomous driving applications.
Method to stabilize Group Relative Policy Optimization (GRPO) for diffusion language models by addressing reward collapse issues in post-training.
DIRECTER: Dynamic rejection steering technique for LLMs to improve instruction following while avoiding oversteering that degrades output quality.
xaitimesynth: Reusable Python package for evaluating time series attribution methods using synthetic ground truth data with class-discriminating features.
OPR: Lightweight mechanism preventing premature policy convergence in deep RL by maintaining buffer of high-performing episodes during optimization.
NEST: Device placement optimization for distributed deep learning that jointly considers parallelism, memory, and network topology without post-hoc feasibility fixes.
Framework for cooperative multi-agent RL with submodular rewards modeling overlapping agent contributions. First formal analysis of diminishing returns in team coordination.
C3 algorithm for multi-agent reinforcement learning with LLMs. Addresses credit assignment problem in sparse reward scenarios through contextual counterfactual analysis.
Gaussian Linear Unit activation function analyzing mathematical relationships among modern transformer alternatives to ReLU.
Stochastic attention mechanism via Langevin sampling on Hopfield energy enabling temperature-controlled retrieval and generation.
Physics-informed surrogate model for ferroelectric NAND retention analysis replacing expensive TCAD simulations.
LLM-based CAD program generation using design procedures and geometric constraints for parametric model synthesis.
Generative models for tabular data using XGBoost as score estimator with denoising diffusion for small and large datasets.
Eigenspectral framework analyzing information flow in LLM feed-forward networks through lightweight spectral metrics.
Mixture-of-experts approach for state space models with expert specialization while maintaining computational efficiency.
Physics-consistent neural networks for Cosserat elasticity modeling deformation and director fields in microstructured materials.
Reinforcement learning formalism for coupled-dynamics environments specifying joint distributions across counterfactual actions.
Study of graph sparsification impact on GNN pipeline performance and scalability for billion-node graphs.
Reinforcement learning approach for chart comprehension in vision-language models using verifiable rewards for symbolic reasoning.
Imitation learning analysis for quadruped locomotion showing effectiveness in small data regimes via limit cycle structure.
Optimal transport framework for conditional generative modeling robust to outliers using unbalanced transport.