OPPO: Accelerating PPO-based RLHF via Pipeline Overlap
Proposes OPPO, a pipeline overlap technique to accelerate PPO-based RLHF training for aligning large language models with human preferences.
Proposes OPPO, a pipeline overlap technique to accelerate PPO-based RLHF training for aligning large language models with human preferences.
Demonstrates control-flow hijacking attacks in multi-agent systems and evaluates LlamaFirewall defense mechanisms for securing inter-agent communications.
LLEMA framework coupling LLM scientific knowledge with evolutionary algorithms and memory for multi-objective materials discovery.
Frequency-domain-guided compression method for multimodal LLM KV cache using outlier-awareness, compatible with FlashAttention.
Guided Flow Policy method for offline RL that distinguishes high-value actions through weighted behavior cloning with flow-matching policies.
RL post-training technique using mixed rewards and canonical action ordering to improve transformer performance on structured problem solving.
Post-training method that sparsifies transformer attention to ~0.4% of edges while maintaining performance for mechanistic interpretability studies.
Uses LLMs to automatically discover proxies for mixed-precision quantization without manual design, reducing training-free DNN optimization costs.
Proposes RePo, a method to improve in-context learning in LLMs by dynamically repositioning context based on Cognitive Load Theory principles.
Research on optimization strategies for scaling LLMs, introducing Spectral Sphere method to address weight drift limitations in μP and Muon optimizers.
ButterflyMoE compresses MoE expert parameters to sub-linear scaling via structured butterfly orbits, enabling deployment on edge devices.
Open-source trillion-parameter MoE LLM with 68.8B activated parameters using Layer-Adaptive Expert Pruning for enterprise and general-purpose tasks.
On-policy self-distillation framework where student LLMs generate their own trajectories while receiving dense token-level supervision from teacher models.
MiTA Attention mechanism using mixture of top-k activations for efficient fast-weight scaling in transformers with extremely long sequences.
VIP allocation strategy optimizing rollout sampling efficiency in reinforcement learning with verifiable rewards, improving upon fixed GRPO allocation.
Position paper proposing agentic time series forecasting with iterative refinement, reasoning-driven inference, and continual adaptation beyond static prediction.
Analysis proving steering vectors in LLMs are fundamentally non-identifiable due to large equivalence classes, challenging behavioral control interpretability.
Framework showing embedding magnitude should be learned independently in contrastive learning, benefiting retrieval and RAG applications.
Study evaluating Kolmogorov-Arnold Networks in physics-informed neural network architectures for discovering unknown terms in oscillatory systems.
Framework enabling masked diffusion models to correct their own tokens iteratively, reducing error accumulation in parallel generation.
Hybrid quantum-classical GAN framework for synthesizing realistic tabular data with heterogeneous features and privacy constraints.
Weight-space sequence modelling technique to improve neural network extrapolation on out-of-support datapoints beyond training distribution.
Three-stage curriculum learning framework for distilling chain-of-thought reasoning from large LLMs into compact student models while preserving interpretability.
cc-Shapley method for measuring multivariate feature importance in ML models using causal context.
Zatom-1 open-source foundation model for unified generative and predictive learning on 3D molecules and materials.
Reference-guided fine-tuning approach for reinforcement learning on mathematical reasoning tasks with sparse rewards.
AOI framework for training LLM agents to diagnose cloud infrastructure failures using failed trajectories as learning signals, addressing safety and data constraints.
Neural solver for vehicle routing problems using distance representation learning for asymmetric real-world scenarios.
Analysis of why linear RNNs achieve transformer-level parallelizability compared to nonlinear RNNs for language modeling.
Learning approach for Dubins Traveling Salesman with neighborhoods using knowledge distillation from expert trajectories.
Theoretical research on game theory attractors and replicator dynamics, characterizing learning equilibria.
Research paper on safety mirage in vision language models: spurious correlations in safety fine-tuning and mitigation via machine unlearning.
Evaluates LLM fault localization capabilities on code changes, assessing semantic program reasoning beyond syntax and lexical features.
Benchmark evaluating LLMs on rigorous causal inference tasks involving statistical pitfalls relevant to medicine, economics, and public policy.
ShIOEnv: Gymnasium environment for grammar-constrained synthesis and modeling of command-line interface behavior with shell input-output data.
Novel framework using multi-kernel Boolean parameters for weight binarization in LLMs to improve efficiency without full-precision latent weights.
SealQA benchmark for evaluating search-augmented LLMs on fact-seeking questions with conflicting or noisy web search results.
EDINET-Bench: New benchmark dataset evaluating LLMs on complex financial document analysis tasks using Japanese financial statements.
MuRating framework transfers English data-quality signals to score documents in 17 languages for multilingual LLM pretraining.
Research generalizes EDM to arbitrary-noise diffusion models, analyzing design space beyond Gaussian noise for image restoration tasks.
Quantum EM algorithm for training quantum Boltzmann machines that circumvents barren plateau problem in quantum machine learning.
Study shows LLM ranking systems are sensitive to small changes in preference data, with top model rankings changeable by dropping few preferences.
Comprehensive evaluation of how weight and activation quantization affects model bias across stereotypes, fairness, toxicity, and sentiment.
Research on optimal alignment of acoustic and linguistic representations in pre-trained models for automatic speech recognition.
BabyHuBERT self-supervised speech model trained on 13,000 hours multilingual child-centered recordings for speaker segmentation.
Framework combining diffusion models with impedance control for robot learning in contact-rich manipulation tasks.
Noise-to-Notes reformulates automatic drum transcription as conditional generative task using diffusion modeling.
BridgeDrive applies diffusion-based planning with expert behavior anchors for closed-loop autonomous driving trajectory planning.
BeyondBench framework uses algorithmic problem generation for contamination-resistant evaluation of reasoning in language models.
SphereAR addresses variance collapse in continuous-token autoregressive image generation by constraining latents to hypersphere.