RAM-Net: Expressive Linear Attention with Selectively Addressable Memory
RAM-Net architecture combines linear attention efficiency with full attention expressivity using selectively addressable memory to reduce information loss.
RAM-Net architecture combines linear attention efficiency with full attention expressivity using selectively addressable memory to reduce information loss.
Manifold-aware temporal domain generalization for LLMs handling distribution shifts via parameter-efficient geometric reformulation.
Momentum LMS theory for non-stationary streaming data analyzing stability and regret without i.i.d. assumptions.
Analysis of differential privacy impact on firing-rate statistics in federated spiking neural networks for neuromorphic learning.
FedGRPO: federated learning method for foundation models using group-relative rewards from domain clients with privacy preservation.
PrefillShare: shared prefill module reducing redundant computation across multiple LLMs in disaggregated serving with KV cache reuse.
Reciprocal-space generative pipeline for crystalline materials using Fourier transforms and diffusion models with symmetry constraints.
Online RL approach training LLMs for HPC code generation using real supercomputer runtime performance (GFLOPS) as rewards.
Automatic soccer event detection from player trajectories without ball tracking using possession path inference.
Empirical GPs: principled framework for learning kernel functions automatically rather than handcrafting from standard functions.
Novel method learning structured latent representations using metric spaces for multimodal state estimation in RL without explicit noise assumptions.
Theoretical analysis of offline RL under Q*-approximation and partial coverage, answering whether Q*-realizability enables sample efficiency.
Few-shot Bayesian optimization framework exploiting auxiliary information from experiments for expensive black-box design problems.
KAN-FIF applies spline-parameterized neural networks for tropical cyclone estimation on resource-constrained meteorological satellite devices.
Meta-Sel: lightweight supervised meta-learning approach for efficient demonstration selection in in-context learning with tight prompt budgets.
Investigates whether LLMs trained with RL spontaneously exploit reward function loopholes without malicious intent, examining alignment risks.
Theoretical analysis of on-policy distillation showing it as special case of KL-constrained RL with reward extrapolation improvements.
Introduces damped harmonic oscillators method for irregular time series modeling as alternative to Transformers and Neural ODEs.
Proposes TIME benchmark addressing limitations in time series foundation models including data composition, integrity, and task formulation issues.
SafeNeuron provides neuron-level safety alignment mechanism for LLMs by targeting safety-critical parameters to prevent alignment bypass attacks.
Group Relative Policy Optimization enables amortized molecular design learning transferable to unseen molecules rather than instance-specific optimization.
Theoretical analysis of sampling and reference policy effects in preference alignment for LLMs through Identity Preference Optimization framework.
Learning to Forget Attention proposes adaptive attention reduction as familiarity increases, reducing compute in hybrid state-space-attention models.
Study shows that downstream adaptation for measuring world models can corrupt latent physics representations, affecting OOD generalization assessment.
Distribution Discriminant Theory enables on-policy supervised fine-tuning for LLMs by bridging computational efficiency gap with RL-based approaches.
Variance Minimisation Policy Optimisation reformulates diffusion model alignment as an SMC process for reward-guided sampling.
Categorical Flow Maps applies flow matching for accelerated few-step generation of categorical data using self-distillation.
Olmix framework addresses practical challenges in data mixing ratios during language model development with principled design choices.
ExtractBench provides a benchmark and methodology for evaluating LLM-based PDF-to-JSON extraction at enterprise scale with diverse schemas.
User study on explainable AI without code examines how non-technical users can understand ML model predictions in no-code platforms.
Study of transformer representations shows direction and magnitude of hidden state vectors serve distinct functional roles in language modeling and syntax tasks.
Research evaluates few-shot temporal reasoning capabilities of LLMs for predicting human activities in smart environments with limited data.
AskBench benchmark and rubric-guided RLVR method evaluates and improves LLMs' ability to request clarification when prompts lack critical details, reducing hallucinations.
SWE-MiniSandbox: container-free reinforcement learning method for scalable training of software engineering agents without isolation overhead.
Self-play framework for vision-language models that actively explores environments to generate tailored visual data for autonomous improvement.
Identifies computational hardness in learning non-trivial mixed-state quantum phases using autoregressive networks and conditional mutual information.
Agentic reinforcement learning approach to optimize proactive LLM agents for multi-turn task completion with learned interaction strategies.
Studies whether self-referential language in LLMs reflects actual internal computation or confabulation through activation analysis.
Method to improve LLM reasoning by identifying critical tokens and verifying consistency through paraphrastic probing.
System-level defenses against indirect prompt injection attacks in AI agents, reducing unsafe actions while maintaining task completion rates.
Training method for LLMs using generalized entropic objectives instead of uniform token-level weighting to improve robustness and sharpening.
Research on budget-constrained LLM agents that plan tool use under monetary constraints. Addresses sequential decision-making with expensive tool executions.
Survey covering multi-agent communication mechanisms in MARL, emergent language, and LLM-based agents across autonomous and collaborative systems.
Approach routing LLM reasoning between latent and discrete spaces to improve efficiency while maintaining reasoning quality and confidence.
Training optimization technique for Mixture-of-Experts models addressing expert load imbalance through dynamic expert relayering.
Method using crosscoders for comparing internal representations across different LLM architectures to identify safety-critical behavioral differences.
Framework for improving feature importance estimation stability by aggregating model explanations across ensemble members for scientific discovery.
9B-parameter LLM architecture combining sparse and linear attention mechanisms for efficient long-context processing and reduced computational costs.
Training method for multi-turn RL of LLM agents using trajectory search rollouts to improve exploration in sparse reward environments.
Decentralized optimization algorithm for distributed machine learning addressing heterogeneous gradient variance across nodes.