Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
Study of L0 hyperparameter effects on sparse autoencoders for LLM feature extraction and interpretability.
Study of L0 hyperparameter effects on sparse autoencoders for LLM feature extraction and interpretability.
FedDAF federated domain adaptation approach addressing domain shifts and limited labeled data across clients.
Theoretical analysis of in-context learning in Mamba models with outliers and generalization guarantees.
Research on adversarial robustness and generalizability of neural operators as PDE surrogates from solver perspective.
Two-phase training strategy enabling per-neuron adaptive activation function selection while maintaining inference efficiency.
LLM-based framework for flight delay prediction integrating aeronautical data and aircraft trajectory representations.
Bayesian online learning approach for adaptive digital twin state transitions in civil engineering.
Transformer architecture for solving complex PDEs on unstructured meshes with geometric feature preservation.
Agentic framework for automated kernel code generation across heterogeneous AI accelerators at scale for recommendation models.
Theoretical analysis of knowledge distillation convergence with Bayesian teacher networks using SGD.
Research on learning discrete activation functions using Gumbel-Softmax for task-specific neural networks.
Constrained optimization framework training transformers as descent algorithms with layerwise loss decrease guarantees.
GraphAllocBench benchmark for multi-objective reinforcement learning with preference-conditioned policies and flexible resource allocation.
TRACE method for causal representation learning handling continuous mechanism transitions across domains.
StepShield benchmark for agent safety measuring detection timeliness of rogue agent behavior on 9,429 code-agent trajectories.
BitLogic training framework for gradient-based neural networks deployable to GPU, FPGA, and ASIC with single codebase.
Meta-learning framework defining practical universality and generalizable learning across arbitrary task distributions.
Study extracting invariant algorithmic cores from transformers across independent training runs to identify necessary computations.
Improvement to TabPFN synthetic tabular data generation by incorporating causal structure and feature dependencies.
LLM compression technique using Leech lattice vector quantization to overcome information-theoretic limits of scalar quantization.
Policy Improvement RL method for post-training LLMs and agents that explicitly verifies policy improvements over baselines.
Analysis showing continual learning performance varies significantly based on fine-tuning regime (trainable parameter subspace).
Reinforcement learning approach for scheduling AIGC workloads and managing energy in distributed data centers using diffusion-based reward shaping.
Streaming reinforcement learning method enabling online learning with partial observability using real-time recurrent backpropagation.
BigMac training pipeline for multimodal LLMs that improves compute-memory efficiency tradeoffs through nested encoder-generator computation.
Framework for continual evolution of agent skills in LLM-based agents by maintaining persistent decision history across task changes.
Reproducibility study of AlphaEdit null-space constrained knowledge editing method for LLMs, validating and extending original results.
Lightweight optimizer exploiting gradient geometry of embedding tables and LM-heads for improved training efficiency across finetuning and pretraining.
Fault-tolerant LLM training system using zero-overhead checkpointing and hot-swapping for resilience across hardware failures.
Credit assignment optimization for RL-based LLM reasoning using fine-grained surrogate entropy for token-level rewards.
Analysis decomposing router-to-oracle performance gap in LLM routing, identifying label noise versus specialist advantage contributions.
Efficient inference system combining quantization and speculative decoding for Qwen3.5-4B LLM on resource-constrained hardware.
Federated learning framework using multimodal LLMs (LLaVA) to address data heterogeneity across distributed clients.
Application of quantum support vector machines to financial data classification using Dhaka Stock Exchange dataset.
Theoretical study of benign overfitting in quantum kernel methods for machine learning on quantum computers.
Human-augmented reinforcement learning approach for 3D bin-packing in logistics, combining RL with human feedback to reduce training time.
Evaluation of neural network graph compilers across heterogeneous hardware platforms, showing how vendor-specific optimizations affect performance comparisons.
Position paper on EU AI Act research exemptions and potential conflicts with academic publication norms at major conferences.
Multi-agent LLM system for credit assessment that mirrors real-world decision-making processes in financial evaluation.
Analysis of reasoning behaviors in thinking language models using Sparse Autoencoders and model diffing techniques.
Production-grade dataset of 2,185 multi-turn examples for training secure code generation models covering OWASP Top 10 and ML security.
LLM-based agentic AI system for insurance underwriting with self-critique mechanisms for high-stakes decision support.
Vision-Language-Action model for robotic control that combines VLM representations with 3D pose understanding for embodied AI tasks.
Cryptographic method to verify fine-tuned neural network models haven't deviated from claimed update procedures without accessing parameters.
arXiv paper on model-based bootstrap for transition kernels in controlled Markov chains with applications to offline reinforcement learning.
arXiv research on CARVE, recurrent model with improved state management for memory-aware chunk-parallel linear attention.
arXiv paper on Wan-Streamer v0.2, audio-visual interaction model achieving higher resolution (640x368) while maintaining 200ms latency.
arXiv paper on TACTIC-KG using small LLM agent teams for constructing cybersecurity threat intelligence knowledge graphs from unstructured text.
arXiv paper introducing Nemotron Audex-30B-A3B, unified audio-text LLM maintaining text performance while adding audio understanding.
arXiv paper on multiplayer world models for dynamic environments with complex physics, conditioning on multiple agent action streams.