Low-Dimensional and Transversely Curved Optimization Dynamics in Grokking
arXiv: Geometric analysis of transformer optimization dynamics revealing low-dimensional manifolds in grokking.
arXiv: Geometric analysis of transformer optimization dynamics revealing low-dimensional manifolds in grokking.
Research paper studying loss-landscape geometry as early-warning signals for grokking in neural networks.
CeRA: parameter-efficient fine-tuning method overcoming LoRA's linear capacity ceiling via non-linear gating and dropout for rank adaptation.
SafeSci: comprehensive benchmark and framework for evaluating LLM safety in scientific domains with multi-domain risk coverage and objective evaluation.
Framework for EEG-to-text decoding addressing semantic bias and signal neglect in neural signal interpretation. Published on arXiv.
DiFlowDubber: discrete flow matching framework for video dubbing with TTS, lip synchronization, and expressive prosody. Published on arXiv.
Qualitative study of 167,000+ AI agents on multiple platforms learning from each other and developing emergent behaviors without researcher intervention.
arXiv: RAG-enhanced diffusion models using adaptive guidance to resolve conflicts between retrieved noisy context and parametric model knowledge.
Studies robustness of medical vision-language models under real clinical workflows using chain-of-distribution attacks and token-space repair techniques.
ArXiv research on parameterized GELU activation for controlled ReLU approximation in deep networks.
ArXiv paper on coarse-to-fine visual processing for efficient document parsing with vision-language models.
ArXiv study on behavioral consistency of LLM agents in SWE-bench comparing multiple models.
ArXiv research analyzing prompt injection attack success stages across five frontier LLM agents.
ArXiv paper on token-level entropy regulation for reinforcement learning in large reasoning models.
ArXiv research on spectral edge thesis controlling phase transitions in neural network training dynamics.
APEX-EM non-parametric framework for LLM agents to accumulate and reuse procedural plans without weight modification.
World model planning for structured origami generation satisfying geometric constraints and kinematic rules via long-horizon reasoning.
Terminal agents executing enterprise tasks via CLI are simpler and more cost-effective than tool-augmented or web agents.
Transfer learning methods for nonparametric Bayesian networks under scarce data with constraint-based and score-based algorithms.
ProdCodeBench evaluates AI coding agents using production-derived tasks reflecting real developer-agent sessions and workflows.
Visual attention inertia in MLLMs causes cognitive hallucinations; proposes mitigation for compositional understanding.
LiME achieves expert specialization in multimodal MoE-PEFT via lightweight modulation instead of separate adapters per expert.
SIEVE enables sample-efficient parametric learning from natural language instructions and feedback without high-quality traces.
Model scheduling for masked diffusion language models uses smaller models at early denoising steps for faster generation.
Process reward models improve LLM mathematical reasoning by providing step-level feedback on intermediate errors, not just final outcomes.
LLM-based compression using domain-adapted LoRA for lossless and lossy text compression achieving 2x improvements.
Systematic characterization of WebGPU dispatch overhead for LLM inference across GPU vendors, backends, and browsers at batch size 1.
UI-Oceanus framework scales GUI agents via synthetic environmental dynamics and self-supervised learning instead of costly human demonstrations.
Benchmark evaluating LLM and embedding performance for drug discovery tasks, assessing advantages over traditional methods.
Contextual RL improves agent generalization by exposing agents to environment characteristics for better zero-shot transfer beyond training distribution.
OPRIDE method for offline preference-based RL reducing human feedback queries through efficient in-dataset exploration strategies.
Differentiable Symbolic Planning architecture combining neural networks with discrete symbolic reasoning for constraint satisfaction problems.
Framework for modeling and controlling ML model reliability under temporal distribution shift during deployment with continuous monitoring.
Contrastive prompt tuning method to optimize LLMs for generating energy-efficient code aligned with Green Software Development goals.
PRISM framework for zero-shot policy transfer in RL using interpretable concept clustering with causal validation across different algorithms.
Entropy-based analysis of combining Chain-of-Thought with RL for text-to-image generation, showing exploration-optimization tradeoffs.
Self-Directed Task Identification framework enabling models to autonomously identify target variables in zero-shot settings without pretraining.
Physics-informed deep generative models for offline RL in spaceflight to mitigate sim-to-real gap with limited real-world training data.
Open-source benchmarking of Matrix Profile methods for time-series anomaly detection on univariate and multivariate datasets.
Study comparing frontier vs smaller LLMs for mathematical proof verification, evaluating whether expensive models are necessary for proof checking.
Analysis of layer-to-layer representation changes in language models, decomposing updates into tokenwise and residual components.
Transfer learning method for RNNs using time-warping rescaling, with theoretical analysis for linear differential equation models.
Framework for validating assumptions in time-series causal discovery through calibrated risk assessment and effect-size diagnostics.
Low-precision training method for LLMs using adaptive Hadamard transforms based on outlier patterns in weights, activations, and gradients.
Research on using LLM-generated synthetic data to warm-start contextual bandits, examining alignment between LLM choices and actual user preferences.
Spectral framework for multi-scale nonlinear dimensionality reduction balancing global-local structure preservation and expressiveness-transparency.
Optimized NF4 dequantization kernels for fast LLM inference on NVIDIA GPUs, addressing FP16 conversion bottleneck.
Communication-efficient distributed learning algorithm with differential privacy using local training and gradient clipping.
VoxelCodeBench platform benchmarking code generation models for 3D spatial reasoning with execution in Unreal Engine.
Analysis of function vectors in LLMs showing they steer behavior beyond logit lens interpretability across 4,032 cross-template transfer pairs.