Provable imitation learning for control of instability in partially-observed Vlasov--Poisson equations
Imitation learning approach for plasma stabilization control in Vlasov-Poisson systems with partial state observability constraints.
Imitation learning approach for plasma stabilization control in Vlasov-Poisson systems with partial state observability constraints.
ORDERED algorithm for unsupervised domain adaptation that reduces variance in discrepancy estimation through optimal data reordering.
Memini system for continual knowledge updating in LLMs using multi-timescale memory dynamics inspired by biological memory mechanisms.
Theoretical framework for analyzing regret distribution in multi-armed bandits and episodic RL with probabilistic guarantees across confidence levels.
RL method for improving sample efficiency in agentic tasks by controlling pass-rate to optimal 50% for maximizing reward signal informativeness.
Theoretical study of signal propagation in finite-width linear recurrent models, analyzing approximation accuracy as sequence depth and width grow jointly.
Analysis of jailbreak attack difficulty on LLMs, showing random search effectiveness challenges assumptions about prompt structure necessity.
Low-cost black-box method to detect LLM hallucinations by treating model as dynamical system and analyzing embedding manifolds.
Case study of AI tools assisting high school students in financial forecasting research project with human-AI co-mentorship.
Mechanistic interpretability study of transformer time series forecasting using sparse autoencoders, questioning superposition hypothesis.
Theoretical analysis of transformer in-context learning for nonlinear regression, explaining how attention acts as a featurizer.
Method to estimate expected output of wide random MLPs without sampling by computing activation distributions analytically.
Information-theoretic analysis showing multi-agent LLM debate preserves answer accuracy but degrades reasoning quality (Reasoning Trap).
FREIA: Free energy-driven reinforcement learning with adaptive advantage shaping for unsupervised LLM self-improvement and reasoning.
Adaptive Power-Mean Policy Optimization method for improving LLM reasoning through adaptive reinforcement learning with verifiable rewards.
arXiv work using authorship attribution machine learning to link anonymous online criminal profiles in trafficking networks.
BOOOM: black-box optimization method over orthonormal manifolds for non-convex and derivative-free machine learning problems.
arXiv security research on membership inference attacks against retrieval-augmented in-context learning in document QA systems.
arXiv paper on tree-conditioned edit flows for ancestral protein sequence reconstruction using deep learning.
JoyAI-Image: unified multimodal foundation model combining spatially enhanced MLLM with diffusion transformer for image understanding and generation.
Evaluation of open-weight LLMs (Gemma, Llama, Mistral, OLMo) on conflict monitoring for West Africa against ACLED benchmark.
ANDRE: neuro-symbolic approach combining attention mechanisms with logic programming for interpretable rule extraction from data.
arXiv paper on supply-chain backdoor attacks hiding sparse perturbations in image classifiers and Vision Transformers.
arXiv theoretical work on unbalanced optimal transport and density control for Gaussian distributions.
arXiv paper on optimal transport methods for curved spaces and Riemannian manifolds in machine learning.
arXiv research on adversarial examples against vision-language models for fact-checking and content moderation.
arXiv dataset paper on synthetic fiber rope degradation for remaining useful life estimation using imagery.
TRIBE v2: tri-modal foundation model predicting human brain activity from video, audio, language across 720 fMRI subjects.
Perturbation-based training framework for LLMs using semantic neighbor prefixes instead of exact tokens for improved extrapolation.
Coral: multi-LLM inference serving framework optimizing heterogeneous GPU allocation across multiple models with adaptive scheduling.
AI-generated image detection using intermediate representations from vision models without retraining.
Coral: adaptive multi-LLM serving system for heterogeneous cloud GPUs with dynamic resource allocation across embedding and KV caches.
Analysis that model-level LLM alignment evaluation is insufficient; deployment-relevant alignment requires system-level evaluation.
SpecPL: prompt learning method for vision-language models using spectral granularity and counterfactual supervision.
SemEval-2026 winning system: ensemble of 7 LLMs with GPT-4o judge for multi-turn faithful response generation with reference passages.
UniVer framework unifying multi-step and multi-draft speculative decoding for LLMs via optimal transport verification.
Evaluation of LLMs on audio embedding benchmark (MSEB), testing multimodal capabilities of models like Gemini on audio understanding tasks.
Research on safety degradation in LLM fine-tuning, analyzing parameter dynamics during training and quantifying sample-level safety risks.
Token-aware gradient optimization attack on audio language models exploiting non-uniform gradient distribution for efficient jailbreaking.
Budget-Aware Optimizer Configurator (BAOC) that reduces GPU memory usage by assigning differentiated optimizer states across network blocks.
Gyan, a neuro-symbolic language model combining transformers with symbolic reasoning for improved interpretability and compositionality.
Local learning approach for LLM post-training reducing computation and memory costs by limiting backward gradient propagation depth.
Neural Rule Inducer (NRI), a pretrained foundation model for zero-shot inductive logic programming enabling rule induction across tasks without retraining.
Study on per-instance algorithm selection for black-box optimization, examining trade-offs between feature computation and algorithm performance.
SLYP agentic pipeline discovers race condition vulnerabilities in Windows COM binaries and generates verified proof-of-concept code.
Piper system for efficient MoE model training via mathematical modeling and pipelined hybrid parallelism addressing memory and communication challenges.
Fine-tunes Gemma models with LoRA for multilingual polarization detection using LLM-generated synthetic data augmentation strategies.
Investigates connection between contrastive learning data augmentation and positive-incentive noise through information theory framework.
Applies quantum-inspired reinforcement learning to synthesizable drug design for molecular optimization with synthetic feasibility constraints.
Proposes dataset-driven channel masks in Transformers for multivariate time series modeling and channel dependency capture.