Optimizing Language Models for Crosslingual Knowledge Consistency
Method using reinforcement learning with structured rewards to improve consistency of LLM knowledge across multiple languages.
Method using reinforcement learning with structured rewards to improve consistency of LLM knowledge across multiple languages.
VDCook platform for automated video dataset construction using natural language queries, retrieval, and synthesis for multimodal LLM training.
Research on credit assignment in cooperative multi-agent LLM systems, showing standard evaluation methods are flawed and proposing exact counterfactual approaches.
Theoretical analysis of log-barrier regularization enabling exploration in stochastic gradient bandit optimization with convergence guarantees.
MARL-Rad uses multi-modal multi-agent RL for radiology report generation, training agentic workflows end-to-end within deployed clinical workflow.
SlopCodeBench evaluates coding agents on iterative 36-problem benchmark with 196 checkpoints where agents extend own solutions over time.
Analysis of emotion vector organization in LLMs via valence-arousal subspace with circular geometry for emotion steering control.
Zero-shot quantization via weight-space arithmetic extracts quantization vectors from donor tasks to improve post-training quantization without receiver training data.
Analysis of Muon optimizer dynamics as spectral Wasserstein flow in mean-field regime, studying gradient normalization for deep learning.
Android Coach improves online RL agent training efficiency by processing multiple actions per state, addressing sample inefficiency in emulator-based learning.
Framework for versioned AI component governance in embodied agent systems with compatibility checking and rollback at deployment lifecycle.
Technique for offloading KV cache to reduce memory and latency bottlenecks in long-context LLM inference while maintaining accuracy.
Lightning OPD enables efficient offline on-policy distillation for LLM post-training by precomputing teacher log-probabilities without live server.
SeaEvo uses LLM-guided evolutionary search for algorithm discovery with persistent population-level state management and structured reasoning.
Neural cellular automaton with discrete bottleneck learns compositional semantic parsing rules without hand-written rules, enabling structural generalization.
Persistent Visual Memory module strengthens visual attention in autoregressive LVLMs to prevent signal dilution during long sequence generation.
Meta's Code World Model preparedness assessment for frontier risks including code generation, reasoning, and misalignment evaluation.
Multi-agent RL approach for tactical deconfliction between heterogeneous unmanned aerial system fleets in dense urban airspace.
Cross-lingual safety alignment framework using self-distillation to transfer safeguards from high-resource to low-resource languages in LLMs.
DGPO improves RL credit assignment for LLM alignment on complex reasoning via distribution-guided policy optimization with finer-grained step isolation.
CoREB introduces a contamination-limited multitask benchmark for code search beyond first-stage retrieval, including reranking and developer-style queries.
MACS improves multimodal MoE inference efficiency using modality-aware capacity scaling to address straggler effects in expert parallelism.
Empirical study evaluating privacy awareness of Vision-Language Models when deployed as autonomous cognitive cores in physical environments.
EGA proposes a residual adapter for vector search systems to handle distribution shift with frozen vision encoders.
Active learning method optimizes communication structure in LLM-based multi-agent systems to improve performance and reduce token usage with limited budgets.
CRAFT proposes continual learning for LLMs via low-rank interventions on hidden representations to prevent catastrophic forgetting during fine-tuning.
VideoRouter addresses efficiency bottlenecks in long video understanding using query-adaptive dual routing to compress visual tokens dynamically.
Safety Anchor defends LLM safety alignment against harmful fine-tuning by introducing geometric bottlenecks in parameter space.
VISD: structured self-distillation for video reasoning combines RLVR and token-level credit assignment in VideoLLMs.
Entropy-regularized adjoint matching for offline RL with flow-matching policies reduces popularity bias in behavior distribution.
NOVA world model framework representing state as neural network weights with latent structural disentanglement for interpretability.
NavOne: top-down vision-language navigation using one-step global planning on maps for efficient path reasoning.
Asymmetric on-policy distillation improves token-level student learning by addressing variance and exploration issues.
Q-MMR framework for off-policy evaluation in MDPs using moment matching and recursive reweighting of trajectories.
AI CFD Scientist: open-source AI agent for autonomous computational fluid dynamics discovery using physics-aware reasoning.
RateQuant applies rate-distortion theory for mixed-precision KV cache quantization with per-head bit allocation.
LKV learns head-wise budgets and token selection for KV cache eviction in long-context LLM inference optimization.
Positive-and-Negative Decoding framework reducing object hallucination in vision-language models via attention rebalancing.
Analysis of strain and vorticity in velocity field Jacobians to understand integration error in flow matching models.
Hierarchical ensemble pipeline for anomaly detection in ESA satellite telemetry using shapelet and statistical features.
Toeplitz MLP Mixer architecture for sequence modeling with O(dn log n) complexity as efficient transformer alternative.
Transformer-based sequence models for wildlife species classification from GPS movement trajectories, outperforming LSTMs and CNNs.
Semantic State Abstraction Interfaces for mapping news text to interpretable coordinates in LLM-augmented portfolio decision systems.
Analysis of model-based RL training on imagined trajectories, quantifying effects of dynamics and reward model errors on policy optimization.
Gauge-aware aggregation method for federated LoRA addressing representation dependence in decentralized LLM adaptation.
Quantum-inspired fast weight programmer combining KAN and quantum circuits for scalable sequence learning on NISQ devices.
Geometric Kolmogorov-Arnold Networks using learned Riemannian metrics for geometry-aware approximation in input space warping.
Closed-form upper bound derivation for admissible learning-rate steps in belief-space dynamics using KL/Bregman geometry.
Gradient Extrapolation-Based Policy Optimization for efficient LLM reasoning improvement, reducing computational cost versus multi-step lookahead.
Sparse attention indexing method for efficient LLM inference that guarantees zero false negatives in KV cache selection during decoding.