EVA: Aligning Video World Models with Executable Robot Actions via Inverse Dynamics Rewards
Video world models for robotics using inverse dynamics rewards to align generated trajectories with executable robot actions.
Video world models for robotics using inverse dynamics rewards to align generated trajectories with executable robot actions.
Systematic analysis of Elastic Weight Consolidation for continual learning showing suboptimal performance and proposing improvements to weight importance estimation.
Benchmark comparing PETNN, KAN, and classical deep learning models on Burmese handwritten digit recognition dataset.
Mi:dm K 2.5 Pro, 32B parameter enterprise LLM supporting multi-step reasoning, long-context understanding, and agentic workflows in Korean and domain-specific applications.
Formal specification for admission control governance of autonomous agents in institutional B2B environments with cryptographic validation.
Graph-based memory system for LLM reward prediction requiring limited labeled data for reinforcement learning post-training.
Research on chain-of-thought faithfulness in LLMs showing measurement methodology significantly affects reported faithfulness scores across 12 open-weight models.
KidGym: 2D grid-based reasoning benchmark evaluating MLLMs on spatial intelligence inspired by Wechsler Intelligence Scales.
CRoCoDiL: Continuous semantic space diffusion model for non-autoregressive language generation with improved coherence.
Industrial-scale RAG framework evaluated on automotive manufacturing requirements engineering with unstructured heterogeneous documentation.
Memory-Keyed Attention: Efficient attention mechanism reducing KV cache memory for long-context LLM inference and training.
TRACE: Multi-agent system using autonomous reasoning for seismological analysis of earthquake mechanisms from geophysical observations.
Unsupervised self-evolution training framework for multimodal LLMs achieving reasoning improvements without annotated data.
DeepXplain: Explainable deep reinforcement learning framework for multi-stage APT cyber defense with provenance graphs.
Case study on LLM-powered workflow optimization for multidisciplinary software development in automotive industry.
mSFT: Iterative algorithm addressing overfitting in multi-task supervised fine-tuning by heterogeneous data mixture optimization.
arXiv paper on safe offline reinforcement learning with budget constraints. Addresses safety-reward trade-offs in sequential decision making.
Research on synthetic data generation using LLMs to improve smaller model fine-tuning. Analyzes diversity and distribution in embedding space.
Uncertainty estimation method for LLMs using intra-layer local information scores from cross-layer agreement patterns.
Sparse Feature Attention method reducing transformer self-attention complexity through feature-level sparsity instead of sequence-level sparsity.
Mathematical framework interpreting LLM hidden states as points on latent semantic manifolds with Riemannian geometry.
Training-free hallucination detector for LLMs using sample transform cost to measure output distribution complexity.
Progressive Quantization method for robust vector tokenization in multimodal LLMs and diffusion models.
Chinese financial news dataset and benchmark for evaluating LLMs as autonomous agents in macro and sector asset allocation.
UniFluids: conditional flow-matching framework using diffusion Transformers to unify learning solution operators across diverse PDEs.
Decision Transformer approach for offline emergency vehicle signal preemption optimization without online exploration.
Delta-Aware Quantization framework for post-training LLM weight compression that preserves knowledge by protecting small-magnitude parameter deltas during quantization.
Classification method for wind power ramp event forecasting addressing severe class imbalance in grid stability systems.
Method for adding trained persistent memory to frozen decoder-only LLMs using memory adapters in latent space.
Distribution-free safety guarantees for wildfire evacuation mapping using conformal prediction with tabular, spatial, and graph models.
Comparative study of LLMs for missing data imputation with analysis of hallucination effects and control mechanisms.
State-space model enhancement combining graph signal processing with Mamba2 for efficient language modeling.
Analysis of systematic biases in Chinchilla scaling law fitting method with implications for compute-optimal LLM allocation.
Cloud-edge collaborative framework using large models for robust photovoltaic power forecasting with latency constraints.
Research on first-mover bias in gradient boosting feature importance explanations under multicollinearity conditions.
Online learning algorithm balancing regret guarantees in adversarial and stochastic settings with safety constraints.
Web-grounded iterative self-play framework for improving LLM reasoning through reinforcement learning with verifiable rewards.
Neural tangent kernel framework for understanding continuous representation full-waveform inversion in geophysical applications.
Research establishing equivalence between classifier-free guidance and alignment objectives in diffusion model training.
Format-aware quantization method for deploying LLMs on edge devices using NVFP4 ultra-low bit precision.
Research on multimodal fusion strategies for time series forecasting using text and vision modalities with constrained approaches.
Study evaluating whether LoRA adapters trained with instruction-tuning objectives actually improve instruction-following capability across different tasks using IFEval benchmarks.
Symbolic Graph Networks framework for discovering PDEs from noisy, sparse observational data using machine learning instead of numerical differentiation.
Reinforcement learning approach for learning optimal decision timing in continuous environments using predictive temporal signals.
Symbolic regression framework using continuous structure search and neural embeddings for interpretable equation discovery.
Offline reinforcement learning with model predictive control using differentiable world models for inference adaptation.
Skill retrieval and ranking system for LLM agents selecting relevant tools from thousands of overlapping options at scale.
Energy-aware gradient pruning framework for federated learning accounting for hardware transmission costs.
Multimodal training framework leveraging unstructured clinical notes to improve structured EHR data deployment.
Foundation model for time-series in-context learning using quantile-regression T5 with instruction conditioning.