GenVid2Robot: From Video Generation to Robot Manipulation via Rigid-Geometric Consistency
Framework converting generated videos to robot manipulation trajectories by enforcing rigid-geometric consistency and kinematic constraints.
Framework converting generated videos to robot manipulation trajectories by enforcing rigid-geometric consistency and kinematic constraints.
Method for initializing language model pretraining using spectral patterns from pretrained checkpoints to improve efficiency.
arXiv paper presenting Tsetlin Machine approach for interpretable PDF malware detection addressing cyberattack vector.
arXiv paper on deep learning system for thoracic disease detection in chest radiographs developed and validated on Thai population data.
arXiv paper studying creativity, honesty and memory in small hyperbolic language models and implications for AI assistants becoming companions.
arXiv paper applying machine learning to automatic thematic indexing of Voltaire's complete works, automating scholarly text categorization at scale.
arXiv paper evaluating NLP approaches for automated keyword extraction from crowdsourced Second World War digital collections.
arXiv paper identifying entity attribution failure in clinical RAG systems where retrieved evidence can be misattributed despite passing standard evaluation metrics.
Soofi S 30B-A3B: open-source MoE hybrid Mamba Transformer foundation model for German and English, 3B active parameters per token, optimized for long-context deployment.
Multimodal framework for retrieving similar autonomous driving scenarios combining visual and motion-based representations.
Evaluates test-time scaling techniques on multilingual visual multiple-choice benchmark comparing self-consistency and beam search strategies on small open VLMs.
Test-time prompt adaptation method for improving Vision-Language Model robustness against adversarial perturbations using distributional structure awareness.
Foveated Dynamic Transformer applies human visual system principles for efficient vision transformer token selection with robustness to adversarial perturbations.
Investigation of counting failures in VLMs reveals internal representations encode count information but fail verbalization; proposes probing-based detection method.
Analysis of representational limitations in AI systems for reasoning, coding, theorem-proving and tool use; identifies vocabulary and verifier gaps in open-ended evaluation.
Physics-constrained ML framework using residual data augmentation and entropy constraints for accelerating turbulent flow simulations via surrogate chemical kinetics models.
ArXiv paper on LLM applications for electronic design automation including HDL generation, testbench construction, and design space exploration.
ArXiv paper on deep Gaussian processes over DAGs for modeling compositions of partially-observed functions in causal and engineering systems.
Privacy-preserving model distillation framework using local differential privacy for remote API access scenarios. Non-adversarial approach.
Bayesian optimization method for bilevel optimization with expensive black-box functions. Theoretical framework without clear AI/ML application focus.
Graph neural network for mixed bundle pricing with self-improving pruning strategy.
Multi-horizon time series forecasting with reduced volatility using neural sequence forking.
Analysis of LLM alignment degradation after fine-tuning beyond black-box evaluation limits.
Multi-objective RL approach for preference-conditioned policy optimization with diversity.
Tensor methods for material design optimization reducing computational search space.
Online optimization of non-monotone DR-submodular functions with linearizability analysis.
Memory-efficient context parallelism for long-sequence Transformer processing via headwise chunking.
ML emulators for climate model acceleration to reduce computational costs in climate science.
RL method for tuning autoregressive image models with instance and distribution-level rewards.
Research paper on logarithmic regret bounds for online convex optimization with bandit feedback.
Probabilistic bias correction improves AI and physics-based subseasonal weather forecasts.
Interpretable multivariate time series classification via anchor-routed MoE for high-stakes applications.
Method for LLMs to self-modify and consolidate in-context knowledge into long-term parameters through memory consolidation.
Contract-based compositional shielding enforces safety in decentralized multi-agent RL through runtime guarantees.
Code correctness is linearly decodable from LLM hidden states before generation; studies repair geometry in Qwen3.
Evolutionary framework discovers developmental reward schedules in RL combining agency, novelty, and reactivity components.
3D masked autoencoders outperform 2D variants for volumetric microscopy cell representations in self-supervised learning.
Continual learning approach for power forecasting in nonstationary energy systems using deep learning.
Temporal link prediction tradeoff analysis: Fisher information bounds on parameter estimation vs predictive accuracy in causal probabilistic temporal graphs.
ECHO: memory-efficient context management for long-horizon language agents via selective turn pruning and tracing; maintains interpretability of decision support.
World models as functional descendants of model-order-reduction techniques; argues modern self-supervised approaches reimplement decades-old control theory concepts.
Training-time information allocation view of implicit bias in neural networks; explains how optimization forms writing patterns for error signals during training.
PeTeR: post-training robustification of probabilistic circuits via distributionally-robust optimization to improve generalization under noise and distribution shift.
Omni-Sleep: foundation model for sleep staging using hierarchical contrastive learning on multimodal polysomnography signals (EEG, ECG, respiration).
Jet-Long: efficient long-context extension for LLMs using dynamic bifocal RoPE, enables zero-shot context extension 10x beyond pretraining window.
Budget-aware test-time selection for routing LLM queries across multiple models; resampling strategy improves cost-quality tradeoff vs single-router approach.
Fourier analytic method for learning Gaussian mixture models with theoretical guarantees on separating mixture components under distance constraints.
Ruby: ML-based tool to detect unsafe Rust regions in stripped binaries without source code access; addresses memory safety analysis gap.
Contrastive learning approach for multimodal EHR analysis combining structured codes and unstructured clinical notes via deep learning.
arXiv paper on accelerated first-order methods for bilevel and minimax optimization.