DIP: Efficient Large Multimodal Model Training with Dynamic Interleaved Pipeline
Training efficiency optimization for large multimodal models using dynamic interleaved pipeline to address stage imbalance and training data heterogeneity.
Training efficiency optimization for large multimodal models using dynamic interleaved pipeline to address stage imbalance and training data heterogeneity.
Lightweight, education-focused implementation of AlphaZero reinforcement learning framework addressing implementation complexity and reproducibility challenges.
Survey of computational persuasion covering AI-driven persuasion in conversational systems, including beneficial applications and ethical risks in politics and marketing.
Analysis of Transformer learning mechanisms through evolutionary biology lens, comparing in-weight learning and in-context learning strategies across environmental timescales.
Comparative analysis of uniform loss versus specialized multi-task optimizers, showing equal-weighted tasks can match specialized approaches with proper hyperparameter tuning.
Framework for robotic manipulation using Vision-Language-Action models with structured supervision for failure diagnosis and recovery in open-world scenarios.
Iterative preference learning method to improve GUI-based mobile agents' reasoning by generating diverse Chain of Action-Planning Thought trajectories with self-training.
Benchmark dataset and evaluation framework for language-prompted molecular structure recognition, editing, and generation tasks combining chemistry and LLMs.
Long-term time-series forecasting model using spectral methods and Wasserstein distance to capture geometric structure in chaotic systems rather than pointwise predictions.
Analysis of LLM-based recommender systems examining whether models understand collaborative signals and proposing diagnostic methods to improve reasoning.
Framework for incremental learning under concept drift with co-evolving label spaces and distributions in non-stationary data streams.
Evaluation framework using Erotetic Theory of Reasoning to characterize whether LLM reasoning errors match established human fallacy patterns across 38 models.
Semi-supervised image segmentation method using vision language model embeddings with domain-specific text anchors to bridge visual-textual semantic misalignment.
Systematic review and meta-analysis of 39 studies measuring LLM-assistant impact on software developer productivity across coding, testing, and debugging tasks.
Knowledge fusion framework integrating dynamic knowledge graphs with static LLMs through bidirectional information aggregation for up-to-date reasoning.
Chain of Retrieval method for scientific paper retrieval using iterative multi-aspect search expansion on full documents beyond abstract embeddings.
Benchmark characterizing state space models and hybrid architectures for long-context processing on local devices versus transformer baselines.
Method for cross-embodiment robot policy transfer using pretrained world models and action representations without embodiment-specific training.
Empirical analysis of model multiplicity in ML systems, examining how equally valid models produce conflicting predictions in medical applications.
Dual-stream temporal-field diffusion method for synthetic DDoS packet-level data augmentation to improve ML-based attack detection.
DMFI framework using LoRA-tuned language models for insider threat detection via dual-modality log analysis combining semantic and behavioral signals.
Investigation of long chain-of-thought reasoning capabilities across 38 models and multiple languages beyond English, analyzing scaling and post-training effects.
Multi-agent visual navigation method combining goal-conditioned RL and conflict-based search with iterative risk allocation for safe coordination.
WorldForge framework for training-free 3D/4D generation from video diffusion models using zero-shot camera control without fine-tuning.
Curriculum learning method for reasoning LLMs using gradient analysis to select training data efficiently, reducing computational waste.
Method reducing LLM overthinking via cumulative entropy regulation to adaptively adjust chain-of-thought reasoning depth.
Video diffusion model approach for generating physically plausible rigid body interactions with object-level control.
Memory-augmented architecture and pretraining strategy separating long-tail and common knowledge for efficient edge deployment.
Method for interpreting weight changes from language model finetuning without access to full training datasets.
Auditing framework using martingale theory to detect token misreporting in cloud-based LLM pay-per-token pricing mechanisms.
Dataset and method for knowledge-based visual QA with explicit reasoning traces using multimodal LLMs without external retrieval.
Investigation of logical consistency failures in video-language models through cross-modal attention analysis.
Analysis showing VAR generative models can function as efficient, interpretable classifiers with benefits over diffusion-based approaches.
Techniques for improving vision-language models' understanding of negation in object detection using chain-of-thought reasoning.
Analysis of Lagrangian methods for safe reinforcement learning, focusing on optimal Lagrange multiplier selection strategies.
Benchmark extending MMLU-Redux to evaluate LLM performance across Arabic dialects beyond Modern Standard Arabic.
Human-centered design approach for creating AI benchmarks with ecological and construct validity for journalism practitioners.
FPGA-based acceleration for LLM inference using lookup table computations and memory-based approaches for efficient hardware deployment.
Method combining multiple LLM models with different cost/accuracy tradeoffs using policy optimization to minimize inference costs.
Algorithm audit evaluating quality and consistency of Google's AI-generated search results on baby care and pregnancy queries.
Framework using LLMs to discover interpretable behavioral rules for autonomous vehicles from real-world data.
Benchmark for evaluating LLM ability to follow complex lexical instructions, addressing limitations of human and LLM-as-judge evaluation methods.
Self-supervised learning method for procedural workflows using Plackett-Luce ranking to improve temporal sequence understanding in activities.
Adversarial fine-tuning evaluation of open-weight genomic language models to assess robustness and misuse potential for pathogenic sequence generation.
Reinforcement learning approach for real-time long-horizon air quality forecasting using group-relative policy optimization in complex terrain regions.
Analysis of LLM price-performance tradeoffs using largest dataset of benchmark costs and model prices, showing progress per dollar differs from benchmark progress.
Deep learning approach for automated post-disaster damage assessment from satellite imagery using semantic segmentation and change detection.
Efficient vision-language model approach with adaptive visual token acquisition that dynamically determines compression ratios based on task requirements.
Basque language encoders capturing linguistic diversity including dialects, historical, and informal text to reduce representational bias in language models.
Novel local explanation method for black-box ML models using multivariate adaptive regression splines and stratified sampling for high-fidelity interpretability.