Meta-weighted online sampling approach for aligning LLMs by reducing distribution mismatch between offline preference data and evolving model policy.
MobileLLM-R1: sub-billion parameter language models with chain-of-thought reasoning and open training recipes, challenging assumptions about model size requirements.
DataMind: scalable data-analytic agent system with open-source training recipes for multi-step reasoning over diverse-format, large-scale data files.
BEV-VLM: trajectory planning approach using vision-language models with bird's-eye-view representations from fused camera and LiDAR data.
VoiceBridge: one-step latent bridge model for general speech restoration from diverse distortions with energy-preserving VAE design.
Max-V1: vision-language model framework for autonomous driving that formulates trajectory planning as next waypoint prediction via language.
FLOP: score-based causal discovery algorithm for linear models using fast parent selection and Cholesky-based updates to find optimal causal graphs.
CMT-Benchmark: 50 expert-level condensed matter theory problems for evaluating LLMs on advanced scientific reasoning and code generation.
Permutation-invariant feature selection method using generative models to capture feature interactions while improving robustness and privacy.
Flow matching variant (Carré du champ flow matching) that improves quality-generalization tradeoff in generative models through geometry-aware noise regularization.
Backdoor attack on vision-language-action models demonstrating vulnerability to behavioral hijacking via hidden training triggers.
Bayesian optimization method using LLM fine-tuning to perform Thompson sampling in large discrete spaces without gradient computation.
Training-free framework for improving vision-language models on information-dense images with text and graphical elements.
Study of continual pre-training for adapting LLMs to low-resource French dialects under tight compute and data constraints.
Investigation of in-context learning across transformer, state-space, and hybrid LLM architectures using behavioral and intervention methods.
Study of misconceptions novice programmers have about LLM-based coding assistants, examining impact of tool capabilities and extensions.
Training method combining supervised learning and reinforcement learning to improve multi-step reasoning in open-source LLMs.
Agentic multimodal model framework enabling tool invocation (code execution, web search) and reasoning integration for vision-language tasks.
Analysis of how LLMs shift moral judgments under persona role-play, introducing benchmark metrics for moral susceptibility and robustness.
Open benchmark for deep learning-based event reconstruction in neutrino telescope data using inverse problem solving.
Diffusion language model using Mamba backbone for efficient inference, achieving higher throughput than transformer-based alternatives.
Benchmark for evaluating embodied AI agents on interaction with physical interfaces (switches, panels, GUIs) in complex environments.
Machine learning system for automating data quality monitoring and anomaly detection in particle physics collider experiments.
SocialNav foundation model for socially-aware embodied navigation with hierarchical architecture trained on 7M samples for human-compliant trajectory generation.
Heterogeneous multi-agent reinforcement learning with attention mechanism for automated feature transformation on structured data.
QKAN-LSTM combining quantum-inspired Kolmogorov-Arnold networks with LSTM for improved sequential modeling with reduced parameter redundancy.
WisPaper end-to-end agent system for academic literature discovery and organization combining semantic search verification with workflow integration.
FRIEDA benchmark evaluating vision-language models on multi-step cartographic reasoning with map interpretation for disaster response and urban planning.
Generalized Primal Averaging optimizer extending Nesterov's method for faster LLM training, unifying DiLoCo and schedule-free approaches with reduced memory requirements.
Trust region masking technique for LLM reinforcement learning addressing off-policy mismatch and approximation errors from implementation divergences in policy gradient optimization.
LIA supervised fine-tuning approach using LLMs for automatic software issue assignment in large open-source projects without heavy project-specific training data.
CSyMR benchmark for compositional music information retrieval testing LLMs on multi-step reasoning over symbolic music scores and natural language queries.
DUET method for LLM unlearning via distillation from a contextualized teacher, removing undesirable knowledge without retraining while avoiding catastrophic forgetting.
LEC-KG framework combining LLM semantic understanding with knowledge graph embeddings for automated domain-specific knowledge graph construction from unstructured text.
Reinforcement learning approach for training humanoid whole-body controllers that generalize across diverse robot embodiments with varied dynamics and degrees of freedom.
Empirical study of 32K LLM agents on Chirper.ai social media platform analyzing collective behaviors, biases, and exclusionary dynamics across 7M posts.
Study on how LLM-generated personalized messages in behavior-change systems affect user perceptions through exposure patterns rather than individual message quality.
Risk-sensitive evaluation framework for LLM hallucinations in medical advice, assessing clinical harm severity beyond factual correctness.
Automated black-box pipeline detecting unverbalized biases in LLM chain-of-thought reasoning without predefined categories using task-specific evaluation.
RooflineBench: Roofline model-based benchmarking framework for characterizing performance of Small Language Models on edge hardware.
Agentic system for respiratory disease diagnosis using multimodal sound generation and active adversarial curriculum learning.
Framework extending Item Response Theory to measure AI model propensities and behavioral tendencies beyond capability metrics.
Agentic system for hierarchical urban geospatial modification using multimodal models to handle dependency-aware city planning changes.
Analysis showing test-time training with KV binding can be expressed as learned linear attention mechanism.
Federated learning aggregation method using gradient-based weighting to address client drift and data heterogeneity.
Safety filtering framework for flow-based generative models providing formal guarantees that generated samples satisfy hard constraints.
Training approach for collaborative AI agents using strategic risk aversion to improve generalization when paired with new partners.
Adversarial dataset and training method to improve multimodal LLM robustness when handling visually complex scenes.
Framework for systematically mapping unsafe regions in LLMs using MAP-Elites quality diversity approach to characterize failure modes.
Improved FSDP implementation supporting structure-aware training and non-element-wise optimizers for large-scale model training.