Deep learning framework using digital network twin for optimizing reinforcement learning training in multi-fidelity 5G networks with antenna tilt adjustment.
Guardian system combines reinforcement learning with LLM-based quality assurance to create spatiotemporal risk surfaces for missing-child search planning from unstructured case documents.
BiCLIP adapts vision-language models to specialized domains via structured geometric transformation, extending canonical transformation theory to domain adaptation.
Research on using machine learning for statistical inference with scientific simulators, focusing on hypothesis testing and model refinement.
Guardian system uses multi-LLM pipeline for intelligent information extraction in missing-person investigations, coordinating end-to-end execution across tasks.
Survey introducing reinforcement learning methods to economists, addressing curse of dimensionality in complex economic models.
Reinterprets generative AI through statistical lens using flow matching, connecting generative models to causal inference and interpretability.
LLM serving system for mobile devices with hardware-based isolation using ARM TrustZone to protect model weights and user data from kernel attacks.
Autonomous AI agent for clinical triage in remote patient monitoring, using 21 medical tools to process vitals data 24/7 without physician bottleneck.
Data curation method for robot learning using influence functions to select high-quality demonstrations from noisy human teleoperation data.
Taxonomy and evaluation framework for latent world models and vision-language-action systems in autonomous driving.
Reinforcement learning approach for dense image captioning using rubric-guided optimization to improve diversity and generalization.
Examination of logical reasoning as mechanistic pathway to situational awareness in advanced AI systems, exploring emergent capabilities risks.
Study of emotion as latent representational factor in LLM reasoning and text processing, beyond sentiment classification tasks.
Vision-language models that self-evolve from zero data without seed images, extending self-improvement paradigms from LLMs to multimodal systems.
TrainDeeploy: framework for hardware-accelerated parameter-efficient fine-tuning of transformer models on extreme-edge devices with memory constraints.
Study of subliminal learning where student LLMs acquire behavioral traits from teacher models via synthetic data training on unrelated domains.
BRACE framework for bandits with noncompliance addressing objective selection between recommendation welfare and treatment learning.
SCDP: sensor-conditioned diffusion policies for humanoid locomotion using only onboard sensors without privileged state estimation.
Multi-DNN inference system for edge devices supporting sparse model variants across heterogeneous processors with improved resource matching.
Study of photonic quantum machine learning under noise, exploring integration of quantum computing with ML for scalable quantum information processing.
EsoLang-Bench: benchmark using esoteric programming languages to evaluate genuine reasoning in LLMs beyond memorization on code tasks.
C2FMAE: coarse-to-fine masked autoencoder for self-supervised visual pre-training combining global semantics with fine-grained detail.
Study on how reasoning affects honesty in LLMs using moral trade-off dataset, finding reasoning increases rather than decreases honesty.
Survey on decentralized federated learning removing central coordinator and using peer-to-peer coordination for collaborative training.
Study of structured lottery tickets in over-parameterized CNNs, connecting pruning to strong lottery ticket hypothesis for computational efficiency.
Sparse Variational Student-t Processes framework extending Gaussian processes for scalable heavy-tailed data modeling.
HYGENE: diffusion-based generative model for creating realistic hypergraphs used in social networks, bioinformatics, and recommender systems.
Research on robust neural network training at arbitrary precision and sparsity, addressing gradient flow issues in quantization and sparse regimes.
ARLBench benchmark for hyperparameter optimization in reinforcement learning agents, addressing evaluation costs and generalizability across domains and algorithms.
Scalable graph neural networks replacing attention with message passing in transformer blocks for large graphs.
Improved RL training method for LLM reasoning handling all-negative sample groups in GRPO framework.
Jailbreak attack on safety-aligned LLMs via self-introspection without requiring model weight access.
Systematic evaluation of quantized LLMs on edge devices, testing models 0.5B-14B with seven PTQ methods.
RL framework using SAT problems to improve LLM reasoning with automatic verification and scalable training.
Large-scale benchmark evaluating ML solvers on combinatorial optimization with real-world industrial datasets.
Semi-supervised extension of conformal prediction using unlabeled data for uncertainty quantification.
Meta-learning approach to rate time series data quality using LLM judgments across diverse domains.
Framework addressing multivariate time series challenges: channel dependencies, asynchronous sampling, and missing values.
Neural latent dynamics model using Langevin equations with VAE framework for capturing spiking activity patterns.
MLES: Multimodal LLM-assisted evolutionary search discovers interpretable programmatic control policies for reinforcement learning.
CTRL: Clustered Transfer Residual Learning for multi-source ML datasets maintaining per-source reliability and differences.
Graph neural networks informed by RF physics principles for accurate, data-efficient prediction of circuit performance.
Iterative in-context learning strategy improves LLM generalization on algebraic reasoning tasks and out-of-distribution examples.
Score-based generative model using Kuramoto dynamics on periodic domains for orientation-rich image generation.
Test-time entropy minimization method using asymmetric learning to adapt models to novel environments without logit inflation.
Shows expert pruning outperforms merging for compressing sparse mixture-of-experts models on generative tasks.
Bradley-Terry policy optimization for extending RL-based training to non-verifiable tasks with only pairwise human preference supervision.
Reinforcement learning framework combining permutation relative policy optimization to improve LLM reasoning on tabular prediction tasks.
Domain-incremental learning framework for graph models that preserves knowledge across multiple graph domains.