SimulCost benchmark evaluates LLM agents on physics simulation tasks with cost-aware metrics, accounting for simulation time and experimental resource usage beyond token costs.
Conversational query rewriting approach for multimodal image retrieval with multi-turn dialogue dataset.
GeoBlock infers optimal block sizes for diffusion language models by analyzing token dependency geometry to enable efficient parallel decoding.
FEMBA: bidirectional Mamba state-space model pre-trained on 21k hours EEG with physiologically-aware objectives for microcontroller deployment.
Enhanced mixture-of-experts architecture using soft nearest neighbor loss to prevent expert collapse and redundant representations.
Cross-lingual evaluation of vision-language models on visual reasoning tasks across Indian languages, revealing performance disparities.
Framework integrating sparse autoencoders with dynamic head pruning in Vision Transformers for interpretable and controllable efficiency.
MotionGPT3 replaces diffusion with rectified flow objectives for efficient text-driven motion generation with improved convergence.
Evolution strategies warm-start reinforcement learning agents for industrial continuous control using CMA-ES-generated demonstrations.
LogicDiff improves reasoning in masked diffusion language models by prioritizing logical connective tokens during inference-time denoising.
Privacy-preserving inference for spiking neural networks using fully homomorphic encryption to enable encrypted computation.
Language-conditioned multi-game procedural level generation using shared neural representations across different game domains.
Study of why minimal GPTs fail at out-of-distribution arithmetic generalization, revealing staged failures from layout barriers to positional encoding limits.
Survey of uncertainty-aware explainable AI methods, examining integration of uncertainty quantification (Bayesian, Monte Carlo, Conformal) into explanatory pipelines.
Controlled evaluation of LLM implementation choices for political text annotation, testing model selection, size, and prompt engineering best practices.
Comparative evaluation of physics-informed neural networks and neural ODEs for modeling nonlinear neuronal dynamics on Morris-Lecar model.
KOMET: model-agnostic framework using Koopman operators to track parameter evolution and handle temporal domain drift in non-stationary environments.
ASTER: agentic toolkit using LLMs for exoplanet research workflows, combining archival queries, literature search, and radiative transfer models.
Online statistical inference framework for sample-averaged Q-learning to reduce variance and improve stability in reinforcement learning.
Analysis of reliability limits in LLM-based multi-agent planning systems modeled as decision networks with language-based communication constraints.
FormalProofBench: benchmark evaluating whether LLMs can produce formally verified graduate-level mathematical proofs using Lean 4.
Lightweight neural network for super-resolution imaging from low-resolution SPAD arrays, reconstructing 256x256 images on embedded devices.
Comparative study evaluating YOLO object detection models on robotics tasks using custom and COCO2017 datasets for workspace object detection.
Study of incentive collapse paradox in AI-assisted task delegation showing accuracy improvements require unbounded payments without intervention mechanisms.
Persona-based LLM approach for simulating diverse human opinions at population scale for social science interventions and consequence modeling.
Theoretical analysis of loss landscape geometry in regularized deep matrix factorization proving unique minimizers under weight decay.
Information-theoretic framework for forecasting measuring mutual information between future observations and information set as predictability limit.
Sovereign Context Protocol defines open runtime attribution layer for human-generated content used in LLM training and inference.
Systematic evaluation of segmentation and geospatial foundation models for global field boundary segmentation using FTW benchmark.
Bayes-MICE extends multiple imputation for time series missing data using Bayesian inference and MCMC sampling.
Conformal Prediction Assessment framework for evaluating conditional coverage validity in distribution-free prediction with finite-sample guarantees.
StretchCast global-regional AI weather forecasting framework using variable-resolution cubed-sphere mesh for refined regional predictions.
Comparative analysis of AI datasets, foundation models, and barriers to achieving general-purpose AI in surgical image analysis.
D-SPEAR dual-stream replay mechanism for stable off-policy reinforcement learning in robotic manipulation with contact-rich dynamics.
Rainbow-DemoRL combines multiple demonstration-augmented reinforcement learning strategies to improve sample efficiency using offline data.
CarbonEdge framework for carbon-aware deep learning inference at network edge, extending model partitioning to optimize environmental impact alongside latency.
High-performance engine for low-bit matrix-vector multiplication enabling efficient inference in neural networks, vector databases, and LLMs.
Open-source benchmark evaluating four AI-powered people search platforms across 119 queries for recruiting, sales, and expert search use cases.
Introduces Hidden Ads backdoor attack class exploiting Vision-Language Models' recommendation behavior to inject unauthorized advertisements through natural triggers.
Extends Bellman Deviation Detection framework for model-free RL to detect man-in-the-middle attacks in cyber-physical systems with refined MDP attack models.
RTLSeek uses multi-stage reinforcement learning to improve LLM-based RTL/Verilog generation with diverse hardware design implementations.
Composer paradigm for test-time instance-specific parameter composition enabling adaptive generative models.
Neural Gaussian mixture model using energy score guidance for predictive uncertainty quantification in machine learning.
LVRPO framework for language-visual alignment in multimodal foundation models using GRPO for understanding and generation.
KAT-Coder-V2 agentic coding model using five expert domains with specialized fine-tuning and unified distillation for software engineering tasks.
GPU-accelerated JAX library for SGP4 orbital propagation of mega-constellations enabling efficient space situational awareness.
ImagenWorld benchmark with 3.6K condition sets for stress-testing image generation models across six core tasks with human evaluation.
Physics-informed neural networks framework (Deflation-PINNs) that identifies multiple distinct solutions to nonlinear PDEs.
Privacy-preserving federated learning framework using flow-matching generation to improve robustness and aggregation in distributed training.
Analysis of prompt injection attacks against LLM agents, tracking attack pipeline stages and defense mechanisms across five frontier models.