Theory of Code Space: Do Code Agents Understand Software Architecture?
Theory of Code Space: Benchmark evaluating whether code agents maintain coherent architectural understanding when exploring multi-file codebases.
Theory of Code Space: Benchmark evaluating whether code agents maintain coherent architectural understanding when exploring multi-file codebases.
Framework enabling AI agents to achieve collaborative context awareness by interpreting concurrent user actions on shared creative artifacts in real-time.
Hybrid LLM architecture fine-tuned for agricultural advisory with domain-specific knowledge to provide accurate, actionable recommendations.
Empirical evaluation of LLM robustness to five types of chain-of-thought perturbations: math errors, unit conversion, sycophancy, skipped steps, and reasoning corruption.
Phys4D: Three-stage pipeline ensuring physics consistency in 4D world models generated from video diffusion models through iterative refinement.
Framework combining LLM common-sense reasoning with robot planning in partially observable environments for task and motion planning.
Analysis of AI R&D automation extent and effects, proposing empirical frameworks beyond capability benchmarks to measure real-world AIRDA impact.
ICR: Framework integrating semiotics and hermeneutics to evaluate meaning in LLM-generated text summaries beyond traditional metrics.
vLLM Semantic Router: Inference framework for intelligent request routing across diverse multimodal LLM deployments using composable signal orchestration.
Traversal-as-Policy: Framework distilling LLM agent execution logs into verifiable Gated Behavior Trees for safe, robust autonomous agents with explicit policy control.
VDCook: A configurable platform for constructing video datasets for multimodal LLMs via natural language queries, combining real retrieval with controlled synthesis.
IntSeqBERT: Transformer encoder with modulo-spectrum embeddings for predicting integer sequences in OEIS, handling out-of-vocabulary large values.
Hybrid heuristic-reinforcement learning approach for optimizing railcar shunting operations in freight yards.
Method for extracting fair and unbiased subnetworks from standard trained models without additional data or complex procedures.
Study of tokenizer pretraining impact on physics foundation models for emulation and simulation of multiphysics phenomena.
Analysis of weak-to-strong generalization in random feature ridge regression, showing improved scaling laws when strong models train on weak labels.
Formal correspondence between Moore machines and state-space models, using automata learning for warm-starting SSM training.
Theoretical analysis of best-of-N sampling for LLM alignment, examining statistical optimality and vulnerability to reward model exploits.
Meta-reinforcement learning approach for multi-objective supply chain optimization in dynamic environments with hierarchical learning structure.
Benchmark for evaluating autonomous data science agents on Kaggle-style tabular ML tasks with time constraints, testing 10 open-source LLMs.
Research on merging task-specific models for domain generalization, analyzing parameter competition and singular value decomposition of task matrices.
Analysis of expert specialization in Mixture of Experts models through routing patterns and early decoding framework to understand inference behavior.
Parameter-efficient fine-tuning method for adapting foundation models to medical imaging tasks with limited data, including automated adapter configuration.
Study on test-time adaptation for LLMs using many-shot prompting, analyzing benefits, limits, and failure modes of in-context learning with large demonstration sets.
Research on using LLM reasoning with reinforcement learning for molecular optimization tasks, addressing limitations of fine-tuning on reference molecules without optimization trajectories.
Weak-SIGReg covariance regularization technique for stabilizing deep learning training in low-data regimes and architectures like Vision Transformers.
Omni-Masked Gradient Descent (OMGD) memory-efficient optimization via mask traversal with improved convergence for training large language models.
EvoESAP non-uniform expert pruning method for Sparse Mixture-of-Experts language models with evolved layer-wise sparsity allocation.
Identifies and addresses performance plateaus in PPO by scaling to 1M parallel environments, showing sample-based loss estimation issues during training.
Analysis of Langevin dynamics and stochastic weight averaging for high-dimensional estimation in tensor PCA and single-index models.
Dynamic Momentum Recalibration method for SGD variants that adaptively adjusts momentum coefficients to balance bias-variance tradeoff in gradient updates.
DQE semantic-aware evaluation metric for time series anomaly detection addressing bias, false alarms, and threshold sensitivity in existing metrics.
Partial Policy Gradients method for reinforcement learning in LLMs that optimizes subsets of future rewards to improve policy gradient accuracy.
Theoretical proof that Predictive Coding Graphs form a mathematical superset of feedforward neural networks, bridging neuroscience and ML.
Federated learning protocol using XGBoost surrogates for distributed health monitoring from wearable sensor data in spinal cord injury patients.
DC-Merge method for merging multiple task-adapted models while maintaining directional consistency in singular spaces and preserving task knowledge.
Synthetic Monitoring Environments (SMEs) benchmark suite for reinforcement learning with configurable tasks and known optimal policies for white-box agent diagnostics.
Learning-based approach (DeCoST) for solving orienteering problems with time windows and variable profits using decoupled discrete-continuous optimization.
Studies how agentic retrieval-augmented reasoning pipelines impact reliability in clinical radiology QA under model variability and heterogeneous LLM deployments.
Stem rethinks sparse attention mechanisms from information flow perspective to reduce quadratic computational complexity during LLM pre-filling for long contexts.
Three-stage pipeline to post-train LLMs for efficient calibrated uncertainty estimation, improving reliability in high-stakes decision-making applications.
ALFCG: adaptive projection-free optimization framework for stochastic composite nonconvex minimization without global smoothness constants or line search.
Adaptive bandit-based scheduling framework for multi-modal LLM inference under heterogeneous budgets, handling variable modality composition and latency constraints.
NOBLE adds nonlinear low-rank branches to transformer linear layers for pretraining from scratch, achieving parameter efficiency different from LoRA/PEFT adapters.
Proposes COLD-Steer, training-free framework for steering LLM activations via in-context one-step learning dynamics without retraining.
Introduces CODEC, sparse autoencoder method for causal interpretation of neural network computations via contribution decomposition.
Presents AllScAIP, attention-based machine learning interatomic potential using all-to-all node attention for long-range interactions.
Studies privacy preservation in sequential multi-agent LLM systems through information-theoretic controls against inference attacks.
Introduces RoboLayout for generating differentiable 3D scene layouts from language instructions feasible for embodied agent interaction.
Analyzes grammar-constrained LLM decoding as coupling between autoregressive distribution and reachability oracle over context-free grammars.