AgenticRecTune: Multi-Agent with Self-Evolving Skillhub for Recommendation System Optimization
AgenticRecTune: multi-agent framework with self-evolving skill hub for optimizing multi-stage recommendation system pipelines.
AgenticRecTune: multi-agent framework with self-evolving skill hub for optimizing multi-stage recommendation system pipelines.
COHERENCE benchmark for evaluating multimodal LLMs on fine-grained image-text alignment in interleaved document-like contexts.
Multi-agent LLM benchmark for negotiation that tests dynamic grounding and communication repair across conversational turns.
Gyan: neuro-symbolic language model combining transformers with symbolic reasoning to improve compositionality, interpretability, and reduce hallucinations.
Framework for evaluating agentic stock prediction systems using LLM judges and closed-loop reinforcement learning feedback on behavioral dimensions.
Asymmetric on-policy distillation method for training student language models with token-level teacher feedback, improving upon standard RL and off-policy approaches.
AI CFD Scientist: open-source physics-aware AI agent for autonomous computational fluid dynamics discovery with LLM-based scientific reasoning loop.
Benchmark for LLM-assisted formal mathematical reasoning using Lean and Mathlib, evaluating pull request merge-readiness for library contributions.
Study on when neural networks fail to extrapolate out-of-distribution, analyzing feature learning versus data-generating-process identifiability.
Mechanistic analysis of hallucination failures in vision-language models, tracing issues to geometric over-alignment in decoder-based VLMs.
FactoryNet: industrial time-series pretraining dataset with 51M datapoints across 23k task executions for zero-shot transfer and anomaly detection.
Research on FP4 quantization for pretraining large language models, investigating stability and convergence issues in full-pipeline low-precision training on Llama 3.1-8B.
Key-Value Means presents block-recurrent attention mechanism achieving O(N) complexity with fixed or growing state for efficient long-context transformer inference.
Tool combining weakest-precondition analysis with agentic Claude Code CLI for specification inference in Move Prover to reduce verification boilerplate.
Metis framework reformulates LLM jailbreaking as inference-time policy optimization using adversarial POMDP for improved red teaming scalability.
CoWorld-VLA multi-expert world model framework for vision-language-action autonomous driving with planning-oriented spatiotemporal representations.
ALAM latent action model for vision-language-action models that extracts action priors from video using algebraically consistent representations.
Formal framework for probabilistic safety shielding in Markov decision processes with conservative guarantees.
DataMaster autonomous system for data-centric ML research automating dataset discovery, adaptation, validation, and knowledge propagation.
TMPO trajectory matching policy optimization for diffusion alignment addressing reward hacking through probability distribution constraints.
MCPShield attack detection framework for LLM agent tool-call traffic via Model Context Protocol, using graph encoding and embeddings.
HEPA self-supervised architecture for event prediction in multivariate time series using horizon-conditioned JEPA pretraining.
LiBaGS lightweight method for selecting informative synthetic training data by scoring boundary proximity, uncertainty, and data density.
ZeNO gradient-free noise optimization method for reward alignment in diffusion and flow models without backpropagation.
Reasoning-prefix masking technique for distilling visual-reasoning capabilities from large VLMs into compact student models.
FAMeX algorithm for AI explainability using feature association maps based on graph-theoretic approaches.
Safe reinforcement learning approach that learns when agents should act via communication-efficient timing decisions under Lyapunov safety constraints.
CAWI initialization method for randomized neural networks using copulas to capture inter-feature dependencies.
Federated multimodal graph learning approach addressing modality heterogeneity and incomplete data across distributed networks.
OceanCBM concept bottleneck model for interpretable ocean forecasting that provides mechanistic explanations aligned with physics.
Study on decision-making with AI assistance, examining how decision-makers interpret model confidence and prediction utility in high-stakes domains.
Theoretical analysis establishing population risk bounds for Kolmogorov-Arnold Networks trained with mini-batch SGD and differential privacy.
Embedding Temporal Logic framework for runtime monitoring of perception-based autonomous systems without expensive learned abstraction modules.
Multi-rollout on-policy distillation method for LLMs that leverages peer successes and failures to provide denser token-level supervision beyond sparse verifier rewards.
FPILOT framework applies inference-time optimization via Model Predictive Control to RL trading agents for portfolio management, incorporating price forecasts at deployment.
Method for robust LLM alignment using ordinal decomposition of discrete rewards in RLHF with stochastic auto-raters for long-form QA and instruction following.
IGT-OMD: implicit gradient transport for decision-focused learning with delayed feedback in online bilevel optimization.
Method for node classification in multiplex graphs with heterophily using adaptive cross-domain learning approach.
UFO: domain-unification-free neural operator framework for learning cross-domain function space mappings.
Research on how upstream training choices affect model robustness when capabilities are retained through subsequent fine-tuning.
RSNet: open-source R package for robust network inference in high-dimensional data using resampling-based framework.
Research on Spectral Energy Centroid metric for improving performance and analyzing spectral bias in implicit neural representations.
Framework measuring layer-wise representation dynamics in LLMs using Frenet geometry, neighborhood retention, and information metrics.
Scaling laws for mixing scarce target data with abundant generic data during LLM pretraining under data constraints.
Analysis of safety probe failures in LLMs where jailbreak evidence distributed across tokens bypasses final-token readout.
Nonparametric framework for learning task-relevant specialist representations from generalist models with identifiability analysis.
Reflection-enhanced self-distillation for LLMs to learn from rare successful interactions with rich environmental feedback.
LoRA initialization method using gradient surgery to mitigate catastrophic forgetting in continual LLM fine-tuning.
Constraint-aware flow matching for physics-informed generative models with strict constraint satisfaction.
Inference-time machine unlearning via gated activation redirection to remove memorized training data from LLMs without retraining.