ActionNex: A Virtual Outage Manager for Cloud
Production-grade agentic system for cloud outage management supporting real-time updates, knowledge distillation, and role-conditioned action recommendations from multimodal operational data.
Production-grade agentic system for cloud outage management supporting real-time updates, knowledge distillation, and role-conditioned action recommendations from multimodal operational data.
Presents energy-based governance framework connecting transformer inference dynamics to constraint-satisfaction models for AI safety and inference-layer control.
Proposes explainable routing architecture for agentic workflows that decomposes tasks to specialized models, recording cost-capability trade-offs with clear rationale for routing decisions.
Develops automated framework using LLMs to compare AI safety policy documents under shared taxonomy, extracting activities and producing summaries with similarity scores.
Introduces Chronos AI Historian, an agentic system enabling historians to extract structured data from primary source image scans via natural language interactions.
Investigates mechanisms of hallucinations in decoder-only Transformers by modeling next-token prediction as graph search, analyzing path reuse and compression in attention patterns.
Studies adaptive reward design in deep reinforcement learning for satellite scheduling, finding static rewards outperform dynamic ones due to PPO stability requirements.
Neuroevolution study of chess agents with learned plasticity, examining Baldwin effect and behavioral diversity emergence.
Selective forgetting approach for Large Reasoning Models to prevent knowledge leakage through chain-of-thought reasoning steps.
Rashomon Memory: Multi-perspective agent memory architecture supporting conflicting interpretations of events for concurrent goals.
Analysis of entropy and attention dynamics in small language models on TruthfulQA, explaining how internal behavior affects factual accuracy.
Foundation model trained on 1.2M spatial transcriptomics profiles combining gene expression mapping with histology for biological discovery.
Comparison of single-agent vs multi-agent Vision Language Models for automated video analysis of collaborative learning behaviors.
Framework addressing RAG limitations in Generative Engine Optimization by modeling confidence decay and proposing deterministic agentic alternatives.
TableVision: Benchmark for multimodal LLM reasoning on complex hierarchical tables, identifying perception bottlenecks in structured data processing.
PRAISE: Training method for agentic search using prefix-based rollout reuse to improve RL efficiency and address reward sparsity in multi-turn QA.
Fuzzy Analytic Hierarchy Process for structured multi-criteria LLM evaluation with confidence-aware uncertainty modeling via triangular fuzzy numbers.
Deep RL framework for optimizing land-use allocation in Lake Malawi Basin to maximize ecosystem service value.
Research on communication with cross-timestep delays in cooperative multi-agent reinforcement learning, analyzing coordination under temporal misalignment.
QualAnalyzer: Open-source Chrome extension for atomistic LLM analysis in qualitative research, preserving prompts and outputs for auditability.
PolySwarm: Multi-agent LLM framework deploying 50 diverse personas for real-time prediction market trading and latency arbitrage on decentralized platforms.
FeynmanBench: Benchmark for evaluating multimodal LLMs on physics diagram reasoning, testing structural logic beyond local information extraction.
discourse_simulator: Open-source framework combining LLMs with agent-based modeling to simulate attitude diffusion on social issues like immigration.
CODE-GEN: Human-in-the-loop RAG-based agentic system with Generator and Validator agents for creating multiple-choice coding comprehension questions.
SkillFoundry: Self-evolving framework that builds agent skill libraries by integrating heterogeneous scientific resources including APIs, scripts, notebooks, and papers.
Framework for quantifying trust in autonomous AI agents through end-to-end outcomes, financial risk management, and operational safety in real-world deployments.
FactReview: LLM-based peer review system that grounds claims in evidence from related work and code, addressing presentation bias in manuscript evaluation.
Framework using generative AI to produce evidence-linked formal arguments for high-stakes system certification and compliance.
InsTraj uses diffusion models with travel intent instructions to generate realistic GPS trajectories for applications.
Profile-Then-Reason framework reduces latency in tool-augmented LLM agents by pre-synthesizing workflows before reactive execution.
LLM poker agents autonomously develop theory-of-mind-like opponent models through extended dynamic interaction.
Philosophical model proposing systematic understanding framework for deep learning systems with internal model tracking.
CoALFake combines active learning with human-LLM co-annotation for cross-domain fake news detection.
Evaluates LLM adaptation in reversal-learning tasks under non-stationary uncertainty with performance criterion triggers.
Study reveals evidence collapse in reasoning VLMs where models lose visual grounding despite high confidence predictions.
TimeSeek benchmark evaluates reliability of agentic LLM forecasters on prediction markets across lifecycle stages.
Four-layer pedagogical safety framework for educational RL systems with Reward Hacking Severity Index metric.
Combee enables scaling prompt learning for self-improving LLM agents through parallelized learning from multiple agent runs.
MC-CPO addresses reward hacking in RL-based adaptive tutoring systems through constrained policy optimization.
Context Engineering methodology for structuring informational context in AI tool usage to improve output quality beyond prompting.
Position paper on failure modes in agentic information retrieval systems, analyzing error cascades in multi-step reasoning workflows.
InferenceEvolve uses LLMs in evolutionary framework to discover and refine causal inference methods automatically.
Study on selecting widened warm starts for language model checkpoints through candidate-selection over training states.
RESCORE uses LLMs to reconstruct executable simulations from control systems research papers, with benchmark of 500 CDC papers.
Soft Tournament Equilibrium: evaluation framework for non-transitive AI agent interactions producing set-valued rankings instead of linear orderings.
Thermodynamics-inspired GeoAI method for modeling spatial heterogeneity and critical transitions in geography/environmental science.
Surrogate goals strategy for LLM agents to reduce bargaining failure risks. Agent prioritizes principal-specified goals over threats during negotiation.
Domain-Contextualized Inference: substrate-agnostic architecture treating domain as first-class parameter. Enables reasoning over symbolic, neural, and hybrid systems.
RoboPhD: LLM-guided evolution of AI agents under tight evaluation budgets. Compares optimization algorithms for improving agent architecture and prompts.
REAM: novel method for pruning experts in Mixture-of-Experts LLMs by merging experts during pruning. Reduces memory requirements for deployment.