A Context Engineering Framework for Improving Enterprise AI Agents based on Digital-Twin MDP
Framework for improving LLM-based enterprise agents via offline RL and digital-twin simulation without extensive real-world data.
Framework for improving LLM-based enterprise agents via offline RL and digital-twin simulation without extensive real-world data.
GSEM: Graph-based memory framework for clinical reasoning agents that organizes experiences with relational structure.
SpecTM: Physics-informed masking approach for foundation models in Earth observation with trustworthiness constraints.
MARCUS: Agentic multimodal vision-language model for automated cardiac diagnosis from ECGs and echocardiograms.
Framework using LLMs and graph analytics to analyze how interdisciplinary research teams develop shared knowledge over time.
Safety enhancement technique exploiting linear separability of harmful and safe query embeddings in LLMs to improve robustness against adversarial prompts.
RedacBench: comprehensive benchmark for evaluating LLM-based redaction systems' ability to remove sensitive information from unstructured text.
KidGym benchmark inspired by Wechsler Intelligence Scales evaluating multimodal LLMs on 2D grid-based reasoning tasks measuring interpretable abilities.
CRoCoDiL: continuous semantic space diffusion model for language addressing token dependencies and semantic coherence in masked diffusion models.
Qualitative study of middle school teachers using LLM chatbots in block-based programming environments examining interaction patterns and affect.
Mixed-method study examining governance of GenAI in academic peer review through social media analysis and area chair interviews.
Locally Coherent Parallel Decoding method improving diffusion language models' token generation by capturing joint dependencies instead of marginal distributions.
Empirical comparison of on-device LLM inference configurations measuring latency-energy-learning tradeoffs for AI tutoring using Phi-3 Mini.
Analysis of expert personas in LLMs revealing why Wharton study found null results; demonstrates structural predictability of outcome through baseline contamination.
Framework evaluating LLMs' ability to simulate Americans' political opinion distributions across policy issues for polling augmentation.
Preordered Multi-Objective MDP framework for autonomous driving addressing conflicting objectives like safety and efficiency without scalar reward collapse.
HR Simulator game studies LLM email generation for workplace communication, comparing human and model outputs on social task performance.
Qualitative study comparing literature reviews generated by LLMs with varying corpus selections, revealing biases and factual gaps in outputs.
Sequence-to-sequence decoder for brain-computer interfaces converting intracortical neural activity to speech with robustness improvements.
SciNav: Agent framework built on LLMs for autonomous scientific coding tasks with objective evaluation through executable benchmarks.
DESRO framework uses LLMs to infer intermediate scientific reasoning steps from experimental outcomes to enable automated scientific discovery and molecule optimization.
Multi-agent reinforcement learning framework for coordinating UAV networks with joint communication and sensing under resource constraints for monitoring missions.
Domain-aware layer pruning analysis of vision-language models with focus on math reasoning and perception coupling.
Fully open-source pipeline for synthesizing long-horizon research trajectories for training deep research agents using reproducible methods.
Multi-agent RL with learned communication protocols for heterogeneous agent coordination in autonomous cyber defense systems.
Empirical analysis of algorithmic collusion fragility in LLM agents with heterogeneous parameters using 2000+ compute hours of experiments.
Multi-agent RL approach for incremental causal DAG discovery from observational data with improved efficiency for online applications.
Collaborative knowledge distillation with adaptive curriculum for heterogeneous distributed learning in edge-based visual analytics.
RAG method using hierarchical code abstraction and architecture-guided retrieval for complex theory-driven codebases in algorithmic game theory.
Argues for redesigning software interfaces from human-centric to agent-centric as LLM agents become primary consumers of systems.
Framework for cooperative autonomous agents to decide what V2X network messages to transmit based on reasoning about receiver benefit.
Natural language-driven LLM agent (kRAIG) for automating data pipeline generation and ETL workflow construction.
Vector-based semantic approach to automatically discover and select relevant tools from large MCP server toolsets for LLM agents.
Comparison of Model Context Protocol (MCP) versus RAG for financial Q&A, showing direct system integration outperforms document retrieval.
Empirical study of how executable tool access changes safety alignment in LLM agents, moving beyond text-centric safety evaluations.
GIP-RAG: Retrieval-augmented framework for interpretable gene interaction and pathway analysis enabling multi-step reasoning across biological networks.
Analysis of multi-agent LLM pipeline contradictions: identifies selection bottleneck threshold determining whether agent diversity improves output quality.
Theoretical analysis of coupled learning dynamics in heterogeneous tri-hierarchical drone swarm systems operating at different timescales.
LLM-driven algorithmic debugging framework for code generation using formal program debugging theory to improve ARC-AGI-2 task performance.
GEM: Native graph-based indexing algorithm for multi-vector retrieval enabling finer-grained semantic matching beyond single-vector approaches.
ContractSkill: Framework for creating repairable multimodal web agent skills with explicit contracts specifying preconditions, steps, and success criteria.
MANA: Agentic multimodal framework using UI navigation and reasoning for detecting mobile ads, combining static and runtime analysis.
Leum-VL technical report on multimodal model for analyzing short video structure and temporal organization beyond scene description.
Analysis of memory poisoning attacks on agentic AI and multi-agent systems, covering semantic, episodic, and short-term memory vulnerabilities.
WebNavigator uses interaction graph retrieval to enable web navigation agents to overcome topological blindness via global environment structure.
ALARA framework for managing multi-agent systems with least-privilege context isolation and composable team coordination.
Study of adversarial attacks targeting cooperative multi-agent reinforcement learning systems in real-world applications.
SymCircuit uses RL to learn probabilistic circuit structures via entropy-regularized search instead of greedy algorithms.
Research on optimizing KV cache memory usage in Transformer-based LLMs to handle longer context windows efficiently during inference.
Meta-learning algorithms for repeated Bayesian persuasion exploiting structural similarity across strategic interaction tasks.