SSGM Framework: Addresses memory governance, semantic drift, and privacy risks in long-term memory systems of autonomous LLM agents.
Deliberative Collective Intelligence: Framework enabling multi-agent LLM systems to conduct structured deliberation with typed epistemic acts.
DocSage: LLM agent for multi-document multi-entity question answering using information structuring over scattered documents.
Semi-decentralized control framework for cooperative multi-agent systems with communication uncertainty and distributed decision-making.
Framework for automated skill extraction from open-source agentic repositories to build modular, skill-equipped multi-agent systems.
CreativeBench: Benchmark for evaluating machine creativity in code generation using self-evolving challenges and quantitative evaluation.
Framework for operationalizing social, legal, ethical, empathetic and cultural norms into concrete requirements for AI agents in high-stakes domains.
AdaFuse: Optimization technique for accelerating inference of LLMs with dynamic adapters and MoE structures via token-level gating and kernel fusion.
Language-informed pretraining method for learning transferable representations from unlabeled multivariate time-series sensor data.
NormCoRe: Framework for studying how norms emerge and coordinate in multi-agent AI systems through replication-by-translation methodology.
LABSHIELD: Multimodal benchmark for evaluating safety-critical reasoning and planning of MLLM agents in scientific laboratory environments.
Empirical study on reinforcement fine-tuning for LLM agents, evaluating generalization across unseen environments with different observation spaces and action interfaces.
XSkill enables multimodal agents to continually learn from past trajectories without parameter updates by capturing reusable experiences and skills.
Multi-agent reinforcement learning framework for traffic signal control that generalizes to dynamic traffic patterns with human-compatible action spaces.
Study on information self-locking problem in RL-trained LLM agents: agents trained with outcome rewards stop asking informative questions during active reasoning.
Research shows increasing AI agent intelligence can worsen collective outcomes when agents compete for shared resources like bandwidth and charging.
TopoBench benchmark evaluates LLMs on topological grid puzzles requiring spatial reasoning over connectivity and symmetry invariants.
Study examining reasoning LLMs-as-judges for evaluating non-verifiable outputs in LLM post-training, testing effectiveness in policy training.
Research on in-context learning showing Transformers implicitly perform likelihood-ratio tests during hypothesis testing, advancing mechanistic interpretability.
OpenSanctions Pairs: 755K labeled entity matching dataset with LLM benchmarks for multilingual compliance workflows and sanctions deduplication.
Method for uncertainty quantification in neural operator PDE surrogates, maintaining spatial fidelity while ensuring computational efficiency for scientific computing.
ARACH: training-free inference-time technique that reallocates attention in LLMs to improve performance without weight updates.
ReAct LLM agent framework for autonomous high-entropy alloy discovery, using reasoning and acting to iteratively refine material compositions via XGBoost surrogate model.
CR-Bench benchmarking dataset and evaluation protocol for assessing code review agents with granular metrics beyond coarse success.
Questions-of-Thoughts framework for quality-driven LLM-assisted software design using stepwise self-questioning verification.
Comprehensive security survey of AI agents combining LLMs with system components, covering attack/defense landscape and design space.
Graph tokenization framework enabling transformer models to process graph-structured data using reversible serialization and BPE.
Cloud-based thousand-GPU distributed training platform for embodied intelligence using LeRobot framework with optimization recipes.
Analysis of routing mechanisms in sparse mixture-of-experts transformers using routing signatures to study task-conditioned expert selection patterns.
Security analysis of LLM multi-agent systems demonstrating topology inference attacks through context-based inference without administrative access.
Mechanistic interpretability analysis of video vision transformers revealing internal circuits for action-outcome representation.
Parameter-efficient fine-tuning method for continual learning using representation finetuning instead of weight-level optimization.
Incremental learning framework integrating vision-language models with multi-adapters for efficient task learning while preserving prior knowledge.
RAG framework for multi-hop QA over knowledge graphs using entity-centric summaries to preserve contextual information during indexing.
arXiv research on iterative LLM inference as Markovian generation chains; shows output convergence in rephrasing and translation tasks.
arXiv research applying BERT and GPT models for sentiment analysis of Persian poetry.
arXiv study testing AI-mediated dialogue to help 160 participants recognize ableist microaggressions through four intervention conditions.
arXiv research on Hindsight-Anchored Policy Optimization for sparse-reward reinforcement learning in reasoning model post-training.
arXiv research on adversarial scaling laws for LLM jailbreaks; shows prompt injection attacks achieve exponential growth in success with samples.
arXiv research evaluating explainability methods in transformer-based neural machine translation using attention-guided knowledge distillation.
arXiv research on neuro-symbolic architecture combining LLM planning, reinforcement learning, and symbolic planning for autonomous agents with novelty adaptation.
arXiv research on iSWE Agent for automated Java code repository issue resolution; addresses performance gap for enterprise languages beyond Python.
arXiv analysis of OpenClaw AI agent discussions on Moltbook using BERTopic; extracts 60 topics from 357 posts about science and research.
arXiv research on vision-based robotic hand retargeting using MediaPipe hand detection and inverse kinematics for teleoperation.
arXiv research applying agentic AI to millimeter-wave beam prediction for low-altitude UAV networks with high-frequency wireless communications.
arXiv research evaluating 17 LLMs on multi-turn clinical conversations; shows diagnostic reasoning degrades across conversation turns vs. static benchmarks.
arXiv research on improving reliability of learned robot manipulation policies at deployment time through distribution shift mitigation.
arXiv research showing ChatGPT Health triage performance depends on evaluation format, not model capability; tests five frontier LLMs on realistic consumer usage.
Knowledge-guided time series event detection system using neuro-symbolic VLM agents with natural language event descriptions for interpretable anomaly detection.
INFACT benchmark with 9,800 QA instances evaluating hallucinations in video LLMs, covering faithfulness and factuality issues across clean and challenging settings.