ABSTRAL framework for automatic multi-agent system design via iterative refinement and topology optimization with measurable coordination metrics.
Neuro-symbolic multimodal reasoning approach for reliable classroom AI addressing privacy and pedagogical challenges.
Dynamic preference inference framework for sequential decision-making when reward preferences shift contextually over time.
Empirical comparison of tool integration vs inter-agent delegation protocols for multi-agent task orchestration systems.
Balanced Direct Preference Optimization method improving safety alignment in LLMs while mitigating overfitting issues.
PhySe-RPO diffusion-based framework for surgical smoke removal optimized via physics and semantic constraints.
CoMaTrack multi-agent reinforcement learning approach for embodied visual tracking using competition-driven game theory.
Chain-of-Authorization framework enabling LLMs to internalize access control via reasoning trajectories for secure data handling.
Analysis of hierarchical reasoning model training dynamics using dynamical systems theory to improve complex reasoning capabilities.
Three-layer framework separating diagnosis from control in agent-based simulations using LLM diagnostics for policy adaptation.
ProGRank defense mechanism against corpus poisoning attacks in RAG systems using probe-gradient reranking.
Ran Score metric for evaluating chest X-ray report generation using LLM-based extraction of clinical findings.
Fine-tuning small language models for NL2SQL tasks, demonstrating cost-effective alternatives to large models for enterprise data access.
PersonalQ framework for efficient serving of personalized diffusion models via intelligent routing and quantization techniques.
Benchmark dataset and computational methods for detecting implicit legal citations in French court decisions using ML.
Benchmark for evaluating LLMs on fault tree analysis and malfunction diagnosis in complex systems using multi-turn dialogue.
Investigates whether LLMs can solve constrained optimization problems, specifically testing reasoning abilities on Optimal Power Flow problems.
Game-playing AI agent that balances performance against human opponents without dominating, improving human-AI interaction and educational value.
MedCausalX framework adding explicit causal reasoning and self-reflection to medical vision-language models for improved clinical reliability.
Dataset and evaluation of 22 LLMs on context-sensitive moral judgment across consequentialist, emotional, and relational dimensions.
Describe-Then-Act framework using distilled language-action world models for proactive agent safety without visual simulation overhead.
PERMA benchmark for evaluating LLM agents with long-term memory using event-driven preferences and realistic task environments.
MemCollab enables memory sharing across heterogeneous LLM agents through contrastive trajectory distillation for collaborative problem-solving.
Analysis of LLM benchmark vulnerabilities and proposal for Olympiad-style sealed exams to prevent benchmark-chasing and improve evaluation integrity.
Framework for studying how LLM-based agents form stable stances and identities in multiagent communities using virtual ethnography methods.
Bilevel autoresearch applying automated research loops to optimize autoresearch systems themselves, improving bottlenecks iteratively.
Using LLMs to detect microservice architecture patterns from Infrastructure-as-Code artifacts for documentation purposes.
Analysis of 1.8M Hugging Face models tracking how multimodal capabilities emerge and propagate across open LLM families.
Systematic evaluation of four prompting strategies across GPT models on chart-based question answering, isolating prompt structure effects.
MERIT system combining memory and retrieval mechanisms with LLMs for interpretable knowledge tracing in educational settings.
TIPS framework for training search-augmented LLMs with reinforcement learning, improving credit assignment and reward shaping for question answering tasks.
Methods for generating high-quality synthetic training data using LLMs to fine-tune smaller models, analyzing diversity and distribution in embedding space.
Mechanistic interpretability study investigating whether LLMs develop genuine emotional representations or merely detect emotion keywords through circuit analysis.
Compact uncertainty estimation method for LLMs scoring cross-layer agreement patterns in internal representations via single forward pass.
Sparse Feature Attention method reducing transformer self-attention cost via k-sparse feature representations instead of sequence-level sparsity.
Mathematical framework interpreting LLM hidden states as points on latent semantic manifolds with Riemannian geometry and Voronoi partitions.
Training-free hallucination detector for LLMs using sample transform cost to measure output distribution complexity without fine-tuning.
Chinese financial news dataset and benchmark for evaluating LLM-based agents in macro and sector asset allocation decision-making.
Decision Transformer approach for optimizing emergency vehicle signal preemption using offline, return-conditioned sequence modeling.
Geometric Mixture-of-Experts framework for graph representation learning using curvature-guided routing on heterogeneous topologies.
Dataset aligning instruction manuals with assembly videos for evaluating multimodal LLMs on real-world technical tasks.
AEGIS infrastructure for governance of adaptive medical AI systems under FDA and EU regulations with continuous improvement.
Multi-task deep learning framework for predicting lithium-ion battery state-of-health and remaining useful life.
Delta-Aware Quantization framework for post-training LLM compression that preserves knowledge from alignment fine-tuning.
Classification approach for wind power ramp event forecasting under severe class imbalance for grid stability.
AgentSLR uses agentic AI to automate systematic literature reviews in epidemiology from retrieval through synthesis.
Method for adding trained persistent memory to frozen decoder-only LLMs without cross-attention mechanisms.
Applies conformal prediction for formal safety guarantees in wildfire spread prediction using tabular, spatial, and graph models.
Comprehensive study of LLM-based data imputation across multiple models and datasets, analyzing hallucination effects and control mechanisms.
Combines graph signal processing with Mamba2 state-space models to create adaptive filter banks for language modeling.