A²utoLPBench is an auto-generated benchmark for testing LLM-driven agents on linear programming problems using inverse-KKT construction to avoid data leakage.
Rubric-based evaluation of frontier LLMs on clinician-authored clinical reasoning tasks, showing performance remains low on difficult open-ended scenarios.
UA-ChatDev adds uncertainty awareness to multi-agent software development frameworks to mitigate hallucination propagation across role-based agent collaboration.
Guard Rail Validation framework intercepts AI agent inference outputs in autonomous telecom networks to validate decisions before triggering live state changes.
Purified OPSD improves on-policy self-distillation for long chain-of-thought reasoning by preserving reflective capabilities while providing token-level supervision.
Multi-agent swarm system for mental health support in low-resource settings using dynamic emotional state calibration and multimodal interfaces.
AgenticSTS is a testbed for evaluating long-horizon LLM agents with bounded memory contracts, isolating effects of memory components on agent decision-making.
HOLA adds hippocampal-inspired exact memory to linear attention models to recover facts that get overwritten in compressed recurrent states, improving needle-in-haystack recall.
Autonomous research agent pipeline for computational physics that handles underdocumented toolchains and validates against external literature to avoid hallucinations.
DRIFTLENS measures how personalization in LLMs changes reasoning trajectories, not just outputs, when user memory is injected into prompts for open-ended questions.
Hardware-enforced semantic coordination framework for safety-critical autonomous systems integrating LLMs, world models, and specialized architectures.
Access control and network policies from human engineering teams transfer to coding agents, providing cheaper oversight than agentic scaffolding.
RFM-AGOP efficiently extracts multi-dimensional refusal subspaces in LLMs for safety and interpretability with reduced computational cost.
SPG-Layout uses LLMs for text-driven 3D indoor scene synthesis in non-Manhattan environments by modeling non-orthogonal spatial relationships.
GPT, Claude, Gemini and GLM are evaluated on grading Linux/bash exam responses, handling partial credit and solution equivalence better than rule-based systems.
EvoPolicyGym evaluates autonomous agents' ability to improve policies through feedback in interactive environments with controlled evaluation methodology.
Dual-channel debate framework studying how social structure affects LLM agent communication with public and off-the-record channels.
ReContext method for long-context reasoning using recursive evidence replay to improve LLM utilization of relevant information in extended contexts.
Real-time safety monitoring method for LLMs using verifier signals with risk-calibrated thresholding to detect unsafe outputs at deployment.
Studies attack surface of persistent-state AI coding agents shipping code iteratively across sessions; introduces Iterative VibeCoding for AI control.
TokenScope interactive tool for token-level explainability in LLM code generation, providing decoding-time signals and uncertainty measures.
Framework for safeguarding LLM agents from misalignment through provenance analysis of tool invocations against user intent.
Kara system for efficient reasoning LLM serving via sliding-window KV cache compression to reduce decoding latency and memory overhead.
SPARCLE method for speaker-aware representations in speech synthesis using contrastive learning with grapheme-based models.
Identifies BPE tokenization fragmentation of safety-critical words as exploitable gap in LLM alignment, tested on five model families.
ErrorBench stress-test protocol showing prompt framing distorts count-based F1 evaluation of LLM error detection without improving span localization.
RAGP method for prompt compression using graph pruning guided by Lévy walks, capturing distributed information across syntactic and semantic relations.
ExPerT framework for personalizing LLM responses using query-wise semantic and keystroke behavioral cues to adapt to user domain expertise.
Introduces Office Comprehension Benchmark for evaluating LLMs on Word, Excel, PowerPoint comprehension across native file formats and variants.
Evaluates six LLM configurations for grading open-ended mathematics exams, assessing reliability and practical usability with partial-credit rubrics.
Practice auditing framework for LLM use in knowledge acquisition, code generation, and automation; introduces collective empiricism concept.
Framework for specifying sociotechnical alignment of AI systems, addressing gap between technical and normative aspects of socially desirable behavior.
Benchmark evaluating federated learning and knowledge distillation for 3D point cloud classification across 504 runs.
Survey of generative AI and federated learning applications for intrusion detection in cyber-physical and IoT environments.
Methods for inferring LLM architectural properties (hidden dimensions, layer counts) through restricted API access with minimal logits exposure.
Mechanistic interpretability analysis of Neural Quantum States using sparse autoencoders for feature extraction and causal steering.
Likelihood-based framework for automatic evaluation of turn-taking naturalness in full-duplex spoken dialogue systems.
Zero-instrumentation monitoring tool for diagnosing GPU training job failures without modifying code or infrastructure.
Empirical study of adoption, retention, and ROI for command-line AI coding agents (Claude Code, GitHub Copilot CLI) at organizational scale.
Training-free method for multimodal attribution in long document QA systems using attention analysis for grounded answers.
Organizational framework for governing agentic AI systems, addressing probabilistic behavior, autonomous actions, and clear accountability.
Benchmark separating reasoning from knowledge retrieval in LLM evaluation using isomorphic cross-domain problem pairs.
Study of pruned Mixture-of-Experts models in biomedical domain, evaluating utility and factual reliability under resource constraints.
Research on gradient geometry of embedding tables in language models; introduces Ember optimizer for efficient finetuning and pretraining with minimal memory.
Framework for reducing LLM hallucinations in resume optimization through temporal validation, contamination detection, and structural checks.
Framework decomposing advantage function variants in RL-based LLM post-training to address training instability and diversity collapse in policy gradient methods.
Multi-Head Recurrent Memory Agents: Addresses reliability degradation in LLM agents with long contexts by decomposing memory capture and retention.
IntentTune: System for resolving ambiguous e-commerce queries by inferring latent user intent attributes using demand and personalization signals.
EFE: Framework using LLM-based evolutionary search to discover preprocessing transformations for structured data as composable Python programs.
X-LogSMask: Explainable multi-scale transformer variant for sparse, structured graph data with improved interpretability over standard transformers.