LLM Readiness Harness: Evaluation, Observability, and CI Gates for LLM/RAG Applications
Readiness harness for LLM/RAG applications combining evaluation, OpenTelemetry observability, and CI gates with Pareto frontier readiness scores.
Readiness harness for LLM/RAG applications combining evaluation, OpenTelemetry observability, and CI gates with Pareto frontier readiness scores.
Defend: Automated peer review rebuttal generation using LLMs with structured reasoning for targeted refutation and factual grounding.
Identity-grounded multi-agent LLM system with debate engine for resilient ethical tutoring addressing semantic drift and logical deterioration.
LLM agents replace random proposal generators in classical optimization algorithms; evaluated on discrete, mixed, and continuous optimization tasks.
AstraAI: CLI framework integrating LLMs with RAG and AST-based analysis for context-aware code generation in HPC scientific codebases.
Novelty Bottleneck model formalizing human-AI collaboration limits: irreducible serial component from tasks requiring human judgment follows Amdahl's Law.
PeopleSearchBench: Open-source benchmark evaluating AI-powered people search platforms across recruiting, sales prospecting, and expert discovery use cases.
Dual-stage LLM framework for scenario-centric semantic risk interpretation in autonomous driving assistance systems with reproducible auditing.
DSevolve: Industrial scheduling framework using LLM to evolve diverse heuristic portfolios for adaptive rule selection in dynamic manufacturing environments.
TianJi: Autonomous AI system using LLMs to discover physical causal mechanisms in atmospheric science beyond statistical weather forecasting.
SkyNet: Extends MuZero model-based RL to partially observable stochastic multi-player games with belief-aware planning under hidden state uncertainty.
Sortify: Agent-based closed-loop ranking optimization for recommendation systems treating ranking as influence allocation problem with online-offline metric calibration.
GAAMA: Graph-based long-term memory system for AI agents maintaining personalized, multi-session behavior through associative memory structure instead of flat retrieval.
GEAKG: Knowledge graphs framework for representing procedural algorithm knowledge as executable, learnable structures for domain-agnostic problem solving.
CARV: Benchmark dataset with 5,500 samples evaluating compositional analogical reasoning capabilities in multimodal LLMs, testing higher-order intelligence composition.
SARL: Label-free reinforcement learning method for improving reasoning in LLMs without requiring verifiable rewards, addressing open-ended domains with ambiguous correctness.
Framework for managing heterogeneous data in multi-embodied agent systems with diverse capabilities operating in dynamic environments.
AI agent uses architecture search across 3,106 experiments to test whether molecular sequences need different transformer designs than language models.
Study of bias in multimodal LLMs on scientific figure QA tasks where answer choices influence model predictions despite contradictory visual evidence.
Research on how LLMs perform scientific reasoning, examining internal heuristics and the effect of prompting on reasoning processes.
arXiv paper: Meta-Harness, automated optimization system searching over LLM harness code to improve system performance beyond model weights.
arXiv paper: SLOW, LLM-based tutoring system with dedicated reasoning workspace for cognitive diagnosis and pedagogical decision-making.
arXiv paper proving reward hacking is structural equilibrium in optimized AI agents under finite evaluation, regardless of alignment method.
arXiv paper: CoT2-Meta, training-free metacognitive reasoning framework combining chain-of-thought with meta-level control over reasoning trajectories.
arXiv paper: PReD, foundation multimodal model for electromagnetic domain covering perception, recognition, and decision-making.
arXiv paper: EpiPersona framework for pluralistic LLM alignment, explicitly coupling stable personas with episode-specific preference factors.
arXiv paper: Energy-Based Reasoning via Structured Latent Planning (EBRM), models reasoning as gradient-based optimization over multi-step latent trajectories.
arXiv study evaluating LLMs for answering student programming questions, addressing balance between pedagogical hints and complete solutions.
arXiv paper: Rhizomatic Research Agent (V3), multi-agent pipeline for non-linear literature analysis in social sciences using Deleuzian ontology.
arXiv paper proposing Collaborative Entropy (CoE), information-theoretic metric for semantic uncertainty quantification in multi-LLM agentic systems.
Survey tracing evolution of LLMs from transformers to agents, analyzing deep research as a prototype vertical application demonstrating progression from QA to agentic tool use.
COvolve co-evolves LLM-generated agent policies and environments as executable Python code using adversarial game dynamics to enable continual learning and improved generalization.
Study evaluates 12 open-weight vision-language models on clinical neuroimaging tasks, revealing that performance gains may reflect prompt artifacts rather than genuine multimodal evidence integration.
MiroEval benchmarks multimodal deep research agents by evaluating both research process and outcomes, addressing limitations of existing benchmarks with real-world query complexity and multimodal coverage.
Entropic Claim Resolution introduces a novel RAG inference method that uses uncertainty-driven evidence selection to handle conflicting information and query ambiguity beyond relevance-based retrieval.
Pilot study comparing t-norm operators in neuro-symbolic reasoning system for EU AI Act compliance classification of AI systems.
Medical AI Scientist: autonomous system generating medical hypotheses, conducting experiments, and drafting manuscripts grounded in clinical evidence.
MonitorBench: comprehensive open-source benchmark studying chain-of-thought monitorability and causal responsibility in LLM reasoning.
Perception-Reasoning Coevolution framework decoupling perception and reasoning optimization in multimodal LLM reasoning with verifiable rewards.
AIGENIE R package automating psychological scale development using LLM text generation with network psychometric methods.
ScanPaper benchmark evaluating MLLMs on scan-oriented academic paper reasoning beyond search-oriented retrieval tasks.
D2Skill: dynamic dual-granularity skill bank organizing reusable experience into task and step skills for reinforcement learning agents.
Study comparing whether smaller and larger LLMs capture culturally diverse moral values from World Values Survey and Pew Research data.
SimulCost: cost-aware benchmark for LLM agents in physics simulations accounting for tool-use costs like simulation time and experimental resources.
M-RAG: improved RAG system addressing retrieval noise and inefficiency through better chunking strategies for long-context LLMs.
Bridge-RAG: novel RAG framework using abstract bridge trees and cuckoo filters to improve retrieval accuracy and efficiency for LLM generation.
ReCQR: conversational query rewriting approach for multimodal image retrieval handling long text and unclear user expressions.
Comparative empirical study evaluating ChatGPT, Gemini, and DeepSeek as teaching agents across pedagogical strategies.
Framework proposing harmful capability uplift as core safety metric for measuring marginal increases in user ability to cause harm with frontier AI models.
Case study of context-aware AI system supporting calculus instruction by answering student questions on discussion forums.