Look Before You Leap: Autonomous Exploration for LLM Agents
Exploration Checkpoint Coverage metric for measuring autonomous exploration in LLM agents in unfamiliar environments.
Exploration Checkpoint Coverage metric for measuring autonomous exploration in LLM agents in unfamiliar environments.
Formal methods combined with ML for auditing and monitoring AI systems throughout development lifecycle.
Cost-performance study of compound LLM agent design dimensions (context, reasoning, hierarchy) in adversarial environments.
Benchmark evaluating 7 LLM tutoring agents on diagnostic precision for student solutions. Shows agents struggle with suboptimal feedback.
Fully Open Meditron: Clinical LLM with complete transparent training pipeline including data provenance and curation procedures.
FORGE: Population-based protocol for evolving natural-language memory in ReAct LLM agents through prompt injection and reflexion loops.
Autonomous LLM-guided tree search system for iterative disease forecasting model generation without manual expert curation.
Sound POMDP synthesis with LTL objectives for autonomous agents navigating uncertain environments with temporal constraints.
Agent4POI: Agentic framework generating context-conditioned multimodal representations for POI recommendation at inference time.
AgentStop: Energy-efficient early termination strategy for local LLM-based agents on consumer devices.
Empirical study showing quantization induces bias emergence in LLMs across model families and precision levels, harming alignment.
RAG-based food recommendation system using Healthy Eating Index and LLMs for personalized diet suggestions.
Universal data mixing method for language model training across pretraining, continual learning, and adaptation phases.
Studies effective harness design for LLM-based algorithm discovery, analyzing trade-offs in token budget allocation for evolutionary search.
LLM and Vision-Language Model workflow for RISC-V supply chain analysis integrating multimodal data extraction and model-driven engineering.
Empirical evaluation of biologically-inspired AI agent frameworks against simpler baselines across three deep benchmarks.
Phoenix-bench: Benchmark of 511 hardware engineering tasks testing whether software-focused agentic AI transfers to realistic EDA workflows.
PBT-Bench: Benchmark of 100 property-based testing tasks evaluating AI agents on deriving semantic invariants from documentation.
A3D: Agentic AI system for autonomous hardware accelerator design using LLMs and tool integration to reduce design labor.
Hydra: LLM code generation system using checkpoint-and-rollback for efficient error detection and repair during decoding.
Group-Query Latent Attention mechanism for efficient LLM decoding across diverse hardware, improving upon DeepSeek-V2/V3 architecture.
AI-driven autonomous testing framework using LLMs for web test automation with integrated security, addressing test suite maintenance failures.
Maximal parameterization update rule for grouped query attention enabling hyperparameter transfer across LLM architectures.
Framework for stabilizing recommendation system predictions through temporal data augmentation and feature pruning.
Discovery agent system for automatic program synthesis from input-output behavior using LLMs.
Security analysis of sleeper memory poisoning attacks on LLM agents with persistent external memory.
Trajectory-level evaluation framework for LLMs in iterative scientific design and autonomous laboratories.
Causal discovery method for DAGs from large-scale interventional data with improved scalability.
Diagnostic evaluation framework for LLM sequential memory, capturing forgetting and transfer patterns.
Study of JEPA principles applied to autoregressive LLM fine-tuning and latent representation learning.
Reinforcement fine-tuning approach for LLM-based discovery of alpha factors in quantitative trading.
Framework for improving LLM confidence estimation and reliability in judgement tasks through margin-adaptive ranking.
Loss function family for training GFlowNets, generative models, and LLMs using off- and on-policy data.
Runtime-structured task decomposition for agentic LLM coding systems improving debuggability and reducing retry costs in software engineering.
DrugSAGE: Self-evolving agent that accumulates and reuses experience for efficient AutoML-style drug discovery model building.
GRLO framework combining RLHF and RLVR paradigms for generalizable LLM post-training across open-ended environments.
Retrieval-augmented LLMs extract structured clinical information from nurse-patient transcripts for medical documentation automation.
LLM-based framework for robot task scheduling using natural language processing to optimize time efficiency and resource allocation.
Ghosted Layers recovers pruned LLM performance by solving boundary activation alignment without retraining pruned decoder blocks.
Theoretical analysis proving RoPE positional embeddings lose locality bias and position discrimination in long-context Transformers.
Demonstrates data attribution values can be manipulated in distributed training while preserving model utility, affecting ML governance.
BetaPRM: Process Reward Model predicting step-level success probability and reliability scores for LLM reasoning verification.
DeltaPrompts improves multimodal distillation by identifying and replacing ineffective prompts that don't change student model outputs.
AstraFlow: dataflow-oriented reinforcement learning system for scaling agentic LLMs with multi-policy training on heterogeneous, elastic compute resources.
Agentic program analysis system for detecting privilege escalation vulnerabilities in polyglot microservices through cross-service interaction modeling.
Offline reinforcement learning method using universal horizon models to reduce compounding errors from repeated self-generated state inference.
PrismLLM: tool for emulating large-scale LLM training on small GPU clusters to enable debugging and optimization without accessing production infrastructure.
Systematic evaluation of latent video prediction models as world models, analyzing robustness of V-JEPA, VideoPrism, and VideoMAEv2 across five axes.
Study of line-tracing ability in vision-language models, diagnosing failures in basic visual path-following operations across controlled tasks.
Interaction-aware influence functions measuring how groups of training examples jointly affect model predictions, capturing redundancy and complementarity.