AutoTool: Automatic Scaling of Tool-Use Capabilities in RL via Decoupled Entropy Constraints
RL method for automatically scaling tool-use capabilities in agents through decoupled entropy constraints addressing overthinking.
RL method for automatically scaling tool-use capabilities in agents through decoupled entropy constraints addressing overthinking.
Study examining whether STAR reasoning prompt technique maintains effectiveness when embedded in complex production system prompts.
Multi-agent LLM orchestration for annotating educational classroom discourse with improved consistency over single-pass annotation.
Analysis of contextual bias in reinforcement learning feedback sources; proposes methods addressing selective truthfulness in evaluators.
Technique to discover and steer category-specific refusal directions in language models for fine-grained safety control at inference time.
Survey analyzing 82 approaches across ARC-AGI benchmark versions; documents consistent 2-3x performance degradation across paradigms.
Examines whether RLHF-trained LLMs face contradictory training objectives analogous to Hofstadter-Mobius loops.
Metric for detecting procedural bias in machine learning models across intersectional demographic groups.
Benchmark environment for evaluating AI agents in enterprise workflows with long-horizon planning, persistent state management, and access controls.
Open source library for constructing and serving LLM-based multi-agent systems with workflow orchestration and tool execution abstraction.
Constraint-based approach formulating LLM routing as MaxSAT problem to match queries with appropriate models based on natural language preferences.
Framework addressing context window limitations in LLMs through dynamic cognitive state management for long-horizon agent tasks.
LLM-based system for extracting Alzheimer's disease phenotypes from clinical notes in electronic health records.
Multi-agent framework using LLMs with self-evolving memory for predicting treatment outcomes in PET theranostics for prostate cancer.
Multi-agent system using LLMs to automatically detect autism intervention strategies in parent-child reading interactions from video.
Transformer-based tokenization method for precipitation nowcasting emphasizing synergistic interactions between meteorological elements.
Benchmarking LLMs against traditional regression for predicting polysulfone membrane mechanical properties from limited experimental data.
GroupGuard framework modeling and defending against collusive attacks where multiple LLM agents coordinate to mislead systems.
Evidence-driven agent framework for radiology report generation using multimodal LLMs with explicit visual evidence tracing and clinical reasoning.
Open source evaluation harness for vision-language-action models decoupling inference from benchmarks via WebSocket protocol and Docker isolation.
Comparative study of supervised fine-tuning versus reinforcement learning as post-training methods for LLMs, examining theoretical connections.
Black-box evaluation of faithfulness in medical reasoning explanations across closed-source LLMs like ChatGPT and Gemini.
Systematic evaluation protocol assessing reliability and robustness of graph-derived signals in tabular machine learning.
GRPO and reflection reward methods for improving mathematical reasoning in LLMs through proactive reflection encouragement during training.
Demand-driven methodology for building enterprise knowledge bases by analyzing LLM agent failures to capture domain-specific tribal knowledge.
Framework for relationship-aware safety unlearning in multimodal LLMs, addressing relational safety failures without collateral damage.
Memory-as-Asset paradigm for human-centric personal memory management as prerequisite for extending LLM knowledge boundaries through self-evolution.
DAG-orchestrated agentic planner for multi-hop question answering over hybrid data lakes combining structured and unstructured sources.
Hierarchical planning framework for analyzing LLM web agent failures across planning, execution, and replanning layers for process-based evaluation.
ScienceClaw + Infinite framework enabling autonomous agents to conduct scientific research through emergent artifact exchange and 300+ interoperable skills.
Economics paper analyzing spillovers and incentive structures in content creation enabled by GenAI reuse capabilities.
Framework for autonomous evolution of data curation strategies across heterogeneous pretraining categories, extending Data Darwinism hierarchy to multi-domain scenarios.
Benchmark for evaluating step-level process quality in tool-using LLM agents, addressing brittleness in long-horizon interactions with irreversible side effects.
Method for resolving proactive interference in LLMs using sleep-inspired memory consolidation to improve context window retrieval accuracy.
System using RAG and LLMs to preserve and query tacit expert knowledge in industrial organizations through structured multimodal capture.
Production-ready job matching platform using transformer embeddings, skill knowledge graphs, and interpretable reranking for candidate matching.
Information-geometric foundations for persistent memory in AI agents, covering retrieval metrics, lifecycle management, and consistency detection.
Algorithm for compiling Bayesian network classifiers into logical formulas for explainability, extending prior work beyond binary classification.
Research on augmenting LLMs with computational argumentation to provide explainable decisions and enable contestation of outputs in high-stakes domains.
Investigation of dynamic Theory of Mind in LLMs as temporal memory problem, measuring ability to represent and update beliefs over time.
Theoretical framework applying punctuated equilibrium to AI development, challenging continuous progress assumptions and predicting capability phase transitions.
Gradient Atoms method for unsupervised discovery of model behaviors via sparse decomposition of training gradients without requiring query behaviors.
RenderMem spatial memory framework for embodied agents treating rendering as interface between memory retrieval and geometric grounding.
GameUIAgent: LLM-powered agentic framework translating natural language descriptions into game UI designs via structured JSON and VLM reflection.
BrainBench: benchmark of 100 brainteaser questions exposing commonsense reasoning failures in LLMs across 20 targeted categories.
OpenHospital arena for evolving and benchmarking LLM-based collective intelligence through physician-patient agent interactions.
Framework for encoding institutional knowledge as AI skills/primitives to enable autonomous agents in enterprise software development.
Planning algorithm using goal recognition-derived heuristics to improve classical planning search efficiency and trajectory prioritization.
Clinical decision support system integrating AI predictive modeling with medical knowledge bases for disease diagnosis using lab results.
Unified world model for remote sensing combining spatiotemporal change understanding with text-guided future scene forecasting.