Method for tracking internal states of LLMs across conversations using self-report-inspired techniques for safety, interpretability, and model welfare without white-box compression.
Manifesto proposing Agentic Business Process Management (APM) framework extending BPM to govern autonomous agents executing organizational processes with agent-oriented abstractions.
Large-scale empirical study analyzing 2,000+ publications on reinforcement learning environments, proposing a taxonomy of RL environment evolution and technological trends.
Examines ethical front-end design choices in conversational AI systems, focusing on user interaction and representation rather than backend algorithmic issues.
AIRA_2 addresses three bottlenecks in AI research agents: synchronous GPU execution, generalization gaps, and fixed LLM operator limitations through improved architectural design.
AutoMS is a multi-agent neuro-symbolic framework using LLMs as semantic navigators for evolutionary search in inverse microstructure design, addressing topology optimization challenges.
CoEvoSkills framework enables LLM agents to self-evolve structured multi-file skill artifacts through co-evolutionary verification without manual authoring.
Deep RL framework optimizes land-use allocation in Lake Malawi Basin to maximize ecosystem service value with ecological constraints.
Position paper on failure modes in agentic IR systems, analyzing error cascades in multi-step reason-act-observe workflows despite linguistic fluency.
Hierarchical multi-agent RL framework for reconfigurable intelligent surfaces removes need for channel state information estimation.
SCMAPR uses self-correcting multi-agent prompt refinement to improve text-to-video generation in complex scenarios with ambiguous prompts.
CuraLight combines LLM-centered control with debate-guided data curation for interpretable and generalizable traffic signal control systems.
EmoMAS uses Bayesian multi-agent framework with small language models for emotionally-aware negotiation in privacy-sensitive edge deployments.
EVGeoQA benchmark evaluates LLM reasoning on dynamic geo-spatial exploration with multi-objective planning and compound constraints.
Rhizome OS-1 is a semi-autonomous operating system deploying multi-modal AI agents as computational and medicinal chemists for drug discovery.
Pre-registered study demonstrating AI safety measures can cause harmful outputs in medical domains, with contextual prompting changing model behavior.
SEARL framework enables self-evolving agents through joint optimization of policy and tool graph memory, reducing reliance on large-scale LLMs.
HiL-Bench evaluates whether coding agents know when to request help with incomplete specifications, exposing judgment gaps in frontier models.
Empirical study of how LLM agents coordinate in multi-agent games, distinguishing baseline action similarity from strategic algorithmic monoculture.
AXIL derives exact instance attribution method for gradient boosting machines, expressing predictions as weighted sums of training targets.
Proposes contrastive learning method for dialogue sentence embeddings using token-level template annotations instead of utterance-level labels.
SciTune framework aligns LLMs with scientific domain knowledge through instruction fine-tuning on multimodal scientific publication data.
MM-LIMA demonstrates multimodal LLM fine-tuning achieves strong results with only 200 high-quality instruction examples, reducing data requirements.
Proposes CROP, a model-based offline reinforcement learning method addressing distribution shift through conservative reward estimation.
Framework using LLMs with philosophical relevance concepts to improve utility-based result ranking in retrieval-augmented generation systems.
MegaFake dataset of LLM-generated fake news for studying mechanisms of misinformation generation and detection methods.
Deep Optimizer States method enables scalable training of transformer models using interleaved offloading to overcome memory constraints.
PoTable framework improves table reasoning in LLMs using plan-then-execute reasoning stages for systematic thinking.
WebLLM inference engine enabling high-performance LLM execution directly in web browsers for on-device deployment without server GPUs.
HumanVBench benchmark for evaluating human-centric video understanding in multimodal large language models with 16 fine-grained tasks.
Three human studies examining whether humans can be influenced to conform to preference models used in RLHF algorithms for LLMs.
Novel curriculum learning approach for sample-efficient reinforcement learning applied to quadrotor stabilization control.
Combines semi-supervised and active learning for semantic segmentation to reduce manual annotation costs and improve model performance.
Proposes using LLMs to help mitigate barren plateaus in quantum neural network training through adaptive parameter initialization.
Study evaluating emergent lifelong learning behaviors in LLMs during multi-turn interactions, proposing new evaluation benchmarks for character-like consistency.
TARAC method addresses hallucinations in vision-language models by improving temporal attention mechanisms during generation without extensive retraining.
Research on energy-efficient optimization techniques for LLM deployment, including quantization and local inference strategies to reduce carbon emissions.
PODS decouples rollout generation from policy updates in LLM RL, addressing compute asymmetry through down-sampling.
LOOPE method learns optimal patch ordering in Vision Transformer positional embeddings for improved spatial information encoding.
RL^V framework unifies LLM reasoners with verifiers using value functions for improved test-time compute scaling during reasoning.
Bayesian approach for Vision Language Models to reduce hallucinations and overconfidence in VQA through selective prediction.
TokUR enables LLMs to self-assess uncertainty at token-level for improved reasoning and response reliability in multi-step tasks.
SpatialScore: comprehensive benchmark for evaluating spatial intelligence of multimodal LLMs with data-driven and agent-based assessment approaches.
GoT-R1: reinforcement learning framework enhancing multimodal LLM reasoning for complex visual generation with precise spatial relationships and attributes.
Fine-tuning approach for LLMs to predict diverse user behaviors, addressing overfitting to frequent behaviors while capturing long-tailed behavior distribution.
World models for interactive video generation with action conditioning and autoregressive decoding to support planning and future prediction.
Framework using LLMs for few-shot code generation to create safety-critical driving scenarios in CARLA simulator for autonomous driving evaluation.
LLM-based autonomous agent for power system voltage control, using experience-driven learning to generate dispatch strategies in distribution networks.
Data Mixing Agent: LLM-based method to automatically re-weight training data domains during continual pre-training, preventing catastrophic forgetting.
PRIX: efficient end-to-end autonomous driving model planning from raw camera pixels without LiDAR, reducing model size and computational requirements.