Adaptive Parallel Monte Carlo Tree Search for Efficient Test-time Compute Scaling
Adaptive Parallel MCTS for test-time compute scaling in LLMs, introducing negative early exit to reduce latency during reasoning.
Adaptive Parallel MCTS for test-time compute scaling in LLMs, introducing negative early exit to reduce latency during reasoning.
Uni-SafeBench: Safety benchmark for unified multimodal large models integrating understanding and generation in single architecture.
BloClaw provides omniscient multimodal workspace for AI scientists, addressing infrastructure limitations in JSON-based tool-calling and fragile execution sandboxes.
Neurosymbolic architecture using ontology-constrained reasoning in Foundation AgenticOS to address LLM hallucination, domain drift, and regulatory compliance in enterprise agents.
Framework for task-level performance prediction in agentic coding benchmarks, moving beyond aggregate pass rates to understand agent challenges across diverse tasks.
CircuitProbe predicts reasoning circuit locations in transformers via activation statistics in CPU milliseconds, achieving 3-4 order speedup over brute-force methods.
UK AISI evaluates whether frontier LLMs sabotage safety research when used as coding assistants, testing alignment and intended behavior in real-world deployment scenarios.
RefineRL enables self-refinement through reinforcement learning for LLM competitive programming, introducing skeptical-agent mechanisms for iterative problem solving.
PARE framework simulates active users to evaluate proactive AI agents, modeling stateful sequential interactions in digital environments for realistic agent development.
Uses multi chain-of-thought voting approach for geometric problem solving in LLMs, combining diagrammatic understanding with symbolic manipulation and logical inference.
Proposes multi-agent RAG system with evolving orchestration and adaptive agent prompts to handle complex multi-hop queries requiring continuous behavioral refinement.
PsychAgent implements experience-driven lifelong learning for psychological counseling, using memory-augmented planning to continuously refine agent behavior through practice.
OmniMem uses autoresearch to design lifelong multimodal memory systems for extended-horizon AI agents, automating exploration of architecture and retrieval strategies.
Evaluates ethical robustness of LLMs under sustained adversarial multi-turn interactions using moral stress testing beyond single-round safety benchmarks.
Introduces NARCBench to detect collusion between LLM agents using multi-agent interpretability and activation analysis, addressing risks in multi-agent deployments.
Analyzes whether LLM reasoning models make decisions before or after chain-of-thought reasoning using linear probes on activation patterns for tool-calling decisions.
HippoCamp benchmark evaluates multimodal AI agents on personal computer file management tasks requiring context-aware reasoning in user-centric environments.
AI agents conduct physics analysis using archived particle collision data, with OpenAI Codex and Claude performing data analysis and report writing autonomously under physicist direction.
Proposes optimizer-aware framework for gradient-based online data selection in LLM fine-tuning, addressing step-dependent utility in sequential data arrival scenarios.
Introduces Olfactory Perception benchmark with 1,010 questions to assess LLM reasoning capabilities about smell across odor classification and related tasks.
Evaluates reliability of LLM and hybrid deterministic-LLM approaches for extracting information from academic course registration documents, comparing three strategies on 140-860 documents.
LinearARD: Self-distillation method restoring RoPE-scaled LLM capabilities on standard benchmarks while maintaining extended context windows.
Dynin-Omni: First masked-diffusion omnimodal foundation model unifying text, image, speech, and video in single architecture.
Study evaluating trustworthiness of LLM-as-Judge for qualitative research workflows and interpretive response assessment.
Eyla: Proposed identity-anchored LLM architecture integrating state-space models and biological priors with implementation failure analysis.
Empirical investigation across 68 tasks and 4 model families showing LLMs fail to estimate task duration, overshooting by 4-7x.
Study quantifying gender bias in LLMs when making hiring decisions and evaluating prompt engineering as mitigation technique.
Research revealing hidden safety mechanisms in post-trained LLMs that degrade during fine-tuning and methods to reactivate them.
MSA-Thinker: Multimodal sentiment analysis framework using hint-guided reinforcement learning for improved interpretability.
Method detecting LLMs posing as humans in online behavioral research by probing memory constraint capabilities.
Entropy-guided decoding strategy improving LLM reasoning by reducing error propagation while maintaining computational efficiency.
RiDiC pipeline generating multilingual entity datasets with controlled popularity for evaluating LLM long-form factuality generation.
Multi-agent simulations across 4 models examining how LLMs internally process ethical instructions in different formats and languages.
Two-phase study validating LLM-as-Judge rubrics against real business conversion outcomes in conversational commerce on major platform.
WHBench: Evaluation suite of 47 expert-validated scenarios for assessing frontier LLMs on women's health topics with 22 models tested.
Study showing larger LLMs underperform smaller ones on 7.7% of problems due to scale-dependent verbosity causing overelaboration errors.
Experimental platform studying behavioral differentiation in multi-agent LLM conversations across 7 models and 208 experimental runs.
DriftScript: Lisp-like DSL that compiles to Narsese for programming non-axiomatic reasoning agents with improved usability.
Personalized federated learning approach for fine-tuning language models on heterogeneous distributed tasks while maintaining individual client performance.
arXiv research comparing energy consumption of retrieval-augmented generation systems versus direct LLM usage for environmental domain applications.
Temporal memory framework for resource-constrained agents using stochastic Bridge Diffusion for continual learning with fixed memory budget.
Perspective on sustainability challenges in AI-driven molecular discovery pipelines including computation and data generation costs.
Empirical validation showing classifier-based safety gates fail to maintain oversight as AI systems self-improve over iterations.
Analysis showing terminal agents suffice for enterprise automation tasks compared to complex tool-augmented or web-based agentic systems.
Curriculum learning framework using LLMs to dynamically generate action sequences for RL agents, tested on Blackjack.
HIVE: Framework for hierarchical pre-training of vision encoders with LLMs to improve vision-language alignment in multimodal models.
Experience report on using customized ChatGPT tutor for teaching domain understanding and design methods in software engineering courses.
Oblivion: Memory control framework for LLM agents using decay-driven forgetting to reduce interference and latency in long-context scenarios.
Study on fault localization granularity impact for repository-scale code repair tasks using machine learning techniques.
Unified architectural metamodel for information systems developed using generative AI and LLMs to standardize code and documentation generation.