When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels
Framework for validating comparative LLM safety scores without labeled benchmarks, enabling safety evaluation across diverse languages and domains.
Framework for validating comparative LLM safety scores without labeled benchmarks, enabling safety evaluation across diverse languages and domains.
Study showing that LLM fine-tuning with the same optimizer as pretraining reduces catastrophic forgetting while maintaining performance on new tasks.
Method for generating valid, challenging mathematical problems using LLM verifiers and self-play, enabling automated problem creation for LLM training.
Training-free bias mitigation method for GUI grounding in agents, improving performance on complex screen interaction tasks.
Knowledge distillation method for multi-modal models focusing on transferring modality-level relationships from teacher to student networks.
Formal game-theoretic framework for safety evaluation of AI deployment protocols through red-teaming exercises.
Framework for aligning LLM-based agents with human preferences through goal inference from multi-turn dialogue, addressing challenges in collaborative agent interactions.
Benchmark for evaluating differentially private text generation methods using LLMs, enabling secure sharing of sensitive datasets across institutions.
Research on how reinforcement learning enables LLMs to develop multi-step reasoning capabilities, using statistical physics framework to explain emergence of slow thinking.
Learning reasoning reward functions for LLMs via inverse reinforcement learning from expert demonstrations, avoiding manual reward specification.
CompassLLM multi-agent approach using LLMs for geo-spatial reasoning on popular path queries from trajectory data without requiring model retraining.
Memory-as-Action (MemAct) framework treating working memory management as learnable policy actions for long-horizon AI agent tasks to mitigate attention dilution.
Analysis showing self-consistency decoding technique yields diminishing returns on modern LLMs like Gemini 2.5, often degrading performance on already-solved problems.
SpatialBench benchmark for evaluating multimodal LLMs on spatial cognition tasks, capturing hierarchical structure of spatial abilities.
ProAgent framework for proactive LLM agents that continuously perceive and assist users in real-world settings using sensory contexts and contextual awareness.
CORE reinforcement learning framework providing fine-grained conceptual signals to teach LLMs genuine mathematical concept application rather than pattern reuse.
SANet framework using specialized AI agents for autonomous decision-making, dynamic adaptation, and cross-layer optimization in 6G networks.
Owen-Shapley policy optimization algorithm for RL-trained LLMs with token-level credit assignment, addressing reward sparsity in recommendation tasks.
E-mem multi-agent episodic context reconstruction framework preserving logical integrity and sequential dependencies for LLM agent memory systems.
BioAgent Bench evaluation suite with curated end-to-end bioinformatics tasks for measuring AI agent performance and robustness with automated assessment.
Process-verifiable thinking data synthesis framework enabling LLMs to perform long Chain-of-Thought reasoning for diverse time series tasks.
Safety harness using capability-safe Scala 3 language to restrict agent tool calls and prevent information leakage, unintended side effects, and prompt injection attacks.
PURE framework addressing preference-inconsistent explanations in LLM-based recommenders through preference-aware reasoning.
Context specification methodology for making AI evaluations operational and relevant to organizational deployment success and decision-making.
MineEvolve framework enabling embodied Minecraft agents to accumulate knowledge through interaction and self-evolution for long-horizon task completion.
Neuro-symbolic framework integrating LLMs into interactive theorem proving to automate proof script generation for systems software verification.
Metacognitive co-regulation framework for LLM design agents to overcome fixation bias and explore alternative solutions in engineering design tasks.
Claw-Eval benchmark suite with 300 human-verified tasks across 9 categories for trustworthy evaluation of autonomous LLM agents in real-world software environments.
Framework for routing NP-hard optimization problems to diverse solvers via polynomial-time reductions, integrated with agentic orchestration.
Autogenesis Protocol for self-evolving LLM-based agent systems, addressing lifecycle management, version tracking, and safe updates to enable modular agent composition.
Critique of current LLM evaluation frameworks identifying systematic failures (distributional, temporal, scope, process) inadequate for deployed agentic systems and RLHF reward hacking.
Deep learning framework for learning lifted action models from visual state sequences without action observation, applicable to AI planning in real-world domains.
FutureWorld live RL environment for training predictive agents on real-world event forecasting with actual outcome rewards.
WaferSAGE uses vision-language models with synthetic data generation for semiconductor wafer defect analysis.
Adaptive Entropy Modulation improves credit assignment in multi-turn LLM agent reinforcement learning tasks.
Position paper arguing agentic AI control layers should be Bayes-consistent for tool/expert selection under uncertainty.
LLM evolutionary search determines Zarankiewicz numbers in discrete mathematics using reinforced optimization.
GR-Ben benchmark evaluates process reward models across diverse reasoning tasks beyond mathematics.
Segment-Aligned Policy Optimization improves credit assignment in multi-modal reasoning RL for language models.
Zero-shot confidence estimation for small LLMs enabling cost-effective local-to-cloud routing without supervised training.
Systematic evaluation of prompting and execution methods for deterministic computation in large language models.
Circuit analysis of LLM agent memory systems revealing internal mechanisms of information management across sessions.
Executor-grounded reward training improves faithful reasoning in LLMs beyond final-answer correctness.
Experience-RAG Skill enables AI agents to orchestrate different retrieval strategies based on task context.
Four pre-trained language models for low-resource Angolan languages using transfer learning and synthetic data.
DeTrigger proposes gradient-centric method to mitigate backdoor attacks in federated learning systems.
CatNet algorithm controls false discovery rate in LSTM feature selection using SHAP importance and Gaussian mirrors with kernel-based independence measures.
LicenseGPT fine-tuned foundation model for dataset license compliance interpretation, addressing legal risks in commercial AI product development.
Evaluates LLM robustness on code understanding against semantics-preserving mutations, assessing reasoning quality beyond accuracy on programming tasks.
Amortized linear-time exact Shapley value computation for product-kernel methods, enabling scalable explainability in kernel-based models.