CADDesigner: Conceptual CAD Model Generation with a General-Purpose Agent
CADDesigner: LLM-powered agent for conceptual CAD modeling accepting text and sketches with interactive refinement dialogue.
CADDesigner: LLM-powered agent for conceptual CAD modeling accepting text and sketches with interactive refinement dialogue.
SkillNav: Modular framework for vision-language navigation using skill-based decomposition to improve generalization on complex spatial-temporal reasoning.
Critical analysis of temporal signals in benchmark contamination detection, showing sensitivity to question construction independent of data memorization.
QuickLAP: Bayesian framework for robot learning that fuses physical corrections and language feedback to infer reward functions.
Prismatic world model for robotic planning that learns compositional dynamics separately for distinct physical modes like contact/impact events.
MobiBench: Multimodal benchmark for mobile GUI agents addressing limitations of existing offline/online benchmarks with multiple valid action paths.
Differentiable Evolutionary RL optimizes reward functions using gradient information to improve policy performance on complex reasoning tasks.
PersonalAlign: GUI agent framework that aligns with implicit user intents by leveraging long-term user records as persistent context for personalized task completion.
PolySHAP improves KernelSHAP approximation of Shapley values for explainable AI by using polynomial regression instead of linear approximation to reduce computational cost.
Study examining how elicitation protocol design affects stated-revealed preference gaps across 24 language models.
THINKSAFE: Safety alignment approach for reasoning models that self-generates safety constraints without external teacher distillation.
M2CL: Multi-LLM context learning method for multi-agent discussion systems addressing discussion inconsistency and context misalignment.
DiscoverLLM: Method enabling LLMs to help users discover intents through interactive exploration rather than just executing stated requests.
SupChain-Bench: Benchmark for evaluating LLMs on real-world supply chain management with multi-step domain-specific orchestration.
AMOR: Hybrid architecture combining recurrence and attention, selectively invoking attention based on predictive uncertainty.
Proxy State-Based Evaluation: LLM-driven benchmark method for multi-turn tool-calling agents that scales beyond deterministic backends.
LLM-based approach for time series question answering using pattern alignment and balanced reasoning across task complexity.
Interactive Benchmarks: Evaluation paradigm assessing model reasoning by testing ability to decide what information to acquire and use.
GAAMA: Graph-augmented memory architecture for AI agents to maintain coherent long-term personalized behavior across sessions.
Trajectory-level safety benchmark (ATBench) for evaluating LLM-based agents across realistic multi-step interactions with diverse failure modes.
Benchmark evaluating whether LLM agents can autonomously design, implement, and execute RL post-training pipelines for model improvement.
Framework for training open-weight language models to simulate student coding behavior for educational tutoring system evaluation.
Industrial evaluation system for LLM-generated meeting summaries with structured ground-truth construction and privacy-bounded monitoring.
Evaluates three LLM agent interaction paradigms (structured tools, computer-use, coding agents) on scientific visualization tasks across 15 benchmarks.
Study on coordinated flow models for offline multi-agent reinforcement learning that balances efficiency and coordination.
Paper arguing automated alignment via research agents risks producing misleading safety assessments without deliberate sabotage.
AI co-mathematician workbench enabling mathematicians to collaboratively leverage AI agents for research including ideation, computation, and theorem proving.
Method to extract and analyze search trees from LLM reasoning traces to quantify planning capabilities and identify myopic decision-making.
Agentick: Unified benchmark enabling fair comparison of RL, LLM, VLM, and hybrid agents on sequential decision-making tasks.
SREGym: High-fidelity benchmark for AI SRE agents with live cloud-native systems and realistic failure scenarios for diagnosis and mitigation.
BoostAPR: Three-stage RL framework for automated program repair using execution feedback and dual reward models for credit assignment.
FORTIS benchmark evaluating privilege escalation vulnerabilities in LLM agent skill layers and access control policies.
SimWorld Studio: System using evolving coding agents to automatically generate diverse 3D environments for embodied agent training.
EpiGraph knowledge graph and benchmark for evaluating knowledge-augmented clinical reasoning in epilepsy diagnosis from heterogeneous evidence.
Expo: Improved RL policy optimization for LLM reasoning with adaptive KL regulation and curriculum sampling over GRPO baseline.
IndustryBench: 2,049-item benchmark testing LLM knowledge on industrial procurement with safety-critical constraints and standards compliance.
Genetic programming approach to symbolic regression with gene editing for discovering mathematical formulas from scientific data.
SORT method for reinforcement learning that adds repair updates for failed rollouts using plan guidance and token probability weighting.
Study on resolving safety-helpfulness trade-offs in LLM alignment through preference dimensional expansion approach.
Research on multi-modal world models integrating tactile and visual feedback for predicting robotic action outcomes in complex environments.
Benchmark for evaluating LLM text-to-SQL performance on complex enterprise databases with intricate schemas and domain knowledge requirements.
AIvaluateXR: evaluation framework for benchmarking LLMs on XR devices, testing 17 models for on-device inference performance and selection.
DeePen: penetration testing methodology for evaluating robustness of machine learning audio deepfake detection classifiers.
Offline reinforcement learning method using value function inconsistency penalties to improve model-based policy learning from fixed datasets.
HiddenBench: 65-task benchmark evaluating multi-agent LLM systems' ability to reason with distributed information in collective settings.
Region-based reinforcement learning approach optimizing LLM performance for table question-answering with structured row-column reasoning.
A3: analytical low-rank approximation framework for transformer attention mechanisms to enable efficient LLM compression and deployment.
Framework combining algorithmic information theory and neural network pruning to improve model generalization through data compression.
Analysis of vision tokens in large vision-language models to reduce hallucinations and improve multimodal decoding with visual semantic guidance.
Dilated Unmasking Scheduler for fast non-autoregressive text generation in masked diffusion language models with improved parallelization.