Decocted Experience Improves Test-Time Inference in LLM Agents
Decocted Experience improves test-time inference for LLM agents by optimizing exploration budgets during reasoning and search without updating model parameters.
Decocted Experience improves test-time inference for LLM agents by optimizing exploration budgets during reasoning and search without updating model parameters.
LLM-powered multi-agent simulation framework for service operation optimization that models human behavior response to design choices via stochastic optimization.
Automatically generates hard math problems from error analysis to evaluate LLM mathematical capabilities, addressing overfitting in math benchmarks.
GUIDE proposes hierarchical diagnosis framework for interpretable evaluation of GUI agents, providing diagnostic insights into long-horizon task failures beyond binary verdicts.
MolDA uses diffusion models instead of autoregressive generation for molecular understanding and generation, addressing sequential generation errors in LLM-based molecular discovery.
ShieldNet addresses supply-chain injection attacks in agentic systems where malicious behaviors are embedded in third-party tools and MCP servers used by LLM agents.
Introduces STEP dataset and STEPPER model for cognitive behavioral therapy counseling agents that identify and address automatic negative thoughts in dialogue.
Proposes metrics to assess consistency of attribution patterns in explainable AI systems under controlled input perturbations for reliable pattern recognition.
Analyzes structural limitations in multimodal AI architectures, arguing that contrastive alignment and cross-attention fusion fail at creative cognition due to modal separability assumptions.
RetailSim simulates end-to-end retail dynamics using LLM agents to evaluate seller persuasion, buyer-seller interaction, and purchase decisions across multiple stages.
Memory Intelligence Agent proposes a novel memory evolution system for deep research agents that improves trajectory retrieval and reduces storage/retrieval costs while enabling autonomous evolution.
SuperLocalMemory V3.3 presents a local-first agent memory system with biologically-inspired forgetting, cognitive quantization, and multi-channel retrieval for AI coding agents without cloud LLM dependency.
Trajectory optimization approach using offline dataset and drifting learned models without requiring forward simulation in unknown dynamics settings.
Architecture for artificial agents with history-dependent perceptual organization through feedback between slow perspective latents and perception processing.
Methods for distilling search agent behaviors from LLMs into smaller language models for knowledge-intensive multi-hop reasoning tasks with lower computational costs.
Persistent runtime system for long-lived LLM agents with auditable execution, case-based memory, deterministic safety gating, and self-perception capabilities.
Pedagogical clarification of the causality step in policy gradient derivations, explaining reward-to-go substitution in REINFORCE estimators with explicit mathematical rigor.
Framework for continuous governance, observability, and compliance of LLM, RAG, and multi-agent AI systems in enterprise environments using zero-trust principles.
ANX protocol-first framework for AI agent interaction with decoupled 3EX architecture, addressing token consumption, security, and fragmentation in agent-native protocols.
MemMachine open-source memory system for LLM agents integrating short-term, episodic, and profile memory with ground-truth preservation across multi-session interactions.
Proves information-theoretic limits of AI safety verification using Kolmogorov complexity, showing incompleteness arises from fundamental constraints beyond computational complexity.
Novel evaluation approach for adaptive AI medical devices using learning, potential, and retention metrics to assess iterative model updates and dataset changes.
QED-Nano trains small models for mathematical theorem proving, achieving performance on complex proofs while remaining reproducible and efficient compared to proprietary systems.
Comprehensive review of LLM applications in healthcare across medical specialties including diagnosis, treatment planning, and patient support with challenges analysis.
Automated LLM-aided system for Universal Verification Methodology testbench construction and stimulus generation, reducing manual IC verification effort.
Identifies Persuasion Paradox: fluent LLM explanations increase user confidence without improving task accuracy and sometimes undermining performance in human-AI teams.
Scales determinantal point processes for RAG to improve diversity in retrieved context, addressing redundancy in standard relevance-ranking retrieval pipelines.
FVRuleLearner uses LLM reasoning and operator-level reasoning trees for formal verification, automating natural language to SystemVerilog assertion translation.
Studies learning rate regularity across 1.8M student interactions on Campus AI platform, automatically generating knowledge components without manual cognitive modeling.
BLK-Assist framework for artist-led co-creation with generative AI using parameter-efficient fine-tuning methods including LoRA and LayerDiffuse for image generation.
Challenges technical objectivity of AI accuracy measurement, showing evaluation depends on context-dependent normative choices affecting error prioritization and risk distribution.
Proposes constrained maximum likelihood estimation approach for rigorous LLM failure rate estimation, addressing tradeoff between expensive gold standards and biased automatic annotation.
SoLA presents training-free compression method for LLMs using soft activation sparsity and low-rank decomposition without special hardware or expensive post-training.
Focus method learns token pair importance through learnable centroids, enabling efficient attention with minimal trainable parameters and zero downstream degradation.
AI Governance Control Stack introduces layered governance architecture for operational stability, reliability, and auditability in high-stakes AI deployments.
LPC-SM proposes hybrid autoregressive architecture separating local attention, persistent memory, and predictive correction for long-context language modeling using Orthogonal Novelty Transport.
Smart FPGA camera platform with optimized AI models for real-time jet flame detection and characterization in industrial fire safety.
Multi-agent reinforcement learning for decentralized EV virtual power plant operation with limited network visibility.
AI agents for code generation in 6G mobile networks, customizing user plane processing through natural language and code generation capabilities.
RAGnaroX: Local ChatOps assistant using small language models, Rust-based, on-premise RAG architecture with function calling for secure deployment.
Research studies scaling trade-offs in multi-agent LLM systems between team size and lifelong learning capability under cost constraints.
3D-IDE method adds implicit 3D depth representation to multimodal LLMs for indoor scene understanding.
XAttnRes proposes cross-stage attention residuals for medical image segmentation, applies LLM attention mechanisms to vision tasks.
Scene dynamic field representation for improving intuitive physics understanding and high-level reasoning in MLLMs.
Generative molecular language models pretrained on chemical data and fine-tuned for discovering new energetic materials.
Technique for reducing hallucinations in MLLMs by enabling active visual interrogation beyond passive observation.
Vision-tactile-language model for material property inference and quality inspection in manufacturing.
Multimodal architecture (CoLoRSMamba) for violence detection using conditional LoRA to couple video and audio representations.
AI-driven system for automating IPv6 communication protocol compliance verification to identify subtle non-compliance.
Method for style-steering symbolic music generation using inference-time composer vectors without requiring labeled datasets.