Agentic Performance at the Edge: Insights from Benchmarking
Empirical study benchmarking agentic AI performance on edge devices constrained to 8B parameters, measuring quality degradation under hardware constraints.
Empirical study benchmarking agentic AI performance on edge devices constrained to 8B parameters, measuring quality degradation under hardware constraints.
Research on safeguarding autonomous driving MLLMs using Markovian safety logic for temporal reasoning in dynamic traffic scenarios.
Research using LLMs to discover efficient branching policies for Mixed Integer Linear Programming solvers, reducing dependence on expert demonstrations.
Research examining reliability gaps in interactive agent benchmarks, showing surface-level outcome checks fail to verify actual agent behavior and task completion.
Research on ASIA, an autonomous agent for system identification that automates model selection, training algorithms, and hyperparameter tuning in dynamical systems learning.
Research on SkillEvolver, a meta-skill framework for online learning and iterative refinement of agent skills from real deployment without static, hand-authored artifacts.
Research on improving LLM structural understanding of graphs by sharpening internal attention mechanisms without external adapters or fine-tuning.
Research establishing statistical methods and U-statistics framework for measuring AI agent reliability and consistency under semantic perturbations and trajectory-level stability.
Research introducing benchmark for continual learning on evolving biomedical knowledge graphs, addressing real-world asynchronous KG updates beyond synthetic splits.
Research on LLM-based storytelling agent for older adults integrating knowledge graphs, argumentation theory, and argument mining to reduce hallucinations and improve transparency.
Agent-First Tool API paradigm redesigns APIs for autonomous agents, addressing architectural mismatches with human-oriented CRUD APIs.
Examines lack of explicit reasoning and interpretability in deep learning models and proposals for understanding representations.
SciAidanBench measures scientific creativity of LLMs through open-ended questions, examining uneven capability progress.
LLARS open-source platform enabling collaboration between domain experts and developers for LLM prompt engineering and evaluation.
Budget-efficient automatic algorithm design using LLMs via code graph representation to reuse algorithmic components.
Argues for calibrated verification over mechanistic interpretability for AI deployment authorization in sensitive domains.
PRISM detection system for preventing secret leakage propagation through shared context in multi-agent LLM pipelines.
Hierarchical causal framework for explainable model predictive control in safety-critical infrastructure.
Evolutionary framework using learned optimization policies as teachers to generate heuristic programs for combinatorial optimization.
Investigation of bias and robustness issues in LLM toxicity evaluation benchmarks used for deployment certification.
Framework for self-evolving agents that distill experience from interactions to adapt to novel tasks at deployment time.
Theoretical framework examining agent cybernetics as foundational science for long-horizon LLM agents using tool loops and reflection.
MATRA threat modeling framework for assessing security risks in LLM-based agent systems with tool access.
Multi-task benchmark for aligning language descriptions with urban trajectory data, focusing on mobility modeling and planning.
ComplexMCP benchmark evaluates LLM agents on interdependent tool use in realistic scenarios with 300+ tests on Model Context Protocol.
Knowledge graph QA method using LLMs with retrieval-augmented generation, proposing informative path supervision for training.
Compares reasoning vs non-reasoning LLM judges showing structured verification benefits vary by task, proposes cost-efficient routing for LLM-as-a-Judge applications.
NanoResearch co-evolves skills, memory, and policy for personalized multi-agent research automation, adapting outputs to individual researcher resource constraints and preferences.
Investigates internal mechanisms and cross-modal information processing in audio-visual LLMs through interpretability probing techniques. Limited practical application focus.
MaD Physics benchmark evaluates agents for scientific discovery under physical measurement constraints, balancing quality and quantity with cost-benefit trade-offs.
Studies nonlinear impact of misleading information on LLM long-context performance, quantifying how distractors degrade retrieval-augmented generation and agentic systems.
Evaluates AI pentesting agents on real-world security targets beyond predefined CTF benchmarks, assessing offensive capabilities in unconstrained environments.
Generalized Turing Test framework for comparing arbitrary agent intelligence through indistinguishability via Turing comparators, dataset and task-agnostic.
BenchCAD comprehensive benchmark for programmatic CAD code generation from visual/textual inputs, requiring understanding of 3D structure, parameters, and engineering operations.
Rate-distortion framework for agent memory optimization preserving decision-relevant distinctions under limited runtime budgets rather than descriptive relevance criteria.
Shepherd runtime substrate formalizes meta-agent operations with typed execution traces and Git-like replay, achieving 5x faster process forking than Docker and 95%+ prompt-cache reuse.
Evaluates open-source LLMs for algorithm generation and ensemble methods for conjecture verification in number theory, demonstrating specialized domain application.
Pair-GRPO unified framework addresses instability in LLM RLHF alignment through preference-based RL optimization with improved gradient interpretability and variance reduction.
Delulu benchmark for detecting code hallucinations in LLM fill-in-the-middle tasks across 7 languages with 1,951 verified samples and 4 hallucination types.
Research evaluating AI companion chatbots for simulated relationships, analyzing emotional dependence risks and psychological harm. No direct developer or technical focus.
MedThink uses knowledge distillation and teacher-guided reasoning correction to compress LLM diagnostic capabilities into smaller models for resource-constrained clinical settings.
BaLoRA extends LoRA fine-tuning with Bayesian uncertainty quantification for large pre-trained models.
Transformer-based causal discovery method for non-stationary time series data with contemporary and lagged relationships.
Research showing product context improves AI coding agent decision compliance by 49%, with benchmark across 8 software engineering tasks.
Safety-Aware Denoiser framework for controlling safety risks in text diffusion models with guidance mechanisms.
Empirical analysis of feature repulsion and spectral properties during two-layer neural network grokking.
Foundation model approach for gene regulatory network inference from single-cell transcriptomic data.
Retrieval-augmented vision-language-action model for autonomous driving addressing long-tail scenario generalization.
Activation reuse technique exploiting token-wise redundancy in diffusion language models for efficient inference.
Empirical study showing weight pruning amplifies bias in compressed LLMs across three models and pruning methods.