Indexed metadata

A Survey on Evaluation of Embodied AI

Liyu Hou, Linyuan Gao, Yuan Wu, Yi Chang

Source record

Source: Crossref

Published: Aug 31, 2026

DOI: 10.20944/preprints202608.2328.v1

Open original source ↗

Source abstract

Embodied AI develops agents that perceive, reason, and act through closed-loop interaction with physical or simulated environments. Standardized evaluation is essential for understanding the capabilities and limitations of these agents and supporting dependable deployment. Nevertheless, existing work still lacks a systematic categorization and thorough analysis of embodied AI evaluation. To address this gap, this survey comprehensively reviews embodied AI evaluation by examining evaluation dimensions, evaluation assets, and evaluation protocols as three complementary aspects. Together, these aspects form a What-Where-How framework. Specifically, for evaluation dimensions, we organize studies across perception and understanding, cognition and reasoning, task planning and decision-making, and action execution and control, while treating safety, robustness, and generalization as cross-cutting properties and synthesizing capability boundaries across the perception-action loop. For evaluation assets, we distinguish simulators that instantiate test conditions, datasets that provide reusable evidence, and benchmarks that standardize tasks and comparisons, while analyzing the trade-offs governing their use. For evaluation protocols, we cover task-performance and physical metrics, multidimensional and hybrid evaluation, and human-in-the-loop assessment, distinguishing terminal outcomes, process signals, transfer evidence, and human oversight. Collectively, this framework establishes a coherent evaluation lifecycle from target definition and evidence construction to behavioral measurement and result interpretation. Finally, we consolidate major challenges in evaluation coverage, physical validity, reproducibility, and failure diagnosis, and advocate an Evaluation-Diagnosis-Enhancement loop for more comprehensive, dynamic, and trustworthy evaluation. Related resources are maintained at https://github.com/EmbodiedAISurvey/Embodied-AI-Eval-Survey.

Evidence graph

No public relationships recorded yet.

Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.