Indexed metadata

The role of reliability in experiments

Jeffrey N. Rouder, Mahbod Mehrvarz, Martin Schnuerch

Source record

Source: Crossref

Published: Mar 16, 2026

DOI: 10.1111/bmsp.70042

Open original source ↗

Source abstract

Abstract We are concerned about an emphasis on reliability for analysis of psychology experiments. Experiments have two elements of sample size: the number of individuals and the number of replicate trials within a task, and that complicates reliability measures. To account for these elements, we distinguish among three levels of analysis: (1) A foundational level that centers task properties without recourse to either element of sample size. An example statistic is intraclass correlation which is the proportion of variances without reference to sample sizes. (2) An intermediate level that centers the number of trials but not the number of individuals. An example statistic on this level is reliability which describes variabilities with reference to numbers of trials but not numbers of individuals. A final level centers both the numbers of individuals and trials. An example quantity is the uncertainty in a correlation coefficient, which, ideally, reflects sample size limits in individuals and trials. Reliability describes an intermediate level – neither useful for communicating foundational task properties nor interpreting correlations. We advocate that researchers consider all three levels and highlight the role of hierarchical models in doing so.

Evidence graph

No public relationships recorded yet.

Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.