Reliability [AL]

Reliability

Reliability refers to the consistency of a measure or procedure — the extent to which it produces the same results under the same conditions. A reliable measure is one that, when applied repeatedly to the same phenomenon, yields consistent results. Reliability is a necessary — though not sufficient — condition for validity: a measure cannot be valid if it is unreliable, but a reliable measure may still be invalid (consistently measuring the wrong thing).

Types of Reliability

Test-retest reliability assesses the consistency of a measure over time. The same participants complete the same measure on two separate occasions (typically separated by two to four weeks — long enough to prevent recall of previous answers, but short enough that the construct being measured is unlikely to have genuinely changed). The scores from the two administrations are correlated; a coefficient of at least +0.80 is generally considered acceptable. High test-retest reliability indicates that the measure produces stable, consistent results across time.

Inter-rater reliability (inter-observer reliability) assesses the consistency of measurement across different observers or raters. Two or more independent raters score or categorise the same material — for example, two observers coding behaviour from the same video recording, or two clinicians diagnosing from the same case vignette. Agreement is assessed by correlating the two sets of scores or calculating percentage agreement. High inter-rater reliability indicates that the measure is sufficiently objective that different raters applying it reach the same conclusions.

Internal consistency assesses the extent to which the items within a scale all measure the same underlying construct. The most common index is Cronbach's alpha (α), which reflects the average intercorrelation between all items on the scale. A value of α ≥ 0.70 is typically considered acceptable; α ≥ 0.80 is considered good. Low internal consistency suggests the items are measuring different constructs and should not be combined into a single score.

Improving Reliability

Reliability of observational research is improved by operationalising behavioural categories precisely and training observers until inter-rater reliability exceeds the criterion threshold. Reliability of questionnaires and interviews is improved by using standardised instructions, clear and unambiguous item wording, and piloting to identify confusing items. In experimental research, reliability is improved by standardising all aspects of the procedure — instructions, timing, stimuli, and recording methods — and using objective rather than subjective measures wherever possible.

 Key Takeaways

  • Reliability is the consistency of a measure — the extent to which it produces the same results under the same conditions.
  • Test-retest reliability: same measure applied twice to same participants; correlation ≥ 0.80 is acceptable — indicates stability over time.
  • Inter-rater reliability: agreement between independent observers scoring the same material — correlation or agreement ≥ 0.80 required.
  • Internal consistency (Cronbach's alpha): extent to which scale items all measure the same construct — α ≥ 0.70 acceptable; ≥ 0.80 good.
  • Reliability is necessary but not sufficient for validity — a reliable measure may consistently measure the wrong thing.
  • Reliability is improved by precise operationalisation, observer training, standardised procedures, and piloting to eliminate ambiguous items.