Observational Design

Observational Design: Behavioural Categories, Event and Time Sampling

When using structured observation as a research method, the researcher must carefully design the observation system before data collection begins. The core decisions concern: how to define and operationalise the behaviours to be observed (behavioural categories); how to record the occurrence of those behaviours over time (sampling method); and how to ensure that the coding system is applied consistently (inter-rater reliability). Well-designed observational systems produce reliable, quantitative data that can be statistically analysed.

Behavioural Categories

Behavioural categories are the pre-defined, operationally specified classes of behaviour that the observer will record. Good behavioural categories should be:

  • Operationally defined: each category must have a clear, unambiguous definition that specifies exactly what counts as an instance of that behaviour — so that any observer can apply the category consistently.
  • Mutually exclusive: any given behaviour should fit into only one category — so that observers make a clear classification without ambiguity.
  • Exhaustive: the set of categories should cover all behaviours likely to be observed — there should be no behaviour that falls outside all categories (an 'other' category can be included as a catch-all).
  • Appropriate in number: too few categories miss important distinctions; too many make coding unreliable (observers struggle to make rapid fine-grained distinctions in real time).

Event Sampling

In event sampling, the observer records each occurrence of a target behaviour whenever it happens during the observation period. A tally mark is added each time the defined behaviour occurs. Event sampling produces a count of the frequency of behaviours and is most appropriate for:

  • Discrete, clearly defined behaviours with a distinct beginning and end (e.g. a physical contact, a verbal aggression, a smile).
  • Behaviours that are relatively infrequent — so that each occurrence can be recorded without the observer missing other events.

Limitation: if behaviours are very frequent or of variable duration, event sampling can miss instances or fail to capture how long each behaviour lasted.

Time Sampling

In time sampling, the observer records behaviour at predetermined time intervals — for example, noting what behaviour is occurring at the end of every 30-second interval. At each interval, the observer records which behaviour category is occurring at that moment (or which categories occurred during that interval, depending on the design).

Time sampling is most appropriate for:

  • Continuous behaviours of varying duration (e.g. proximity to a caregiver, engagement with a task).
  • Situations where multiple behaviours could occur simultaneously — time sampling provides a snapshot rather than a continuous record.

Limitation: behaviours occurring between sampling intervals are missed. The interval length must be chosen carefully — too long and many behaviours are missed; too short and the observer cannot code adequately between intervals.

Inter-Rater Reliability

Whatever sampling method is used, inter-rater reliability must be assessed: two independent observers code the same observation session and the degree of agreement is measured (typically using a correlation coefficient or percentage agreement). High inter-rater reliability (conventionally r > 0.80 or >80% agreement) indicates that the behavioural categories are clearly operationalised and consistently applied. Low agreement indicates that the coding system needs revision before the main study proceeds.

 Key Takeaways

  • Behavioural categories: pre-defined, operationalised classes of behaviour — must be mutually exclusive, exhaustive, operationally defined, and manageable in number.
  • Event sampling: records each instance of a target behaviour when it occurs — produces frequency counts; best for discrete, infrequent behaviours.
  • Time sampling: records behaviour at preset intervals (e.g. every 30 seconds) — provides snapshots; best for continuous or simultaneous behaviours.
  • Event sampling limitation: misses very frequent or long-duration behaviours; time sampling limitation: misses behaviours between intervals.
  • Inter-rater reliability: two independent observers code the same session — agreement measured by correlation or percentage. Conventionally ≥ 0.80 or 80%.
  • A well-designed observational system produces reliable, quantitative data — the quality of the coding system determines the quality of the data.