Cognitive Approach
Operant Conditioning
Whilst classical conditioning explained how organisms learn involuntary, reflexive responses to stimuli, it could not account for the vast range of voluntary, purposeful behaviours that make up most of everyday life. Operant conditioning, developed principally by B.F. Skinner (1938), extended the behaviourist account to cover the learning of voluntary behaviours through their consequences. The core principle is that behaviour is controlled by its outcomes: responses that produce favourable consequences are strengthened; responses that produce unfavourable consequences are weakened.
Thorndike's Law of Effect
Skinner's work built on Edward Thorndike's (1911) Law of Effect: behaviours followed by satisfying consequences are more likely to recur; behaviours followed by unsatisfying consequences are less likely to recur. Thorndike demonstrated this with puzzle box experiments — cats placed in boxes learnt over successive trials to operate the release mechanism to escape, as escape (a satisfying consequence) strengthened the effective behaviour through trial and error learning.
Skinner's Operant Conditioning
Skinner developed a highly controlled experimental apparatus — the Skinner box — in which animals (typically rats or pigeons) could perform simple, measurable responses (lever pressing, key pecking) whose consequences were automatically controlled and recorded. This allowed precise, objective measurement of learning under various reinforcement schedules.
Skinner identified three key types of consequence that shape behaviour:
- Positive reinforcement: the presentation of a desirable stimulus following a behaviour, increasing the probability of that behaviour recurring. A rat receiving a food pellet for pressing a lever is positively reinforced.
- Negative reinforcement: the removal or avoidance of an aversive stimulus following a behaviour, increasing the probability of that behaviour recurring. A rat that presses a lever to switch off an electric shock is negatively reinforced. Negative reinforcement always increases behaviour — it is not punishment.
- Punishment: a consequence (either the application of an aversive stimulus or the removal of a desirable one) that decreases the probability of the preceding behaviour. Skinner distinguished positive punishment (adding something aversive, e.g. a shock) from negative punishment (removing something desirable, e.g. confiscating a privilege).
Schedules of Reinforcement
Skinner showed that the pattern or schedule with which reinforcement is delivered has powerful effects on the rate and persistence of behaviour. Continuous reinforcement (rewarding every correct response) produces rapid learning but rapid extinction when reinforcement stops. Partial reinforcement schedules — particularly variable ratio schedules (reinforcing after an unpredictable number of responses) — produce the highest response rates and the greatest resistance to extinction. This explains why gambling behaviour (rewarded on a variable ratio schedule) is highly persistent and difficult to extinguish.
Evaluation
Operant conditioning has extensive empirical support from animal research and has been applied to education (token economy systems, programmed learning), behaviour therapy (behaviour modification), and the analysis of workplace motivation. Skinner's careful experimental work exemplifies scientific rigour. However, like classical conditioning, operant conditioning is criticised for ignoring cognitive processes — the behaviourist insistence that behaviour is entirely a function of reinforcement history cannot account for latent learning (Tolman's rats learning mazes without reinforcement), insight learning, or the role of expectations and representations in human and animal behaviour. The approach is also criticised for its mechanistic view of humans as passive responders to environmental consequences.
Key Takeaways
- Operant conditioning (Skinner, 1938) explains the learning of voluntary behaviour through consequences: favourable outcomes strengthen behaviour; unfavourable outcomes weaken it.
- Thorndike's Law of Effect (1911): behaviours followed by satisfying consequences become more likely; behaviours followed by unsatisfying consequences become less likely.
- Positive reinforcement: adding a desirable stimulus, increasing behaviour. Negative reinforcement: removing an aversive stimulus, increasing behaviour. Punishment: a consequence that decreases behaviour.
- Negative reinforcement is NOT punishment — both negative reinforcement and positive reinforcement increase the probability of behaviour; only punishment decreases it.
- Variable ratio reinforcement schedules produce the highest response rates and greatest resistance to extinction — explaining the persistence of gambling behaviour.
- Operant conditioning is criticised for ignoring cognitive processes — latent learning and insight learning demonstrate that behaviour can be shaped by expectations and representations, not only by reinforcement history.