Stimulus equivalence: how learners derive relations you never taught

By INTERLAZA

Teach a learner that the spoken word “dog” goes with a picture of a dog. Then teach that the same spoken word goes with the written word dog. If, without any further teaching, the learner can now match the picture to the written word — and the written word back to the spoken one — something remarkable has happened. You taught two relations. The learner walked away with far more than two.

That is stimulus equivalence, and it is arguably the closest thing behaviour analysis has to a theory of how symbolic meaning gets built. It is also, for anyone teaching language and concepts, the single largest source of unclaimed efficiency in a teaching programme.

This article covers what equivalence is, the relations that define it, how to arrange teaching to produce it, and — the part most often done wrong — how to test for it without accidentally destroying the test.

The phenomenon

Murray Sidman first documented this in the early 1970s while working with a young man with an intellectual disability who could match spoken words to pictures, but could not read. Sidman taught him to match the spoken word to the written word. He then tested relations nobody had trained — and found the learner could now match written words to pictures, and pictures to written words. Reading comprehension, in a limited but real sense, had emerged without being taught.

The important claim is not that the learner memorised more pairs. It is that a set of physically unrelated stimuli — a sound, an image, a printed string — had become functionally interchangeable. Behave toward one, and you behave toward the others the same way. That set is called an equivalence class.

The three defining relations

For a set of stimuli to qualify as an equivalence class, three properties must hold when tested. Using the convention of labelling stimuli A, B and C:

Reflexivity — matching a stimulus to itself, with no training. Given A, the learner selects A. Sometimes described as generalised identity matching; in practice this is what identity match-to-sample establishes.

Symmetry — reversibility of a trained relation. You taught A→B (given the spoken word, select the picture). Symmetry means B→A emerges untaught: given the picture, the learner selects the spoken word.

Transitivity — combination across a shared element. You taught A→B and A→C. Transitivity means B→C emerges: the two stimuli that were never presented together are now related through the one they share.

When symmetry and transitivity are both demonstrated together — the learner derives C→B as well as B→C — the relation is called equivalence proper. The class is closed and bidirectional.

A useful way to hold this: reflexivity is sameness, symmetry is reversal, transitivity is linkage. Equivalence is all three at once.

Why this multiplies your teaching

The efficiency argument is arithmetic, and it is stronger than most people expect.

Consider a three-member class — spoken word (A), picture (B), written word (C). The total number of possible relations among three stimuli is 3 × 3 = nine. You directly train two of them: A→B and A→C. If equivalence emerges, the learner ends up with all nine.

Two taught relations. Seven derived for free.

Now extend to a four-member class — add, say, the written word in a second language, or a sign. Sixteen possible relations. You train three. Thirteen emerge untaught. The ratio improves as classes grow, which is why equivalence-based instruction scales so unusually well: each stimulus you add to a class multiplies the derived relations rather than adding to them.

This is the mechanism behind the observation that typically-developing children acquire vocabulary far faster than anyone explicitly teaches it. They are not being taught more words. They are deriving relations.

For a learner who is not deriving relations, this is precisely the deficit worth targeting — and it is targetable. Equivalence responding can be established through training rather than waited for.

Choosing a training structure

Which relations you train directly, and in what arrangement, measurably affects whether equivalence emerges. Three structures dominate the literature:

One-to-Many (OTM), also called sample-as-node: train A→B, A→C, A→D. One stimulus serves as the hub, related outward to each of the others.

Many-to-One (MTO), also called comparison-as-node: train B→A, C→A, D→A. Every stimulus is related inward to a shared comparison.

Linear Series (LS): train A→B, then B→C, then C→D. Each relation chains to the next.

Research comparing these — Arntzen and colleagues in particular have run this comparison repeatedly — generally finds MTO and OTM produce equivalence more reliably than linear series, with linear arrangements tending to yield the weakest outcomes, especially as class size grows. The intuition is that linear structures require the longest chain of derivation to connect the endpoints, and give the learner the fewest opportunities to contact the shared node.

Practical guidance: default to many-to-one or one-to-many. Reach for linear series only when the content genuinely has a sequential structure. If equivalence fails to emerge, restructuring the training arrangement is often more productive than simply running more trials of the same arrangement.

Testing without destroying the test

This is where equivalence programmes most often go wrong, and the error is subtle enough to be worth stating plainly:

A derived-relations probe must be run without reinforcement or corrective feedback.

The moment you deliver feedback on a probe trial, you are no longer testing whether the relation emerged — you are training it. The data become uninterpretable, because a correct response on trial five could be derivation, or it could be a response you just shaped on trials one through four.

A defensible probe therefore looks like this:

  • No feedback of any kind on probe trials — no praise, no correction, no differential consequence. This is uncomfortable to run, because the learner may experience an unusual stretch without reinforcement.
  • Mixed in among maintained baseline trials, which do carry reinforcement, so the overall density of reinforcement does not collapse and the probes are not obviously marked.
  • Symmetry and transitivity probed separately where possible, so you can see which property failed rather than only that “equivalence didn’t emerge.”
  • Criterion set in advance — typically high, since chance performance on a three-comparison array is already 33%.

If a class fails its probes, the useful diagnostic questions are: was the baseline genuinely mastered before testing, or merely at criterion on the last block? Was the training structure linear? Were the class members too perceptually similar, allowing a discrimination shortcut? Did the probe context differ so much from training that it disrupted responding?

Common mistakes

Testing before baseline is solid. Derived relations rest on trained relations. A baseline at 80% is not a baseline; it is a source of noise that will appear as “equivalence failure.”

Reinforcing during probes. Covered above, and worth repeating because it is the most common single error.

Confusing generalisation with derivation. A learner selecting a new picture of a dog is generalising along a physical dimension. A learner relating a spoken word to a printed word shares no physical dimension at all — that is derivation. They are different phenomena with different implications.

Treating equivalence as all-or-nothing. Partial class formation is informative. A learner showing symmetry but not transitivity is telling you exactly where to intervene.

Ignoring the nodal distance. Relations that require passing through more nodes are typically weaker. If you must use a longer chain, expect to strengthen the intermediate relations more thoroughly.

From theory to a running session

The gap between the literature and a Tuesday-afternoon session is real. Running equivalence-based instruction properly means maintaining baseline relations while interspersing unreinforced probes, tracking which specific derived relations have emerged for which learner, and restructuring training when a class fails — all while the person in front of you is a child with a limited tolerance for unreinforced stretches.

This is the work Interlaza was built to carry. Equivalence training and no-feedback probes are first-class exercise types rather than something to improvise: training structure is selectable, probe trials are generated without reinforcement by design, and symmetry and transitivity results are tracked per relation and per learner, so a failed class tells you which property failed. The learning history is recorded trial by trial, which is what makes the diagnostic questions above answerable rather than speculative.

The science here is fifty years old and well replicated. What has been missing is not evidence but a practical way to run it — which is the gap worth closing, because the arithmetic at the top of this article is the difference between teaching a learner nine things and teaching them two.

Further reading

  • Sidman, M. (1994). Equivalence Relations and Behavior: A Research Story. Authors Cooperative.
  • Sidman, M., & Tailby, W. (1982). Conditional discrimination vs. matching to sample. Journal of the Experimental Analysis of Behavior, 37(1), 5–22.
  • Arntzen, E. (2012). Training and testing parameters in formation of stimulus equivalence. European Journal of Behavior Analysis, 13(1), 123–135.

For the wider relational framework this sits inside, see our introduction to match-to-sample, and the science behind the platform.