Stokes and Baer’s old complaint still holds: programs teach a skill, measure it where it was taught, and call it learned. The child who names the apple card in the therapy room and not the apple in the kitchen has learned something real, and it is not what the program says.
Four kinds, and they fail separately
Across materials. The skill holds with a different picture, a different exemplar, the real object rather than the photo. This is the one a matching program most often fails, because it is so easy to teach one picture per concept.
Across people. It holds with a parent, a sibling, a different Instructor. A skill available only to the person who taught it is a fact about that relationship.
Across settings. Kitchen, classroom, car, park. Rooms carry an enormous amount of control that nobody plans.
Over time. The skill is still there in three weeks with no practice in between. This one is usually called maintenance and is usually the least measured of the four.
Passing one says nothing about the others, which is why “he generalizes well” is not a sentence a record can support.
Figure — What changes when checking generalization
- Reference situation
- A different exemplar
- A different person
- A different place
Plan it in from the first session
The reliable methods are unglamorous.
Teach with several exemplars, not one. Three photographs of a dog, not the same photograph thirty times. Early on this looks slower and it is not: what a single exemplar buys in speed it takes back in a skill that only works on that image.
Vary the irrelevant while you teach. Sit on the other side of the table. Change the room. Let someone else run a session. Much of what you hold constant through teaching can end up as part of what the child learned, without anyone planning it that way.
Recruit the natural consequence. The skills that survive are the ones the world responds to on its own. A request that gets the thing needs nobody to maintain it; a label that nobody reacts to needs the program to keep supporting it.
Testing it without destroying it
The protocol below rests on one central rule: no feedback, no prompting, no correction. Praise the effort afterwards if you like, but during the probe the child’s answer must not be told whether it was right — otherwise you have run a teaching trial with an unusual name, and the number that comes out cannot be compared to anything.
What counts as untrained material depends on what you are testing. If you are probing whether a skill generalizes across materials, the exemplar itself has to be genuinely new — reserve a couple of exemplars per concept at the start, hold them out of every teaching session, and use them only for probing. If instead you are probing across people, settings or time, previously-trained material is expected: what has to be new is the dimension you are testing (a different person, a different room, or the gap since it was last practiced), not the material itself.
Interlaza does exactly this: reserved exemplars are pinned as explicit items before the session pool is trimmed, so the same ones stay held out session after session, and probe trials are excluded from the mastery streak, from the accuracy window and from the mastery estimate. That exclusion is the whole point — a probe that fed the mastery calculation would let the test decide its own result.
What a mastery score is and is not
Worth stating plainly, because it is the claim most easily overstated: a child matching pictures reliably has demonstrated they can tell those things apart and pair them. That is visual discrimination, not the same thing as understanding language. It is not evidence that they understand the word, that they will use it, or that they will recognize the thing anywhere else — that last one is precisely what a generalization probe is for, which is why it deserves to be in the plan rather than in the hopes.
Further reading
- Stimulus equivalence — relations that emerge untaught, and how to test them without training them.
- ABA data collection methods — choosing the measure before you need it.