Stokes and Baer’s old complaint still holds: programs teach a skill, measure it where it was taught, and call it learned. The child who names the apple card in the therapy room and not the apple in the kitchen has learned something real, and it is not what the program says.
Four kinds, and they fail separately
Across materials. The skill holds with a different picture, a different exemplar, the real object rather than the photo. This is the one a matching program most often fails, because it is so easy to teach one picture per concept.
Across people. It holds with a parent, a sibling, a different Instructor. A skill available only to the person who taught it is a fact about that relationship.
Across settings. Kitchen, classroom, car, park. Rooms carry an enormous amount of control that nobody plans.
Over time. The skill is still there in three weeks with no practice in between. This one is usually called maintenance and is usually the least measured of the four.
Passing one says nothing about the others, which is why “he generalizes well” is not a sentence a record can support.
Plan it in from the first session
The reliable methods are unglamorous.
Teach with several exemplars, not one. Three photographs of a dog, not the same photograph thirty times. Early on this looks slower and it is not: what a single exemplar buys in speed it takes back in a skill that only works on that image.
Vary the irrelevant while you teach. Sit on the other side of the table. Change the room. Let someone else run a session. Anything you hold constant through teaching becomes part of what the child learned.
Recruit the natural consequence. The skills that survive are the ones the world responds to on its own. A request that gets the thing needs nobody to maintain it; a label that nobody reacts to needs a program forever.
Testing it without destroying it
A generalization probe has one rule and it is absolute: no feedback, no prompting, no correction. Praise the effort afterwards if you like, but during the probe the child’s answer must not be told whether it was right — otherwise you have run a teaching trial with an unusual name, and the number that comes out cannot be compared to anything.
The second rule: the material has to be genuinely untrained. Reserve a couple of exemplars per concept at the start, hold them out of every teaching session, and use them only for probing. Anything the child has practiced has stopped being evidence.
Interlaza does exactly this: reserved exemplars are pinned as explicit items before the session pool is trimmed, so the same ones stay held out session after session, and probe trials are excluded from the mastery streak, from the accuracy window and from the mastery estimate. That exclusion is the whole point — a probe that fed the mastery calculation would let the test decide its own result.
What a mastery score is and is not
Worth stating plainly, because it is the claim most easily overstated: a child matching pictures reliably has demonstrated they can tell those things apart and pair them. That is receptive discrimination. It is not evidence that they understand the word, that they will use it, or that they will recognize the thing anywhere else — that last one is precisely what a generalization probe is for, which is why it deserves to be in the plan rather than in the hopes.
Further reading
- Stimulus equivalence — relations that emerge untaught, and how to test them without training them.
- ABA data collection methods — choosing the measure before you need it.