Here are two tasks. Both take a few seconds. Both would be described by anyone watching as easy.
In the first, a photograph of a dog appears at the top of the screen. Below it, two photographs: the same dog, and a chair. The child touches the dog.
In the second, the tablet says the word “dog”. Below, the same two photographs. The child touches the dog.
Same child, same afternoon, same two pictures, both at 95%. A percentage does not say what the child was responding to: that is Ribes’ distinction between recognizing and comparing, and out of it come the four decisions Interlaza uses to force comparison. Both of these tasks sit on the same side of that line: both are coupling, recognizing and repeating. Neither requires comparing.
So by the structure of the task they are the same thing. And yet anyone who has sat with a child through both knows perfectly well that they are not. This article is about what the difference actually is, because naming it changes what you do next.
What makes them the same
Ribes’ criterion is structural, and it is strict. A contact is coupling when the contingencies of occurrence stay constant, when nothing permutes from trial to trial. That is not the same as saying the correct answer never changes: switch the sample to a cat and the correct comparison switches to the cat too — what stays constant is the rule (this sample’s correct comparison is always this sample’s match), never which particular item is correct on a given trial. In both of our tasks the dog photograph is always the right answer for the dog sample specifically, and there is no rule in force that could change and make that same pairing wrong on a later trial. That is why both land in the same box — by that structural test, not by how advanced either task looks.
What makes them worlds apart
Ask a different question: what holds the correspondence together?
In the identity task, the physical match is right there on the screen. The photograph at the top and the photograph below it share their form — the technical word is isomorphism. The formal support the child needs is present in the field, here and now, even though noticing it still rests on a learning history of its own: perceiving that two photographs share a form is a discrimination the child has learned to make, not a given. What the task asks of them is to tell apart the perceptual features that make up that sameness, which is exactly the adjustment criterion Ribes assigns to this contact: differentiality.
In the word→picture task, nothing of that sort is available. There is nothing dog-shaped about the sound “dog”. The two are not similar, not analogous, not connected by any physical property either of them has. What holds them together is convention — an agreement made by a community of speakers long before this child was born — and reaching it requires a conventional medium of contact in the first place.
One correspondence is held up by a physical property present in the field. The other is held up by nothing physically shared between the sound and the picture at all — the sound itself is very much present, it just carries no formal resemblance to what it names.
Two axes, not one ladder
The usual way to read Ribes is as a single ladder: contacts get more complex as you climb, each one demanding more than the last. Read that way, our two tasks are at the same rung, and the difference above simply disappears.
We find it more useful to read it as two independent axes — our own reading for this purpose, not a distinction Ribes himself draws in these terms.
Recursion (our term) — how much has to be held at once. How many things must be kept in play for a response to be correct: one datum, a rule, a rule that depends on another rule. On this axis our two tasks genuinely are at the same point. Simple pairings, nothing permuting, nothing to keep in suspension.
Anchoring (also our term), who holds the correspondence. At one end, what is present: the field itself supplies the reason the answer is right. At the other, convention: to be right, the child must already have learned that the sound “dog” names that animal, and that is not on the screen. On this axis our two tasks are at opposite extremes.
The two tasks look equally simple, and they do not demand the same learning. That is why a single ladder cannot see the difference, and why we keep the second axis.
A distinction worth not blurring
It is tempting to say the word and the photograph “differ in dimension”. They do not, and the sloppiness costs something later.
A dimension is the respect in which things are being compared — color, size, taste. It is the in what way. A word and a photograph of a dog are not two values of one dimension; they are the same content arriving through a different channel and in a different format. Different axes of the model entirely. Keeping them apart is what makes it possible to say precisely what a probe is testing — a handful of trials on material the child has not practised, with no help and no feedback.
The part that matters at the table
Here is the consequence, and it is the reason any of this is worth an article.
The same word→picture task can be two completely different contacts, depending on the child’s history.
If “dog” is functioning for the child as an acoustic pattern — a particular noise, in a particular voice, that has come to be followed by touching a particular picture — then the task is pure coupling. Recognize a sound, repeat a gesture. Nothing about it is language.
If “dog” is functioning as a word, something else is going on entirely.
The task on screen is identical in both cases. The child’s accuracy is identical. And the score cannot separate them, which is the same lesson as the first article, arriving from a new direction: performance tells you how often the child was right, never what made them right.
How you would actually tell
Only one thing distinguishes them, and it does not run through a harder version of the same task: a probe, change something the acoustic pattern depends on but the word does not, and see whether the answer survives.
If “dog” is working as a word, it should tend to survive:
- another exemplar — a different dog, one the child has not seen;
- another voice — the same word said by someone else.
If it is working purely as an acoustic pattern, the first may survive and the second is less likely to — though not a sure thing either, since different voices still share considerable acoustic similarity (the same phonemes, a broadly similar pitch contour), so some generalization across voices can happen even without the word having become a word for this child. The written word is not part of this test at all, and it would be a mistake to add it: understanding spoken language does not require reading, and at this product’s age range most children have not learned to read yet. A child who fails to respond to the printed word “dog” has told you nothing about whether the spoken word functions as a word for them — only that they cannot read, which is expected and unrelated.
We should be straight about which of these Interlaza can run today. The exemplar probe is implemented: the early stages close by redrawing the answer in its other variant, realistic photograph to pictogram, with no feedback and no bearing on progression. The other-voice probe is not: the library holds one recording per concept per language, so it needs recordings made, not code written. A word→picture exercise, separately, already exists in the app (the symbolic type, built on each concept’s stored word text) — it is a genuinely different task from a voice-generalization probe, not a substitute for one, and we mention it here only to be precise about what already exists rather than to suggest it answers the question this section is about.
What to do with this
Mostly, stop letting “simple” do so much work.
Two stages that both look easy, that both sit early in a route, that a child passes at the same rate, can be resting on completely different foundations. One is asking them to see a sameness that is in front of them. The other is asking them to bring something to the situation that is not in it at all, and a child who does that has done something considerably more interesting than the accuracy suggests.
Which returns to where the first article ended. The criterion is never in the cards. Somebody holds it — the instructor, the family, the community that agreed what each word names. What this second axis adds is that how far away that somebody is standing is itself part of what makes a task hard, and it is not visible in any score.
Interlaza draws on the interbehavioural tradition (Kantor; Ribes and López), Varela and Quintana’s Competence Transfer Matrix (1995), Sidman’s stimulus equivalence, and Relational Frame Theory. Full references are on our science page.