95% of What? The Question a Percentage Cannot Answer

By INTERLAZA 10 min read

A child finishes the week at 95%. Good news — but 95% of what?

Here are two children who both hit that number this month.

Marta matched twelve photographs to their identical twins, over and over, and got nearly all of them right. Show her a different photograph of the same dog and she stalls.

Diego hit the same 95% on the same kind of screen. Show him a different photograph of the same dog and he is still right. Show him a drawing of it, and he is still right.

Marta and Diego produced the same number. They did not do the same thing. And nothing in a percentage — not the total, not the trend line, not the streak — will tell you which of the two you have.

This matters because it changes what you do on Monday. Marta needs new exemplars before anything else; more repetitions of the same twelve will raise her score and teach her nothing. Diego is ready to move on. Read only the percentage and you would give them both the same next session.

Figure — The same score, twice

Practised Probe — new material, no feedback
Practised material with feedback New material no feedback Marta 95% collapses Recognising the same thing again Diego 95% holds Responding to the relation

One performance collapses, the other barely moves. The score never said which was which; the probe does.

Two children at 95%. The percentage is identical because it counts correct answers, and both are answering correctly; what differs is what they are answering to. Only a no-feedback probe with material the child was never trained on tells the two apart — which is why Interlaza reports those probes separately from everything else.

The distinction, and where it comes from

The Mexican behavioural psychologist Emilio Ribes spent a career on precisely this problem: that the same visible performance can be produced by qualitatively different kinds of interaction. In Teoría de la Conducta he sets out five, each a different way an individual and their environment can be organised into a single episode.

The two that concern ordinary practice are the first and the third.

Coupling — the interaction is organised around something that is simply there. Recognising, matching, repeating. What holds it together is present in front of the child: this photograph and that photograph share their form. Nothing swaps roles between trials; the dog picture is always right for the dog picture.

Comparison — no element carries its function on its own. What matters is a relation — bigger than, the same as, the opposite of — and the relation itself is what must be responded to. Invert the terms and the correct answer inverts with them.

Marta is coupling. Diego, on the evidence, is doing something more.

The other three — alteration, extension, transformation — matter enormously in the theory and rarely in a tablet session, so they are at the end of this article rather than the middle.

The uncomfortable part

Here is the finding that makes this more than vocabulary: you cannot tell which contact you are looking at from what the task looks like.

A word-to-picture task feels more advanced than photo-to-photo matching. Structurally it is not: nothing permutes, the word is always right for its picture, the contingencies stay constant. It is coupling — a different kind of coupling, held together by convention rather than by resemblance, but coupling.

And a task that genuinely does point higher can still be solved by coupling. Carpio places matching-to-sample at the level of relating; Ribes warns that those tasks are usually solved by recognising. Formally they point up. Functionally they can stay down. A correct answer does not distinguish the two, which is exactly why a score cannot.

So: not from the task, and not from the percentage. Then from what?

From what happens when you take the support away

There is one arrangement that separates them, and it is not clever — it is just strict.

Remove the feedback. Change the material. See what survives.

If the performance was recognition, it collapses: the child was responding to these pictures, and these pictures are gone. If the child was responding to the relation, the new material makes little difference, because the relation is still there.

That is the whole test. It is why Interlaza runs no-feedback probes with material the child has not been trained on, and why those probes are reported separately from everything else.

Reading it in your own results

On a child’s Results tab you will now find two statements, and they are deliberately different in kind.

What kind of practice this was. Something like:

Recognising the same thing again — 94% · 146 trials Responding to a relation — 6% · 9 trials

This describes the tasks you set, not the child. It is knowable, because it is a property of the exercise. And it is often the more useful of the two, because it makes visible something that is otherwise invisible: two months of practice that was almost entirely recognition, without anyone deciding that.

What the no-feedback checks say. Something like:

Practised material — 71% New material, no help — 33%

The score comes from recognising material already seen. With new material and no help it drops sharply, so what is happening is recognition rather than relating.

That is the sentence a percentage cannot produce. And when the probes hold up instead, it says the opposite — that something transferred, and the child is responding to the relation rather than to the particular items.

Three things it deliberately refuses to do:

  • It never gives the child a level. The same task is a different contact depending on what the child already knows, so a label read off the exercise type would be wrong a good share of the time — while carrying the app’s authority. What is classified is the task.
  • It stays silent when it cannot know. Too few checks, or trained material that is not solid yet, and there is no verdict — because there is no established performance for a check to confirm or contradict.
  • It always shows both figures. Where the line falls between 71% and 33% is a judgement, not a standard. You should be able to disagree with it.

What to do with it

If practice is 90%-plus recognition and the probes collapse, the answer is not more trials. It is more exemplars — different photographs, different voices, the drawing as well as the photo — and probes to check.

If the probes hold, the recognition work has done its job and the child is ready for tasks where the relation is what varies.

And if the probes say nothing yet, that is a real answer too: run some.

The deeper layer, for the curious

Everything above is the part you can act on. Underneath it there is a research literature that goes considerably further, and it is worth knowing it exists.

Ribes’ five contacts each have their own criterion of adjustment — differentiality, effectiveness, precision, congruence, coherence — and the first three have published indices with formulas. They are not accuracy in disguise: they count omissions as failures to adjust, two of them measure time rather than trials, and the precision index is a product across two conditions, so succeeding under one rule and failing the other collapses it toward zero.

One property worth stating because it trips people up: that precision index cannot exceed 0.25. Both its factors share a denominator, so a flawless run reads 0.25 — while the interpretation bands published alongside it run from 0 to 1 and would call that “incipient”. The bands are described by their own author as arbitrary. That is a good reminder that a number from a paper still needs reading, not just reporting.

Two of the five contacts also cannot be arranged on a tablet at all: extension is a contact between two people, not between a person and an object, and transformation — talking about how one talks — is far outside the age range this product serves. Naming them and leaving them undone is more honest than a version that pretends.

None of that is needed to use what is on your Results tab. It is there because the distinction on that tab is not something we invented; it comes from somewhere, and the somewhere is checkable.


References

Andrade-González, D. E., León, A., & Hernández Eslava, V. (2020). Tarea de transposición y contactos funcionales de comparación: una revisión metodológica y empírica. Acta Comportamentalia, 28(4), 539–565. — The six criteria a task must meet to measure the comparison contact — and with them the reason colour will not do: red is not “more” than blue.

Carpio, C. A. (1994). Comportamiento animal y teoría de la conducta. In L. J. Hayes, E. Ribes & F. López (Eds.), Psicología interconductual: contribuciones en honor a J. R. Kantor (pp. 45–68). Guadalajara, Mexico: Universidad de Guadalajara. — The alternative naming of the five criteria — ajustividad, efectividad, pertinencia, congruencia, coherencia — and the note placing matching-to-sample at the pertinencia level.

Ribes Iñesta, E. (2018). El estudio científico de la conducta individual: una introducción a la teoría de la psicología. Mexico: Editorial El Manual Moderno. — Chapter 10, Tabla 10-1: the five functional contacts. This is the table the distinction in this article comes from.

Ribes Iñesta, E. (2021). Teoría de la psicología: corolarios. Granada, Spain: Co-presencias Editorial. — The naming we use here and in the app: differentiality and precision where Carpio says ajustividad and pertinencia.

Ribes, E., & López, F. (1985). Teoría de la conducta: un análisis de campo y paramétrico. Mexico: Trillas. — The original taxonomy, which everything above is a reformulation of.

Ribes, E., Vargas, I., Luna, D., & Martínez, C. (2009). Adquisición y transferencia de una discriminación condicional en una secuencia de cinco criterios distintos de ajuste funcional. Acta Comportamentalia, 17(3), 299–331. — Training and transfer tests running through all five criteria, with 24 participants.

Serrano, M. (2009). Complejidad e inclusividad progresivas: algunas implicaciones y evidencias empíricas en el caso de las funciones contextual, suplementaria y selectora. Revista Mexicana de Análisis de la Conducta, 35(monographic issue), 161–178. — The formulas for the three indices — differentiality, effectiveness, precision — and the interpretation bands. Footnote 2 on page 168 is where the author says the band ranges are arbitrary.