When a child stalls, the first question is not about the child

By INTERLAZA 9 min read

There is a particular kind of flat line every instructor recognizes. Three weeks on the same handful of concepts. Accuracy hovering around 60%. The child is cooperative, the sessions run, the data gets recorded — and nothing moves.

The real question at that moment is not why can’t this child learn this? It is: how long have I been running the same thing, and what exactly am I going to change?

That question has had a rigorous answer since the 1990s, and it comes from a place most of us never look: a school system. CABAS — Comprehensive Application of Behavior Analysis to Schooling, developed by R. Douglas Greer and colleagues at Columbia University Teachers College — is an attempt to make teaching a measurable practice rather than a craft. Its two books, Designing Teaching Strategies (2002) and Verbal Behavior Analysis (2008, with Denise Ross), spend most of their pages on something almost nobody measures: the decisions of the adult.

The rule for when to decide

Greer’s decision protocol starts somewhere unglamorous: how many data points before you are allowed to conclude anything?

The answer is to count segments, not points. The first session is an origin — it has nothing to compare against. The second gives you one segment, the third gives you two, the fourth gives you three. At three segments a trend exists and you can read it. If the direction wobbles, extend to five and read again.

Then a rule with no wiggle room in it:

  • Ascending — keep going. Whatever you are doing is working.
  • Flat or descending — a decision is due. Now.

And then the sentence that makes the whole system different from every “insights panel” ever shipped: every instructional session run without the change that was due is counted as another error. Not the child’s error. The teacher’s. The reasoning is blunt — a session spent under teaching that isn’t working is educational time gone, and it may be actively compounding the difficulty.

Greer also counts the correction: a decision finally made three sessions late does not go in the ledger as a good decision. It goes in as the correction of an error.

If you find that harsh, notice what it does. It converts “I should probably change something soon” — a feeling — into a number that someone can look at. Keohane and Greer (2005) taught the protocol to three instructors and tracked six children over eighteen months; the children needed measurably fewer teaching trials to reach the same objectives once the adults were using it. The intervention was on the adults. The effect was measured on the children.

Once a decision is due, the second half of the protocol says where to look — and the order is the point, not the list.

Figure — Where to look, in order

Four flat or falling sessions in a row

  1. Was the teaching itself intact?

    What you check:
    Every trial had a clear sample, a real chance to answer, and a consequence that matched the answer — reinforcement or a correction the child actually looked at.
    If the answer is no:
    Fix the presentation. Nothing below this rung can be read until this one is clean.
  2. What was happening around the session?

    What you check:
    Sleep, illness, hunger, a reward the child has had enough of, the ten minutes before the tablet came out.
    If the answer is no:
    Change the conditions, not the program. Same route, different moment.
  3. Are the prerequisites really there?

    What you check:
    Not "was it taught" but "is it there now" — mastered, recently, and in conditions like today's.
    If the answer is no:
    Insert the missing step before the current one and come back.
  4. Is there something physical in the way?

    What you check:
    Hearing, vision, motor access to the screen.
    If the answer is no:
    Adapt the materials, and get the professional opinion the app cannot give.

The question that is never on the ladder

“Can this child learn?” It has no answer that changes what anyone does tomorrow. The question that replaces it — “what does this child need from the way we are teaching?” — has four, and they are above.

The order matters as much as the list. Each rung is only readable once the one above it is clean: a prerequisite gap and a session run with unclear presentations produce the same flat line, and checking the child first is how the flat line gets blamed on the child. Adapted from the decision protocol in Greer (2002).

The first rung is the one that gets skipped. Before anything about the child is analyzed, Greer requires that you establish the teaching itself was intact. In his vocabulary a learn unit is one complete interlock: the adult secures attention, presents an unambiguous sample, leaves a real window to answer — about three seconds — and delivers a consequence that matches what the child did. If any piece is missing, the interaction happened but the learn unit did not, and there is nothing to diagnose yet.

That sounds like bookkeeping. It isn’t. Across five studies, replacing adult–child interactions that were not learn units with interactions that were raised correct responding by a factor of four to seven. Read that for what it is: a comparison of classroom instruction with and without a complete teaching interaction, not a promise about any particular app or child. What it establishes is narrower and more useful — the difference between teaching and almost-teaching is large enough that you cannot skip past it on the way to conclusions about a learner.

One component of that interlock is easy to get wrong and worth naming on its own. When a child answers incorrectly, the correction only counts if the child looks at the sample again while giving the right answer. Hogin (1996) tested exactly this: children who saw only their own answer and the consequence did not master the operation; children who also saw the stimulus did. The correction is not there to announce the verdict. It is there to put the right answer next to the thing it belongs to.

The trap in “he knows it”

There is a second idea in these books that lands harder on a matching app than on a classroom, and leaving it out would be dodging it.

Both volumes repeat a warning: pointing is not naming. A child who reliably selects the red card when you say “red” has a listener repertoire. A child who says “red” when shown the card has a speaker repertoire. These are functionally independent — they do not come as a pair, and assuming they do means a child gets marked as knowing a concept and never receives the instruction they were actually missing. Greer states it flatly: identifying an item among choices is not the same as naming it.

That is a real limit on what any selection-based screen can tell you, ours included. A tablet is very good at the listener half. The speaker half happens between a child and an adult in a room, which is why Interlaza’s off-tablet modes exist and why a route that only ever lives on the screen is an incomplete route.

The same books also give the flip side, and it is the more surprising half: matching to sample is not just a vocabulary exercise. In Greer’s developmental sequence, matching is the procedure used to induce “the capacity for sameness” — the pre-listener milestone that discrimination itself is built on — and visual matching is the step that sets up naming. A child who cannot yet match is not a child failing an easy task. They are a child working on the thing that comes before the task.

What Interlaza does with this, and what it cannot

Interlaza records what this method runs on: every trial, what was presented, what was chosen, how long it took, how much help was active, and whether the session followed the procedure it was configured with. Three things are built on top of that record, one per idea above.

Rung 1 is a gate, not a panel. The Program Advisor will not propose restructuring a child’s program until the recent sessions are established as having run the way they were configured — if they did not, it says so first and holds the rest back. The order is the whole point: teaching first, child second. It deliberately fails open, because a gate that shuts for lack of data blocks the advisor exactly when nobody has information.

What the teaching cost is now a number. Alongside accuracy, the Results tab reports the median trials it took to reach mastery, per concept. “Still getting them right” cannot tell you whether a change of tactic made learning cheaper; this can.

And the count Greer puts at the center now exists. When a route’s trend has gone flat or downward and nothing in the record has changed since, the app says how many sessions that has been going on for. It is about the adult, not the child — which is why it stays silent whenever the child is climbing. A number that only ever meant “sessions since you last touched this” would read as a reproach aimed at a child who is doing fine.

Now the realistic half, and it is not a to-do item. That count means “sessions since anything we can see changed” — and the app cannot see everything. An instructor who changes how they present the sample, moves the session to a better hour, or swaps a reinforcer by hand has changed the program in the way that matters, and left no trace the software can read. Neither can it attribute an edit to a route several children share. So the number is a prompt to look, never a verdict about whether anyone was paying attention, and it says as much on screen.

The ladder above works on paper regardless. The flat line was going to be there either way.