Reading the data

Single-case design: how one learner can answer a research question

How single-case designs work: each learner is their own control, the common designs and what each one can conclude, and what a graph alone cannot settle.

For: Instructors Researchers

By INTERLAZA 7 min read

An Instructor changes how prompts are faded and, over the next two weeks, the child’s independent responses climb. Something worked — but what? The child also turned three, started sleeping better, and got a new person on Thursdays.

Group research answers that with a different logic: random assignment and controls built into the design are what let averaging over many people cancel out the unrelated things — averaging by itself, without those, does not do that work. Group designs are also not available to somebody working with one child, and most practice questions are not about that anyway: not does this work on average but is this working, for this learner, now.

Single-case designs are built for exactly that. They are not a weaker version of a trial — they are a different design with its own logic, standard in applied behavior analysis since the field started.

The idea: the learner is their own control

Instead of comparing one group against another, you compare conditions within one person, measuring repeatedly under each.

What makes it evidence rather than a story is replication: the effect showing up more than once, tied to the condition rather than to the calendar. A behavior that rises when you introduce the change, falls when you withdraw it, and rises again when you bring it back has told you something no single before-and-after can.

That is also the reason the baseline matters more than people expect. A baseline is not dead time before the interesting part — it is the prediction. You measure until the data would let you say what the next few sessions look like without the change; only then does a departure from it mean anything. Introducing the change while the line is still climbing on its own throws away the comparison you were about to make.

The designs, and what each can conclude

Which one fits depends on the question being asked, whether the effect is expected to reverse, whether several learners or behaviors are available to stagger, and whether withdrawing something that is helping is ethically acceptable — no single design is the right choice for every case.

A-B. Baseline, then intervention. Easy, and the weakest: it shows that something changed at the same time as your change. Everything else that happened that week has the same claim. Useful for monitoring, not for concluding.

A-B-A-B (withdrawal). Baseline, intervention, back to baseline, intervention again. The effect has to appear, go, and come back — which is hard for a coincidence to imitate. Where reversing the effect is both expected and ethically acceptable, it buys one of the strongest conclusions among the common designs — a conditional strength, not an automatic one — and it has two real costs: withdrawing something that is helping is not always acceptable, and some learning does not reverse. Note that this design alternates exactly two conditions and returns to the first: four differently-named phases in a row is not an A-B-A-B, it is an A-B-C-D.

Multiple baseline. Start the change at a different time for each of several learners, behaviors or settings, while everyone stays in baseline until their turn. If each line moves when its own change arrives and not before, the staggering itself is the replication — and nothing has to be withdrawn.

Alternating treatments. Two conditions alternate rapidly, often session by session, and you compare them directly. It answers “which of these two is better here?” quickly, and it needs conditions the learner can tell apart.

Changing criterion. The bar moves in planned steps and the behavior is expected to track it. Useful for something built gradually, where a withdrawal makes no sense.

Reading the graph

Judgment here is visual and structured, not a p-value. What is being asked of the chart:

  • Level — did the value move?
  • Trend — which way was it going in each phase, and did the direction change?
  • Variability — how noisy is each phase? Judging a change against that noise is a matter of degree, not a hard rule — a shift barely bigger than the existing spread is weak evidence on its own, and what strengthens it is the same pattern repeating across the design’s replications, not a single phase read alone.
  • Immediacy — how fast did it move after the phase line?
  • Overlap — how much do the two phases’ points share a range?
  • Consistency — do similar phases (two baselines, or two intervention phases in a withdrawal design) look like each other? Data that repeats its own pattern across the design is more convincing than a single instance of anything (What Works Clearinghouse standards).

The trap is overlap in reverse: a striking rise after a baseline that was already rising says much less than the same rise after a flat one. That is why the baseline gets read first.

What the chart cannot settle

Two questions sit outside it, and a design without them is a design that can be wrong quietly.

Would a second person have scored it the same? If only one observer ever scored the behavior, the record describes one person’s judgment as much as the child’s behavior. The answer is a second observer scoring some portion independently, and a figure for how much they agree.

Did the procedure actually run as written? A flat line has two explanations that look identical on a graph: the procedure did not help, or the procedure was not what happened. Checking that is a separate measurement, and without it a study of the learner may be a study of the adults.

Both of these have well-established methods and neither is optional in published work. Interlaza records the phase markers, the agreement between two observers and the fidelity checks alongside the sessions themselves — but the design is still a decision a person makes, and the app will not make it.

Further reading