ABA data collection methods: which measure fits which behavior

By INTERLAZA 4 min read

Applied behavior analysis stands on one habit: decisions follow data. Which makes the unglamorous question — how exactly is this behavior being recorded? — the most consequential one in a program. This is a field guide to the main methods, the rule for choosing between them, and the two checks that decide whether the numbers deserve to drive anything.

Figure — Which measure fits which behavior

The behavior looks like The adult presents something; the child responds once. Matching, receptive labels, imitation
Measure Accuracy per trial Correct / incorrect / prompted, with the prompt level attached. How it lies: A session percentage hides that most correct responses were prompted, or that every error sat on one target.
The behavior looks like The child emits it freely; you can count it. Requests, vocalizations, a target sound
Measure Frequency → rate Occurrences divided by a recorded observation window. How it lies: A count without its window cannot be read at all, and a window too short to matter still produces a per-minute number.
The behavior looks like It lasts, and the length is the point. Engagement, tantrums, time on task
Measure Duration Total duration and duration per occurrence, reported separately. How it lies: One twenty-minute episode and twenty one-minute episodes share a total and need different programs.
The behavior looks like The response comes, but you are watching the wait. Fluency goals, emerging prompt dependency
Measure Latency Instruction to response, summarized by the median. How it lies: A fast error is still an error — speed is evidence alongside accuracy, never instead of it.

Whatever the row, two questions decide whether its numbers can be trusted: would a second observer have recorded the same thing, and did the session run as the program prescribes?

Four shapes of behavior, the measure each one calls for, and the specific way that measure misleads when it is read carelessly. Choosing the row is the professional judgment the software cannot make; recording it faithfully once chosen is the part it can.

Trial-by-trial accuracy

For discrete, teacher-presented tasks — matching, receptive labels, imitation — the natural record is one scored response per trial: correct, incorrect, prompted, at what prompt level. From it come percent correct, prompted-versus-independent breakdowns, and error patterns.

Its failure mode is aggregation: a session’s “85%” can hide that most correct responses were prompted, or that all the errors sat on one target. Keep the prompt level attached to every response, and read per-target numbers before session-wide ones.

Frequency and rate

For behavior that the child emits freely — requests, vocalizations, a target sound — counting occurrences is the measure, and dividing by observation time turns a count into a rate you can compare across days. Two disciplines keep it trustworthy: record the observation window (a count without its window is uninterpretable), and refuse rates from windows too short to mean anything — thirty seconds of observation does not license a per-minute claim.

A time line carrying seven marks, one for each occurrence, with a bracket underneath reading '10 minutes observed' and, beside it, the result: 7 responses in 10 minutes.
A count on its own is half a measure. The same seven marks spread over a whole morning describe a different child, which is why the observation window belongs inside the record rather than in the notes around it.
Two staircases climbing from the same origin against axes of time and responses so far: one steep, labeled high rate, and one almost flat, labeled low rate.
Dividing by the window turns a count into a rate you can compare across days: the steeper the climb, the higher the rate. It also explains the trap — a window far too short to mean anything still produces a per-minute number.

Duration

For behavior that lasts — engagement, tantrums, time on task — count and rate miss the point; the clock is the measure. The literature separates total duration from duration per occurrence, and both matter: one twenty-minute episode and twenty one-minute episodes have the same total and call for different programs. Decide in advance what happens to an episode still open when the session ends — closing it at the boundary, rather than discarding it, keeps the record accurate.

A time line carrying three blocks of clearly different lengths, marked 1.5, 3.8 and 0.7 minutes, with a summary reading total 6 min and 2 min per episode.
Total and per occurrence are reported separately on purpose. Both numbers come from the same three blocks, and only together do they say whether you are looking at one long episode or several short ones.

Latency

Time from instruction to response. It is the early-warning channel accuracy cannot provide: responding that stays correct while latency stretches usually means control is shifting to something other than the material — often to the adult. It also anchors fluency goals: “correct within three seconds” is a different achievement from “correct eventually”.

A time line with two posts on it — the first labeled instruction, the second response starts — and an arrow of 2.6 seconds measuring the gap between them.
Latency measures the wait, not the answer. The rest of the line is deliberately empty: what is being timed is everything that happens before the response, which is exactly what an accuracy score cannot see.

Permanent product

Sometimes the behavior leaves an artifact — a completed puzzle, a written word — and the artifact can be scored after the fact. Cheapest of all methods, applicable only where the product genuinely proves the behavior.

The two checks behind every method

Whatever the method, two questions decide whether its numbers can be trusted. Would a second observer, scoring independently, have produced the same record? That is interobserver agreement (IOA), and it is measured, not assumed. And did the session actually run as the program prescribes? That is procedural fidelity. A beautiful graph fails both silently; nothing in the plot tells you the record behind it was reproducible. We wrote a full piece on both — see below.

The procedural fidelity card in the app: a percentage over the checks that ran as configured, and below it a list of items headed 'not checked', each with the reason the trial log cannot answer it
Procedural fidelity as the app reports it, read off the trial log with no extra data entry. The part that matters is the lower half: what could NOT be checked is named with its reason and kept out of the percentage — a fidelity figure that quietly drops what it could not measure is worse than none.

What Interlaza records, and what it deliberately leaves to you

Because the app presents the trials, it records what happens in them as a side effect of running the session: every response with its prompt level and latency, trial by trial, with no clipboard involved. Its free-operant modes record counts over a measured window (refusing to report a rate under a minimum of observation), and its duration mode reports total and per-occurrence separately, closing still-open episodes at the boundary. For a second observer, it can stream a session live so the observer scores independently — without seeing the primary record first — and computes the agreement.

What it does not do is decide what to measure. That choice — the pairing of behavior and measure this article is about — stays with the professional, where it belongs.

Further reading

  • Ledford, J. R., & Gast, D. L. (2018). Single Case Research Methodology (3rd ed.). Routledge — chapters 5–8 on measurement and dependent variables.
  • Cooper, J. O., Heron, T. E., & Heward, W. L. (2020). Applied Behavior Analysis (3rd ed.). Pearson.