An instructor finishes a session and the screen says 92% correct. Good week, apparently. But the child was working with the correct card faded in beside the sample for most of those trials, three of the concepts were ones they had already mastered in March, and the two that were actually being taught scored 4 out of 12 between them. The 92% is arithmetic. It is not a fact about the child.
That gap — between a number a tool can produce and a number you can defend to a supervisor, a family, or an ethics committee — is what this checklist is about. It is written for instructors evaluating any tool of this kind, including ours.
1. Does it separate independent responses from prompted ones?
Ask to see a report that shows both. A prompted correct response tells you the prompt worked; only an unprompted correct response tells you the child can do it. A tool that pools them into one accuracy figure produces its highest numbers exactly when the child is being helped the most, which is the opposite of what you need.
What a good answer looks like: the report shows independent trials as their own count, and the mastery decision is made from those alone.
2. Does it check what it taught, or only what it practiced?
A probe is a small set of trials run with no help and no feedback, on material the child has not practiced — that combination is what makes it a check rather than more teaching. Practicing and checking answer different questions. Almost every tool reports practice. Ask specifically whether it can run trials with the feedback switched off and report them separately.
What a good answer looks like: the tool can present untrained material, withholds feedback on it by design, and never lets a probe count toward the mastery streak.
3. Can it say what the criterion was, before the session ran?
“Mastered” means nothing on its own. Three correct in a row is a different claim from 80% over the last ten trials, which is a different claim again from a probability estimate crossing a threshold. Ask where the criterion is written down, whether it was set in advance, and whether the report prints it next to the verdict.
What a good answer looks like: every “mastered” label is shown together with the rule that produced it.
4. Could a second person score the same session the same way?
This is interobserver agreement, and it is the standard evidence that a measure describes the child rather than the observer. Ledford and Gast (2018) treat agreement of 90% or higher — the share of trials where two independent observers recorded the same outcome — as preferred, and below 80% as unacceptable. Ask whether the tool can produce that number at all.
What a good answer looks like: a second observer can score a session without seeing how the first one scored it, and the report gives agreement per session and per concept, not one averaged figure.
5. Does it record whether the procedure ran as designed?
Greer (2002) makes the case for checking the teaching before analyzing the learner: a missing prerequisite and a session run with unclear presentations produce the same flat line on a graph. Procedural fidelity is the record of what the session actually did — prompts delivered, feedback given, intervals honored — and most tools have the data and never report it.
What a good answer looks like: the tool reads fidelity off the session it already recorded, names what it could not check, and does not fold the unmeasurable parts into the percentage.
6. What happens when there is no internet?
A therapy room in a school basement, a home visit, a center with one shared connection. Ask what the app does when the connection drops mid-session: whether the trials already run are kept on the device, whether the session can be finished, and whether anything is lost when it reconnects.
What a good answer looks like: the session runs to the end offline and the data is on the device before any server is involved.
7. Where does the child’s data live, and can you get it out?
You are handling health data about a minor. Ask three separate things: what leaves the device and when, whether cloud storage is something you switch on rather than something already on, and what an export actually contains. An export you cannot open in a spreadsheet or hand to a committee is not an export.
What a good answer looks like: the storage answer is specific, the cloud copy is opt-in, and the export is a documented file format you can read without the vendor.
The pattern in all seven
| A vague answer sounds like | A checkable answer sounds like |
|---|---|
| ”Our reports are comprehensive" | "Here is the field, here is where it appears in the report" |
| "The AI adapts to the child" | "Here is the rule, here is what it changes, here is where the change is logged" |
| "It’s fully compliant" | "Here is what is stored, here is what is transmitted, here is who controls it" |
| "Data is secure" | "Here is the mechanism, and here is what it does and does not guarantee” |
The questions are not a scoring rubric and there is no passing total. What they do is move a demo from what the interface looks like to what its numbers commit to — and a vendor who can answer six of the seven precisely, and says plainly that the seventh is not built yet, is telling you more than one who answers all seven smoothly.