Teaching without errors

When a child stalls, the first question is not about the child

Four flat sessions in a row is a decision point, not a diagnosis. A teaching system has a rule for when to decide and an order for where to look.

For: Instructors

By INTERLAZA 11 min read

There is a particular kind of flat line every instructor recognizes. Three weeks on the same handful of concepts. Accuracy hovering around 60%. The child is cooperative, the sessions run, the data gets recorded, and nothing moves.

The real question at that moment is not why can’t this child learn this? It is: how long have I been running the same thing, and what exactly am I going to change?

That question has had a rigorous answer since the 1990s, and it comes from a place most of us never look: a school system. CABAS — Comprehensive Application of Behavior Analysis to Schooling, developed by R. Douglas Greer and colleagues at Columbia University Teachers College — is an attempt to make teaching a measurable practice rather than a craft. Its two books, Designing Teaching Strategies (2002) and Verbal Behavior Analysis (2008, with Denise Ross), spend most of their pages on something almost nobody measures: the decisions of the adult.

The rule for when to decide

Greer’s decision protocol starts somewhere unglamorous: how many data points before you are allowed to conclude anything?

The answer is to count segments, not points. The first session is an origin — it has nothing to compare against. The second gives you one segment, the third gives you two, the fourth gives you three. At three segments a trend exists and you can read it. If the direction wobbles, extend to five and read again.

Then a rule with no wiggle room in it:

  • Ascending — keep going. The line does not prove which change is responsible, but there is no decision to make while it climbs.
  • Flat or descending — a decision is due. Now. (That’s for a measure where more means more learned. For one where less is better — errors, latency — read the direction the other way.)

And then the sentence that makes the whole system different from every “insights panel” ever shipped: every instructional session run without the change that was due is counted as another error. Not the child’s error. The teacher’s. The reasoning is blunt — a session spent under teaching that isn’t working is educational time gone, and it may be actively compounding the difficulty.

Greer also counts the correction: a decision finally made three sessions late does not go in the ledger as a good decision. It goes in as the correction of an error.

If you find that harsh, notice what it does. It converts “I should probably change something soon”, a feeling, into a number that someone can look at. Keohane and Greer (2005) taught the protocol to three instructors and tracked six children over eighteen months; the children needed measurably fewer teaching trials to reach the same objectives once the adults were using it. The intervention was on the adults. The effect was measured on the children.

Once a decision is due, the second half of the protocol says where to look, and the order is the point, not the list. That order decides which explanation to test first for a flat line — it is not a rule about when an adult may act on pain, illness or distress. That gets checked the moment it’s noticed, at any rung.

Figure — Where to look, in order

Four flat or falling sessions in a row

  1. Was the teaching itself intact?

    What you check:
    Every trial had a clear sample, a real chance to answer, and a consequence that matched the answer — reinforcement or a correction the child actually looked at.
    If the answer is no:
    Fix the presentation. Nothing below this rung can be read until this one is clean.
  2. What was happening around the session?

    What you check:
    Sleep, illness, hunger, a reward the child has had enough of, the ten minutes before the tablet came out.
    If the answer is no:
    Change the conditions, not the program. Same route, different moment.
  3. Are the prerequisites really there?

    What you check:
    Not "was it taught" but "is it there now" — mastered, recently, and in conditions like today's.
    If the answer is no:
    Insert the missing step before the current one and come back.
  4. Is there something physical in the way?

    What you check:
    Hearing, vision, motor access to the screen.
    If the answer is no:
    Adapt the materials, and get the professional opinion the app cannot give.

The question that is never on the ladder

“Can this child learn?” It has no answer that changes what anyone does tomorrow. The question that replaces it — “what does this child need from the way we are teaching?” — has four, and they are above.

This is a guide for reading a flat line — which explanation to test first, not a diagnostic exam the child has to pass. Each rung is only readable once the one above it is clean: a prerequisite gap and a session run with unclear presentations produce the same flat line, and checking the child first is how the flat line gets blamed on the child. Wellbeing is never gated by this order — pain, illness or distress get acted on the moment they're noticed, at any rung. Adapted from the decision protocol in Greer (2002).

The first rung is the one that gets skipped. Before anything about the child is analyzed, Greer requires that you establish the teaching itself was intact. In his vocabulary a learn unit is one complete interlock: the adult secures attention, presents an unambiguous sample, leaves a real window to answer (about three seconds) and delivers a consequence that matches what the child did. If any piece is missing, the interaction happened but the learn unit did not, and there is nothing to diagnose yet.

That sounds like bookkeeping. It isn’t. Across five studies, replacing adult–child interactions that were not learn units with interactions that were raised correct responding by a factor of four to seven (Greer & McDonough, 1999). Read that for what it is: a comparison of classroom instruction with and without a complete teaching interaction, not a promise about any particular app or child. What it establishes is narrower and more useful — the difference between teaching and almost-teaching is large enough that you cannot skip past it on the way to conclusions about a learner.

One component of that interlock is easy to get wrong and worth naming on its own. When a child answers incorrectly, the correction only counts if the child looks at the sample again while giving the right answer. Hogin (1996) tested exactly this: children who saw only their own answer and the consequence did not master the operation; children who also saw the stimulus did. The correction is not there to announce the verdict. It is there to put the right answer next to the thing it belongs to.

The trap in “he knows it”

There is a second idea in these books that lands harder on a matching app than on a classroom, and leaving it out would be dodging it.

Both volumes repeat a warning: pointing is not naming. A child who reliably selects the red card when you say “red” has a listener repertoire. A child who says “red” when shown the card has a speaker repertoire. These are functionally independent — they do not come as a pair, and assuming they do means a child gets marked as knowing a concept and never receives the instruction they were actually missing. Selecting correctly doesn’t rule out that naming will emerge on its own later; it just can’t be assumed alongside it. Greer states it flatly: identifying an item among choices is not the same as naming it.

That is a real limit on what any selection-based screen can tell you, ours included. A tablet is very good at the listener half. The speaker half happens between a child and an adult in a room, which is why Interlaza’s off-tablet modes exist: a route that stays on the screen leaves out half the work.

The same books also give the flip side, and it is the more surprising half: matching to sample is not just a vocabulary exercise. In Greer’s developmental sequence, matching is the procedure used to induce “the capacity for sameness”, the pre-listener milestone that discrimination itself is built on, and visual matching is the step that sets up naming. That is Greer’s own developmental sequence, not a claim that matching is the only route in — a child can also build a listener or speaker repertoire through gesture, AAC or plain exposure with no formal matching task at all. A child who cannot yet match is not a child failing an easy task. They are a child working on the thing that comes before the task.

What Interlaza does with this, and what it cannot

Interlaza records what this method runs on: every trial, what was presented, what was chosen, how long it took, how much help was active, and whether the session followed the procedure it was configured with. Three things are built on top of that record, one per idea above.

Rung 1 is a gate, not a panel. The Program Advisor withholds its restructuring recommendations only when there is enough evidence — at least 20 checked trials — that recent sessions did not run the way they were configured. Below that floor, or with no evidence at all, it fails open and still proposes: a gate that shuts for lack of data would block the advisor exactly when nobody has information. A closed gate here always means “we looked, and the record deviated” — never “we don’t know yet.”

What the teaching cost is now a number. Alongside accuracy, the Results tab reports the median trials it took to reach mastery, per concept. “Still getting them right” cannot tell you whether a change of tactic made learning cheaper; this can.

And the count Greer puts at the center now exists. When a route’s trend has gone flat or downward and nothing in the record has changed since, the app says how many sessions that has been going on for. It is about the adult, not the child, which is why it stays silent whenever the child is climbing. A number that only ever meant “sessions since you last touched this” would read as a reproach aimed at a child who is doing fine.

Now the realistic half, and it is not a to-do item. That count means “sessions since anything we can see changed”, and the app cannot see everything. An instructor who changes how they present the sample, moves the session to a better hour, or swaps a reinforcer by hand has changed the program in the way that matters, and left no trace the software can read. Neither can it attribute an edit to a route several children share. So the number is a prompt to look, never a verdict about whether anyone was paying attention, and it says as much on screen.

The ladder above works on paper regardless. The flat line was going to be there either way.


References

Greer, R. D. (2002). Designing Teaching Strategies: An Applied Behavior Analysis Systems Approach. San Diego: Academic Press. — The learn unit with its full experimental base, and the decision protocol: count segments, not points; flat or falling means decide now.

Greer, R. D., & Ross, D. E. (2008). Verbal Behavior Analysis: Inducing and Expanding New Verbal Capabilities in Children with Language Delays. Boston: Pearson. — Verbal capabilities as inducible cusps, the warning that pointing is not naming, and the half that surprised us: matching to sample is the procedure used to induce the capacity for sameness.

Keohane, D. D., & Greer, R. D. (2005). Teachers’ use of a verbally governed algorithm and student learning. International Journal of Behavioral Consultation and Therapy, 1(3), 252–271. — cited here as it appears in Greer & Ross (2008); a 2002 book cannot itself cite a 2005 article, so we no longer attribute it to Greer (2002). We have not read the article directly. — The experiment behind the protocol: three instructors, six children, eighteen months — and measurably fewer teaching trials to the same objectives once the adults used it.

Greer, R. D., & McDonough, S. H. (1999). Is the learn unit a fundamental measure of pedagogy? The Behavior Analyst, 22(1), 5–16. — The five-study review behind the four-to-seven factor: interactions that were complete learn units versus ones that weren’t.

Hogin, S. (1996). Doctoral dissertation, Columbia University. — cited here through Greer (2002, pp. 26–27), not read directly. — Children who saw only their own answer and the consequence did not master the operation; those who also saw the stimulus did. It is why a correction only counts if the child looks at the sample again.