Teaching without errors

Errorless teaching and the ABA prompt hierarchy: fading and transferring stimulus control

The ABA prompt hierarchy in practice: why you start with the most intrusive prompt, and how to fade it without creating prompt dependency.

For: Instructors

By INTERLAZA 12 min read Updated

Ask a practitioner what errorless teaching is and most will say “you don’t let the learner make mistakes.” True, but it hides the part that actually matters. Errorless teaching is not the absence of errors — it is a deliberate arrangement of stimulus control, in which help is present from the first trial and then withdrawn on a schedule, so that responding transfers from the prompt to the thing that is supposed to control it. The ordered set of helps you fade through is the prompt hierarchy, and which end of it you start from decides whether the procedure is errorless at all.

Get the arrangement right and acquisition is fast and durable. Get the withdrawal wrong and you produce a learner who performs beautifully with you and not at all without you. This article is about the difference.

What it’s also called

The literature is not consistent, which makes searching for it harder than it should be. You will encounter the same family of procedures under:

  • Errorless learning and errorless teaching — used interchangeably
  • Errorless discrimination training — the older experimental term, from Terrace’s work with pigeons in the early 1960s
  • Stimulus fading and stimulus shaping — two distinct techniques within the family (extra-stimulus and within-stimulus prompting, in Etzel & LeBlanc’s 1979 terms), often collapsed together in practice
  • Prompt fading — strictly the withdrawal step, not the whole procedure
  • Transfer of stimulus control — the outcome the procedure exists to produce

One term that does not belong here: no-no-prompt. That is an error-correction procedure. It permits two errors before prompting, which is the opposite arrangement.

Why errors are worth engineering out

The standard justification is emotional — errors are frustrating, and a frustrated learner disengages. That’s true and it matters, particularly with young children.

But there’s a stronger behavioral argument. An error is not neutral information; it is a practiced response. A learner who responds incorrectly three times has not had three opportunities to learn the right answer, and repeating the error is not by itself proof that anything reinforced it — practice alone does not strengthen a response without a reinforcing consequence behind it. But where the incorrect response has any history of reinforcement, or where the array is small enough that guessing pays off intermittently, those repeated errors do become three opportunities to strengthen the wrong one, and error patterns consolidate quickly and are then expensive to undo.

For learners with restricted attention or a history of failure, both problems compound: errors are practiced and they occasion escape.

The prompt hierarchy: where to begin, and why most-to-least

Choosing the prompt is where practitioners most often go wrong, and it is the question exam candidates ask most:

When you begin teaching a response errorlessly, you start with the prompt most likely to produce a correct response — and “most effective” is not automatically the same thing as “most intrusive,” even though the two often line up in practice. A highly reinforcing gesture can be as reliably effective as full physical guidance for a particular child, and less intrusive besides; the hierarchy is organized by effectiveness, intrusiveness is just the usual proxy for it.

That is a most-to-least prompt hierarchy. Begin with whatever guarantees success — a full physical prompt, a model, or in a match-to-sample array, a comparison field where the correct choice is the only viable one — then systematically reduce intrusiveness as responding stabilises. It is the most common route into errorless teaching, but not the only one: an arrangement built entirely from stimulus prompts, with no response prompt at any point, is errorless too, provided the learner never actually gets the chance to answer wrong.

Least-to-most does the reverse: minimum help first, escalating after the learner responds incorrectly or does not respond at all. It has its uses, particularly in assessment and with learners who already have partial skills. But it is not errorless, by construction: it produces an error (or a failure to respond) before it delivers help. Choosing least-to-most and calling the program errorless is the most common confusion between these two procedures.

A third option worth knowing: progressive time delay, where the prompt stays maximally effective but is delivered progressively later after the instruction, giving the learner an expanding window to respond independently. It’s often the gentlest route to independence when intrusiveness is hard to grade.

Response prompts versus stimulus prompts

Two categories, and they fade differently.

Response prompts act on the learner’s behavior — physical guidance (full or partial), a model, a gesture, a verbal cue. They’re direct and effective, and often the only option when the target is a physical action rather than a discrimination among materials that can themselves be manipulated. They carry the higher risk of dependency because the learner can come to wait for them.

Stimulus prompts act on the materials — altering the correct choice or its competitors so the discrimination is temporarily easier. In a match-to-sample array this includes:

  • Intensity or opacity — the incorrect comparisons start faint and strengthen as the learner improves
  • Size — the correct comparison starts larger and returns to parity
  • Highlighting — a border or marker on the correct comparison that fades out
  • Positional prompting — the correct comparison occupies a fixed, predictable location before being randomised
  • A directional cue — an arrow or pointer that reduces in salience

For teaching a discrimination specifically, stimulus prompts have an advantage worth knowing about: they operate on the same materials the learner will eventually face unaided, which shortens the distance the control has to travel — though that is a strength for discrimination targets, not a general claim that stimulus prompts outperform response prompts everywhere; the two solve different problems more often than they compete for the same one. The classic distinction here is stimulus fading (gradually changing an added dimension — a color cue, a size difference — until it disappears) versus stimulus shaping (gradually morphing the form of the stimulus itself into the target). Fading is simpler to program; shaping tends to survive better when the discrimination is genuinely difficult.

Fading: the actual mechanics

Here is the opacity form of the arrangement across the default three phases, with the two criteria that govern movement drawn under it:

Figure — The same trial, three phases of help

Sample

Phase 1 · 10 %

Failing is unlikely by design. It is still a choice among three options, not a pairing with no response required — the correct one is simply the only visually obvious answer, so the sample and the correct comparison sit side by side as far as the child can tell.

Sample

Phase 2 · 62 %

The child keeps succeeding, so the distractors keep gaining presence. Withdrawal is gradual on purpose: no phase is a jump.

Sample

Phase 3 · 100 %

All options at full contrast. The child answers with no help — this trial is unprompted, whatever came before it in the session.

Each success moves help one step down.

One error moves it one step back — immediately.

The 10% floor in phase 1 is deliberate and never rises: with the distractors barely visible, the correct answer is the one visually obvious option, and that arrangement starts teaching before the child has had to choose under any real uncertainty.

The three opacities are computed with the same fading formula the app runs (floor 10%, gentle curve, final phase always 100%), for its default of three phases per concept.

Errorless fading across the default three phases: distractors start at 10% opacity and reach full contrast only after the child has been succeeding. The rule under the panels is what separates the technique from simply making things easy — help returns the moment an error appears.

Fading is where programs quietly fail, because the criteria are usually left implicit. Make them explicit:

Advance criterion. How many consecutive correct responses at the current level before help reduces? A single correct response is often enough with a well-graded hierarchy and a young learner — it keeps the schedule moving and helps stop the prompt becoming a fixture, though it is not a guarantee on its own: a learner who keeps regressing on errors can still spend a whole session near the top of the hierarchy whatever the advance criterion says. More conservative programs use two or three.

Regress criterion. What happens on an error? Returning to the previous level immediately, on the first error, is the standard and defensible choice. Delaying the regression allows the error to repeat, which is precisely what the procedure exists to prevent. It is also what stands between the previous two criteria and their intended effect: a well-tuned advance criterion and level count still cannot promise the learner reaches full independence if enough regressions eat into the trials available.

Number of levels. Enough to make each step small, few enough that the learner has a realistic chance of reaching full difficulty within the session. A hierarchy with more levels than the learner has trials per target is a hierarchy the learner will never finish — they simply never experience the unprompted condition. This is a surprisingly common and invisible failure, and even a well-sized hierarchy is not a guarantee by itself: it makes reaching full difficulty reachable, not certain, since regressions can still use up the budget first.

Per-target tracking. Fading level belongs to the target, not the session. A learner may be at full independence on one item and maximum support on another; collapsing them to a single session-wide level guarantees that one of the two is wrong.

The point is transfer, not the absence of errors

The purpose is that the natural discriminative stimulus ends up controlling the response — the spoken word, the printed word, the object — and not the arrow, the instructor’s hand, or the position on the screen.

The characteristic failure is prompt dependency: near-perfect responding whenever help is present, collapsing the moment it is withdrawn. What’s happened is that control transferred to the prompt and stayed there. Related, and easier to miss, is when the learner discriminates the prompt rather than the target — responding to “the big one” or “the bright one” rather than to what the item actually is. Performance looks excellent while nothing intended has been learned.

Two safeguards, and they are not the same check:

  • Check unprompted responding as a routine part of teaching, not only at the end. An independent correct response, scored and fed back to like any other, is what tells you whether help is still needed at all. If independence isn’t emerging by mid-program, the hierarchy is too slow or the steps are too large.
  • Run no-feedback probes on untrained material separately. That is a different question — not “can they do it unprompted” but “does the response hold up away from this exact material” — and answering it needs feedback withheld during the probe itself, which an ordinary independence check does not require.
  • Watch latency alongside other signs, not as a standalone diagnostic. A learner who is correct but slow is not conclusively still under prompt control on latency alone — slowness has other causes too. It is the combination that is telling: correct, slow, and hesitating or glancing toward the instructor.

Common mistakes

Calling least-to-most errorless. Covered above; it’s the most frequent one.

Fading on the clock rather than on the data. Reducing help because it’s session three, rather than because the advance criterion was met.

Removing several dimensions at once. If size, position and highlighting all fade on the same trial, a failure tells you nothing about which was carrying the control.

Never reaching full difficulty. If the learner always finishes the session with some help still present, they have never practiced the target condition.

Treating an error as a scoring event. In an errorless arrangement an error is a signal that the hierarchy is wrong — too few levels, steps too large, advance criterion too loose. It’s diagnostic information about the program, not about the learner.

In practice

Running this properly means tracking a separate fading level per target, applying advance and regress criteria consistently trial by trial, capping the hierarchy so full difficulty is actually reached, and watching latency alongside accuracy — while teaching a child whose tolerance for all of this is finite.

This is the part Interlaza automates. Prompt type is selectable across the five stimulus-prompt forms described above; fading level is tracked per concept rather than per session; advance and regress criteria are explicit and configurable; and the hierarchy is sized against the trials available per target so the unprompted condition is reachable within a normal session — a learner who keeps regressing on errors can still use up that budget first, the same limit any hierarchy has. An adaptive mode adjusts the advance criterion using the learner’s estimated mastery, recent errors and response latency, which is the same judgment an experienced instructor makes, applied consistently on every trial.

Disclosure: Interlaza is our product. The procedures described above are standard in the literature and do not depend on any particular software.

Further reading

  • Terrace, H. S. (1963). Discrimination learning with and without “errors.” Journal of the Experimental Analysis of Behavior, 6(1), 1–27.
  • Touchette, P. E., & Howard, J. S. (1984). Errorless learning: Reinforcement contingencies and stimulus control transfer in delayed prompting. Journal of Applied Behavior Analysis, 17(2), 175–188.
  • Etzel, B. C., & LeBlanc, J. M. (1979). The simplest treatment alternative: The law of parsimony applied to choosing appropriate instructional control and errorless-learning procedures for the difficult-to-teach child. Journal of Autism and Developmental Disorders, 9(4), 361–382.

Related: our introduction to match-to-sample, and stimulus equivalence for what these discriminations build toward.