A child finishes the week at 95%. Good news, but 95% of what?
Here are two children who both hit that number this month. They are composites written to make the point — not two learners from our records, and not data.
Marta matched twelve photographs to their identical twins, over and over, and got nearly all of them right. Show her a different photograph of the same dog and she stalls.
Diego hit the same 95% on the same kind of screen. Show him a different photograph of the same dog and he is still right. Show him a drawing of it, and he is still right.
Marta and Diego produced the same number. They did not do the same thing. And nothing in a percentage — not the total, not the trend line, not the streak — will tell you which of the two you have.
This matters because it changes what you do on Monday. Marta needs new exemplars before anything else; more repetitions of the same twelve will raise her score and teach her nothing. Diego is ready to move on. Read only the percentage and you would give them both the same next session.
Figure — The same score, twice
One performance collapses, the other barely moves. The score never said which was which; the probe is evidence toward an answer.
The distinction, and where it comes from
The Mexican behavioral psychologist Emilio Ribes spent a career on precisely this problem: that the same visible performance can be produced by qualitatively different kinds of interaction. In Teoría de la Conducta he sets out five, each a different way an individual and their environment can be organized into a single episode.
The two that concern ordinary practice are the first and the third.
Coupling — the interaction is organized around something that is simply there. Recognizing, matching, repeating. What holds it together is present in front of the child: this photograph and that photograph share their form. Nothing swaps roles between trials; the dog picture is always right for the dog picture. Coupling can still be dressed up as a conditional task — “match this way when the border is blue, that way when it’s red” — without becoming comparison; what keeps it coupling is that the rule itself, once set, does not swap which item is correct for which sample.
Comparison — no element carries its function on its own. What matters is a relation — bigger than, the same as, the opposite of — and the relation itself is what must be responded to. The same physical comparison item can be the right answer against one sample and the wrong answer against another, because its function comes from the relation to whatever sample is on screen, not from anything fixed about the item itself. Invert the terms and the correct answer inverts with them.
Marta’s performance points toward coupling. Diego’s, on the evidence, points toward something more — though neither is a verdict about the child, only about what this one arrangement showed on this occasion; the app itself never goes further than that, for the reason the next sections explain.
The other three — alteration, extension, transformation — matter enormously in the theory and rarely in a tablet session, so they are at the end of this article rather than the middle.
The uncomfortable part
Here is the finding that makes this more than vocabulary: you cannot tell which contact you are looking at from what the task looks like.
A word-to-picture task feels more advanced than photo-to-photo matching. Structurally it is not: nothing permutes, the word is always right for its picture, the contingencies stay constant. It is coupling — a different kind of coupling, held together by convention rather than by resemblance, but coupling.
And a task that genuinely does point higher can still be solved by coupling. Carpio places matching-to-sample at the level of relating; Ribes warns that those tasks are usually solved by recognizing. Formally they point up. Functionally they can stay down. A correct answer does not distinguish the two, which is exactly why a score cannot.
So: not from the task, and not from the percentage. Then from what?
From what happens when you take the support away
There is one arrangement that gives real evidence, and it is not clever — it is just strict.
Remove the feedback. Change the material. See what survives.
If the performance was pure recognition, it tends to collapse: the child was responding to these pictures, and these pictures are gone. If the child was responding to the relation, the new material tends to make little difference, because the relation is still there. Neither is a guarantee — recognition can generalize too, along whatever physical similarity the new material happens to share with the old, so a performance that mostly holds up is not on its own proof of relating; and a single probe, however clean, is one piece of evidence rather than a definitive test on its own.
That is not the only possible test, but it is a real and useful one. It is why Interlaza runs no-feedback probes with material the child has not been trained on, and why those probes are reported separately from everything else.
Reading it in your own results
On a child’s Results tab you will now find two statements, and they are deliberately different in kind.
What kind of practice this was. Something like:
Recognizing the same thing again — 94% · 146 trials Responding to a relation — 6% · 9 trials
This describes the tasks you set, not the child. It is knowable, because it is a property of the exercise. And it is often the more useful of the two, because it makes visible something that is otherwise invisible: two months of practice that was almost entirely recognition, without anyone deciding that.
What the no-feedback checks say. Something like:
Practiced material — 71% New material, no help — 33%
The score comes from recognizing material already seen. With new material and no help it drops sharply, which is the pattern recognition rather than relating produces.
That is the sentence a percentage cannot produce. And when the probes hold up instead, it says the opposite — that something transferred, and the child is responding to the relation rather than to the particular items.
Three things it deliberately refuses to do:
- It never gives the child a level. The same task is a different contact depending on what the child already knows, so a label read off the exercise type would be wrong a good share of the time — while carrying the app’s authority. What is classified is the task.
- It stays silent when it cannot know. Too few checks, or trained material that is not solid yet, and there is no verdict — because there is no established performance for a check to confirm or contradict.
- It always shows both figures. Where the line falls between 71% and 33% is a judgment, not a standard. You should be able to disagree with it.
What to do with it
If practice is 90%-plus recognition and the probes collapse, the priority is more exemplars, not more trials — different photographs, different voices, the drawing as well as the photo — and probes to check. That is where to start, not necessarily the whole answer; a child who still doesn’t generalize once exemplar variety is genuinely broad may need a different kind of task entirely, not just more of this one.
If the probes hold, the recognition work has done its job and the child is ready for tasks where the relation is what varies.
And if the probes say nothing yet, that is a real answer too: run some.
Two articles carry on from here. Repetition is not relation sets out the four decisions Interlaza uses to make a correct answer reachable only by comparing, and what to expect from a probe. Two simple tasks, two different worlds looks inside coupling itself: matching a photograph and touching a dog on hearing the word are both recognition, and they still do not ask the same thing.
The deeper layer, for the curious
Everything above is the part you can act on. Underneath it there is a research literature that goes considerably further, and it is worth knowing it exists.
Ribes’ five contacts each have their own criterion of adjustment — differentiality, effectiveness, precision, congruence, coherence — and the first three have published indices with formulas. Worth being precise about where those formulas come from: they were worked out (Serrano, 2009) under the vocabulary of Ribes’ earlier model — the functions of the 1985 taxonomy (Ribes & López) — rather than directly under the later five-contact naming (coupling, comparison, and so on, from Ribes 2018/2021) this article otherwise uses. The two vocabularies describe the same line of theory at different points in its development, and the correspondence between them is not something this article is asserting as settled beyond what those sources themselves establish. What the indices are not, whichever vocabulary names them: accuracy in disguise. They count omissions as failures to adjust, two of them measure time rather than trials, and the precision index is a product across two conditions, so succeeding under one rule and failing the other collapses it toward zero.
One property worth stating because it trips people up: that precision index cannot exceed 0.25. Both its factors share a denominator, so a flawless run reads 0.25 — while the interpretation bands published alongside it run from 0 to 1 and would call that “incipient”. The bands are described by their own author as arbitrary. That is a good reminder that a number from a paper still needs reading, not just reporting.
Two of the five contacts also have no way to be arranged in this app, as it is built today: extension is a contact between two people, not between a person and an object, which a screen cannot stand in for, and transformation — which reaches well beyond simply talking about how one talks, into language turned reflexively on itself — is far outside the age range this product serves in any case. Naming them and leaving them undone is more realistic than a version that pretends.
None of that is needed to use what is on your Results tab. It is there because the distinction on that tab is not something we invented; it comes from somewhere, and the somewhere is checkable.
References
Andrade-González, D. E., León, A., & Hernández Eslava, V. (2020). Tarea de transposición y contactos funcionales de comparación: una revisión metodológica y empírica. Acta Comportamentalia, 28(4), 539–565. — The six criteria a task must meet to measure the comparison contact — and with them the reason color will not do: red is not “more” than blue.
Carpio, C. A. (1994). Comportamiento animal y teoría de la conducta. In L. J. Hayes, E. Ribes & F. López (Eds.), Psicología interconductual: contribuciones en honor a J. R. Kantor (pp. 45–68). Guadalajara, Mexico: Universidad de Guadalajara. — The alternative naming of the five criteria — ajustividad, efectividad, pertinencia, congruencia, coherencia — and the note placing matching-to-sample at the pertinencia level.
Ribes Iñesta, E. (2018). El estudio científico de la conducta individual: una introducción a la teoría de la psicología. Mexico: Editorial El Manual Moderno. — Chapter 10, Tabla 10-1: the five functional contacts. This is the table the distinction in this article comes from.
Ribes Iñesta, E. (2021). Teoría de la psicología: corolarios. Granada, Spain: Co-presencias Editorial. — The naming we use here and in the app: differentiality and precision where Carpio says ajustividad and pertinencia.
Ribes, E., & López, F. (1985). Teoría de la conducta: un análisis de campo y paramétrico. Mexico: Trillas. — The original taxonomy, which everything above is a reformulation of.
Ribes, E., Vargas, I., Luna, D., & Martínez, C. (2009). Adquisición y transferencia de una discriminación condicional en una secuencia de cinco criterios distintos de ajuste funcional. Acta Comportamentalia, 17(3), 299–331. — Training and transfer tests running through all five criteria, with 24 participants.
Serrano, M. (2009). Complejidad e inclusividad progresivas: algunas implicaciones y evidencias empíricas en el caso de las funciones contextual, suplementaria y selectora. Revista Mexicana de Análisis de la Conducta, 35(monographic issue), 161–178. — The formulas for the three indices — differentiality, effectiveness, precision — and the interpretation bands. Footnote 2 on page 168 is where the author says the band ranges are arbitrary.