Someone looks at your graph and asks the only question that matters: how do you know this is true?
Not “did the child improve”. Something earlier and harder: how do you know the line on that graph is about the child, and not about who was holding the clipboard, or what got typed in afterwards?
There are (at least) three separate ways a record can be wrong, and they need three separate answers.
Figure — Three checks between the room and the graph
What happened in the room
-
Would a second person have written the same thing?
Interobserver agreement
- Proves:
- Two adults, scoring independently, recorded the same events.
- Does not prove:
- That either of them was scoring the right thing.
Agreement and Cohen’s kappa, per session and per concept — and the number of trials scored twice, which is what decides what the other two mean.
-
Did the session run the way it was designed?
Procedural fidelity
- Proves:
- Prompts, probes, reinforcement and error correction followed the plan.
- Does not prove:
- That the plan was the right plan.
Read off the trial record, not from a second adult with a clipboard. What cannot be read from it — the interval, the number of comparisons — is named on screen and left out of the figure.
-
Could the record have been changed afterwards?
A record that cannot be edited quietly
- Proves:
- The file is the one written at the time, or a partial edit to it is visible.
- Does not prove:
- That the file is impossible to alter, and not that every possible alteration shows — a chain with no copy kept outside the device can, in principle, be rewritten consistently from a point forward.
Each trial is hashed over the one before it, over a fixed set of core fields, and a finished session is locked against new trials. That catches a stray edit to one trial; it does not by itself rule out someone with full access rewriting a whole session end to end.
The line on the graph
What none of the three prove
That the child learned anything. Together they buy something narrower and more useful: the graph becomes a claim about the child rather than a claim about the recording.
1. Would a second person have written the same thing?
That question is interobserver agreement, and it is the one that decides whether a dependent variable can be reported at all. If only one person ever scored a trial, there is no evidence that the scoring was consistent — only that it was confident. High agreement on its own is not the same as accuracy: two observers applying the same wrong criterion still agree with each other, every time. What agreement establishes is consistency; whether the criterion itself was the right one is a separate question this check does not answer.
The procedure is unglamorous: a second adult watches the same child and scores the same trials, independently, and afterwards the two records are compared. The word doing the work is independently. Two observers who can see each other’s answers are one observer with a witness.
Two numbers come out of it:
- Percentage agreement. Ledford and Gast (2018) put the bar at 90% or above for applied research, and treat anything below 80% as unacceptable. Those are their thresholds, and they do not become friendlier because a session went badly.
- Cohen’s kappa. Raw agreement inflates when one category dominates. If the child gets almost everything right, two observers agree constantly by accident. Kappa does two things to correct for that: it subtracts the agreement you would expect from chance alone, and then normalizes by the agreement that was actually still possible beyond chance — so what is left is a proportion of the available agreement, not just a raw difference.
There is a third number people forget, and it changes the meaning of the other two: how many trials were scored twice. Agreement of 100% over three trials and agreement of 100% over eighty are not the same claim. Only the denominator tells you which one you are reading.
In Interlaza, you arm the next session for a second observer before it starts, from the child’s Results tab. This live-scoring feature specifically cannot be arranged after a session has already run — the second adult scores on their own device while the trials happen, and the screen does not show them how the first observer scored a trial until they have committed their own answer. That is a limit of this one feature, not of interobserver agreement as a concept: a session recorded on video, watched independently afterward by someone who did not see the primary scoring, can still support a computed IOA outside the app. Afterwards the report gives you agreement and kappa, broken down by session and by concept, with the number of trials it was computed over.
The breakdown is not decoration. Ledford and Gast warn specifically against reporting a single averaged agreement figure, because averaging hides the condition where the two observers were pulling apart. So the overall number never appears without its tables.
2. Did the session actually run the way it was designed?
That second question is procedural fidelity, and it is treated as a risk-of-bias criterion, not as an administrative nicety: Ledford and Gast name it as one of the things a risk-of-bias assessment weighs, and a study that measures and reports it has cleared that one item — measuring fidelity alone does not, by itself, classify an entire study as low risk of bias, since other criteria (like the agreement check above) still have to hold too.
On paper this is expensive. It means a second adult with a checklist, ticking boxes: was the prompt delivered at the planned level, was the interval respected, was the reinforcer delivered on the right schedule, was the correction procedure run when it should have been.
Software knows most of that without asking anyone. The engine that ran the session already wrote down, on every single trial, what it did, so the checklist is a reading of the record, not a second person in the room. Interlaza checks four things:
- Prompt fading stayed inside its configured range, never rose after an error, and never jumped more than one step. What this checks is the fading phase — the discrete step in the schedule — not how strong the prompt on screen actually looked; it doesn’t see which prompt type (opacity, size, a highlight) was in use.
- Unprompted opportunities. Two different kinds of trial that happen to share one requirement: a probe, used to test transfer to material the child was never trained on, and an error-correction trial, the child’s next attempt right after a miss, must both arrive with no help showing. A probe carrying a prompt measures the prompt instead of transfer, and Carroll’s “repeat until independent” is only independent if nothing is on screen either.
- Reinforcement matched the response, and thinning never started before the concept had earned it.
- Error correction ran the configured procedure and stayed inside its ceiling.
Only invariants that survive the app’s own adaptivity are counted. If the intervention monitor holds a fading phase, or adaptive fading changes how many correct responses a phase needs, or a correction procedure stops early because the child is showing distress — none of those are deviations, and none of them are counted as one.
And the panel says out loud what it cannot check. The trial record stores when a response arrived, not when the trial began, so the inter-trial interval is not verifiable from it; the number of comparisons on screen is not stored either. Both are listed on screen, with the reason, and neither is folded into the percentage. A fidelity figure that quietly drops the items it could not measure is worse than no figure at all.
3. Could the record have been changed afterwards?
The first two checks are about how the data was produced. This one is about what happened to it since.
Every trial in Interlaza carries a hash computed over the trial before it, so the trials of one session form a chain. What is hashed is a fixed set of ten core fields — which concept, which response, whether it was correct, its latency, its timestamp, and a few more — not the whole row a trial might carry. Editing one of those fields after the fact breaks the chain from that point on, in a way that shows up when the chain is verified. When a session finishes it is locked, and the server refuses to insert new trials into a session that is already closed.
The precise wording matters here, and there is a real limit worth stating rather than glossing over. This does not make the data impossible to alter. It makes a partial alteration evident — someone editing one trial without recomputing everything after it. The chain itself lives on the device with no independently held copy elsewhere, so it has no external anchor: someone with full access to the database could, in principle, rewrite an entire session’s trials and recompute every hash after the edit, and the chain would verify as consistent again. What the chain buys you is protection against a casual, partial edit, not a cryptographic guarantee against a determined rewrite with no outside copy to check against.
What none of this proves
None of the three tells you the child learned anything. Agreement says two people saw the same thing. Fidelity says the procedure ran as designed. The chain says the record is the one that was written at the time.
Together they give you something narrower and more useful than confidence: they make the graph a claim about the child, rather than a claim about the recording. Whether the child is learning is the next question, and it is worth asking only once these three are answered.
Agreement thresholds and the procedural-fidelity criterion are from Ledford, J. R., & Gast, D. L. (2018). Single Case Research Methodology: Applications in Special Education and Behavioral Sciences (3rd ed.), chapter 5.