What is treatment fidelity in ABA?
Treatment fidelity - also called procedural fidelity or treatment integrity - is the degree to which a plan is delivered as it was written, and it determines whether the data collected during a program can be interpreted at all.
Every clinical decision in ABA assumes a link between the procedure on paper and what happened in the room. Fidelity is that link, measured. When it is high, the data answers the question the plan asked; when it is low, the data describes something else entirely, and no amount of careful graphing recovers the difference.
That makes low fidelity uniquely expensive, because it corrupts both possible outcomes. A plan that appears not to work may simply never have been run, and revising it discards a procedure that was never tested. A plan that appears to work may have worked for reasons not written anywhere, which is why it collapses when the person implementing it changes. Either way the team ends up making confident decisions on evidence that does not support them.
Measuring it is deliberately concrete. The procedure is broken into its component steps - deliver the instruction, wait the specified delay, prompt at the planned level, reinforce on the stated schedule, respond to errors as written - and an observer scores which steps were implemented correctly, usually as a percentage of steps. Checks are done periodically, across every person implementing the plan and across the settings where it runs, since fidelity is a property of an implementation rather than of a document.
Drift has predictable sources. Plans written as prose instead of steps leave room for interpretation; procedures that are difficult to run alongside actual supervision demands get quietly shortened; new implementers learn the plan from whoever is nearest rather than from the plan; and small in-the-moment adjustments accumulate until the procedure has changed without anyone deciding to change it. Steps that are unpleasant to implement, extinction in particular, erode fastest.
The fix is usually training and design rather than reprimand. Behavioral skills training - explain, model, rehearse, give feedback - reliably raises fidelity, and persistent low scores on a specific step are better read as information about the plan than about the person: a step nobody can run at full caseload is a design problem, and simplifying it will do more than restating it. Fidelity checks also give feedback something specific to attach to, which is what makes supervision actionable instead of impressionistic.
Fidelity sits alongside interobserver agreement rather than substituting for it. Agreement establishes that the measurement is trustworthy; fidelity establishes that what was measured is the procedure the plan describes. A program can have excellent agreement on data faithfully recorded from a procedure nobody ran correctly, and the resulting graph will be precise and misleading in equal measure.
Practically, fidelity data lives with supervision records, since supervision is when most observation happens anyway. Keeping the checklists, the scores, and the feedback attached to the plan they belong to is what lets a clinician tell the difference between a treatment that failed and a treatment that was never delivered - and that distinction, more than any single graph, is what makes a program's conclusions defensible.