How to Tell If the Change Didn't Work, or Simply Wasn't Implemented
Why it's worth checking implementation before ruling out the hypothesis
Shani Shimshian, social worker and therapist, founder of ALNA · Updated September 2026
"We tried it, and it didn't work"
This is one of the most common sentences heard in wrap-up meetings for improvement processes, and it's often said with full confidence. But behind that sentence sits an assumption worth examining: that what was actually tried truly reflected what was planned.
An intake team moves to a new step in the enrollment process, chosen to shorten wait times. Everyone was enthusiastic about the idea in the team meeting, no one objected, and a start date was set. A month later, the key metric is checked, the wait time hasn't improved. The immediate, common conclusion: "We tried it, and it didn't work; the hypothesis was wrong."
Before rushing to that conclusion, it's worth asking one more question, almost embarrassingly simple: was the new step actually implemented as planned, by everyone on the team, in every case that came in? In many cases, the answer turns out to be "only partially", some of the team adopted the new step, some kept doing things the old way out of habit, and some weren't sure exactly what was expected of them. And that changes the correct conclusion completely.
Two entirely different questions
"Was our hypothesis about what's causing the problem correct?" and "Did the change we planned actually happen the way we planned it?" are two completely separate questions, requiring different ways of checking and leading to different conclusions, yet on the ground they constantly get blended into one. In the program-evaluation literature, this confusion is sometimes called a "Type III error", evaluating the outcome of an intervention that, in practice, was never carried out as designed. The result: good programs get wrongly ruled out, and problematic programs keep going because no one checked whether they were implemented correctly in the first place.
This distinction matters especially when a change depends on dozens of people acting without close oversight, case coordinators, therapists, or frontline staff. Even if leadership believes the directive is clear and binding, there's no way to know for certain it's actually being carried out in every interaction unless that's explicitly checked.
The professional literature suggests: always separate the measurement of implementation fidelity, whether the action happened as planned, from the measurement of the outcome itself.
What to check before checking the outcome
Measuring implementation can be light and proportionate. A short conversation with several staff members, a sample of case files, or one question in a weekly team meeting may be enough. Before drawing a conclusion from the outcome, answer four basic questions:
Did whoever is supposed to carry out the change actually start doing it?
Sometimes the directive went out, but in practice only part of the team started acting on it, and no one paused to check that before measuring an outcome.
Did it happen at the frequency and quality planned, or only partially and occasionally?
A change that happens once a week instead of in every case gives a misleading picture, the outcome is measured as if the change were fully in effect.
Did the team understand exactly what was expected of them, or did the directive stay general?
A vaguely worded directive leads to execution that varies from person to person, and when the average outcome is measured, it's hard to know exactly what was measured.
Were there technical or operational barriers quietly blocking execution?
A system that doesn't support the new field, a form that wasn't updated, or a step that requires an extra approval, all of these can stop implementation without anyone explicitly noticing.
When it turns out the problem is implementation, not the hypothesis
If the change was not implemented as planned, first address what prevented it. The next action may involve clearer ownership, adjusted workload, additional training, or a more precise statement of what is expected. The organization can then examine the same hypothesis with a more consistent implementation. Ruling out a hypothesis that was never properly tested wastes time and can erode the team's trust in future improvement efforts.
There's also a practical advantage here: fixing an implementation problem is usually much faster and cheaper than going back to the diagnosis stage and searching for an entirely new explanation. Before investing in testing another mechanism, it's worth making sure the first one actually got a real chance to be tested.
At ALNA, here's what we do: before checking whether the outcome changed, we check whether the change happened at all, following organizational behavior management (OBM) principles, which explicitly distinguish between measuring implementation behavior and measuring outcome.
And if the change was implemented and still didn't work
When implementation measurement confirms the change was carried out fully, at the planned frequency and quality, and the outcome still hasn't changed, that's exactly the moment to return to the mechanism hypothesis itself. It's possible the factor identified wasn't the main driver, and it's worth examining another possible mechanism that was presented as an alternative from the start. This is a completely different process from deciding too early to abandon a direction that never got a real chance.
Distinguishing between the two possibilities, failed implementation versus an unsupported hypothesis, is exactly what turns "it didn't work" from a frustrating feeling into information you can act on. In the first case, the way forward is fixing execution. In the second, it's going back a step earlier, reconsidering the possible explanations in light of what's now known.
In both cases, what this process doesn't allow is the most common thing in practice: throwing up your hands and assuming there's no way to know. There's always a way to reduce uncertainty by one more step, before moving on.
Why it's so easy to skip this question
There's a psychological logic to jumping straight to "it didn't work." Checking implementation requires a critical look inward, asking whether our own team, tools, or process were up to the task. It's much easier, and sometimes less uncomfortable, to conclude the idea itself wasn't good and move on. There's also real organizational pressure in that direction: with ten other things in the queue, it's easier to try something new than to stop and figure out why the old thing didn't move.
Without this check, an organization can move through several plausible hypotheses and rule them out even though none received a fair test. One simple question before moving on can prevent that cycle.
Sources & Further Reading
- Basch, C.E., Sliepcevich, E.M., Gold, R.S., Duncan, D.F., Kolbe, L.J. (1985). "Avoiding Type III Errors in Health Education Program Evaluations: A Case Study." Health Education Quarterly, full article
- Carroll, C., Patterson, M., Wood, S. et al. (2007). "A conceptual framework for implementation fidelity." Implementation Science, full article
- OBM Network, the professional association for organizational behavior management. obmnetwork.com
- Consolidated Framework for Implementation Research (CFIR), a widely used framework for evaluating intervention implementation. cfirguide.org
Tried a change that didn't produce results?
In a short introductory call, we can work through together whether this is an implementation gap or a hypothesis worth revisiting.
Also worth reading: Why a Good Recommendation Can Still Fail at Implementation