Write the revisit criteria first
Everyone writes success criteria. Some write failure criteria. Almost nobody writes the third thing, which is the one that survives a pivot.
Before you run anything, write down three things. What success looks like. What failure looks like. And the conditions under which you would revisit the test itself. Everyone writes the first. Some people write the second. Almost nobody writes the third.
The reason the third matters is that success against a threshold you set too low is its own kind of failure. A test can go overly well, and you will read that as a green light rather than as evidence that you set the bar in the wrong place.
Related, and worth saying plainly: a metric is only a metric if it is time-bound. Time is what qualifies the number. Without it you do not have a target, you have a wish.
The example I use is founders going zero to one on a consumer app. Someone says, if we can get to twenty users that would be great. Fine. Twenty users is meaningless on its own. Not getting to twenty is one problem. Getting to twenty in six months is a different problem. Sitting at twenty-five a year later is not a problem at all, it is a verdict.
What the revisit criterion actually captures is not the number. It is the assumptions underneath the number. What did I believe about the world when I set this, and what would have to change for this to stop being the right test? Founders pivot all the time and carry the same success criteria straight through the pivot without noticing. The metric survived. The world it described did not.
The obvious objection, which I have taken more than once, including from my wife: pre-writing the conditions under which you would revisit a target is just giving yourself permission to move the goalposts.
My answer is that sometimes the goalpost does need to move. Sometimes it is too close, or too far, or too big, or too small, or pointing in the wrong direction entirely. Moving it is not the problem. Moving it without knowing that is what you are doing is the problem.
And the alternative to written revisit criteria is not holding firm. Nobody holds firm. The alternative is moving the goalposts anyway, unrecorded, usually months later, and usually after the person who held the tacit knowledge of why that number was chosen has left. The team then re-derives the target from first principles as though it is a brand new test, having lost everything the previous one already accounted for.
Which reframes what you are writing. The revisit criterion is not a hedge. It is a knowledge-transfer artefact.
One smaller habit in the same family, and I forget it as often as anyone. People test that a guard blocks the bad case. They almost never test that it permits the good one. Same blindness, which is instrumenting one side of the coin and calling it covered.
“Sometimes the goalpost does need to move. The problem is never moving it. It is moving it without knowing that is what you are doing.”