Traction is the absence of surprise, not the presence of customers.
A week that beats the last one feels like proof. At seed volumes it usually is not. Below roughly fifty meaningful attempts a week, output is dominated by variance, which means a good week and a bad week can be produced by the same process with nothing changed in between. A founder who reads the good week as signal and the bad week as noise is not measuring anything. They are selecting.
Runway is not really cash. It is a finite number of controlled attempts, and a scaling decision spends several of them at once. Hiring, spend and roadmap commitments raise coordination costs that do not come back down afterwards. Making those commitments on top of a process you cannot predict does not produce growth. It produces the same variance, more expensively.
What a spike is evidence of
A spike is evidence that something happened. It is not evidence that you can make it happen again, and the second claim is the only one a scaling decision rests on.
Founder heroics produce spikes. So does a well-timed introduction, a post that travelled further than usual, or a competitor’s outage. Each of those is real revenue and none of them is a loop. A spike you bought stops when the spending stops. A spike you were given does not repeat on request. The test that separates a loop from an event is not size. It is whether output moves when the inputs move, and holds still when they do not.
Which is why a loop with flat output and low error is worth more than a loop with rising output and high error. The first can be tuned and then multiplied. The second cannot be multiplied at all, because you do not know which part of it to multiply.
Loop B produced more over the ten weeks. Only Loop A tells you anything about the eleventh.Dr. Hafiz Muhammad Ali
What the measure has to survive
Measuring this is where most attempts come apart, and they come apart for a documented reason.
The obvious approach is percentage error: take what you expected, take what happened, express the gap as a proportion. It is the standard measure and it is the wrong one here. Hyndman and Koehler showed that percentage error is undefined when the actual value is zero, and severely skewed when it is anywhere near zero, which makes it unusable for what they call intermittent data: small counts that often include zeros.
A seed-stage loop is intermittent data. Two qualified calls one week and none the next is not an edge case, it is an ordinary fortnight. A measure that goes to infinity on a zero week will report a catastrophe in the one week that was always going to happen.
So state a tolerance in whole units and measure against that instead. Not “we were thirty per cent out” but “we said three, we got one, and we had agreed that one either way was acceptable”. Hyndman and Koehler’s own alternative scales the error against what a naive forecast would have produced, which is the same instinct in more formal dress: judge the loop against the dumbest available prediction, and if it cannot beat that, it is not yet a loop.
What the number is for
The point of the measure is not the measure. It is that it gates exactly one decision.
Inside tolerance, the loop is behaving. Raise the inputs, change nothing else, and watch whether output follows. If it does, you have something that scales. If it does not, you have learned that the loop was smaller than it looked, which is worth knowing before the hire rather than after it.
Outside tolerance, hold the inputs where they are and change exactly one thing. Change two and you will not know which one moved the error, and you will have spent a week to learn nothing.
When the error refuses to come down, the cause is usually not the script or the list. It is that the loop is being fed by a boundary that has never excluded anyone, so the attempts inside it were never comparable to begin with. Averaging incompatible segments produces a number that cannot be forecast because it is not measuring one thing.
There is no threshold worth quoting here. Any specific error rate offered as a benchmark is somebody’s convention rather than a finding, and adopting it will tell you less than a tolerance you set yourself and then hold to. What is not arbitrary is the ordering. Control first, then scale. A loop that cannot predict its own output under steady inputs will not begin predicting it under larger ones.
Three questions, asked before the week starts.
-
What did you predict last week, in writing, before it began?
If there was no number written down in advance, last week produced no information about the loop. It produced a result, which is a different thing, and results assessed after the fact are always explicable.
-
Which of last week’s wins could you reproduce on purpose?
Go through them one at a time and name the controllable input behind each. The ones with no controllable input behind them are not part of your loop, whatever they did to the total.
-
What would have to be true before you raised inputs by a fifth?
Write the condition now, while nothing is at stake. A condition written during a good week will be written to permit the thing you already wanted to do.
Growth that surprises you is not yours yet. The loop you can predict is the only one you can decide anything with.
References
- Hyndman, R. J., & Koehler, A. B. (2006). Another look at measures of forecast accuracy. International Journal of Forecasting, 22(4), 679–688.