Beating zero is not an achievement, and most comparisons are made against the wrong alternative
The short answer
A test compares a result against what an alternative would have produced, and the alternative is a choice. Four are in common use, in increasing severity: zero, a price index, a total return index, and a matched random strategy holding exposure and holding period constant. Each one removes an explanation that is not skill. The first removes none. Against a price index rather than a total return index, an Indian comparison hands over roughly 6.9 per cent over five years and 31 per cent over twenty at a 1.35 per cent dividend yield, before anything is decided. And against a properly matched null, a strategy with modest exposure produces a strongly positive year one time in twenty by chance alone. Most results that survive a zero null do not survive a matched one.
This is the least glamorous idea in research method and the one that decides the most. A strategy can be described accurately, tested carefully and reported honestly, and still mean nothing, because the thing it was compared against was never the real alternative.
A test is a comparison, and you chose what it compares to
The mechanics of a significance test are usually taught from the arithmetic end, which hides the part that matters. The arithmetic asks how surprising an observed result is under an assumed alternative. That alternative is not given by the data. It is supplied by whoever is running the test, and the entire content of the conclusion depends on it.
This creates a specific and very comfortable failure. Choosing a weak alternative does not make a test look weak. It leaves every visible feature of a test intact: there is still a comparison, still a number, still a conclusion. What has been removed is the only part that carried information. A result reported as beating its benchmark tells you nothing at all until you know what the benchmark was and what it controlled for.
| Null | Controls for | Still lets through |
|---|---|---|
| Zero | Nothing | The entire market move, plus exposure, plus timing |
| A price index | The market's price move | The distributions, plus exposure, plus timing |
| A total return index | The price move and the distributions | The exposure choice, plus timing |
| A matched random strategy | All of the above, plus exposure, turnover and holding period | Selection and timing, which is what was claimed |
The first rung, and why it is still in use
Comparing against zero asks whether the strategy made money. In a market that rose, holding anything at all made money, and so did a randomly timed position in it. The comparison is therefore not between the strategy and its alternative but between the strategy and not participating, which is a question about whether to invest rather than a question about the strategy.
It persists because it is the comparison a profit and loss statement naturally produces. Nobody has to choose it. It is simply what is in front of you, which is exactly why it needs to be displaced deliberately.
The second rung, and the amount it quietly concedes
Comparing against a price index is a real improvement and it contains a specific, quantifiable gift. A price index tracks the level of its constituents and nothing else. It does not include the distributions those constituents pay. An investor who simply held the constituents would have received them.
So a comparison against a price index is a comparison against a version of buying and holding that threw the distributions away. The strategy is credited with an advantage it did not create.
| Comparison horizon | Advantage conceded before anything is decided |
|---|---|
| 1 year | 1.4 per cent |
| 3 years | 4.1 per cent |
| 5 years | 6.9 per cent |
| 10 years | 14.4 per cent |
| 15 years | 22.3 per cent |
| 20 years | 30.8 per cent |
Over one year this is small enough to ignore and usually is. Over the horizons at which long term claims are made, it is not small. A twenty year record measured against a price index begins roughly 31 per cent ahead of where it should, which is enough to convert an unremarkable result into an impressive one without anything having happened.
The yield is not a constant, so these figures show the shape of the effect at a current yield rather than a measurement of any historical period. The direction, however, is never in doubt: the gift always runs the same way, toward the strategy.
The third rung, and what it still cannot see
A total return index fixes that. It puts the distributions back and gives the honest passive comparison. It is the right benchmark for the question "should I have done this instead of holding the market", and for a fund that is fully invested it is usually the correct null.
It is not the right null for a strategy that chooses when to be in the market, and the reason is worth being precise about. If a strategy is invested only part of the time, it is taking less market exposure than the index, so beating the index is a harder test than it should be. If it is invested with leverage, it is taking more, so beating the index is an easier test than it looks. In both cases the comparison is contaminated by the exposure decision, and the exposure decision is not usually the thing being claimed.
The rung that actually tests the claim
The demanding null holds constant everything about the strategy except the part being claimed as skill. Take the real strategy's exposure, its turnover and its typical holding period, and build a version that makes those same commitments at random times. That is the matched null, and it is what the claim should be measured against, because the only difference between the two is selection and timing.
Running it produces a distribution rather than a single number, and the width of that distribution is the finding.
| Exposure profile | Median | 90th percentile | 95th percentile | 99th percentile |
|---|---|---|---|---|
| 30 per cent of the year, 5 day holds | -0.00 | +0.67 | +0.85 | +1.20 |
| 30 per cent of the year, 20 day holds | -0.00 | +0.69 | +0.88 | +1.26 |
| 60 per cent of the year, 20 day holds | +0.01 | +0.85 | +1.09 | +1.54 |
| fully invested all year | +0.01 | +1.30 | +1.65 | +2.32 |
The median is zero, as it must be, because random timing has no edge by construction. The percentiles are the point. A strategy exposed 30 per cent of the year in short holds lands above +0.85 annual volatilities one year in twenty with no skill at all. To make that concrete, if a broad index had an annual volatility of 15 per cent, a figure used here purely to convert the units and not a measurement of any index, that is a result of roughly 13 percentage points in a year, produced by nothing.
Notice also what happens down the table. The more of the year a strategy is exposed, the wider its null distribution becomes, because there is more market variation inside it. A fully invested strategy has the widest null of all, which is the formal version of a familiar observation: the more market you take, the more of your outcome the market explains, in both directions.
The start date is a parameter too
A comparison has a beginning and an end, and both are choices that nobody thinks of as choices. They behave exactly like any other free parameter: there are many available, they produce different answers, and the one that gets reported is the one that made the result legible.
The mechanism is straightforward once it is named. Over any long record there are periods in which a given approach did well and periods in which it did not. A comparison beginning at a market low flatters a long strategy. One beginning at a high flatters a defensive one. Neither start date was chosen dishonestly. Both were chosen because the data happened to begin there, or because a round number of years was convenient, or because that is when the record starts.
The defence is the same as for any other searched parameter, and it is worth stating because it is cheap. Report the result across several start dates rather than one, and report the range. If the conclusion survives moving the start by a few months in either direction, it was not a start-date artefact. If it does not survive, that is the finding, and it is a more useful finding than the original number was.
The same applies at the other end. A comparison that ends at a favourable moment has the identical problem, and it is harder to notice because the end of a record feels like a fact rather than a decision. It is a decision whenever the record could have been cut anywhere and was cut there.
Building the matched null for your own strategy
None of this is useful as a principle. It is useful as a procedure, and the procedure is short enough to carry out on any strategy that has a trade log.
| Step | What to take from the real strategy | What to randomise |
|---|---|---|
| 1. Exposure | The fraction of the period with a position open | Nothing yet |
| 2. Holding period | The distribution of holding lengths, not just the average | Nothing yet |
| 3. Direction | The proportion of long and short positions | Nothing yet |
| 4. Instrument | The same instrument or the same universe | Nothing yet |
| 5. Entry timing | Nothing | Start each position at a uniformly random point |
| 6. Repetition | Nothing | Repeat several thousand times to get a distribution |
The discipline is entirely in the first four rows. Everything about the strategy that is not the claim must be copied across exactly, because anything left uncopied becomes a second explanation for the difference and the test loses its edge. A common slip is matching the average holding period while ignoring its spread, which produces a null with quite different variance from the strategy and therefore the wrong percentiles.
The direction row matters more than it looks for anything that shorts. A long-short strategy compared against a long-only random null is being credited for the entire short book, which is an exposure decision rather than a selection one. The null has to carry the same mixture.
What comes out is not a verdict but a percentile: the fraction of random versions that did at least as well. That number is the honest summary of the result, and it is usually a good deal less impressive than the raw comparison it replaces.
Comparing two strategies is a different problem again
Everything above concerns testing one result. Ranking two introduces a separate error, and it is the most common one in practice: comparing outcomes without matching risk.
A strategy that takes twice the variability should produce roughly twice the swing in both directions. Over a favourable period it will look better. Over an unfavourable one it will look worse. Neither observation says anything about which is the better method, and both will be reported as though they did.
| Comparison | Valid when | Fails when |
|---|---|---|
| Raw outcome | Risk and exposure are genuinely matched | Either differs, which is almost always |
| Scaled to a common variability | You state the scaling and apply it before looking | The scaling is chosen after seeing the result |
| A ratio that already divides by variability | The sample is long enough for the ratio to be estimated | The sample is short, which makes the ratio itself noise |
What even a matched null cannot control for
It would be a poor article on this subject that presented its own strongest method without its limits, so here is the one that matters most.
A matched null tests whether this strategy's timing beat random timing over this period. What it cannot test is whether the strategy was chosen for having done well over this very period. If a hundred variants were tried and the best one is now being compared against its matched null, the comparison is being run on a result that was already selected for looking good. The null distribution is correct, the percentile is computed correctly, and the conclusion is still wrong, because the strategy did not arrive at the test at random.
That failure is invisible from inside the test. Nothing about the trade log records how many alternatives were discarded on the way to it, and the percentile that comes out will look exactly as convincing either way. This is the point at which a good null runs out and the only remaining defence is a record of what was tried, kept at the time.
Two smaller limits are worth naming alongside it. A matched null says nothing about whether the cost model was honest, because both the strategy and the null are charged whatever costs you assumed, so a shared error cancels and leaves the comparison looking clean. And it says nothing about capacity: a result that depends on trading sizes the market would not absorb will beat its null comfortably and still be unavailable to anyone.
| Question | Does the matched null answer it? |
|---|---|
| Did the timing beat random timing this period? | Yes, and this is its job |
| Was the market exposure doing the work? | Yes, that is what matching removes |
| How many variants were tried first? | No, and nothing in the data can tell you |
| Was the cost model honest? | No, a shared error cancels out |
| Would it work at size? | No, capacity is a separate question entirely |
| Will it work next period? | No, and no test of past data can |
Choosing the null before you see the result
All of this collapses into one procedural rule, and it is the only part of this page that has to be remembered.
Write down the null before you look at the result. Not the strategy, not the period, not the metric: the alternative explanation, stated specifically enough that somebody else could construct it. Doing this first makes the comparison a test. Doing it afterwards makes it a description, because a null chosen once the result is known will be chosen, without any dishonesty being involved, to be one the result beats.
It is worth being clear about how ordinary that failure is. Nobody sits down intending to select a flattering benchmark. What happens is that several comparisons are available, one of them is more familiar or more readily to hand, and the one that gets reported is the one that made the result legible. The defence against it is not integrity, which was never missing. It is sequence.
The payoff is real, and it runs in the direction people do not expect. Most results do not survive a properly matched null, which sounds discouraging until you notice what it means: the small number that do survive are worth the whole of your attention, and you will now be able to find them.
Frequently asked questions
What is a null in this context?
The alternative explanation a result is being tested against. Every comparison implicitly has one, and the question a test answers is whether the observed result is surprising under that alternative. Change the alternative and the same result can go from remarkable to entirely ordinary without a single number about the strategy changing.
Why is beating zero not enough?
Because a long position in a rising market beats zero without any skill being involved, and so does a randomly timed one. Zero is the alternative in which markets do not go up, which is not the alternative anyone was actually choosing between. It is the weakest possible comparison and it is the most commonly used.
What is the difference between a price index and a total return index?
A price index tracks only the level of its constituents. A total return index adds the distributions those constituents pay, reinvested. The second is what an investor holding the constituents would actually have experienced, which is why it is the honest comparison and the first is not.
How large is that difference in India?
It compounds at roughly the index dividend yield, which the Nifty 50 factsheet dated 29 May 2026 gives as 1.35 per cent, and which has historically sat in the region of 1 to 2 per cent. Compounded, that is a low single digit figure over a few years and a substantial one over fifteen or twenty. The yield changes over time, so the figures on this page show the shape of the effect at a current yield, not a historical measurement.
What is a matched random strategy?
A strategy that takes positions at random while holding constant everything about the real strategy except its selection and timing: the same fraction of the year exposed, the same typical holding period, the same instrument. Comparing against it isolates the decisions the strategy claims to be good at, because everything else is identical by construction.
Why is that null so much harder to beat?
Because it already contains the market exposure. A strategy that is invested most of the time will beat a total return index comparison if it is invested at slightly better moments, but a matched random strategy is invested just as much, so the comparison strips out the exposure and leaves only the timing. Most of what looks like skill under a weaker null is exposure.
Does a result beating a matched null prove the strategy works?
No. It clears one specific alternative explanation, which is a real thing to have done and is more than most results manage. It says nothing about how many variants were tried before this one, whether the period was representative, or whether costs were modelled honestly. A good null is necessary and not sufficient.
How do I compare two strategies with different risk?
You cannot compare them directly on outcome. A leveraged or concentrated result and an unleveraged one are different risks, and the one taking more risk should produce more of both outcomes. Either scale one to the other's variability before comparing, or compare on a measure that already divides by variability, and say which you did.
Why is one year of results not enough even against a good null?
Because the null distribution is wide. In the simulation on this page, a matched random strategy with modest exposure has a one in twenty chance of producing a strongly positive year with no skill whatsoever. A single year that looks good is therefore consistent with luck at a probability nobody would accept as evidence in any other context.
What should I actually do with this?
Before looking at a result, write down what the alternative is and what it would have produced. Doing that first is what makes the comparison a test rather than a description, and doing it afterwards, once the result is known, is how a null gets chosen to be beaten rather than to be informative.
How these numbers were produced. The dividend gap compounds the stated index dividend yield over each horizon. It is the shape of the effect at a current yield, not a measurement of any historical period; the yield varies, so recompute it for the period you care about. The null distribution comes from 20,000 simulated runs per exposure profile over 246 trading days with a fixed random seed, taking positions at uniformly random start points for the stated holding period, with results expressed in units of the market's own annual volatility so that no volatility figure has to be assumed. The 15 per cent volatility used once in the text is a unit conversion offered to make the scale concrete and is not a measurement of any index. All simulation figures are illustrative of the structure of the problem and are not a prediction about, or a measurement of, any actual strategy.
The position is stated as at September 2026. Index yields, constituents and methodology change; confirm the current factsheet and methodology directly before relying on any figure here, and take advice on your own circumstances.
Ready to go deeper than this article?
Bharath Shiksha is a 90-volume curriculum across 6 stages, from chart reading at ₹14,999 through capital raising, or the full bundle at ₹1,49,999. Choosing the alternative before looking at the result is a habit rather than a technique, and it is the habit that separates research from recollection.
Take the free diagnostic →