The best of fifty backtested rules is a statistic about your search, not about the market
The short answer
When you test one rule, its result describes the rule. When you test fifty and report the winner, the result describes the maximum of fifty draws, and the maximum of a set behaves nothing like a single draw. On simulated data engineered to contain no edge at all, one three-year candidate clears an annualised Sharpe of 1.0 about 4 percent of the time, while the best of fifty such candidates clears it 88 percent of the time. So the claim that a rule beat a benchmark carries no information until the number of candidates searched is stated alongside it, and that number is almost never recorded, including by the person who did the searching.
This is the failure that quietly invalidates most retail strategy research, and it is not a matter of care or of computing power. A researcher who is scrupulous about survivorship, about transaction costs and about look-ahead can still produce a result that means nothing, because the defect enters through the act of choosing rather than through the act of testing.
Everything below is computed rather than asserted. Two of the three instruments are simulations on random data, which is stated at each one, and the third runs a real rule family on a series built from the exchange archive. The parameters are given so the numbers can be reproduced.
A backtest result is a statistic about a search
Consider what a performance figure is supposed to answer. Run one rule, get an annualised Sharpe, and the honest question is: if this rule had no edge, how often would three years of data flatter it this much? That is a well posed question with a known answer.
Now run fifty rules and report the best one. The question the number answers has silently changed. It is now: if none of these fifty had an edge, how often would the best of fifty look this good? Those two questions have very different answers, and the gap between them widens with every candidate added.
The reported figure did not lie. It simply stopped being about the rule and became a statistic about the search that produced it. Nothing on the page tells the reader which of the two questions was answered, because the only thing that distinguishes them is a number that was never written down.
The distribution that matters is the distribution of the maximum
The mechanism is easiest to see where the answer is known in advance, so here it is built deliberately. Generate N return series of 756 trading days each, drawn independently with a true mean of exactly zero. There is no edge anywhere in this experiment by construction. Compute each candidate's annualised Sharpe, keep the highest, and repeat two hundred thousand times.
| Candidates searched | Median best Sharpe | 90th percentile | Best clears 1.0 | Best clears 1.5 |
|---|---|---|---|---|
| 1 | 0.00 | 0.74 | 4.1% | 0.5% |
| 5 | 0.65 | 1.17 | 19.1% | 2.3% |
| 20 | 1.05 | 1.48 | 57.4% | 9.1% |
| 50 | 1.27 | 1.66 | 88.2% | 21.2% |
| 200 | 1.56 | 1.90 | 99.98% | 61.7% |
| 1000 | 1.85 | 2.15 | 100% | 99.2% |
Read the first row and the fourth together. A single candidate reaches an annualised Sharpe of 1.0 in 4.1 percent of runs, which is roughly the rarity a naive reader assumes when they see that figure. The best of fifty reaches it in 88.2 percent of runs. A number that would have been surprising becomes the expected outcome, and the market did not participate in the change.
The same arithmetic runs through the significance test. With one true-null test at a 5 percent threshold you get a false positive 5 percent of the time. With fifty independent ones the probability that at least one clears the bar is 92.3 percent. With two hundred it is 99.997 percent. So a winner described as significant at 5 percent, drawn from a search of fifty, is not evidence of anything. It is what the search was always going to produce.
Two instruments were used here and they were made to agree. The main figures come from the exact sampling distribution of the Sharpe on an independent sample, and a slower brute force pass that simulates actual return paths reproduces them: at two hundred candidates the exact method gives a median best of 1.564 and the brute force pass gives 1.565. An identity that is asserted rather than checked is still an assertion.
The trial count you never wrote down
The natural defence is that this describes a factory running mass searches, not a person with one idea. It does not. A single idea is a grid, and the grid is enumerated during development whether or not anyone counts it.
Take a crossover rule. You try a few fast lengths and settle on one. You try slow lengths, discard the ones whose equity curve looks wrong, keep the survivor. You test a stop, then a wider stop. You add a volatility filter, look at the result, and keep the filter. You check the rule on a second and a third instrument. That is a grid of three thousand, and the write-up will describe it as one strategy with one set of parameters.
What makes this so resistant to correction is that the discards do not feel like tests. Rejecting a parameter because the curve looked poor feels like craftsmanship. But the criterion for rejection was the data, which makes it a test, and the arithmetic of the maximum has no interest in what the researcher called it.
There is a second, larger search most people never count at all: other people's. A rule adopted because somebody published it carries their entire trial count, invisibly. If a hundred people each test a hundred rules on the same market history and only the winners get published, a reader who takes a published rule at face value has inherited an effective count in the thousands while believing it to be one.
Two different things to be wrong about
Once the count is acknowledged, the question becomes what to do about it, and the literature offers two different objectives that are often confused.
Family-wise error rate is the probability that even one of the things you accepted is noise. Controlling it at 5 percent means there is only a one in twenty chance that even one entry on your shortlist is a false discovery. The Bonferroni correction achieves it by dividing the threshold by the number of tests, which for two hundred candidates means demanding a p-value below 0.00025.
False discovery rate is the expected proportion of your accepted set that is noise. Controlling it at 10 percent means accepting that roughly one entry in ten is wrong, in exchange for a shortlist that contains something. The Benjamini-Hochberg procedure achieves it by ranking p-values and comparing each to a threshold that rises with its rank.
Here is the comparison run on a universe where the truth is known. Two hundred candidate rules, of which exactly ten carry a genuine annualised Sharpe of 0.8 and the other one hundred and ninety are pure noise, each with three years of daily data. The planted count is the ground truth no live research ever has, which is precisely why the comparison has to be made in simulation.
| Selection rule | Real edges found, of 10 | Noise accepted | Share of the shortlist that is noise |
|---|---|---|---|
| No correction, 5 percent | 3.91 | 9.55 | 70.4% |
| Bonferroni, 5 percent family-wise | 0.16 | 0.05 | 4.7% |
| Benjamini-Hochberg, 10 percent false discovery | 0.38 | 0.19 | 9.7% |
The first row is the one worth sitting with. Applying an uncorrected 5 percent threshold to two hundred candidates produces a shortlist that is seventy percent noise, even though ten genuinely good rules were present and waiting to be found. The procedure is not failing to find them. It is drowning them.
The third row shows the false discovery rate procedure delivering what it advertised: a realised noise share of 9.7 percent against a target of 10. The second row shows Bonferroni delivering what it advertised too, a shortlist that is 95 percent clean, at the cost of surfacing one of the ten real edges only about once in every six searches.
Which one a trader wants follows from what a shortlist is for. A regulator approving a drug is making one irreversible decision where a single false positive is a disaster, and family-wise control is right. A trader is assembling a set of candidate rules to allocate to in small sizes, watch, and cut. Insisting that not a single false rule ever enters that set is an expensive guarantee, and the price is refusing almost every real one as well.
What the correction costs, priced in years of data
Both procedures in that table recovered almost nothing, and the reason is not the procedures. It is that three years of daily data cannot resolve a Sharpe of 0.8 against a bar set for two hundred candidates. The honest way to express the cost of a correction is therefore not in statistical power but in years of history.
| True annualised Sharpe | One test named in advance | Best of 20 candidates | Best of 200 candidates |
|---|---|---|---|
| 0.5 | 10.8 then 24.7 | 31.5 then 53.3 | 48.5 then 74.8 |
| 0.8 | 4.2 then 9.7 | 12.3 then 20.8 | 19.0 then 29.2 |
| 1.0 | 2.7 then 6.2 | 7.9 then 13.3 | 12.1 then 18.7 |
| 1.5 | 1.2 then 2.8 | 3.5 then 5.9 | 5.4 then 8.3 |
The closed form was checked against a simulation at the 0.8 and two hundred candidate cell: the analytic power at 19.0 years is 0.502 and twenty thousand simulated samples give 0.503.
Read across the 0.8 row. A rule that genuinely carries an annualised Sharpe of 0.8, named in advance as a single hypothesis, needs about four years of daily data for a coin-flip chance of clearing the bar. The identical rule, arrived at as the winner of a two hundred candidate search, needs about nineteen. The difference is not a matter of statistical taste. It is fifteen years of Indian market history that cannot be bought, borrowed or waited for.
That reframes pre-registration entirely. Writing the hypothesis down before testing it is not a bureaucratic courtesy borrowed from academia. It is the single largest lever available to a researcher with a finite amount of history, and on these numbers it is worth more than every other methodological improvement combined.
Read the 0.5 row for the harder message. A true annualised Sharpe of 0.5 is a respectable rule, and as the winner of a two hundred candidate search it needs roughly half a century of daily data before it can be distinguished from the noise it was selected against. For a great many real rules, the honest conclusion is not that they fail but that the available history cannot answer the question, and a research process that never returns that verdict is not testing anything.
Ninety-six rules on a real Indian series
Simulations settle the mechanism. The next question is what the mechanism does to a real search on real Indian data, so here is one, run end to end.
The series is a composite built from the exchange security archive: for every trading day, the two hundred counters with the highest turnover on the previous day, equally weighted. Ranking on the same day's turnover is a look-ahead, because heavy volume arrives alongside large moves, and the first version built that way produced an annualised drift above 150 percent, which is how the bug announced itself. Membership is recomputed daily from each day's own file, so nothing has to survive the window and no survivorship filter is applied. It is a breadth composite of liquid counters, not an index, and it is not investable. Over the window it carried an annualised drift of 3.5 percent and annualised volatility of 18.6 percent.
The rule family is every moving-average crossover from ten fast lengths and ten slow lengths where the fast is shorter, which is 96 rules. Each rule holds the composite when the fast average is above the slow one and holds nothing otherwise, acting on the following day's return so no rule can trade on information it did not have. Costs are excluded, which biases every rule upward. The split was fixed at 60 percent before the data was loaded, giving 246 in-sample days and then 164 held-out days beginning 20 January 2026.
| First window | Held-out window | |
|---|---|---|
| Best rule by first-window Sharpe | 1.21 | -0.98 |
| Median rule in the family | -0.21 | -0.39 |
| Simply holding the composite | -0.11 | 0.53 |
| Rules beating the composite | 38 of 96 | 0 of those 38 |
| Top decile by first window, mean in the other | -0.98, against -0.45 for the family as a whole | |
| Rank correlation between the two windows | -0.43 | |
Three readings. First, thirty-eight of ninety-six rules beat simply holding the composite in the first window, and not one of those thirty-eight managed it again in the second. Second, the rank correlation between the two windows is minus 0.43, so the first window's ordering did not merely fail to carry over. It pointed the wrong way, and not slightly. The ninety-six rules share most of their signals, so that figure is not ninety-six independent votes, but on balance what the first window rewarded, the second punished. Third, and most awkward, the rules that looked best in the first window did worse in the second than the family average, -0.98 against -0.45, which is what selection produces when whatever it selected on does not last.
The winner's own numbers make the point sharpest. Its first-window Sharpe of 1.21 carries a naive one-sided p-value of 0.12, unremarkable even before the search is counted. But the right null is not a single test. Shuffling the composite's daily returns preserves the return distribution exactly and destroys the time ordering every one of these rules depends on, so the best Sharpe across the family on each shuffle is a draw from the best-of-96 under the null. Two thousand shuffles give a median best of 0.89 and a 95th percentile of 2.19. The observed best of 1.21 sits above the median of pure noise but well inside its range: 32 percent of the shuffles produced a best at least that high, a permutation p-value of 0.32.
Put plainly: shuffled data, which holds no time structure for any rule to use, produces a best at least as good from the same search about one time in three, and the winner then lost money in the held-out window. The split was also re-run at 50 and 70 percent. The rank correlation stayed negative at all three, at minus 0.27 and minus 0.05; at no split did a single first-window winner beat the composite again; and the permutation p-value never fell below 0.24. The window is short at 610 trading days and nothing here settles whether crossover rules work. What it does show is the selection arithmetic operating on real Indian prices exactly as the simulation says it must.
A correction applied after the results is not a correction
This is the failure mode general explanations leave out, and it defeats researchers who have understood everything above.
The obvious remedy looks like this: run the two hundred rules, see which won, then apply a Bonferroni threshold for two hundred tests. It has the shape of the right procedure and it does not work, for three separate reasons.
The first is that the count has to include the tests you would have run. If the first two hundred had all failed you would have tried a different family, and a different one after that. Your search stopped because something worked, which means the denominator is not two hundred but the unknowable size of the search you were prepared to conduct.
The second is that the family itself was chosen after looking. Correcting for ninety-six crossover rules while quietly ignoring that crossovers were picked over breakout rules, mean-reversion rules and volatility filters you had already glanced at counts ninety-six when the real number is far larger. The correction is applied to the last stage of a search whose earlier stages were also data driven.
The third is the one that cannot be engineered around. Once you have seen which rule won, every subsequent choice is made by someone who knows the answer. Where the hold-out boundary falls, which metric is reported, what cost assumption is used, whether an outlier day is a data error or a real event: all of these are now decided by a person with an interest in one outcome. This is not dishonesty. It is how attention works, and the only defence is that the choices were made before the results existed.
The same logic governs the hold-out period, which is why holding data back is weaker protection than it seems. A hold-out works exactly once. Consult it, see the rule fail, adjust the rule and consult it again, and it has become part of the search. The second look is in-sample with extra steps. A hold-out is a consumable, and the discipline that makes it worth anything is writing down in advance what result would make you abandon the rule, because that sentence is impossible to write honestly afterwards.
What verification reaches, and what it cannot
India now has real machinery for checking performance claims, and as of this year it is running. Under a framework operationalised by a SEBI circular dated 29 April 2026, a recognised credit rating agency was designated as the first Past Risk and Return Verification Agency, with the National Stock Exchange acting as the data centre behind it, and services went live on 4 May 2026 after a pilot. Registered investment advisers, research analysts and algorithmic trading service providers who want to communicate past performance enrol with it, and the enrolment window, first set at 3 August 2026, was extended to 3 September 2026. Certified performance data from before the framework may be communicated only up to 3 May 2028, after which verified metrics are the only ones permitted.
One detail in that framework is more statistically serious than it looks. Verification runs prospectively, from the date an entity opts in. It does not retro-validate a history someone arrives holding. That is pre-registration in all but name, and it is the correct design, because a record that accumulates after you have declared yourself cannot have been selected from a search of records.
But notice precisely what verification establishes and what it cannot. It establishes that the numbers are real: these recommendations were made, on these dates, and this is what followed. It says nothing about how many candidate approaches were examined before the one on display was adopted, because that quantity leaves no audit trail anywhere. Nobody can verify the ideas you discarded. A verified track record with an unstated trial count is an accurate number about an unmeasured search.
The contrast with a neighbouring corner of the same market is instructive. Fund houses running mid-cap and small-cap schemes have published standardised liquidity stress test results every fortnight since 15 March 2024, on a defined methodology, covering a named risk. That disclosure works because the quantity is observable and the method can be fixed in advance. There is no equivalent for a published backtest, and the reason is not regulatory neglect. The disclosure that would matter is the trial count, and the trial count is not observable by anyone except the researcher, who usually does not know it either.
The practical consequence for a reader in India is narrow and useful. The framework covers registered intermediaries communicating performance. The great bulk of strategy material, courses, screenshots of equity curves, template backtests and channel content, sits entirely outside it. For all of that, the only correction available is the one the reader applies: treat every result as the maximum of an unstated number of attempts, and ask for the number.
The four honest fixes, and what each one costs
None of these makes a weak rule work. Each of them makes it harder to mistake a search artefact for a discovery, and each has a price that explains why they are rare.
| Fix | What it removes | What it costs |
|---|---|---|
| Pre-register the hypothesis | The ability to choose the question after seeing the answer | You must commit before you know anything, and most committed hypotheses fail |
| Count and report N | The reader's assumption that N was 1 | Your result looks far weaker, correctly, and looks weaker than everyone else's uncounted one |
| Hold out data and consult it once | Iterative fitting to the whole sample | You lose that data from the search, and you get exactly one attempt |
| Test the winner against a shuffled null | The single-test p-value that ignores the contest | Most winners stop being significant, including ones you liked |
The second row is the one that fails socially rather than technically. A researcher who reports having tested three thousand candidates and found one that survives correction has produced better work than a researcher who reports one strategy with a clean-looking curve, and will be read as having produced worse work. That asymmetry is why trial counts stay unpublished, and it is a reason to read an unstated count as a large one rather than a small one.
Walk-forward testing deserves a note because it is often offered as a complete answer. Re-fitting parameters on a rolling window and evaluating on the next does remove one kind of fitting, and it is a genuine improvement. But the rule family is still chosen once, before any of the folds, and usually chosen after looking at the data. Walk-forward protects the parameters and leaves the larger selection untouched.
What survives
Very little, and that is the finding rather than a failure of the method. The honest position after this arithmetic is that most of what a retail researcher can establish from the history available is negative: this rule cannot be distinguished from noise on the data I have. That verdict feels like nothing and is worth a great deal, because the alternative is allocating capital to the winner of an uncounted contest.
Three questions do the bulk of the work. How many candidates were examined, including the ones abandoned on the evidence. Was the hypothesis written down before the result was known. Would the same conclusion have been reached had the search been stopped one candidate earlier. Any result that cannot answer all three is a question rather than a finding, and treating it as a question is the whole of the discipline.
This is also why research method and judgement cannot be separated in practice. The arithmetic above says what a sample can and cannot support; deciding what to do when the honest answer is "this history cannot tell you" is the part no procedure settles. That combination, a method that will return an uncomfortable verdict and the judgement to act on it, is what the curriculum is built around.
Frequently asked questions
If a strategy beat the index in a backtest, why is that not evidence?
Because the sentence is incomplete until it says how many strategies were tested. Beating a benchmark is a comparison; being selected as the best of a set is a maximum. On simulated data with a true mean of exactly zero, a single three-year candidate clears an annualised Sharpe of 1.0 about 4 percent of the time, while the best of fifty such candidates clears it 88 percent of the time. Nothing about the market changed between those two numbers.
What is the trial count, and why does nobody report it?
It is the number of distinct candidates examined before the reported one was chosen, including every parameter setting tried and abandoned. It goes unreported because discarding a parameter feels like development rather than like a test. A grid of ten fast lengths, ten slow lengths, five stop settings, a filter switch and three instruments is three thousand candidates, and it would normally be described as one strategy.
I only ever built one strategy. Does this apply to me?
Yes, and usually more than to someone running a formal search, because at least a formal search knows its own size. A discretionary research process tries settings, looks at the equity curve, adjusts and tries again. Each of those cycles is a test against the same data. The count is unrecorded rather than small.
What is the difference between family-wise error and false discovery rate?
Family-wise error is the probability that even one thing you accepted is noise. False discovery rate is the expected share of the things you accepted that are noise. Controlling family-wise error at 5 percent gives you a shortlist that is almost certainly clean and almost certainly empty. Controlling false discovery rate at 10 percent gives you a shortlist where roughly one entry in ten is wrong and the rest are worth examining.
Which of the two should a trader use?
Usually false discovery rate, because a trader is not making one irreversible decision. A shortlist of candidate rules is allocated to in small sizes, monitored and cut. In that setting, demanding that not a single false rule ever enters is an expensive guarantee, because the price of it is rejecting almost every real rule as well.
How much data does a real edge need before it can be certified?
More than most people have. Computed from the non-central t distribution, a rule with a true annualised Sharpe of 0.8 needs about 19 years of daily data for a coin-flip chance of clearing a Bonferroni threshold set for 200 candidates, and about 29 years for an 80 percent chance. The same rule, named in advance as a single hypothesis, needs about 4 and about 10. Deciding in advance what you are testing is worth roughly fifteen years of history you cannot obtain.
Is a hold-out period enough on its own?
Only if it is used once. A hold-out consulted a second time has become part of the search, because the adjustment you make after a failure was informed by it. Treat it as a consumable with one use, and write down in advance what result would make you abandon the rule, because that sentence is impossible to write honestly after the result is known.
Can I apply a correction after I have seen which rule won?
Not meaningfully, and this is the failure mode most explanations skip. The count you correct for has to include the tests you would have run had the first batch failed, and once you have stopped searching because something worked, that number is unknowable. Worse, every remaining choice, the hold-out boundary, the metric, the cost assumption, is now being made by someone who already knows the answer.
Does SEBI require anyone to disclose how many strategies were tested?
No. India now has a verification framework for past performance, operational since 4 May 2026, under which a recognised agency verifies risk and return metrics for registered investment advisers, research analysts and algorithmic trading service providers, with the exchange acting as data centre. Verification runs prospectively from the date an entity opts in. It establishes that a track record is real. It does not reach the trial count, and most strategy content sits outside the framework altogether.
What does this leave a retail researcher able to do?
Write the hypothesis down before testing it, count and report every candidate examined, hold out a period and consult it once, and treat any result that has not survived those three things as a question rather than a finding. None of that makes a rule work. It stops a search artefact being mistaken for one, which is the more common outcome by a wide margin.
Simulated figures are from random data with the parameters and seeds stated beside each one and contain no market information of any kind. The 96-rule study is a historical exercise on a non-investable composite with costs excluded, presented to demonstrate selection arithmetic; it is not a strategy, not a recommendation, and no conclusion about future results follows from it. The regulatory position is stated as at 19 September 2026. Verify the current framework, the applicable circulars and the enrolment position directly with SEBI and with a registered intermediary before relying on anything here, and take advice on your own facts.
The figures in the 96-rule study were corrected on 23 September 2026. The first version counted every cached bhavcopy file as a trading session, but when the exchange archive is asked for a date on which there was no session it returns the previous session's file, so 23 files in this window repeated a session already present and the composite carried each of those returns twice. Sessions are now keyed by each file's own trade date, one file per session: the NSE full security bhavcopy files named from 1 April 2024 to 18 September 2026 number 634 and hold 611 distinct sessions, the first of which supplies only the previous day's turnover, leaving 610 daily returns. The Saturday special session of 18 May 2024, which exists only under the file name of 20 May, is kept. The weekend budget-day sessions of 1 February 2025 and 1 February 2026 have no file in the cache and are missing; every return is taken within one file from its own previous close, so none spans two sessions, and a counter that changes its symbol drops out on the day of the change. The correction moved the composite's drift from 8.5 to 3.5 percent a year, the winning rule from 20 and 30 days to 20 and 40, the rank correlation between the windows from minus 0.17 to minus 0.43 and the winner's permutation p-value from 0.67 to 0.32. It made one earlier sentence false, that the p-value never fell below 0.48 at the other splits, and that sentence has been rewritten. The finding that the first window's ranking did not carry over to the second is unchanged.
Ready to go deeper than this article?
Bharath Shiksha is a 90-volume curriculum across 6 stages, from chart reading at ₹14,999 through capital raising, or the full bundle at ₹1,49,999. Counting your own trials, and accepting the verdict when the history cannot answer, is taught as the first stage of research method rather than as an advanced topic.
Take the free diagnostic →