The best of fifty backtested rules is a statistic about your search, not about the market

The short answer

When you test one rule, its result describes the rule. When you test fifty and report the winner, the result describes the maximum of fifty draws, and the maximum of a set behaves nothing like a single draw. On simulated data engineered to contain no edge at all, one three-year candidate clears an annualised Sharpe of 1.0 about 4 percent of the time, while the best of fifty such candidates clears it 88 percent of the time. So the claim that a rule beat a benchmark carries no information until the number of candidates searched is stated alongside it, and that number is almost never recorded, including by the person who did the searching.

This is the failure that quietly invalidates most retail strategy research, and it is not a matter of care or of computing power. A researcher who is scrupulous about survivorship, about transaction costs and about look-ahead can still produce a result that means nothing, because the defect enters through the act of choosing rather than through the act of testing.

Everything below is computed rather than asserted. Two of the three instruments are simulations on random data, which is stated at each one, and the third runs a real rule family on a series built from the exchange archive. The parameters are given so the numbers can be reproduced.

A backtest result is a statistic about a search

Consider what a performance figure is supposed to answer. Run one rule, get an annualised Sharpe, and the honest question is: if this rule had no edge, how often would three years of data flatter it this much? That is a well posed question with a known answer.

Now run fifty rules and report the best one. The question the number answers has silently changed. It is now: if none of these fifty had an edge, how often would the best of fifty look this good? Those two questions have very different answers, and the gap between them widens with every candidate added.

The reported figure did not lie. It simply stopped being about the rule and became a statistic about the search that produced it. Nothing on the page tells the reader which of the two questions was answered, because the only thing that distinguishes them is a number that was never written down.

The distribution that matters is the distribution of the maximum

The mechanism is easiest to see where the answer is known in advance, so here it is built deliberately. Generate N return series of 756 trading days each, drawn independently with a true mean of exactly zero. There is no edge anywhere in this experiment by construction. Compute each candidate's annualised Sharpe, keep the highest, and repeat two hundred thousand times.

The best annualised Sharpe found rises with the number of rules tested, on data with no edge at all Measured bars for the median best annualised Sharpe among N candidate strategies whose returns are pure noise with a true mean of exactly zero, over three years of daily data, from two hundred thousand repetitions. One candidate gives a median of zero. Fifty candidates give one point two seven. A thousand candidates give one point eight five. The dashed line above each bar marks the ninetieth percentile. The percentage below each bar is how often the best of that many clears an annualised Sharpe of one. 1.0 2.0 0 0.0014.1%0.65519.1%1.052057.4%1.275088.2%1.5620099.98%1.851000100% Number of candidate rules tested, and how often the best of them clears an annualised Sharpe of 1.0 Every series here has a true mean of exactly zero. None of them has any edge. Simulation on random data, 756 trading days per candidate, 200,000 repetitions, seed 20260919
Computed, not illustrative. The only thing that changes across the six groups is how many candidates were searched. The market is identical in all of them, and it has nothing in it.
The best of N candidates, on simulated data with a true mean of exactly zero. 756 trading days per candidate, 200,000 repetitions, seed 20260919.
Candidates searchedMedian best Sharpe90th percentileBest clears 1.0Best clears 1.5
10.000.744.1%0.5%
50.651.1719.1%2.3%
201.051.4857.4%9.1%
501.271.6688.2%21.2%
2001.561.9099.98%61.7%
10001.852.15100%99.2%

Read the first row and the fourth together. A single candidate reaches an annualised Sharpe of 1.0 in 4.1 percent of runs, which is roughly the rarity a naive reader assumes when they see that figure. The best of fifty reaches it in 88.2 percent of runs. A number that would have been surprising becomes the expected outcome, and the market did not participate in the change.

The same arithmetic runs through the significance test. With one true-null test at a 5 percent threshold you get a false positive 5 percent of the time. With fifty independent ones the probability that at least one clears the bar is 92.3 percent. With two hundred it is 99.997 percent. So a winner described as significant at 5 percent, drawn from a search of fifty, is not evidence of anything. It is what the search was always going to produce.

Two instruments were used here and they were made to agree. The main figures come from the exact sampling distribution of the Sharpe on an independent sample, and a slower brute force pass that simulates actual return paths reproduces them: at two hundred candidates the exact method gives a median best of 1.564 and the brute force pass gives 1.565. An identity that is asserted rather than checked is still an assertion.

The trial count you never wrote down

The natural defence is that this describes a factory running mass searches, not a person with one idea. It does not. A single idea is a grid, and the grid is enumerated during development whether or not anyone counts it.

One strategy is a grid of candidates, and only the survivor is written down A single idea branches into a grid of parameter choices: ten fast lengths, ten slow lengths, five stop settings, a filter on or off, and three instruments. The product is three thousand candidate rules. All but one branch is discarded during development and never recorded, so the trial count reported afterwards is one. One idea fast length10slow length10stop setting5filter2instrument3 = candidates 3,000 2,999 discarded during development Never logged, because rejecting them felt like building, not testing 1 written up reported as if N were 1 The arithmetic does not care what you called the discards Every parameter you looked at and rejected on the evidence is a test that entered the count
An example grid. The product is exact arithmetic on those five choices. Most researchers would describe the result as one strategy, and the number they would quote is the maximum over three thousand.

Take a crossover rule. You try a few fast lengths and settle on one. You try slow lengths, discard the ones whose equity curve looks wrong, keep the survivor. You test a stop, then a wider stop. You add a volatility filter, look at the result, and keep the filter. You check the rule on a second and a third instrument. That is a grid of three thousand, and the write-up will describe it as one strategy with one set of parameters.

What makes this so resistant to correction is that the discards do not feel like tests. Rejecting a parameter because the curve looked poor feels like craftsmanship. But the criterion for rejection was the data, which makes it a test, and the arithmetic of the maximum has no interest in what the researcher called it.

There is a second, larger search most people never count at all: other people's. A rule adopted because somebody published it carries their entire trial count, invisibly. If a hundred people each test a hundred rules on the same market history and only the winners get published, a reader who takes a published rule at face value has inherited an effective count in the thousands while believing it to be one.

Two different things to be wrong about

Once the count is acknowledged, the question becomes what to do about it, and the literature offers two different objectives that are often confused.

Family-wise error rate is the probability that even one of the things you accepted is noise. Controlling it at 5 percent means there is only a one in twenty chance that even one entry on your shortlist is a false discovery. The Bonferroni correction achieves it by dividing the threshold by the number of tests, which for two hundred candidates means demanding a p-value below 0.00025.

False discovery rate is the expected proportion of your accepted set that is noise. Controlling it at 10 percent means accepting that roughly one entry in ten is wrong, in exchange for a shortlist that contains something. The Benjamini-Hochberg procedure achieves it by ranking p-values and comparing each to a threshold that rises with its rank.

Family-wise error and false discovery rate ask different questions of the same shortlist Two panels over the same shortlist of selected rules. On the left, family-wise error asks whether even one selected rule is noise, and controlling it at five percent means the whole shortlist is almost certainly clean but almost empty. On the right, false discovery rate asks what proportion of the shortlist is noise, and controlling it at ten percent accepts that roughly one in ten entries is wrong in exchange for a shortlist that contains something. Family-wise error False discovery rate Is ANY of these wrong? WHAT SHARE of these is wrong? Shortlist is near certainly clean and it is near certainly empty 0.16 of 10 real edges recovered About one entry in ten is noise and the shortlist is not empty 0.38 of 10 real edges recovered Right when one wrong answer is catastrophic and final one irreversible decision Right when the wrong ones are sized small, watched and cut which is what a trader is doing Neither is the careful answer and the other the sloppy one. They answer different questions. Figures from a simulation of 200 candidate rules of which 10 carry a true annualised Sharpe of 0.8, three years of daily data each, 2,000 repetitions, seed 20260920
Computed, not illustrative. At three years of data both corrections recover almost nothing, which is the sample speaking rather than the method failing.

Here is the comparison run on a universe where the truth is known. Two hundred candidate rules, of which exactly ten carry a genuine annualised Sharpe of 0.8 and the other one hundred and ninety are pure noise, each with three years of daily data. The planted count is the ground truth no live research ever has, which is precisely why the comparison has to be made in simulation.

What each rule recovers from a planted universe. 200 candidates, 10 with a true annualised Sharpe of 0.8, 756 trading days, 2,000 repetitions, seed 20260920. Simulated data.
Selection ruleReal edges found, of 10Noise acceptedShare of the shortlist that is noise
No correction, 5 percent3.919.5570.4%
Bonferroni, 5 percent family-wise0.160.054.7%
Benjamini-Hochberg, 10 percent false discovery0.380.199.7%

The first row is the one worth sitting with. Applying an uncorrected 5 percent threshold to two hundred candidates produces a shortlist that is seventy percent noise, even though ten genuinely good rules were present and waiting to be found. The procedure is not failing to find them. It is drowning them.

The third row shows the false discovery rate procedure delivering what it advertised: a realised noise share of 9.7 percent against a target of 10. The second row shows Bonferroni delivering what it advertised too, a shortlist that is 95 percent clean, at the cost of surfacing one of the ten real edges only about once in every six searches.

Which one a trader wants follows from what a shortlist is for. A regulator approving a drug is making one irreversible decision where a single false positive is a disaster, and family-wise control is right. A trader is assembling a set of candidate rules to allocate to in small sizes, watch, and cut. Insisting that not a single false rule ever enters that set is an expensive guarantee, and the price is refusing almost every real one as well.

What the correction costs, priced in years of data

Both procedures in that table recovered almost nothing, and the reason is not the procedures. It is that three years of daily data cannot resolve a Sharpe of 0.8 against a bar set for two hundred candidates. The honest way to express the cost of a correction is therefore not in statistical power but in years of history.

Years of daily data before a genuinely good rule can clear the bar. Computed from the non-central t distribution, one-sided, at a Bonferroni threshold of 5 percent divided by the number of candidates. Shown as years for a 50 percent chance, then for an 80 percent chance.
True annualised SharpeOne test named in advanceBest of 20 candidatesBest of 200 candidates
0.510.8 then 24.731.5 then 53.348.5 then 74.8
0.84.2 then 9.712.3 then 20.819.0 then 29.2
1.02.7 then 6.27.9 then 13.312.1 then 18.7
1.51.2 then 2.83.5 then 5.95.4 then 8.3

The closed form was checked against a simulation at the 0.8 and two hundred candidate cell: the analytic power at 19.0 years is 0.502 and twenty thousand simulated samples give 0.503.

Read across the 0.8 row. A rule that genuinely carries an annualised Sharpe of 0.8, named in advance as a single hypothesis, needs about four years of daily data for a coin-flip chance of clearing the bar. The identical rule, arrived at as the winner of a two hundred candidate search, needs about nineteen. The difference is not a matter of statistical taste. It is fifteen years of Indian market history that cannot be bought, borrowed or waited for.

That reframes pre-registration entirely. Writing the hypothesis down before testing it is not a bureaucratic courtesy borrowed from academia. It is the single largest lever available to a researcher with a finite amount of history, and on these numbers it is worth more than every other methodological improvement combined.

Read the 0.5 row for the harder message. A true annualised Sharpe of 0.5 is a respectable rule, and as the winner of a two hundred candidate search it needs roughly half a century of daily data before it can be distinguished from the noise it was selected against. For a great many real rules, the honest conclusion is not that they fail but that the available history cannot answer the question, and a research process that never returns that verdict is not testing anything.

Ninety-six rules on a real Indian series

Simulations settle the mechanism. The next question is what the mechanism does to a real search on real Indian data, so here is one, run end to end.

The series is a composite built from the exchange security archive: for every trading day, the two hundred counters with the highest turnover on the previous day, equally weighted. Ranking on the same day's turnover is a look-ahead, because heavy volume arrives alongside large moves, and the first version built that way produced an annualised drift above 150 percent, which is how the bug announced itself. Membership is recomputed daily from each day's own file, so nothing has to survive the window and no survivorship filter is applied. It is a breadth composite of liquid counters, not an index, and it is not investable. Over the window it carried an annualised drift of 3.5 percent and annualised volatility of 18.6 percent.

The rule family is every moving-average crossover from ten fast lengths and ten slow lengths where the fast is shorter, which is 96 rules. Each rule holds the composite when the fast average is above the slow one and holds nothing otherwise, acting on the following day's return so no rule can trade on information it did not have. Costs are excluded, which biases every rule upward. The split was fixed at 60 percent before the data was loaded, giving 246 in-sample days and then 164 held-out days beginning 20 January 2026.

96 crossover rules on the composite. NSE security bhavcopy, 2 April 2024 to 18 September 2026, 610 trading days of which the first 200 are warm-up. Annualised Sharpe, costs excluded. Not a recommendation and not a forecast.
 First windowHeld-out window
Best rule by first-window Sharpe1.21-0.98
Median rule in the family-0.21-0.39
Simply holding the composite-0.110.53
Rules beating the composite38 of 960 of those 38
Top decile by first window, mean in the other-0.98, against -0.45 for the family as a whole
Rank correlation between the two windows-0.43
In-sample rank against holdout result for 96 rules on a real Indian series A scatter of ninety-six moving-average crossover rules run on a composite of the two hundred most traded counters on the National Stock Exchange. The horizontal axis is the annualised Sharpe each rule achieved in the first window and the vertical axis is what the same rule achieved in the held-out window. The cloud tilts downward. Every rule that scored above 0.4 in the first window, marked in red on the right, sits below the zero line in the holdout. holding the composite the winner 0 +0.5 -1.0 -2.0 Holdout Sharpe Annualised Sharpe in the first window, which is the number a backtest report would show Rank correlation between the two axes: minus 0.43. The in-sample ordering pointed the wrong way. NSE security bhavcopy, 2 April 2024 to 18 September 2026, 246 in-sample and 164 holdout trading days, costs excluded
Measured on real data, not illustrative. If the first window's ranking carried over to the second, this cloud would tilt upward from left to right. It tilts the other way, and every rule marked in red, the ones that scored highest first, lost money in the held-out window.

Three readings. First, thirty-eight of ninety-six rules beat simply holding the composite in the first window, and not one of those thirty-eight managed it again in the second. Second, the rank correlation between the two windows is minus 0.43, so the first window's ordering did not merely fail to carry over. It pointed the wrong way, and not slightly. The ninety-six rules share most of their signals, so that figure is not ninety-six independent votes, but on balance what the first window rewarded, the second punished. Third, and most awkward, the rules that looked best in the first window did worse in the second than the family average, -0.98 against -0.45, which is what selection produces when whatever it selected on does not last.

The winner's own numbers make the point sharpest. Its first-window Sharpe of 1.21 carries a naive one-sided p-value of 0.12, unremarkable even before the search is counted. But the right null is not a single test. Shuffling the composite's daily returns preserves the return distribution exactly and destroys the time ordering every one of these rules depends on, so the best Sharpe across the family on each shuffle is a draw from the best-of-96 under the null. Two thousand shuffles give a median best of 0.89 and a 95th percentile of 2.19. The observed best of 1.21 sits above the median of pure noise but well inside its range: 32 percent of the shuffles produced a best at least that high, a permutation p-value of 0.32.

Put plainly: shuffled data, which holds no time structure for any rule to use, produces a best at least as good from the same search about one time in three, and the winner then lost money in the held-out window. The split was also re-run at 50 and 70 percent. The rank correlation stayed negative at all three, at minus 0.27 and minus 0.05; at no split did a single first-window winner beat the composite again; and the permutation p-value never fell below 0.24. The window is short at 610 trading days and nothing here settles whether crossover rules work. What it does show is the selection arithmetic operating on real Indian prices exactly as the simulation says it must.

A correction applied after the results is not a correction

This is the failure mode general explanations leave out, and it defeats researchers who have understood everything above.

The obvious remedy looks like this: run the two hundred rules, see which won, then apply a Bonferroni threshold for two hundred tests. It has the shape of the right procedure and it does not work, for three separate reasons.

The first is that the count has to include the tests you would have run. If the first two hundred had all failed you would have tried a different family, and a different one after that. Your search stopped because something worked, which means the denominator is not two hundred but the unknowable size of the search you were prepared to conduct.

The second is that the family itself was chosen after looking. Correcting for ninety-six crossover rules while quietly ignoring that crossovers were picked over breakout rules, mean-reversion rules and volatility filters you had already glanced at counts ninety-six when the real number is far larger. The correction is applied to the last stage of a search whose earlier stages were also data driven.

The third is the one that cannot be engineered around. Once you have seen which rule won, every subsequent choice is made by someone who knows the answer. Where the hold-out boundary falls, which metric is reported, what cost assumption is used, whether an outlier day is a data error or a real event: all of these are now decided by a person with an interest in one outcome. This is not dishonesty. It is how attention works, and the only defence is that the choices were made before the results existed.

The same logic governs the hold-out period, which is why holding data back is weaker protection than it seems. A hold-out works exactly once. Consult it, see the rule fail, adjust the rule and consult it again, and it has become part of the search. The second look is in-sample with extra steps. A hold-out is a consumable, and the discipline that makes it worth anything is writing down in advance what result would make you abandon the rule, because that sentence is impossible to write honestly afterwards.

What verification reaches, and what it cannot

India now has real machinery for checking performance claims, and as of this year it is running. Under a framework operationalised by a SEBI circular dated 29 April 2026, a recognised credit rating agency was designated as the first Past Risk and Return Verification Agency, with the National Stock Exchange acting as the data centre behind it, and services went live on 4 May 2026 after a pilot. Registered investment advisers, research analysts and algorithmic trading service providers who want to communicate past performance enrol with it, and the enrolment window, first set at 3 August 2026, was extended to 3 September 2026. Certified performance data from before the framework may be communicated only up to 3 May 2028, after which verified metrics are the only ones permitted.

One detail in that framework is more statistically serious than it looks. Verification runs prospectively, from the date an entity opts in. It does not retro-validate a history someone arrives holding. That is pre-registration in all but name, and it is the correct design, because a record that accumulates after you have declared yourself cannot have been selected from a search of records.

But notice precisely what verification establishes and what it cannot. It establishes that the numbers are real: these recommendations were made, on these dates, and this is what followed. It says nothing about how many candidate approaches were examined before the one on display was adopted, because that quantity leaves no audit trail anywhere. Nobody can verify the ideas you discarded. A verified track record with an unstated trial count is an accurate number about an unmeasured search.

The contrast with a neighbouring corner of the same market is instructive. Fund houses running mid-cap and small-cap schemes have published standardised liquidity stress test results every fortnight since 15 March 2024, on a defined methodology, covering a named risk. That disclosure works because the quantity is observable and the method can be fixed in advance. There is no equivalent for a published backtest, and the reason is not regulatory neglect. The disclosure that would matter is the trial count, and the trial count is not observable by anyone except the researcher, who usually does not know it either.

The practical consequence for a reader in India is narrow and useful. The framework covers registered intermediaries communicating performance. The great bulk of strategy material, courses, screenshots of equity curves, template backtests and channel content, sits entirely outside it. For all of that, the only correction available is the one the reader applies: treat every result as the maximum of an unstated number of attempts, and ask for the number.

The four honest fixes, and what each one costs

None of these makes a weak rule work. Each of them makes it harder to mistake a search artefact for a discovery, and each has a price that explains why they are rare.

What each correction buys and what it costs
FixWhat it removesWhat it costs
Pre-register the hypothesisThe ability to choose the question after seeing the answerYou must commit before you know anything, and most committed hypotheses fail
Count and report NThe reader's assumption that N was 1Your result looks far weaker, correctly, and looks weaker than everyone else's uncounted one
Hold out data and consult it onceIterative fitting to the whole sampleYou lose that data from the search, and you get exactly one attempt
Test the winner against a shuffled nullThe single-test p-value that ignores the contestMost winners stop being significant, including ones you liked

The second row is the one that fails socially rather than technically. A researcher who reports having tested three thousand candidates and found one that survives correction has produced better work than a researcher who reports one strategy with a clean-looking curve, and will be read as having produced worse work. That asymmetry is why trial counts stay unpublished, and it is a reason to read an unstated count as a large one rather than a small one.

Walk-forward testing deserves a note because it is often offered as a complete answer. Re-fitting parameters on a rolling window and evaluating on the next does remove one kind of fitting, and it is a genuine improvement. But the rule family is still chosen once, before any of the folds, and usually chosen after looking at the data. Walk-forward protects the parameters and leaves the larger selection untouched.

What survives

Very little, and that is the finding rather than a failure of the method. The honest position after this arithmetic is that most of what a retail researcher can establish from the history available is negative: this rule cannot be distinguished from noise on the data I have. That verdict feels like nothing and is worth a great deal, because the alternative is allocating capital to the winner of an uncounted contest.

Three questions do the bulk of the work. How many candidates were examined, including the ones abandoned on the evidence. Was the hypothesis written down before the result was known. Would the same conclusion have been reached had the search been stopped one candidate earlier. Any result that cannot answer all three is a question rather than a finding, and treating it as a question is the whole of the discipline.

This is also why research method and judgement cannot be separated in practice. The arithmetic above says what a sample can and cannot support; deciding what to do when the honest answer is "this history cannot tell you" is the part no procedure settles. That combination, a method that will return an uncomfortable verdict and the judgement to act on it, is what the curriculum is built around.

Frequently asked questions

If a strategy beat the index in a backtest, why is that not evidence?

Because the sentence is incomplete until it says how many strategies were tested. Beating a benchmark is a comparison; being selected as the best of a set is a maximum. On simulated data with a true mean of exactly zero, a single three-year candidate clears an annualised Sharpe of 1.0 about 4 percent of the time, while the best of fifty such candidates clears it 88 percent of the time. Nothing about the market changed between those two numbers.

What is the trial count, and why does nobody report it?

It is the number of distinct candidates examined before the reported one was chosen, including every parameter setting tried and abandoned. It goes unreported because discarding a parameter feels like development rather than like a test. A grid of ten fast lengths, ten slow lengths, five stop settings, a filter switch and three instruments is three thousand candidates, and it would normally be described as one strategy.

I only ever built one strategy. Does this apply to me?

Yes, and usually more than to someone running a formal search, because at least a formal search knows its own size. A discretionary research process tries settings, looks at the equity curve, adjusts and tries again. Each of those cycles is a test against the same data. The count is unrecorded rather than small.

What is the difference between family-wise error and false discovery rate?

Family-wise error is the probability that even one thing you accepted is noise. False discovery rate is the expected share of the things you accepted that are noise. Controlling family-wise error at 5 percent gives you a shortlist that is almost certainly clean and almost certainly empty. Controlling false discovery rate at 10 percent gives you a shortlist where roughly one entry in ten is wrong and the rest are worth examining.

Which of the two should a trader use?

Usually false discovery rate, because a trader is not making one irreversible decision. A shortlist of candidate rules is allocated to in small sizes, monitored and cut. In that setting, demanding that not a single false rule ever enters is an expensive guarantee, because the price of it is rejecting almost every real rule as well.

How much data does a real edge need before it can be certified?

More than most people have. Computed from the non-central t distribution, a rule with a true annualised Sharpe of 0.8 needs about 19 years of daily data for a coin-flip chance of clearing a Bonferroni threshold set for 200 candidates, and about 29 years for an 80 percent chance. The same rule, named in advance as a single hypothesis, needs about 4 and about 10. Deciding in advance what you are testing is worth roughly fifteen years of history you cannot obtain.

Is a hold-out period enough on its own?

Only if it is used once. A hold-out consulted a second time has become part of the search, because the adjustment you make after a failure was informed by it. Treat it as a consumable with one use, and write down in advance what result would make you abandon the rule, because that sentence is impossible to write honestly after the result is known.

Can I apply a correction after I have seen which rule won?

Not meaningfully, and this is the failure mode most explanations skip. The count you correct for has to include the tests you would have run had the first batch failed, and once you have stopped searching because something worked, that number is unknowable. Worse, every remaining choice, the hold-out boundary, the metric, the cost assumption, is now being made by someone who already knows the answer.

Does SEBI require anyone to disclose how many strategies were tested?

No. India now has a verification framework for past performance, operational since 4 May 2026, under which a recognised agency verifies risk and return metrics for registered investment advisers, research analysts and algorithmic trading service providers, with the exchange acting as data centre. Verification runs prospectively from the date an entity opts in. It establishes that a track record is real. It does not reach the trial count, and most strategy content sits outside the framework altogether.

What does this leave a retail researcher able to do?

Write the hypothesis down before testing it, count and report every candidate examined, hold out a period and consult it once, and treat any result that has not survived those three things as a question rather than a finding. None of that makes a rule work. It stops a search artefact being mistaken for one, which is the more common outcome by a wide margin.

Simulated figures are from random data with the parameters and seeds stated beside each one and contain no market information of any kind. The 96-rule study is a historical exercise on a non-investable composite with costs excluded, presented to demonstrate selection arithmetic; it is not a strategy, not a recommendation, and no conclusion about future results follows from it. The regulatory position is stated as at 19 September 2026. Verify the current framework, the applicable circulars and the enrolment position directly with SEBI and with a registered intermediary before relying on anything here, and take advice on your own facts.

The figures in the 96-rule study were corrected on 23 September 2026. The first version counted every cached bhavcopy file as a trading session, but when the exchange archive is asked for a date on which there was no session it returns the previous session's file, so 23 files in this window repeated a session already present and the composite carried each of those returns twice. Sessions are now keyed by each file's own trade date, one file per session: the NSE full security bhavcopy files named from 1 April 2024 to 18 September 2026 number 634 and hold 611 distinct sessions, the first of which supplies only the previous day's turnover, leaving 610 daily returns. The Saturday special session of 18 May 2024, which exists only under the file name of 20 May, is kept. The weekend budget-day sessions of 1 February 2025 and 1 February 2026 have no file in the cache and are missing; every return is taken within one file from its own previous close, so none spans two sessions, and a counter that changes its symbol drops out on the day of the change. The correction moved the composite's drift from 8.5 to 3.5 percent a year, the winning rule from 20 and 30 days to 20 and 40, the rank correlation between the windows from minus 0.17 to minus 0.43 and the winner's permutation p-value from 0.67 to 0.32. It made one earlier sentence false, that the p-value never fell below 0.48 at the other splits, and that sentence has been rewritten. The finding that the first window's ranking did not carry over to the second is unchanged.

Related guides

The backtesting mistakes that survive every rewrite

Read →

Building a first trading system, in order

Read →

Regime detection in Indian markets

Read →

Ready to go deeper than this article?

Bharath Shiksha is a 90-volume curriculum across 6 stages, from chart reading at ₹14,999 through capital raising, or the full bundle at ₹1,49,999. Counting your own trials, and accepting the verdict when the history cannot answer, is taught as the first stage of research method rather than as an advanced topic.

Take the free diagnostic →