A Sharpe ratio that came out of a search is not a Sharpe ratio
The short answer
A Sharpe ratio computed from data is an estimate with a standard error, governed first by sample length and second by the skew and fat tails of the returns. The deflated Sharpe ratio changes the question from is this above zero to is this above what the best of N no-skill attempts would have produced anyway. Measured on 495 sessions of real NSE data, a search over 170 rules returned a best annualised Sharpe of 1.71. As a single test that is a real result: 2.54 standard errors from zero, with a 95 percent interval of 0.39 to 3.03. Re-running that entire search 2,000 times on the same returns with the time order destroyed, the best rule averaged 1.16 and matched or beat 1.71 in 8.1 percent of runs. Against a null that keeps ten-day blocks intact it did so in 36.4 percent. Deflated for the 170 rules actually run, the probability that its true Sharpe clears the hurdle is 0.79, short of the 0.95 a conventional test asks for, and at about 17,400 trials it is a coin toss. Deflation needs an honest N and nobody has one, so it does not certify anything. It is a destructive test, and it is very good at that.
Two different objects are routinely called a Sharpe ratio. The first is a property of a strategy, which is unobservable. The second is a number computed from a finite sample of returns, which is an estimate of the first and carries all the sampling noise that implies. Reporting the second and discussing the first is the single most common error in quantitative trading, and it survives because the estimate has no error bar printed next to it.
There is a second, larger problem sitting on top of that one. When a number is the best of many that were computed, it is not an estimate of anything. It is the maximum of a set of noisy draws, and the maximum of a set has its own distribution that sits well to the right of the thing being estimated. The deflated Sharpe ratio is the correction for both problems at once, and the arithmetic is simple enough to do on the inputs below.
The Sharpe ratio is an estimate, and it arrives with an error bar
Write the estimate as the sample mean of returns divided by their sample standard deviation, both measured per period, then multiply by the square root of periods per year to annualise. Both inputs are statistics. The ratio inherits their noise.
For returns that are independent and identically distributed and normal, the standard error of the estimate in per-period units is the square root of one plus half the squared per-period Sharpe, all over the number of observations less one. Two things follow immediately. The sample length sits under a square root, so halving the error bar takes four times the data. And the Sharpe itself barely enters, because at daily frequency the per-period Sharpe is a small number and its square is negligible.
That leaves sample length as very nearly the only thing that governs how much a reported Sharpe can be trusted.
The rows above use the measured skew and kurtosis of the strategy return series from this article's search. At one year of daily data, an estimate of 1.71 carries a 95 percent interval running from just below zero to above three and a half. At two years the lower end clears zero, which is roughly where this article's own sample of 495 sessions sits. Getting the annualised standard error down to 0.20, which is still wide enough to be uncomfortable, needs 5,594 sessions, which is about 22.2 years.
That figure is worth holding against what Indian data can actually supply. Index futures started trading in June 2000 and index options in June 2001, so exchange traded Indian index options have existed for about 25 years and index futures for about 26. The entire run of that market, traded without a single change of rule, is about enough to pin one Sharpe estimate to a standard error of 0.20. And that is for one strategy, tried once.
One arithmetic slip is worth naming because it is everywhere. The standard error formula is in per-period units, and the Sharpe that goes into it must be the per-period Sharpe, not the annualised one. Putting an annualised Sharpe of 1.5 into a formula built for daily data inflates the correction term by a factor of 252 and produces intervals that are nonsense. Annualise at the end, by multiplying both the estimate and its standard error by the square root of periods per year.
Skew and fat tails move the error bar, and rarely by as much as you have been told
Trading strategies do not produce normal returns. A trend rule takes many small losses and a few large gains. An option selling book does the reverse. Both make the normal formula wrong, and the correction is a single extra term inside the same square root: one, minus skewness times the per-period Sharpe, plus kurtosis less one over four times the squared per-period Sharpe, all over observations less one. Set skewness to zero and kurtosis to three and the normal case comes straight back out, which is the check that the two formulas are the same object.
The direction is set by the sign of the skew. A negatively skewed strategy with a positive Sharpe gets a wider standard error than the normal formula reports, and excess kurtosis widens it further. So for the strategies most likely to be sold on their Sharpe, short volatility structures and mean reverting books above all, the naive figure understates the uncertainty. A positively skewed trend rule gets a slightly narrower one.
The magnitude is where most published treatments overstate the case. Both correction terms scale with the per-period Sharpe, which is the annualised figure divided by the square root of periods per year. On daily data that is a small number and the correction is usually a rounding error. On monthly or quarterly data it is not.
| Sampling frequency | Sharpe per period | Skew -1.0, kurtosis 10 | Skew -0.5, kurtosis 6 | Normal | Skew +1.0, kurtosis 10 |
|---|---|---|---|---|---|
| Daily | 0.0630 | 1.034 | 1.017 | 1.000 | 0.972 |
| Weekly | 0.1409 | 1.083 | 1.041 | 1.000 | 0.946 |
| Monthly | 0.2887 | 1.190 | 1.095 | 1.000 | 0.929 |
| Quarterly | 0.5000 | 1.354 | 1.179 | 1.000 | 0.972 |
On the series measured for this article the effect is mostly that small, with one instructive exception. The composite's daily skew is -0.22 and its kurtosis 4.64, and the moment adjustment moves the standard error by 0.4 percent. Across all 170 rules the skew ranged from -2.30 to 1.78 and the kurtosis from 5.7 to 33.7, and all but two of them moved their standard error by less than 5 percent. Those two are the two best rules in the search, breakouts whose returns are strongly positively skewed, and on them the adjustment works in the rule's favour: the best rule, with skew of 1.64 and kurtosis of 23.3, gets a standard error 6 percent narrower than the normal formula reports, because its daily Sharpe is high enough for the skew term to bite.
This matters because the non-normality correction is the part that gets the attention, and it is the small part. The large part is next.
The comparison against zero is the wrong comparison
A conventional test asks whether one strategy with no edge could have produced this Sharpe by luck. If the answer is no at the usual threshold, the result is called significant. That test is correct for one strategy examined once, which describes almost no backtest ever run.
What actually happened is that many candidates were computed and the best was kept. The right question is therefore whether the best of N no-skill candidates could have produced this figure, and the answer to that question is different because the maximum of N draws is not a draw. Its expectation grows with N, roughly with the square root of twice the logarithm of N for normal draws, and the practical version of that formula is the one used for deflation: the cross-trial standard deviation of the Sharpe estimates, multiplied by a weighted combination of two normal quantiles at one minus one over N and one minus one over N times e, weighted by the Euler constant of about 0.5772.
| Trials | Expected best Sharpe | Bought by the tenfold increase |
|---|---|---|
| 1 | 0.00 | a single attempt, by definition |
| 10 | 0.67 | plus 0.67 |
| 100 | 1.08 | plus 0.41 |
| 1,000 | 1.39 | plus 0.31 |
| 10,000 | 1.65 | plus 0.26 |
| 100,000 | 1.88 | plus 0.23 |
Read the last column, because it contains the part people get wrong. The climb is logarithmic, so the damage is front loaded. The first tenfold increase, from one honest attempt to ten, buys 0.67 of hurdle out of nothing. The tenfold increase from 10,000 to 100,000 buys 0.23. The defence that only a handful of variations were tried is therefore not the defence it sounds like. A handful is where the hurdle climbs fastest.
What a 170-rule search does to two years of real Indian data
The demonstration below is run on real market data and a simulated null, and the two are kept clearly apart. The data is the NSE full security bhavcopy for 2024-09-18 to 2026-09-18, which gives 495 sessions once the archive's copies of earlier sessions, filed under holiday dates, are set aside. The series is an equal weighted composite of the 100 most traded symbols in the EQ series that were present on every session, with daily moves beyond 25 percent set aside as corporate action artefacts, which removed 19 observations out of 49,500.
It is a real return series with real higher moments, which is all the demonstration needs. It is not an index, it is not a portfolio anybody could hold at those weights without cost, and no costs are charged anywhere in what follows. The financing rate is set to zero, which flatters every figure here: at a stated 6 percent rate against the measured annualised volatility of 17.18 percent, a fully invested series would lose 0.35 of annualised Sharpe.
| Property | Measured | What it governs |
|---|---|---|
| Sessions in the sample | 495 | The standard error of every Sharpe below |
| Constituents | 100 symbols, equal weighted | Most traded by median turnover, present every session |
| Observations set aside | 19 | Daily moves beyond 25 percent, treated as corporate action artefacts |
| Annualised volatility | 17.18 percent | The denominator of every Sharpe here |
| Skewness | -0.2240 | Enters the standard error linearly in the per-period Sharpe |
| Kurtosis, not excess | 4.6425 | Enters the standard error in the square of the per-period Sharpe |
| Holding the composite throughout | Annualised Sharpe 0.56 | The reference every rule below has to beat to be worth anything |
The search is the kind anybody would run: 114 moving average crossover pairs and 56 breakout rules with a separate exit lookback, each long or flat, with the state set on one session trading the next so that nothing sees its own future. That is 170 trials.
| Annualised Sharpe | Note | |
|---|---|---|
| Holding the composite throughout | 0.56 | Total return 17.24 percent over the window |
| Median of the 170 rules | 0.28 | 125 of 170 finished above zero |
| Worst of the 170 rules | -0.77 | The same search, the other tail |
| Best of the 170 rules | 1.71 | In the market 24.0 percent of sessions, 24 completed round trips |
| Rules beating the composite | 41 of 170 | The other 129 did worse than simply holding it |
| Spread of Sharpe across the search | 0.43 | This is the input the expected maximum is computed from |
| Simulated best, order destroyed | mean 1.16, 95th percentile 1.83 | 8.1 percent of no-skill searches beat the real one |
| Simulated best, ten-day blocks kept | mean 1.50, 95th percentile 2.64 | 36.4 percent of no-skill searches beat the real one |
The best rule returned an annualised Sharpe of 1.71 against 0.56 for simply holding the composite. Its standard error is 0.67, so the 95 percent interval runs from 0.39 to 3.03. Against zero it stands 2.54 standard errors clear, which most reports would call a result, and here the interval excludes zero. Taken as the only rule anyone tested, it would be a genuine finding. It was the best of 170.
Costs have not been charged anywhere, and they are not a rounding error on a rule like this one. It is in the market on less than a quarter of sessions, so its annualised volatility is only 5.95 percent. A Sharpe of 1.71 on that volatility is a small return divided by a small risk, which makes it unusually sensitive to anything charged per round trip, and there were 24 of those in 1.96 years.
| Cost per round trip | Annualised Sharpe | Against holding the composite at 0.56 |
|---|---|---|
| 0.00 percent | 1.71 | ahead |
| 0.10 percent | 1.50 | ahead |
| 0.20 percent | 1.30 | ahead |
| 0.40 percent | 0.89 | ahead |
Even at a stated 0.40 percent per round trip the winning rule stays ahead of simply holding the composite, 0.89 against 0.56. What costs erase is its lead over the null. At 0.10 percent it is level with the average best that the same search finds when ten-day blocks are kept intact, 1.50 against 1.50, and at 0.40 percent it falls below the average best the same rules produce on returns with the time order destroyed, 0.89 against 1.16. The cost levels are stated inputs; the Sharpe in each row is computed from the rule's own trade record.
Now the null. The whole search, all 170 rules, was re-run 2,000 times on resampled versions of the same returns. Two nulls were used rather than one, because a single null is an untested instrument. The first shuffles the daily returns into a random order: identical mean, identical volatility, identical skew and kurtosis, and no time structure whatsoever, so no rule can possibly have an edge. Real returns are not independent though, and a trend rule feeds on short run persistence, so a shuffle-only null is generous to the rule. The second draws ten-day blocks with replacement, which keeps momentum and volatility clustering intact inside each block and destroys it across blocks. Neither is the truth. The pair brackets it, and a conclusion has to survive both.
Under the shuffle, the best of 170 no-skill rules averaged 1.16 and reached 1.83 at the 95th percentile. 8.1 percent of those searches produced a best rule at least as good as the real one, so against this null the real result is unusual, better than about 92 searches in 100, but short of the 95 in 100 a conventional test asks for. Under the block resampling the average best was 1.50, below the 1.71 the real search found, but 36.4 percent of no-skill searches matched or beat it. Against the null that keeps short-run persistence, the best rule out of a 170-rule search on two years of real Indian equity data is an ordinary outcome of the identical search run on resampled data.
The session count flatters the sample as well. The winning rule was in the market on 24.0 percent of sessions and produced 24 completed round trips, so 495 rows of data stand behind far fewer independent bets: on a per-trade basis the Sharpe is 0.50 with a standard error of 0.17, and a single best episode accounts for 25 percent of the total log growth. Sample length should be counted in independent bets, not in rows of the data file.
The deflation arithmetic, with every input on the page
The deflated Sharpe ratio is the normal cumulative distribution evaluated at the observed Sharpe less the hurdle, times the square root of observations less one, divided by the square root of the moment adjusted variance term already given. It returns a probability that the true Sharpe exceeds the hurdle, not a p-value against zero. Every input for the rows below is measured: observed per-period Sharpe 0.10761, 495 observations, skewness 1.6428, kurtosis 23.2888.
| Hurdle used | Hurdle, annualised | Distance in standard errors | Deflated Sharpe ratio |
|---|---|---|---|
| Against zero, the usual test | 0.00 | 2.54 | 0.99 |
| Against the formula at N equal to the 170 rules run | 1.16 | 0.81 | 0.79 |
| Against the simulated best-of-search, order destroyed | 1.16 | 0.82 | 0.79 |
| Against the simulated best-of-search, ten-day blocks kept | 1.50 | 0.31 | 0.62 |
| Against the formula at 1,000 trials | 1.39 | 0.47 | 0.68 |
| Against the formula at 10,000 trials | 1.65 | 0.08 | 0.53 |
The observation did not change between the first row and the last. The sample did not change, the estimate did not change, the standard error did not change. Only the reference point moved. A figure that reads as near certain at 0.99 against zero reads as 0.79 against the 170 rules actually run, short of the 0.95 a conventional test would demand, and as a coin toss, 0.53, if the search behind it was really 10,000 trials. At about 17,400 trials it reaches 0.50: the observed Sharpe is then exactly what that many no-skill attempts produce on average, and beyond that count the same number becomes evidence against the rule. That is the whole mechanism of deflation in one table.
One input deserves a note, because it is the one most often taken from thin air. The spread that goes into the expected maximum is the cross-trial standard deviation of the Sharpe estimates from the search itself, measured here at 0.43 annualised. That is the right quantity because it already contains the sample length, the correlation between the rules and the shape of the returns as they actually are, none of which a textbook value would know. Two things break it. Trials run over samples of different lengths are not comparable draws. And a search in which some trials genuinely do have an edge inflates the spread and therefore the hurdle, which is one more reason to compute the null by re-running the search rather than by reading a formula.
The same inputs answer a more useful question than the probability does: how long a sample would have to be for this result to clear its own hurdle at the 95 percent level. Against zero, the observed Sharpe needs about 208 sessions, under a year, which this sample already has. Against the hurdle for 170 trials it needs about 2,000 sessions, roughly 8 years, four times the history it has. Against the hurdle for 1,000 trials it needs about 6,100 sessions, or 24 years, roughly the whole history of exchange traded Indian index options, and against 10,000 trials about 186,000 sessions, on the order of 740 years. Run the question the other way and the sample it actually has clears the 95 percent bar only against a hurdle of 0.60 or less, which is what the best of about seven no-skill trials is expected to produce. The grid alone was 170. The price of having searched is paid in sample length, and past about a thousand trials it is not a price anybody can pay with more data.
Notice also that the formula and the simulation agree without being told to. The analytic expected maximum at N of 170 is 1.160; the simulated best-of-search with the order destroyed averaged 1.156. Two instruments, one built from a closed form and one from 2,000 re-runs, landing within 0.0036 of each other, less than the 0.0085 standard error of the simulated average itself. The agreement is closer than it has any right to be, and why is the subject of the next section.
Nobody has an honest N, and the errors run in both directions
Deflation needs the number of trials. That number is unknowable, and it is unknowable in two opposite ways at once.
It is too small because the grid is not the search. Every variant coded and abandoned, every re-run after a bug that changed the answer, every parameter nudged because the equity curve looked wrong, every idea taken from a source that had itself searched before publishing: all of it belongs in N, and almost none of it is written down. A trader who reports a grid of 170 and has been working on the problem for a year has run thousands of trials, and the difference between 1.16 and 1.65 on that hurdle takes this article's best rule from a deflated Sharpe ratio of 0.79 to 0.53, from weak evidence to a coin toss.
| Source of trials | Belongs in N | Usually recorded |
|---|---|---|
| The parameter grid in the final run | Yes | Yes, it is the only one people count |
| Variants coded and deleted before that run | Yes | Almost never |
| Re-runs after fixing a bug that changed the result | Yes | No |
| Parameters nudged by eye between runs | Yes, one each | No |
| Filters and exits swapped in and out | Yes | No |
| Rules inherited from published material | Yes, plus the search behind them | Not even in principle |
| Rules the same community already tested on the same series | Yes | Not knowable |
| Correlation between the trials | It lowers the effective count | Measurable, and rarely measured |
It is too large because the trials are not independent. A crossover with a 20-day fast leg and one with a 25-day fast leg are very nearly the same bet, and the expected-maximum formula assumes something closer to independence. This one is measurable rather than guessable, but on this search the measurement comes with a trap. Taken at face value it finds no redundancy at all: the formula at N of 170 predicts a hurdle of 1.16, the simulation that carries the real correlation structure produced 1.16, and solving back, the 170 rules behaved like about 166 independent ones. The redundancy is there, and a second error is hiding it. The shuffle keeps the series' upward drift, and a long-or-flat rule collects part of that drift with no skill at all, which lifts every rule in the null, while the formula measures from zero. Remove the drift, shuffle 2,000 more times, and the best of 170 averages 0.87, which solves back to about 27 independent trials. So the formula fed a raw grid size makes two errors here, counting correlated rules as independent and ignoring the drift a long-or-flat rule collects for nothing, each worth about 0.29 of annualised Sharpe, and on this series they happen to cancel. A raw grid size lands on the right hurdle here by coincidence, not by design.
Both errors in the count are real, the trials never recorded and the correlated trials counted as independent. They run in opposite directions, and the one that is easy to measure is the smaller one. That is why the honest use of the instrument is one sided. A strategy that fails deflation at a conservative N is finished, and the finding is safe because the untracked trials only push the hurdle higher. A strategy that passes has not been certified by anything except your own account of your own record keeping.
The practical response is to stop treating N as something to be recovered after the fact. Declare it before the search. Log every run, including the ones that end in a shrug. A trial counter that increments automatically is worth more than any amount of care applied afterwards, because afterwards the count is already lost.
India verifies the return now, and still not the search
This is the part that dates most of what is currently published on Indian strategy performance. Since 4 May 2026, past performance shown by investment advisers, research analysts and algorithmic trading services in India goes through a verification agency. The framework, operationalised by a circular dated 29 April 2026 carrying the reference HO/38/14/(4)2026-MIRSD-POD/I/10557/2026, recognised a credit rating agency as the first Past Risk and Return Verification Agency and the exchange as the data centre behind it. Close to fifty risk and return metrics are computed from transaction data taken directly from exchanges and clearing corporations. Advisers and research analysts were required to enrol by 3 August 2026, and figures from before the framework can be shown only until 3 May 2028.
The link to algorithmic strategies is direct. The circular dated 4 February 2025 on safer participation of retail investors in algorithmic trading, reference SEBI/HO/MIRSD/MIRSD-PoD/P/CIR/2025/0000013, requires providers of black box algorithms to register as research analysts and to hold a research report for each algorithm. Implementation timelines under that circular were extended by a later circular, so check the current dates before acting on any of it. The effect of the two together is that the performance figure attached to a retail algorithmic strategy in India is on its way to being a verified figure computed from executed trades.
That is a genuine improvement, and it fixes a different problem from this one. Verification establishes that a claimed track record is arithmetically real. It cannot establish how many candidate strategies were searched before that one was put in front of you, because realised transactions do not contain that information and no pipeline built on realised transactions ever could. There is no field for N.
So the Indian retail market is about to have verified Sharpe ratios that are still not deflated Sharpe ratios. Worse, verification is computed on live records, which are short. A provider with eighteen months of verified trading has about 378 sessions behind its figure, which puts the annualised standard error on its Sharpe at 0.77 by the arithmetic above. The number will be true. It will also be a point estimate with an error bar nobody prints, selected from an unknown number of attempts, and a reader who treats it as a property of the strategy will be making exactly the error this article is about.
What deflation is actually for
It is not a certificate and it should never be used as one. Given an honest N it would be a proper test, and there is no honest N. What survives that limitation is a destructive instrument, which is worth more than it sounds.
Most candidate strategies are not real. They are the right tail of a search, and the entire cost of discovering that comes later, in capital, in the weeks it takes to be sure the drawdown is not just variance, and in the confidence spent on the next one. Deflation moves that discovery to the cheapest possible point available. The search in this article takes a few thousandths of a second to run, and the 4,000 simulated searches that test it take under a minute on an ordinary laptop.
The workable order is short. Fix the hypothesis and the trial budget before touching the data. Prefer few motivated trials to many mechanical ones, because the hurdle climbs fastest at the start and a smaller N is a real asset. Compute the null by re-running the entire search on resampled returns rather than testing the survivor alone, because the survivor is not the unit of analysis, the search is. Compare against the best of that null rather than against zero. Then treat a pass as permission to keep testing, never as a conclusion.
None of this makes a strategy work. It makes the failures cheap and early, which is the only part of the process that compounds. Judgement about what to test at all, and the discipline to count the tests you ran, do more for a systematic book than any single statistic, and neither can be bought from a backtest report.
Frequently asked questions
What is the deflated Sharpe ratio, in one sentence?
It is the probability that a strategy's true Sharpe ratio is above a hurdle, where the hurdle is not zero but the Sharpe that the best of N no-skill attempts would be expected to produce on the same data. It adjusts for three things at once: the number of trials behind the result, the length of the sample, and the skew and kurtosis of the returns.
Why does a Sharpe ratio have a standard error at all?
Because the mean and the standard deviation in it are sample statistics, not known quantities. The ratio is an estimate of something unobservable, and like any estimate it has a sampling distribution. To first order the standard error falls with the square root of the number of observations, which is why sample length dominates everything else.
Does non-normality make the standard error bigger or smaller?
It depends on the sign of the skew. Negative skew with a positive Sharpe widens the standard error, and excess kurtosis widens it further. Positive skew narrows it. The size of the effect scales with the Sharpe measured per period, so on daily data it is usually small and on monthly or quarterly data it is material. On the series measured for this article the daily adjustment moved the composite's standard error by 0.4 percent, and the best rule's by about 6 percent, narrower rather than wider, because its returns are strongly positively skewed and its daily Sharpe is unusually high.
How many trials should I use for N?
There is no honest answer, and that is the point rather than an evasion. N is the size of the search that produced the survivor, which includes every variant abandoned, every re-run after a fix, every parameter nudged by eye, and every idea inherited from someone else who searched first. Most people do not record any of it. The workable discipline is to declare an N before the search starts and treat any figure that fails at that N as dead.
My backtest shows a Sharpe of 2. Is that good?
The number on its own carries no information without the sample length and the number of trials behind it. In the search measured for this article, a best of 2.0 or more came out of 2.2 percent of the no-skill searches with the time order destroyed, which would make it unusual, and out of 21.1 percent of those that keep ten-day blocks intact, which would make it an ordinary outcome of a 170-rule search on data with no lasting edge in it. Which null is nearer the truth decides the answer, and a backtest report will not say.
Does a longer backtest fix the problem?
It shrinks one part of it. The standard error falls with the square root of sample length, so getting an annualised standard error down to 0.20 at the Sharpe measured here needs about 5,594 sessions, which is roughly 22 years. It does not fix the trials problem at all, because a longer sample also lets you try more things on it.
Is a holdout sample or walk-forward testing an alternative?
They are complements, not substitutes, and both leak. A holdout stops being a holdout the second time you look at it, and walk-forward re-optimisation is itself a search whose trials belong in N. Deflation is useful precisely because it puts a number on the cost of the searching that those methods try to manage by procedure.
Does this apply to a live track record, or only to backtests?
It applies to any Sharpe that was selected. A live record that is shown to you because it was the best of several the manager was running has been selected in exactly the same way a backtest is, and the same arithmetic applies. What a live record does buy is the absence of look-ahead and the presence of real costs, which a backtest never has.
Does the verification of past performance in India cover this?
No. The verification framework operational since 4 May 2026 computes risk and return metrics from transaction data sourced from exchanges and clearing corporations, so it establishes that a claimed record is arithmetically real. It carries no information about how many candidate strategies were tried before that one was shown, because realised transactions cannot contain that information.
What should I do with a strategy that fails deflation?
Stop working on it and record the trial. Failing at a conservative N is the cheapest possible result, because the alternative is finding out with capital. A strategy that passes has not been certified either, since the N you supplied was your own estimate of your own record-keeping.
Every market figure here was computed for this article by tools/measure-article-90-dsr.py in this repository, from NSE full security bhavcopy files for 2024-09-18 to 2026-09-18. The simulation is fixed seed and reproduces the same figures on a re-run. The null distributions are simulated and labelled as such. No costs, taxes or financing charges are applied anywhere except the stated cost levels in the one table that says so, and applying them would lower every figure shown. Regulatory positions are stated as at 19 September 2026: verification and algorithmic trading timelines in India have already been revised once and should be checked against the current circulars before anyone acts on them. Nothing here is advice, and no figure in it is a forecast.
The market figures were corrected on 23 September 2026. The first version counted every cached bhavcopy file as a session, but when the exchange archive is asked for a date on which there was no session it returns the previous session's file, so 17 of the 512 files named inside this window repeated a session already present and each of those returns was counted twice. Sessions are now keyed by each file's own trade date, one file per session, which gives 495. The two Saturday special sessions that exist only under Monday file names fall before this window. The weekend budget-day sessions of 1 February 2025 and 1 February 2026 have no file in the cache and are missing; each return is taken within one file from its own previous close, so none spans two sessions. The constituents are the symbols present in the EQ series on every one of the 495 sessions, so a company that changed its symbol inside the window is left out. The correction raised the best rule's Sharpe from 1.24 to 1.71 and moved its 95 percent interval from -0.15 to 2.62, which contained zero, to 0.39 to 3.03, which does not, and the number of no-skill trials whose expected best equals it rose from about 311 to about 17,400. The sections on the null distributions, the cost table, the deflation arithmetic and the effective number of trials were rewritten on the corrected figures, including where those figures weaken the case the first version made against the winning rule.
Ready to go deeper than this article?
Bharath Shiksha is a 90-volume curriculum across 6 stages, from chart reading at ₹14,999 through capital raising, or the full bundle at ₹1,49,999. Counting the trials behind a number, and testing the search rather than the survivor, is the part of systematic work that decides whether the rest of it was worth doing.
Take the free diagnostic →