A five per cent test on Indian index returns was measured rejecting six times too often, and it is blind to the change that would actually hurt you

The short answer

A two sample distribution test asks whether two batches of returns could have come from the same distribution. It is the right question when a system stops working, and three measured facts decide how much its answer is worth. Power: half of a thirteen year daily index record, 1,686 observations, can only reliably detect a shift of about 0.11 standard deviations in the daily average, which is a change of roughly 1.8 in the ratio of annualised average to annualised variability. Independence: on a stationary simulated series with nothing changed at all and the volatility clustering the Nifty 50 actually shows, the nominal five per cent test rejected 31.4 per cent of the time. On the real record, testing adjacent one year windows in their historical order rejected 32.8 per cent of the time against 4.2 per cent for the identical numbers shuffled. Blind spots: a change confined to the worst two per cent of days cannot move the statistic past 0.02, and the threshold at this sample size is 0.0468, so it cannot be rejected at all.

Every figure below was computed in the build script for this page, on 3,385 real NSE sessions from 2013-01-01 to 2026-09-18, with the method and the seeds stated. The test was implemented rather than imported, and it was proven in both directions before anything was built on it.

What the test asks, in plain terms

Take two batches of numbers. For each one, walk along the number line and record the fraction of the batch that falls at or below each point. That running fraction is the batch's distribution function, and it rises from zero to one. Lay the two on top of each other and measure the largest vertical distance between them anywhere. That distance is the statistic. The question the test answers is how often two batches drawn from a single distribution would happen to be that far apart.

The appeal is that it assumes nothing about the shape. A comparison of averages tests the average and a comparison of variability tests the variability, whereas this responds to a change in the middle, the spread, the skew or the shape without being told where to look. That is exactly what you want when someone says a system stopped working, because the statement is a distributional claim and its two rivals are distributional too: either the distribution of outcomes has moved, or it has not and this is the ordinary bad stretch an unchanged distribution delivers from time to time.

The generality is also the source of every difficulty that follows. A test that looks everywhere has to spend its sensitivity everywhere, and it turns out to spend almost none of it in the place a trader cares about most.

Proving the instrument before believing it

A test you have never fed a known answer is not an instrument, it is a function that returns numbers. It has to be proven in both directions, and the second is the one people skip: it must reject when there is genuinely something to find, and it must hold its stated false positive rate when there is not. A test that rejects everything passes the first check perfectly and is useless.

So: fed pairs of samples of 1,686 each, drawn from distributions that genuinely differ, and then fed 10,000 pairs drawn from the same distribution.

The instrument proven in both directions, at 1,686 observations a side, five per cent level
What it was fedAverage statisticHow often it rejected
Average higher by half a standard deviation0.2077100.0 per cent
Spread one and a half times wider0.1111100.0 per cent
Same average and spread, heavier tails0.070199.4 per cent
Two samples from the same distribution0.02975.00 per cent

The last row is the one that licenses everything else on this page. Against two samples from an identical distribution the test rejected 5.00 per cent of the time at a nominal five per cent, and 0.90 per cent at a nominal one. It is calibrated. Two concrete runs at the same sample size: two samples from the same law gave a statistic of 0.0297 and a p value of 0.45; two samples differing only in spread, by a quarter, gave 0.0973 and a p value below a millionth.

The implementation was also checked against an established one, in a script that does not import this page's build code. Over two hundred random sample pairs the statistic agreed to the last digit, the p values differed by at most about 0.013 in the region that decides a verdict, and the two agreed on the verdict at five per cent in 3,985 of 4,000 pairs. Every failure described below is therefore a property of the data, not of a broken tool.

The power problem, which is the spine of the whole subject

The requirement scales as one over the square of the difference being looked for. A change half as large needs roughly four times the data. That is not a convention, it is the same square root law that governs every average, and it is why the arithmetic runs away so fast.

The size of a change, in the units this test works in, is the largest gap it opens between the two distribution functions. For a shift of the average by a given number of standard deviations, that gap has a closed form and it is smaller than people expect: a shift of a tenth of a standard deviation opens a gap of only 0.0399. The constant relating the gap to the sample needed was measured here by bisection, not assumed, and then validated at two other shifts where the predicted eighty per cent power came out at 82 and 81 per cent.

Observations needed against the size of the change A steeply falling curve on a logarithmic vertical scale. The number of observations per sample needed to detect a change falls sharply as the change gets larger. A horizontal line marks the number of observations one half of the real index record actually contains, and the curve crosses it only at fairly large changes. Observations per sample needed, at five per cent and eighty per cent power 1001,00010,000100,000 1,686, one half of the real record 893512,1788,70734,821 Shift in the daily average, in standard deviations 0.025 0.10 0.50 Halve the change and the requirement roughly quadruples. The constant was measured, not assumed.
Computed by simulation. The crossing point is the honest reading: with one half of a thirteen year daily record in hand, only a change of about 0.11 standard deviations or larger is reliably visible.
Observations per sample needed to detect a shift in the daily average, five per cent level, eighty per cent power. Years at 246 trading days.
Shift, in standard deviationsLargest gap it opensObservations a sideYearsIn annualised ratio termsPower at 1,686 a side
0.5000.1974890.47.84100 per cent
0.2500.09953511.43.92100 per cent
0.1000.03992,1788.91.5771 per cent
0.0500.01998,70735.40.7824 per cent
0.0250.010034,821141.50.3911 per cent

Read the last two columns together and the difficulty is plain. With 1,686 observations a side, which is every trading day from 2013-01-01 to 2019-11-19, the smallest shift reliably visible is about 0.11 standard deviations a day. Converted into the units people quote, that is a change of roughly 1.8 in the ratio of annualised average to annualised variability. A change that large is not a regime shift, it is a different instrument.

The change that actually ends a system is nothing like that size. Suppose a method averaged 0.05 standard deviations a day and that edge halved. In distributional terms the shift is 0.025 standard deviations, and the table says detecting it takes 34,821 observations a side, or about 142 years of daily data. The trader knows within a year that something is wrong. The test will not know within a lifetime. The general arithmetic of that gap is worked through in how many trades it takes to tell an edge from luck, and this test sits on the unfavourable side of it.

Returns are not independent, and the correlation people check is the wrong one

The p value the test reports is computed from how fast a distribution function settles as observations accumulate, and that rate assumes every observation is a fresh piece of information. Returns are not like that, and the usual defence against the objection is checking the wrong quantity.

On this record the correlation between consecutive daily returns is -0.0044, which is effectively nothing, and that is the number most people look at before declaring the independence assumption safe. The correlation between consecutive absolute returns is 0.232, it is still 0.202 five days out, and 0.124 twenty days out. Large days arrive next to large days. The dependence is in the size of the moves, not their direction, and the size is exactly what a distribution test reads.

Whether that matters is measurable. Below, two kinds of stationary series in which nothing changes at all: the two halves are drawn from one distribution by construction, so every rejection is a false positive and the column should read five per cent throughout.

False positive rate at a nominal five per cent, on series of 3,373 observations split in half, where nothing has changed. For comparison, the real index record measures 0.232 at one day and 0.124 at twenty.
The processAbsolute values, correlated at one dayAt twenty daysRejections at a nominal five per cent
Consecutive returns correlated at 0.00-0.0010.0005.30 per cent
Consecutive returns correlated at 0.100.008-0.0007.52 per cent
Consecutive returns correlated at 0.200.035-0.00010.12 per cent
Consecutive returns correlated at 0.300.079-0.00012.15 per cent
Volatility clustering0.0600.0228.60 per cent
Volatility clustering0.1570.10229.68 per cent
Volatility clustering, closest to this index0.2050.13231.36 per cent
Volatility clustering0.2500.15733.72 per cent
Volatility clustering0.2890.17536.84 per cent
Volatility clustering0.2950.11623.84 per cent

Correlation between consecutive returns does modest damage: even at 0.30, which no equity index comes close to, the rate reaches only 12.2 per cent. Volatility clustering does severe damage. At the calibration whose absolute return correlations best match this index, 0.205 at one day against the measured 0.232, the nominal five per cent test rejected 31.4 per cent of the time with nothing whatsoever having changed. That is 6.3 times the rate it advertises, and the neighbouring calibrations put the plausible range at roughly 24 to 37 per cent.

The third column is the reason, and it is worth reading before the fourth. The rate is not driven by how strong the clustering is at one day but by how long it persists. The last row has the highest one day correlation in the table and a distinctly lower false positive rate, because by twenty days its clustering has largely decayed. What breaks the test is a series that spends months in a quiet state and months in a violent one, which is a fair description of an equity index.

The same numbers, in a different order

A simulation is a model, and a model can be argued with. The cleanest version of this measurement uses the real series and changes exactly one thing about it.

Take the 3,373 real daily returns. Walk a pair of adjacent windows along the record and test every position, so at one year a side that is 2,874 separate tests. Then shuffle the identical 3,373 numbers into a random order and do the same. The shuffled series has the same values, the same histogram, the same average, the same spread, the same skew and the same worst day. A distribution test cannot see any difference between the two, because everything it looks at is unchanged. Only the sequence differs.

The same returns, in order and shuffled Three pairs of bars. In each pair the left bar is how often the test calls the distribution changed when the real returns are left in their historical order, and the right bar is how often it does so when exactly the same numbers are shuffled. The left bars are far taller, and the right bars sit on the five per cent line the test claims. How often a five per cent test says the distribution changed 5 14.54.2125 per side32.84.2250 per side52.15.2500 per side historical order same numbers, shuffled Identical values, identical histogram. Only the order differs.
Every adjacent pair of windows in the real record was tested. Shuffling destroys the order and nothing else, and it takes the false alarm rate back to the five per cent the test promises.
How often a five per cent test calls the distribution changed, on adjacent windows of the real index record
Window lengthPairs testedHistorical orderSame numbers shuffledInflation
125 a side3,12414.50 per cent4.16 per cent3.5 times
250 a side2,87432.85 per cent4.21 per cent7.8 times
500 a side2,37452.11 per cent5.22 per cent10.0 times

Shuffled, the rate lands between 4.2 and 5.2 per cent against a nominal five, which is the calibration again. In historical order, at one year a side, the test calls the distribution changed on 32.8 per cent of adjacent window pairs, and at two years a side on 52.1 per cent. Ordering is the only thing that changed.

One honesty note, because it matters for what the measurement proves. Part of that excess is genuine slow change in the real distribution rather than dependence alone, and the test cannot separate the two. The simulated rows in the previous table isolate dependence, because there the distribution provably does not move. Both effects are present in the real record, both inflate the rejection rate, and neither is what the reported p value is describing. Whether a series can be treated as having one fixed distribution at all is the subject of a price series has a history, not a distribution.

What the test structurally cannot see

A test that looks at the whole distribution finds the largest gap wherever it happens to be. The consequence is a blind spot with a hard edge, and it sits precisely where risk lives.

There is only so much probability in a tail. If a change touches only the worst two per cent of observations, it can move the distribution function by at most 0.02 at any point, because two per cent is all the mass there is to move. That is a ceiling, not a tendency. Below is the same normal distribution with its worst two per cent of days made progressively worse, tested at 1,686 observations a side.

A change confined to the worst two per cent, tested at 1,686 observations a side
The worst two per cent becomeStandard deviation ratioShortfall in that tailAverage statisticHow often it rejected
1.25 times worse1.0331.250.02935.7 per cent
1.50 times worse1.0721.500.03035.6 per cent
2.00 times worse1.1642.000.03014.9 per cent
3.00 times worse1.3953.000.03026.0 per cent
5.00 times worse1.9575.000.03005.0 per cent

Read the last column down. It does not fall away gradually, it simply never moves: 5.7 per cent when the tail is a quarter worse and 5.0 per cent when it is five times worse, both indistinguishable from the 5.0 per cent the test rejects at when nothing has changed. The average statistic is pinned at about 0.030 regardless. The test is not weak against this change, it is blind to it.

A change confined to the tail runs into a ceiling A curve that rises out of the far left of the chart, levels off at a low ceiling, and falls back to zero well before the middle of the range. A horizontal line above the ceiling marks the level the curve would have to reach for the test to reject. The curve cannot reach it, because the change affects only two per cent of observations. The worst two per cent of days made three times worse rejection needs 0.0468 this change can never exceed 0.02 Return level, in standard deviations minus 8 minus 4 0 The worst one in fifty days is three times worse, and the test cannot see it below 9,223 a side.
The ceiling is exact rather than empirical. A change touching only the worst two per cent of observations moves the distribution function by at most 0.02 anywhere, so below 9,223 a side the test rejects no more often than it does when nothing has changed, however severe the change is.

The threshold arithmetic makes the edge exact. Rejection at five per cent needs a statistic above 1.3581 times the square root of two over the sample size a side, and for that to fall below the 0.02 ceiling the sample must be at least 9,223 observations a side, which is about 37 years of daily data. Measured against a tripled tail, the test rejected 4.8 per cent of the time at 1,686 a side, 52 per cent at 9,223, which is exactly the coin flip the arithmetic predicts when the true gap sits on the threshold, and 100 per cent at 20,000.

This is the single most important limitation on the page, because the change a trader most needs to detect is almost always a tail change. Standard deviation already understates it, as set out in standard deviation is a second moment, and the whole distribution test understates it further.

The real split, and what it does not say

Now the obvious thing to do with the data. Take the 3,373 daily log returns, split them at the midpoint, 2019-11-19, and test 1,686 against 1,686. The statistic is 0.0302 against a threshold of 0.0468, for a p value of 0.42. It does not reject.

The gap between the two halves of the index return record A near flat wandering line close to zero, drawn against two horizontal reference lines far above and below it. The line is the difference between the two halves' distribution functions at every return level. The reference lines are the level the difference would have to reach for the test to call the distribution changed. The line crosses neither, and at its furthest from zero it gets 65 per cent of the way to one of them. Nifty 50 daily returns, first half against second half threshold 0.0468 resampled threshold 0.0688 largest gap 0.0302 minus 3 pcminus 1.5 pc0plus 1.5 pcplus 3 pc Daily return level at which the two halves are compared Real data, 3385 sessions. The record straddles a crash and the test returns a p value of 0.42.
Computed from the cached index closes, not illustrative. The statistic is the largest vertical distance this line reaches from zero. It gets 65 per cent of the way to the line it would have to cross.

Anyone reading that as evidence the two periods were alike would be badly wrong, and the way in which is instructive.

The two halves of the Nifty 50 daily return record, in per cent
 AverageStandard deviationMedian absolute deviationInterquartile rangeWorst dayWorst one per cent
First half+0.04510.89850.49590.9913-6.10-3.05
Second half+0.03911.12330.51991.0410-13.90-5.37
Second as a multiple of first0.871.251.051.052.281.76

The second half is 25 per cent more variable by standard deviation, its worst day is 2.3 times deeper than anything in the first, and the average of its worst one per cent of days is 1.8 times worse. Those are not small differences to live through. Yet the median absolute deviation rose only 4.8 per cent and the interquartile range only 5.0 per cent.

That is the whole article in one table. The middle of the distribution barely moved and the tail moved enormously, so the standard deviation moved while the distribution function did not. A useful counterfactual: if the second half genuinely were the first half rescaled by 1.250, the largest gap between them would be 0.0538, comfortably above the 0.0468 threshold, and the test would very likely have caught it. The observed statistic was 0.0302. The halves are not a rescaling of one another, and the test read the part that did not change.

Running the same split on two more indices gives the other half of the lesson.

The same midpoint split on three indices, 1,686 observations a side
IndexStatisticp valueStandard deviation ratioVerdict at five per cent
Nifty 500.03020.4231.250does not reject
Nifty Bank0.02850.5021.138does not reject
Nifty IT0.05400.0151.295rejects at five per cent

One of three rejected, at a p value of 0.015. Reporting that one and not the other two is the trap, though the usual multiple testing arithmetic is the weaker objection here, because three index return series are nowhere near independent of one another. The stronger objection is the threshold itself. Resampled from that index's own history in 250 day blocks, the honest five per cent level is 0.0599 against an observed 0.0540. Once the dependence the series actually has is allowed for, it does not reject either.

Testing again every month guarantees a rejection

The natural way to use this test is as monitoring: run it each month as new data arrives and act when it fires. That converts a single test into a sequence of them, and a sequence of five per cent tests is not a five per cent procedure.

Chance of at least one rejection when nothing has changed, rolling 250 observation windows advanced by 21
How long you monitorIf the tests were independentMeasured, independent observationsMeasured, clustered like the index
6 monthly tests26.5 per cent13.3 per cent47.6 per cent
12 monthly tests46.0 per cent22.1 per cent68.8 per cent
24 monthly tests70.8 per cent37.3 per cent84.1 per cent
36 monthly tests84.2 per cent48.7 per cent92.7 per cent

Both measured columns disagree with the naive bound, in opposite directions, and both reasons are worth knowing. With independent observations the measured rate is lower than the bound, because overlapping windows share most of their data, so consecutive tests are nearly the same test and you get far fewer independent chances than you paid for. With clustered observations it is higher, because each individual test is already rejecting several times too often before any repetition is counted.

Three years of monthly monitoring on a series where nothing changes produced at least one rejection in 93 per cent of runs, with the first false alarm arriving at a median of month 5. The structure of this trap, and what to do about it, is the subject of the best of fifty backtested rules.

What to do instead, measured rather than recommended

None of the above is an argument for abandoning the test. It is an argument against the textbook threshold, which is the part that is wrong. Three changes, each measurable.

Take the threshold from your own series. Resample the record in contiguous blocks, so short range dependence survives and no real change can, split each resampled series in half and test it, and read off the ninety fifth percentile of the statistic. That is the level your data produces when nothing has changed.

Five per cent threshold taken from the index record itself by block resampling, against the textbook 0.0468
Resampling blockThresholdAgainst the textbookVerdict on the real split
no blocks, single days0.04630.99 timesdoes not reject
20 day blocks0.05161.10 timesdoes not reject
60 day blocks0.05751.23 timesdoes not reject
125 day blocks0.06111.31 timesdoes not reject
250 day blocks0.06881.47 timesdoes not reject

The single day row is the check on the method: with blocks of one there is no dependence left and the threshold lands on the textbook value. As the blocks lengthen the honest threshold rises to 1.47 times it. Which block length is right is a judgement about how long the dependence lasts, and reporting the verdict across several rather than picking one is the honest presentation. This is the same resampling logic used for interval estimates in bootstrap intervals on an equity curve.

Remove the clustering first, and accept that it changes the question. Dividing each return by a trailing estimate of volatility cuts the absolute return correlation from 0.232 to 0.111, and the effect on the false alarm rate is real.

Rejection rate on adjacent real windows, before and after dividing by a 60 day trailing volatility
Window lengthRaw returnsVolatility standardised
125 a side14.50 per cent11.05 per cent
250 a side32.85 per cent20.95 per cent
500 a side52.11 per cent18.86 per cent

At two years a side the rate falls from 52.1 per cent to 18.9 per cent. That is a large improvement and it is not a fix, because the rate is still several times the nominal five. It also narrows what is being asked: a standardised test can no longer see a change in the level of volatility, only a change in the shape that remains once volatility is divided out. That is often the more interesting question, but it is a different one and it should be stated as such.

Decide the split before you look, and test the quantity you care about. A split chosen after seeing where the results turned is guaranteed to produce a gap, and no threshold repairs that. And since the test is blind to the tail, a targeted comparison is worth far more: on the same samples where the whole distribution test rejected a tripled tail 4.8 per cent of the time, a direct comparison of the average of the worst two per cent rejected essentially always. Sensitivity is not free, but it can be aimed.

The question that is usually better

Put the three limits together and what this test can do is narrower than its reputation. It rules out a dramatic break well, because a dramatic break opens a large gap and a large gap shows up in a few hundred observations. It confirms a subtle one poorly, because subtlety is what the sample size cannot reach. And it is worst at the change that matters most, because a tail change cannot open a large gap by construction.

A rejection therefore tells you something worth having: the two periods are further apart than dependence alone explains, which is a reason to look harder. A non rejection tells you very little, and the real split on this page is the demonstration, returning a p value of 0.42 across a boundary with a day more than twice as bad on one side.

Which leads to the reframing. When a system stops working, the instinctive question is has the distribution changed, and at the sample sizes anyone has it is usually unanswerable. The answerable question is was it ever what I thought it was. If a method was established over a period short enough that a distribution test cannot now resolve a change in it, the original evidence was already too thin to support the claim, and nothing broke because nothing was ever demonstrated. That conclusion is reachable today, from the record you already hold.

None of this argues against measuring. It argues for knowing what a measurement can carry, which is the difference between a research process and a ritual, and it is why this is taught as method and as the honest limits of method rather than as a list of tests to run. Running the test is easy. Knowing which of its answers you are entitled to use is the skill, and the related discipline of publishing a number someone else can check is set out in publishing a result someone else can check.

Frequently asked questions

What exactly does a two sample distribution test ask?

Whether two batches of numbers could plausibly have come from the same underlying distribution. It builds the distribution function of each batch, measures the largest vertical distance between them, and asks how often two samples from a single distribution would happen to be that far apart. It assumes nothing about the shape, which is why it suits returns, where parametric assumptions fail.

Why is that the right question when a system stops working?

Because the competing explanations are a distributional pair: either the distribution of outcomes changed, or it did not and this is the ordinary bad stretch an unchanged distribution produces from time to time. A test of means compares one feature and can miss everything else, while this asks about the whole shape at once, which is what you want when you do not know which part moved.

How large a sample does it actually need?

Far more than people expect, and the requirement grows as the square of how small the change is. Measured here, a shift of a tenth of a standard deviation in the daily average needs about two thousand observations a side and a shift of a fortieth needs about thirty five thousand, which is more than a century of daily data. The sample size question is treated at length in the guide on detecting an edge.

My returns show almost no autocorrelation. Does the independence problem still apply?

Yes, and that check is the wrong one. On the index record used here the correlation between consecutive daily returns is about minus 0.004, which is effectively nothing, while the correlation between consecutive absolute returns is about 0.23 and is still above 0.12 twenty days out. The dependence lives in the size of the moves, not their direction, and the size is what a distribution test reads.

Why does dependence make the p value wrong rather than just imprecise?

Because the p value is calculated from how fast a distribution function settles as observations accumulate, and that rate assumes each observation is fresh information. When observations arrive in clusters it settles more slowly, so the gap between two halves is routinely larger than the calculation allows, and the test reports a small p value for a gap that is entirely ordinary.

Can the test miss a change that matters for risk?

It can miss it completely, and the reason is structural rather than a question of sample size. A change confined to the worst two per cent of observations cannot move the distribution function by more than 0.02 anywhere, because there is only two per cent of probability there to move. If the threshold at your sample size is above 0.02, the change itself can never push the statistic over the line, so the test rejects no more often than it would if nothing had happened.

What does a non rejection actually license me to say?

Only that the samples are not far enough apart to rule out a single distribution at your chosen level, which is much weaker than saying nothing changed. On the real index split here the p value is 0.42, far above any usual cut off, while the second half contains a day more than twice as bad as anything in the first.

If I run the test every month as new data arrives, what happens?

You eventually reject, whether or not anything changed. Measured here on a series with nothing changing and the clustering the index actually shows, thirty six monthly tests produced at least one rejection in the large majority of runs. Repeated testing on accumulating data is the same trap as searching many rules for the best one, which is covered in the guide on multiple testing.

Is there a version of this that is actually usable?

Yes, with three changes. Take the threshold from the series itself by resampling it in blocks rather than from the textbook formula. Decide the split before you look, so the comparison is not chosen by the answer it gives. And test the quantity you actually care about, usually something in the tail, because a targeted comparison at the same sample size can be far more sensitive.

If the answer is usually inconclusive, what should I be asking instead?

Whether the result was ever what you thought it was. A system that appeared to work over a period short enough that a distribution test cannot now resolve a change was never established in the first place, and the honest reading is that the original evidence was thin rather than that something broke. That question is answerable with the data you have, and the change question generally is not.

How these numbers were produced. The data are the exchange's daily index close files, 3,385 sessions from 2013-01-01 to 2026-09-18, including 14 weekend special sessions (budget days, muhurat trading and disaster-recovery drills), which are real sessions and are kept. The archive holds no file for 12 weekday sessions between 2013-10-09 and 2016-06-20, each found because the next file's own reported change does not match the previous close; the change across each such gap spans two or more sessions and is left out of every sample of daily returns, so 3,373 returns remain, and no autocorrelation pairs two returns across a gap. The file for 2023-03-13 reports its change against the wrong prior session; its return is computed from consecutive closes like every other day. The statistic is evaluated at every observed value of the pooled sample and the p value comes from the limiting Kolmogorov distribution summed to one hundred terms. The critical value constants are re-derived by bisection at build time and asserted against the values used. The instrument was proven in both directions first: against three genuinely different alternatives at 1,686 a side it rejected in 99 per cent of trials or more, and against 10,000 pairs from an identical distribution it rejected 5.00 per cent of the time at a nominal five and 0.90 per cent at a nominal one. The sample size constant was measured by bisection at a shift of 0.25 standard deviations and validated at 0.10 and 0.05, where a predicted eighty per cent power measured 82 and 81 per cent. Dependence results use a first order autoregressive process and a stationary volatility clustering process whose absolute value correlations are reported beside the measured ones. The ordered against shuffled comparison tests every adjacent window pair in the real record and an equal number of shuffles of the identical values. Block thresholds are the ninety fifth percentile across 1,200 circular block resamples. All draws are seeded, so every figure on this page is identical on every rebuild. Computed on Python 3.9.6.

What could not be verified this session. No external source could be fetched or searched while this page was built, so nothing here rests on a third party claim. Every quantitative statement is computed from the cached NSE index closes named above or from a stated simulation, and both are reproducible from the description. Simulation results describe the behaviour of a statistical procedure and are not a measurement of, or a prediction about, any trading system or index. The position is stated as at September 2026. Verify the data and the arithmetic against your own before relying on either, and take advice on your own circumstances.

Related guides

How many trades before you can tell an edge from luck

Read →

The best of fifty backtested rules

Read →

A price series has a history, not a distribution

Read →

Ready to go deeper than this article?

Bharath Shiksha is a 90-volume curriculum across 6 stages, from chart reading at ₹14,999 through capital raising, or the full bundle at ₹1,49,999. Knowing which questions your data can actually answer, and refusing the ones it cannot, is the part of research method that no software supplies.

Take the free diagnostic →