A price series has a history, not a distribution
The short answer
Stationarity is the assumption that the process producing your data is not itself changing. A price series fails it outright: on 3,385 real sessions of the Indian broad index, 2013-01-01 to 2026-09-18, the unit root statistic on the log level is -0.99 against a five per cent critical value of -2.86. Returns fail a weaker version, because the size of a move predicts the next at 0.23 and still at 0.12 twenty sessions out. The consequences are arithmetic: the trailing one year mean ranged from -40 to 66 per cent, days beyond four standard deviations occurred 17 times where a normal law expects 0.21, and the banking to technology correlation ranged from -0.09 to +0.64. And the series changed its own construction on 3 August 2026, when the official closing price of derivative eligible Indian stocks began to be set by a closing auction rather than by an average of the last half hour.
Every figure here was computed from the exchange daily index close files, not quoted. The archive is 3,385 sessions over 13.7 years with no calendar gap longer than 6 days, and the method is stated at the end so the numbers can be reproduced rather than believed.
Stationarity is a claim about the generator, not about the chart
The strict definition is short. A series is stationary when the joint distribution of any group of observations depends on how far apart they are, not on where in the record they sit. A hundred days from 2014 and a hundred from 2025 are then two samples of the same thing, with no special dates and no before and after.
Almost no financial series satisfies that and almost no method needs it. The working version, sitting underneath the tools people actually use, is weaker and has three parts: the mean does not change, the variance does not change, and the covariance between two observations depends only on the gap between them. That is the assumption inside every sample average, standard deviation, correlation matrix and regression coefficient computed on market data. It is nearly true often enough to be dangerous, and rarely written down, which is why it is rarely checked.
A price has no centre to return to
A price series fails the first condition in the most basic way. A shock to today's price is carried into every future price rather than decaying, so the best forecast of tomorrow is today, and the variance of where the series will be grows without limit with the horizon. That property is a unit root.
Run the standard test on the log of the broad index and the statistic is -0.99, with 28 lagged differences and 3,356 usable observations, against a five per cent critical value of -2.86. Not close. Take first differences and the same test on daily log returns gives -11.15, past the one per cent value of -3.43 by a wide margin. That pair is the textbook picture and where most treatments stop, which is unfortunate, because the second result is routinely over-read.
Returns are closer, and still not there
Differencing removed the wandering level. It did not make the variance constant, which the weaker definition also needs.
The signed return is close to unpredictable: its correlation with the previous session is -0.003, inside the ordinary noise band of 0.034 for a sample of 3,373. The size of the return, ignoring direction, is another matter. It is 0.23 at one session, 0.19 at five, still 0.12 at twenty. Large days follow large days and quiet days follow quiet days, for weeks.
That is volatility clustering, and it means the variance is not a constant to be estimated but a quantity with a history of its own. Over trailing sixty sessions, annualised volatility ran from 7.2 per cent around 2025-11-25 to 58.0 around 2020-06-01, a factor of 8.0. A single number called the volatility of this index summarises that range.
One layer down, the variability of volatility is not constant either. Taking the log of trailing sixty session volatility and splitting the sample in half, its standard deviation is 0.27 in the first half and 0.41 in the second. The quantity you would hold fixed to model the changing variance is itself changing, which is why the search for a stable transformation keeps uncovering one more unstable layer.
| Statistic | Whole sample | Rolling low | Rolling median | Rolling high |
|---|---|---|---|---|
| Annualised mean return | 10.42 per cent | -40.2 per cent | 10.5 per cent | 65.8 per cent |
| Annualised volatility | 15.98 per cent | 8.8 per cent | 13.9 per cent | 32.2 per cent |
| Correlation with one session earlier | -0.003 | -0.211 | +0.044 | +0.208 |
| Correlation of return size with one session earlier | +0.231 | -0.161 | +0.139 | +0.402 |
Read the top row across. The whole sample annualised mean is 10.42 per cent; the trailing year mean ranged from -40.2 to 65.8, with 15 per cent of rolling years below zero. The whole sample number is not a level the series visits. It is an average of a quantity that spent its time elsewhere.
An average is a weighted average over conditions
Label each session by the realised volatility of the twenty sessions before it, using only information available before that session opened, and sort the sample into three equal groups.
| Preceding condition | Sessions | Annualised mean | Annualised volatility |
|---|---|---|---|
| Calm | 1,118 | +6.25 per cent | 10.40 per cent |
| Middling | 1,117 | +2.75 per cent | 13.45 per cent |
| Stressed | 1,118 | +22.06 per cent | 21.90 per cent |
The pooled figure of +10.36 per cent is exactly the weighted average of those three, weighted by how often each condition occurred. That is an identity, not an approximation, so the pooled figure inherits the sample's mix, and the mix is not a constant of nature.
| Year | Calm, per cent of sessions | Middling | Stressed | Pooled mean under this mix, per cent |
|---|---|---|---|---|
| 2013 | 10 | 20 | 70 | +16.70 |
| 2014 | 25 | 64 | 11 | +5.72 |
| 2015 | 8 | 43 | 48 | +12.40 |
| 2016 | 10 | 49 | 41 | +11.05 |
| 2017 | 81 | 19 | 0 | +5.60 |
| 2018 | 49 | 31 | 20 | +8.38 |
| 2019 | 36 | 38 | 27 | +9.12 |
| 2020 | 2 | 32 | 66 | +15.54 |
| 2021 | 33 | 22 | 46 | +12.69 |
| 2022 | 17 | 18 | 65 | +15.88 |
| 2023 | 65 | 35 | 0 | +5.04 |
| 2024 | 39 | 37 | 25 | +8.91 |
| 2025 | 59 | 24 | 16 | +8.00 |
Holding every conditional average fixed and changing only the mix, the pooled figure runs from +5.04 per cent under the 2023 mix to +16.70 under the 2013 mix, a gap of 11.66 points, none of it from a change in behaviour. When a long run average is quoted and the future assumed to deliver it, the assumption is not about returns. It is about the mix recurring, and the table says it does not.
A standard deviation understates what it is usually asked to describe
The second condition fails in a way that is easy to quantify. Standardise the daily returns by their own standard deviation and count the days landing far out.
| Threshold | Days observed | Days a normal law expects | Implied frequency under a normal law |
|---|---|---|---|
| Beyond 3 standard deviations | 44 | 9.106 | once in 1.5 years |
| Beyond 4 standard deviations | 17 | 0.214 | once in 64 years |
| Beyond 5 standard deviations | 11 | 0.002 | once in 7,062 years |
| Beyond 6 standard deviations | 7 | under 0.001 | once in 2.1 million years |
Excess kurtosis is 17.6 and skewness -1.15. The worst session was -13.90 per cent, or 13.7 standard deviations. Under a normal law of this sample's standard deviation that has probability near ten to the power -42, about once in ten to the power 40 years. The universe is not that old. The number does not say how unlikely the day was. It says how badly the standard deviation describes this series at the edges.
That matters wherever a standard deviation is a risk input rather than a description. A position sized so a three standard deviation day is survivable is sized against an event this sample produced 44 times in 14 years, which is the arithmetic behind position sizing from first principles.
A correlation is a period, not a property
| Pair | Whole sample | Rolling low | Rolling high | Range |
|---|---|---|---|---|
| Nifty Bank and Nifty IT | +0.31 | -0.09 | +0.64 | 0.73 |
| Nifty 50 and Nifty FMCG | +0.65 | +0.33 | +0.89 | 0.55 |
| Nifty 50 and Nifty Pharma | +0.52 | +0.19 | +0.74 | 0.55 |
| Nifty Bank and Nifty FMCG | +0.48 | +0.11 | +0.80 | 0.70 |
Every pair here ranges wider than most would accept as estimation error. The banking and technology pair spans 0.73, from -0.09 around 2014-12-04 to +0.64 around 2020-05-15. A hedge ratio built on the whole sample figure is built on the midpoint of a quantity that spent years far from it, and these correlations rise exactly when the diversification was supposed to be working.
Differencing removes the trend, and takes something with it
The standard instruction is to difference until the series passes a test. What it leaves out is what differencing discards, which is the ordering. A return series is a set of numbers plus the sequence they arrived in, and every distributional statistic ignores the second half.
Shuffle the real daily returns into a random order ten thousand times. Every shuffle has the same mean, standard deviation, skewness, kurtosis and final level, because log returns add and addition does not care about order.
| Property | The real ordering | Random reorderings of the same returns |
|---|---|---|
| Mean, volatility, skewness, kurtosis, final level | as measured | identical by construction |
| Deepest peak to trough fall | 38.4 per cent | median 30.5 per cent, ninety fifth percentile 43.4 |
| Longest stretch below a previous high | 492 sessions | median 777, ninety fifth percentile 1,457 |
The two path results move in opposite directions, which is the part worth sitting with. The real path fell further from its high than 86 per cent of shuffles, because the large declines arrived together. Yet it spent far less time below a previous high, 492 sessions against a median of 777, because the recovery was concentrated too. Clustering made the worst moment worse and the waiting shorter.
Neither fact is recoverable from the return distribution. Both live in the ordering, and differencing throws the ordering away. For a method needing only the distribution that is a fair price. For anything path dependent, which includes drawdown limits, margin, leverage, stop placement and every decision about whether to keep going, the differenced series has discarded the thing being asked about. Calling it the stationary version of the data is true and incomplete.
Failing to reject is not evidence of stationarity
Tests for a unit root are the usual instrument, and their output is the most misreported number in applied work. The test puts a unit root in the null hypothesis. Rejecting it is evidence against a unit root. Failing to reject is not evidence for stationarity, in the way failing to convict is not a finding of innocence. Before quoting results the instrument has to be proved both ways.
| Series | Test statistic | What it supports |
|---|---|---|
| Log of the index level, 3,385 sessions | -0.99 | does not reject a unit root |
| Daily log returns, 3,373 observations | -11.15 | rejects a unit root decisively |
| Log of 20 session realised volatility | -5.63 | rejects a unit root |
Now the two controls, both simulated. On random walks of the same length, where a unit root is present by construction, the test rejected in 5.3 per cent of 1,000 runs against its nominal five, so its size is right. Mean reverting series, stationary by construction, give the other direction.
| Persistence per session | Half life of a shock | Unit root rejected |
|---|---|---|
| 0.980 | 34 sessions, about 0.1 years | 100.0 per cent |
| 0.990 | 69 sessions, about 0.3 years | 99.8 per cent |
| 0.995 | 138 sessions, about 0.6 years | 71.4 per cent |
| 0.999 | 693 sessions, about 2.8 years | 7.9 per cent |
Read the last row. A series returning halfway to its mean in about 2.8 years is stationary by construction, and this test on 3,373 daily observations identified it as such only 7.9 per cent of the time. Over 14 years of daily data a slowly mean reverting series and a random walk are not distinguishable here, so a non rejection tells you the series is one or the other, and that is all.
A test result is evidence about persistence, not a licence to treat what follows as stable. The same distinction makes pair trading on Indian equities harder than it looks, where a test is read as establishing a relationship it only failed to rule out.
A series that spans a rule change is two series
Everything above treats non stationarity as market behaviour. The sharper version is that the measurement itself changes, on published dates that have nothing to do with the market.
The most recent example is six weeks old. From 3 August 2026, the official closing price of Indian stocks with derivative contracts is set by a closing auction rather than by the volume weighted average price of the last thirty minutes of continuous trading, the method for years before. Since index levels are computed from constituent prices, an index built on those stocks now closes by a different mechanism than before that date. This archive holds 34 such sessions, far too few to measure anything, so the right response is to record the date and treat any comparison crossing it as crossing a boundary. SEBI issued a further consultation paper on 12 September 2026 on expiry day settlement prices in light of the new session, comments invited until 3 October 2026, so the boundary is still moving.
It is not the only one. On 1 September 2025, index and stock derivative expiry moved from Thursday to Tuesday on one Indian exchange and to Thursday on the other, following a circular of 26 May 2025. Every day of week statistic on a window spanning that date pools two weekly structures.
| Weekday | 2024-09-01 to 2025-09-01 | 2025-09-01 to 2026-09-18 | Change, percentage points | In standard errors |
|---|---|---|---|---|
| Monday | 0.757 per cent | 0.665 per cent | -0.091 | -0.73 |
| Tuesday | 0.572 per cent | 0.547 per cent | -0.026 | -0.25 |
| Wednesday | 0.392 per cent | 0.668 per cent | +0.275 | +2.73 |
| Thursday | 0.632 per cent | 0.475 per cent | -0.158 | -1.44 |
| Friday | 0.728 per cent | 0.615 per cent | -0.113 | -1.08 |
Only one of the five weekdays moved by more than two standard errors, so nothing there establishes an expiry effect and nothing is meant to. That is the point. The reason not to pool these windows is not that the table is significant. It is that a rule published in advance makes them measurements of two weekly structures, and no quantity of data repairs that. A test would have to find the break. The circular already told you where it is.
The cleanest demonstration is older and fully measurable. On 31 March 2021, following a revision to the index maintenance methodology announced the previous month, the exchange moved the earnings basis in its published valuation ratios from standalone to consolidated. That session the index closed at 14,690.70 against 14,845.10, a move of -1.04 per cent, while the published price to earnings ratio went from 40.43 to 33.20, a move of -17.88 per cent. Price alone implies 40.01, so the definition change accounts for 6.81 points and an implied jump of +20.5 per cent in the earnings base. Published dividend yield moved -10.3 per cent the same session, the book value ratio -0.2 per cent.
That session is the largest one day move in the published valuation ratio in 3,385 sessions, larger than any during the sharpest declines in the sample. The largest apparent event in the series is not an event. Anyone averaging valuation across that date, or fitting a mean reverting model to it, is fitting a model to an administrative decision.
A quieter version sits in the same files. The broad index appears under 3 published names, the current one covering 2,689 sessions since 2015-11-09. A script selecting rows by that name silently drops the other 696, about 21 per cent of the record, and raises no error. The series was renamed, not interrupted, and only a row count reveals it.
What to assume, and what to check
The usual advice is to hunt for a transformation that makes the data stationary. On the evidence above that hunt does not terminate. Differencing fixes the level and leaves the variance clustered, modelling the variance leaves its variability moving, and any long sample carries dated breaks no transformation addresses, because they are not features of the process at all.
The workable position is narrower and more demanding. Do not ask whether the series is stationary. Name the property your method needs stable, name the horizon it must hold over, and test that property over that horizon.
A volatility forecast for next month needs one month volatility to be predictable from recent one month volatility, which is checkable, and the 0.23 correlation in return size is why it works at all. A correlation sizing a quarterly hedge needs stability over quarters, and the range above says size against the range, not the point. A long run average used for planning needs the mix of conditions to recur, which the reweighting table shows is the load bearing assumption and is rarely stated. A backtest crossing 3 August 2026, 1 September 2025 or 31 March 2021 needs the measurement constant across it, and it is not.
Put this way the question stops being a technical preliminary and becomes the substance of the work. Each assumption is specific, each is checkable, and each can be reported as holding, failing, or untestable at the sample size available. That is a more useful output than a single test statistic, and it separates a research process from a search for the transformation that finally makes the warning go away. On identifying conditions rather than pooling them, see regime detection in Indian markets.
Frequently asked questions
What does stationary actually mean?
That the process generating the numbers does not change over time. In the strict form, any group of observations has the same joint distribution wherever you take it from, so no date is special. Almost nothing in finance satisfies that, so the working version is weaker: constant mean, constant variance, and a covariance depending only on the gap between two observations.
Why is a price series not stationary?
Because its level has no fixed centre. A shock is carried into every later price rather than decaying, so the mean and variance of the level both depend on how long you have watched. Over 3,385 sessions of the Indian broad index the unit root statistic on the log level is -0.99, nowhere near the five per cent critical value of -2.86.
If returns are stationary, is the problem solved?
No, returns are closer, not there. Direction says almost nothing about the next return, but size does: the correlation between the size of one day's move and the next is 0.23 and is still 0.12 twenty sessions later. Variance arriving in clusters is not constant variance, so the weaker definition fails too.
What is wrong with an average computed on such a series?
It is a weighted average over conditions, and the weights are whatever happened to occur in your sample. Grouping the Indian index sample by the volatility of the preceding month and reweighting the conditional averages by each year's own mix gives a pooled figure ranging from +5.04 to +16.70 per cent annualised. Only the mix changed.
Does differencing fix it?
It removes the trend and takes the ordering with it. Shuffling the real Indian daily returns preserves mean, volatility, skewness, fat tails and final level exactly, yet the deepest fall moves from 38.4 per cent to a median of 30.5 across ten thousand shuffles, and the longest stretch below a previous high from 492 sessions to a median of 777. Everything a distribution describes is unchanged. The experience of holding it is not.
If a unit root test does not reject, is the series stationary?
No. Failing to reject says the data did not supply enough evidence against a unit root, a statement about the evidence and not the series. Simulating series stationary by construction makes it plain: when mean reversion has a half life of about 2.8 years, the test rejected in only 7.9 per cent of runs at this sample size.
How do I know the test itself is working?
Run it where the answer is known, in both directions, before trusting it on real data. On simulated random walks, where a unit root is present by construction, it rejected in 5.3 per cent of runs against a nominal five, so its size is right. With fast mean reversion it rejected in 100 per cent, so it has power. A test only ever shown to agree with you has not been checked.
What counts as a structural break, and how do I find one?
A dated change in the rules, the instrument or the measurement that makes observations before and after it observations of different things. Indian examples include the expiry weekday change of 1 September 2025 and the closing auction for the official closing price of derivative eligible stocks from 3 August 2026. You find them by reading exchange and regulator circulars for your sample period, not by searching the data.
Should I just use a shorter, more recent window then?
A shorter window reduces the chance conditions changed inside it and raises the chance that what you measured is noise. No length escapes both. The choice is not between a stationary sample and a non stationary one but between two kinds of error, and stating which you chose is the part usually left out.
So what should I actually check before using a statistic?
Name the property you need stable, name the horizon it must hold over, and test that rather than the series. A volatility forecast for next month needs one month volatility to have been predictable from the preceding month. A correlation sizing a hedge needs its range over windows the length of the hedge. The general question has a known answer and it is no.
How these numbers were produced. The archive is every daily index close file published for 2013-01-01 to 2026-09-18, checked for calendar holes first; the largest gap between sessions is 6 days. The broad series is assembled across all 3 names it has carried, and includes the 14 weekend special sessions the archive holds, such as budget days and muhurat trading. Returns are daily log differences of the closing level, annualised at 247 sessions, this archive's measured average. A return is accepted as one session only when the later file's own change column agrees with the two closes. In all, 11 returns between 2013 and 2016 fail that test, each where the archive has no file for the session or sessions in between; they stay in the price path, which really did make those moves, and in the shuffles, but are left out of every per-session statistic, which therefore rests on 3,373 daily returns. The file for 13 March 2023 carries a wrong change column, so that day's return is taken from its consecutive closes. Rolling statistics use trailing 250 session windows except volatility, shown at 20 and 60 sessions where stated. Volatility states come from the twenty sessions ending the day before, so no label uses information from the day it labels. The unit root test is the augmented Dickey and Fuller regression with a constant and 28 lagged differences, the lag order set by the standard rule for this sample size. Shuffles use ten thousand reorderings with a fixed seed. Simulated results, labelled simulated throughout, use 1,000 runs per row of first order autoregressive series the same length as the real sample.
The position is stated as at 19 September 2026. Exchange methodologies and microstructure rules change on dated effective days, and several here changed within the last eighteen months; confirm the current position with the exchange and the regulator before building on any of it. Nothing here is investment advice or a statement about future returns.
Ready to go deeper than this article?
Bharath Shiksha is a 90-volume curriculum across 6 stages, from chart reading at ₹14,999 through capital raising, or the full bundle at ₹1,49,999. Knowing which stability assumption a method is making, and checking that one rather than the series, is taught here as method rather than as a result to memorise.
Take the free diagnostic →