The extreme observations in market data are usually real, and a rule that deletes them moves every risk estimate in the dangerous direction
The short answer
A cleaning rule is a model of what an error looks like, and on market data it is usually wrong in one direction. Measured on 3,395 sessions of the broad Indian index from 2013-01-01 to 2026-09-18, every one of the 44 sessions beyond three standard deviations was matched by at least 8 of ten sector indices moving the same way, so all of them were market moves, not recording errors. Deleting the 11 sessions that moved more than 5 per cent, 0.32 per cent of the record, cuts the worst drawdown from 38.4 to 19.6 per cent. Across ten common rules and two treatments the same data gives annualised volatility from 12.5 to 16.1 per cent, excess kurtosis from 0.2 to 22.2 and a worst drawdown from 16.0 to 44.0 per cent. Against closes deliberately planted five standard deviations wrong, the detector that caught nearly all of them (98 per cent) with the fewest real sessions flagged (2) was a cross check against the sector indices. And the number being cleaned changed definition on 3 August 2026, when a closing auction began setting the official close of every stock with derivatives: in the exchange's own file the last trade now equals the close on 99.3 per cent of their sessions, against 1.8 per cent before.
Every figure below was computed from the exchange's own daily files, and the method is stated in full at the end so the work can be redone. The page makes one argument and measures it several ways: an extreme observation in market data is guilty only when something outside the series says so, and a rule applied without that evidence changes the answer in a direction the rule chose.
An outlier is a question about the recording, not about the number
Two different things produce an extreme value. A bad tick is a recorded value that does not describe what traded: a mistyped price, a decimal slip, a value carried forward from a stale feed, a print from a different instrument, or two sessions merged into one. An extreme move is a value that describes exactly what traded, however improbable it looked. A daily return of minus 13 per cent looks the same in both cases. The number carries no mark of its origin.
A single series offers one signature. An isolated bad print enters and leaves: the close is wrong on one session and right on the next, so the return series shows a move followed by its mirror image. A real move usually persists. Usually is the problem. A rule that flags any move beyond three standard deviations that is at least half undone the next session flags 12 sessions of this record, and every one was matched by the sector indices. They include the rally of 3 June 2024 after the general election exit polls and the fall of 5.93 per cent on counting day, 4 June 2024, of which 54 per cent came back the next session. By the textbook signature, both were bad ticks.
The series can also hide a real extreme entirely. On 13 March 2020 the market wide circuit breaker halted trading for 45 minutes after the broad index fell 10 per cent in the morning. That day's file records a low 10.79 per cent below the previous close and a close 3.81 per cent above it. In a close to close series the session is an ordinary up day, and the extreme lives only in the low column. Whether an observation is an outlier depends on the resolution of the record as much as on the market.
What separates the two cases is not in the broad index at all. That is the idea the rest of this page tests: corroboration from series that share a cause, rather than the size of the number.
The close is already a cleaned number, and on 3 August 2026 its rule changed
The official close of a stock is not its last trade, and was not before the auction either. Until 31 July 2026 the exchange set the close of every stock in continuous trading as the volume weighted average price of its trades in the last 30 minutes, the method the regulator itself describes in its consultation paper of 12 September 2026. That average is a cleaning rule applied before anyone downloads anything: a stray print at 3:29 pm moves a 30 minute average only in proportion to its share of the half hour's volume. It is one reason bad ticks are rare in a daily close series and common in tick data.
From 3 August 2026 the rule is different for every stock with derivatives. SEBI's circular of 16 January 2026 (SEBI/HO/47/11/11(3)2025-MRD-POD2/I/2765/2026) introduced a closing auction session in the equity cash segment. The exchange's operating circular NSE/CMTR/73362 of 18 March 2026 sets the mechanics: the close is the equilibrium price discovered in the auction, the auction runs inside a band of 3 per cent either side of a reference price, and the reference price is the volume weighted average of trades between 3:00 and 3:15 pm, which becomes the close itself if no equilibrium price is found. Circular NSE/CMTR/75479 of 30 July 2026 made it live from 3 August. Stocks without derivatives keep the 30 minute average.
The change is visible in the exchange's own security file, and it breaks a common cleaning check. The daily bhavcopy reports both the last traded price and the close, and comparing the two is a routine test for a late erratic print.
| Group | Sessions | Stock sessions | Last price equals close, per cent | Median gap, basis points | 99th percentile gap, basis points |
|---|---|---|---|---|---|
| With derivatives, before 3 August | 63 | 13,230 | 1.8 | 13.2 | 84.3 |
| With derivatives, from 3 August | 34 | 7,140 | 99.3 | 0.0 | 0.0 |
| All other stocks, before 3 August | 63 | 139,341 | 6.1 | 25.8 | 241.8 |
| All other stocks, from 3 August | 34 | 80,537 | 8.0 | 24.4 | 248.4 |
For the 210 stocks with derivatives the median gap between the last price and the close fell from 13.2 basis points to 0.0. For every other stock nothing moved. The check did not fail on 3 August. It stopped measuring anything for those stocks, silently, on a date no column header records.
Three consequences follow for anyone cleaning a close series. The close of a stock with derivatives now carries a hard clamp of 3 per cent around its reference price, normally the average of the 15 minutes before the auction, which is winsorisation performed by the exchange, by rule, before the data reaches you. An index built from those closes inherits the auction: the operating circular computes the indicative index close during the auction from the constituents' indicative auction prices, under its paragraph 6.5.3. And the definition is moving again. The consultation paper of 12 September 2026 proposes new auction timings, two options for the expiry day settlement price, one of them a blended average across the last 30 minutes of continuous trading and the auction, and an end to the indicative index value during the auction, with comments invited until 3 October 2026. A guide that describes the Indian close as a plain 30 minute average is now describing only the stocks without derivatives.
Most anomalies in a real archive are missing sessions, and the file reports them itself
Before any outlier rule there is a decision about what counts as a session, and it is usually made by a download script. The daily all index file carries the exchange's own points change for each index against the previous session's close, so each file contains a check on its neighbour: the close minus the reported change must equal the previous file's close. Across the 3,371 files a weekday downloader collected for this record, the check failed 26 times, and every failure had a cause that the archive itself could establish.
| Cause | Disagreements | Largest session hidden | Largest composite created | Resolution |
|---|---|---|---|---|
| Budget day sessions on a weekend | 4 | -2.51 on 2020-02-01 | -2.13 | Special session file fetched and merged |
| Diwali muhurat sessions | 5 | +0.52 on 2023-11-12 | +1.75 | Special session file fetched and merged |
| Other Saturday sessions: two short ones in 2013 and 2014, the replacement session of 20 January 2024, the recovery site drills of 2 March and 18 May 2024 | 5 | -0.23 on 2024-01-20 | -1.88 | Special session file fetched and merged |
| Weekday sessions the archive answers with not found, 2013 to 2016 | 11 gaps, 12 sessions | -2.15 on 2015-09-04 | -3.38 | Close recovered as the next close minus its reported change; one two session gap leaves a single composite |
| A change column computed against the wrong session | 1 | none missing | reports -2.47 against a true -1.49 | Files trusted over the change column on 2023-03-13 |
Fourteen are weekend sessions, the most recent the Sunday session of 1 February 2026, held because the Union Budget was presented that day. A downloader that requests weekdays only never asks for any of them. Twelve more are weekday sessions between 2013 and 2016 for which the archive answers not found, which a downloader that treats not found as a holiday silently skips.
The damage is to individual sessions. In the downloaded series 25 returns span more than one session, and one became a false outlier: 2015-09-07 shows -3.38 per cent, beyond three standard deviations, and is the sum of -2.15 per cent on 2015-09-04 and -1.26 per cent on 2015-09-07, neither of which is. The budget session of 1 February 2020 fell 2.51 per cent and does not exist in the download at all; it is folded into a Monday of -2.13 per cent.
The file also supplies the repair. The next file's close minus its change recovers the missing close, and the method can be validated wherever both numbers exist: on the fourteen special sessions it matched the special file's own close in 14 of 14 cases, to within 0.05 of an index point. The reconciled record holds 3,396 sessions and 3,395 returns, with one two session composite left where two consecutive weekdays are missing. In aggregate the repair is small: annualised volatility moves from 16.20 to 16.13 per cent and excess kurtosis from 17.37 to 17.53. It is the session by session record that changes, and a cleaning rule acts on sessions. Security level data has a third category, the split or bonus that is neither a bad tick nor a move, which the guide to backtest integrity covers.
Ten detectors, tested on the real record and on planted errors
Each detector below is a standard way of finding outliers, run with its textbook threshold on the reconciled record, and each is measured twice. First, what it flags in the real record, which shows no sign of a recording error once reconciled: every session any detector flagged moved the same way as at least 6 of the ten sector indices, and every one beyond three standard deviations as at least 8, so each flag is an observation a deletion rule would destroy rather than an error it would remove. Second, what it catches when an error is known to be there. Every eligible session in turn, 3,321 of them, had its close replaced by one recorded three, five or eight standard deviations too low, correct again the next session. The experiment is exhaustive rather than random, so it needs no seed and carries no sampling error.
| Detector | Rule | Real sessions flagged | 3 sd print caught | 5 sd print caught | 5 sd caught, stressed third | 8 sd print caught |
|---|---|---|---|---|---|---|
| Fixed threshold | a move beyond 5 per cent either way | 11 | 2 | 43 | 42 | 99 |
| Fixed threshold | a move beyond 10 per cent either way | 1 | 0 | 0 | 0 | 1 |
| Standard deviations, whole record | beyond 3 standard deviations of the whole record | 44 | 48 | 98 | 97 | 99 |
| Standard deviations, repeated | 3 standard deviations, recomputed after each removal until nothing more goes | 85 | 80 | 99 | 98 | 99 |
| Standard deviations, trailing | beyond 4 standard deviations of the previous 63 sessions | 19 | 41 | 85 | 60 | 97 |
| Median absolute deviation | beyond 3 robust standard deviations of the whole record | 98 | 84 | 99 | 99 | 99 |
| Median absolute deviation, rolling | beyond 3 robust standard deviations of the 21 sessions centred on it | 94 | 69 | 94 | 86 | 99 |
| Reversal signature | beyond 3 standard deviations and at least half undone the next session | 12 | 43 | 97 | 94 | 99 |
| Cross check | beyond 4 robust standard deviations and the median sector move under 50 per cent of it | 2 | 49 | 98 | 96 | 99 |
| Cross check, direction only | beyond 4 robust standard deviations and fewer than 7 of the 10 sectors moving the same way | 0 | 21 | 66 | 67 | 68 |
Read across a row and each detector's trade is visible. The fixed 5 per cent threshold caught only 43 per cent of prints five standard deviations wide, because a planted error adds to the day's real move and a down print on an up day stays inside the line; the 10 per cent threshold caught none at that size. The whole record standard deviation rule caught 98 per cent and flagged 44 real sessions to do it. The median absolute deviation caught 99 per cent and flagged 98. The cross check caught 98 per cent and flagged 2. Only one family tells an error from a move, because only one looks at evidence the error cannot fake.
The standard deviation rule defeats itself
A standard deviation rule asks whether a value is far from the centre in units of the spread, and it estimates the spread from the same values it is testing. On fat tailed data the extremes inflate the yardstick meant to catch them. The whole record standard deviation of the broad index is 1.016 per cent a day. Clip everything beyond three of those, recompute, and repeat until nothing more goes, and the core that remains has a standard deviation of 0.799 per cent. The 85 sessions removed on the way, 2.5 per cent of the record, had widened the yardstick by 27 per cent. The procedure took 5 passes because each removal shrinks the scale and exposes the next layer, which is why repeated clipping of a fat tailed series does not converge on the errors. It converges on a thinner market.
A window that includes the value being tested makes the problem absolute. In a sample of n observations no single value can sit more than n minus 1, divided by the square root of n, sample standard deviations from the sample mean, a classical inequality. For a 10 session window that ceiling is 2.85, so a three standard deviation rule computed over a trailing window that includes today cannot flag anything, on any data; on this record the largest value it produced was 2.69. For a 20 session window the ceiling is 4.25, so a five standard deviation rule built that way never fires. The usual rolling window functions in dataframe libraries include the current row unless told otherwise, so this is not a contrived construction.
Excluding the tested value fixes the arithmetic and not the masking. A trailing rule compares each session with the 63 before it, and every real shock raises its threshold for the next three months.
The planted errors measure the cost. The trailing rule caught 99 per cent of five standard deviation prints in the calmest third of sessions and 60 per cent in the most turbulent third. Across March and April 2020 it caught 26 per cent, and the centred robust filter, whose window was full of real crash sessions, caught 5 per cent. A detector built on the series' own spread is least able to see an error exactly when an error does most damage to a risk estimate, because that is when the spread is widest. The measurement of how far a standard deviation understates the tail is the same fact seen from the other side.
Remove, winsorise or flag: three different records
Once a session is flagged there are three things to do with it. Removal deletes it, so the path the returns imply skips that move. Winsorisation keeps the session and clamps its size to the rule's own boundary. Flagging changes nothing and keeps the list. Each produces a different record, and the statistics below describe the record, not the market.
| Rule | Sessions affected | Volatility, removed | Volatility, winsorised | Excess kurtosis, removed | Excess kurtosis, winsorised | Worst drawdown, removed | Worst drawdown, winsorised |
|---|---|---|---|---|---|---|---|
| As recorded, flagged only | 0 | 16.13 | 17.53 | -38.4 | |||
| 5 per cent threshold | 11 | 14.59 | 15.27 | 2.51 | 4.06 | -19.6 | -28.6 |
| 10 per cent threshold | 1 | 15.68 | 15.94 | 8.31 | 11.21 | -33.2 | -36.3 |
| 3 standard deviations, whole record | 44 | 13.44 | 14.44 | 0.73 | 1.39 | -17.2 | -22.0 |
| 3 standard deviations, repeated | 85 | 12.69 | 13.90 | 0.24 | 0.60 | -16.0 | -20.0 |
| 4 standard deviations, previous 63 sessions | 19 | 14.72 | 15.49 | 5.21 | 7.79 | -17.6 | -31.5 |
| 3 robust standard deviations | 98 | 12.50 | 13.74 | 0.16 | 0.43 | -17.7 | -19.3 |
| 3 robust, centred 21 sessions | 94 | 15.04 | 15.62 | 22.19 | 18.85 | -35.4 | -37.7 |
| Reversal signature | 12 | 15.67 | 15.91 | 18.71 | 17.68 | -44.0 | -40.4 |
| Cross check | 2 | 16.01 | 16.05 | 17.63 | 17.48 | -41.0 | -38.4 |
The spread is the finding. The same data gives annualised volatility from 12.5 to 16.1 per cent, excess kurtosis from 0.2 to 22.2, and a worst drawdown from 16.0 to 44.0 per cent. None of the rules is unreasonable on its face; each is a standard construction. Removing the 11 sessions beyond 5 per cent roughly halves the worst drawdown, and after the robust rule removes 98 sessions the worst episode in the record is no longer the pandemic crash at all but a decline from 2024-09-26 to 2025-03-04.
The direction is not random, and the mechanism is the skew. The record's largest moves lean down: of the 44 sessions beyond three standard deviations 25 were falls and 19 rises, and of the 98 robust flags 59 were falls and 39 rises. A symmetric rule applied to an asymmetric tail deletes more losses than gains, so the mean rises while the volatility falls, and every risk statistic moves toward calm.
Two rows run the other way, and they matter more than the ones that flatter, because they show that a cleaning rule does not simply lower risk estimates. It replaces them with the rule's own shape. The centred robust filter deleted 94 sessions and raised excess kurtosis from 17.5 to 22.2: its window could see an isolated shock or an ordinary move in a calm fortnight, but not a crisis session surrounded by other crisis sessions, so it removed the shoulders of the distribution and kept the clustered extremes that dominate the tail. Deleting the 12 sessions with a bad print's reversal signature deepened the worst drawdown to 44.0 per cent, because the rebound sessions it deleted, 13 and 20 March 2020 among them, sat inside the worst fall in the record.
Flagging is the only treatment under which a risk estimate remains an estimate of what happened. It costs nothing but discipline: the statistic is computed on the full record, the flagged list is kept beside it, and any sensitivity is reported as a second figure rather than substituted for the first. The worst drawdown is already a noisy sample statistic, and a cleaning rule adds a second source of variation to it that has nothing to do with the market.
A move that appears everywhere is a market move
A bad print is a property of one record. A market move is a property of many. On the real record the evidence is unanimous: all 44 sessions beyond three standard deviations were matched by at least 8 of the ten sector indices moving the same way, 32 of them by all ten, and the median sector move ran between 0.40 and 1.54 times the broad index's. The planted prints were matched by nothing, which is why the cross check caught them.
The converse does not hold. A move that appears alone is suspect, not guilty. Scanning the ten sector indices for sessions beyond six robust standard deviations of their own history, on which the broad index moved less than 1 per cent, finds 10, and every one of them is real.
| Sector index | Session | Move | Broad index | Median of the other nine | Turnover |
|---|---|---|---|---|---|
| Public sector banks | 2017-10-25 | +29.63 | +0.86 | +0.53 | 24.9 times |
| Media | 2019-01-25 | -16.37 | -0.64 | -0.55 | 7.2 times |
| Media | 2021-09-14 | +14.40 | +0.14 | +0.24 | 17.3 times |
| Media | 2021-09-22 | +13.57 | -0.09 | +0.82 | 17.5 times |
| Information technology | 2013-01-11 | +9.33 | -0.29 | -1.50 | 6.1 times |
| Consumer staples | 2017-07-18 | -6.73 | -0.90 | +0.03 | 5.1 times |
| Pharmaceuticals | 2015-07-21 | -6.99 | -0.86 | -1.48 | 4.1 times |
| Metals | 2022-05-23 | -8.14 | -0.32 | -0.34 | 2.2 times |
| Realty | 2014-10-14 | -9.02 | -0.26 | -0.03 | 2.7 times |
| Information technology | 2026-02-04 | -5.87 | +0.19 | +0.92 | 3.1 times |
The public sector bank index rose 29.6 per cent on 25 October 2017, the day after the government announced a ₹2.11 lakh crore recapitalisation plan for public sector banks. The metals index fell 8.1 per cent on 23 May 2022, after export duties on iron ore and steel were notified on 21 May. No other sector corroborates either, because no other sector shared the cause. The corroboration lives in the series that did: the banking index, which includes the largest public sector bank, rose 3.4 per cent that October day, and on every one of the 10 sessions the sector index's own turnover ran at least 2.2 times its recent median. A bad print arrives with ordinary turnover, because nothing traded at it.
The cross check has a threshold too, and it produced two false flags. On 25 and 26 March 2020 every sector rose, but the broad index rose more than twice the median sector because banking led, at 8.0 and 6.1 per cent, and banking weighs heavily in a capitalisation weighted index. The rule as specified, that the median sector must reach half the broad move, flagged both. A version that asks only whether seven of ten sectors moved the same way flags no real session and misses 34 per cent of the planted prints, the sessions on which most sectors happened to drift in the print's direction. Every detector here trades one error for another, and the best of them still needs a person to read its list.
A cleaning rule chosen after the result is a fitted parameter
The statistics table above holds 19 versions of one record. Take the simplest statistic anyone reports, the ratio of mean daily return to its volatility, annualised, for merely holding the index. As recorded it is 0.63. Across the versions it runs from 0.59 to 1.29, and the highest comes from removing the 98 robust flags: a doubling produced by editing 2.9 per cent of the sessions, without one assumption about the market changing. These are gross measurements of a published index, not of anything tradable, and they describe the past only.
A researcher who tries several cleaning rules and keeps the one that makes the result look best has run a search, and the reported figure is the maximum of that search. It is the same defect as choosing the null after seeing the result, and it enters through the gap the measurement of multiple testing describes: once the result is known, whether an outlier day is a data error or a real event is decided by someone with an interest in one answer. No dishonesty is required. The rule that made the result legible is simply the one that gets written up.
The defence is sequence, not virtue. Fix the definition of a session, the detector, its threshold, its window and the treatment before looking at any performance statistic. If the rule changes afterwards, which is sometimes right, report both versions and say which came first.
What a stated cleaning rule contains
The session. Whether weekend special sessions are included, how sessions missing from the source are treated, and how many returns span more than one session. On this record that decision alone changed 25 returns and created one false outlier.
The detector. The statistic, the threshold, the window, and whether the window includes the value it tests. A rule stated without its window cannot be checked, because the window decides what it can see.
The treatment. Removed, winsorised to which boundary, repaired from which related series, or flagged only.
The list. Every affected session, with its size and the evidence for calling it an error: the related series that failed to move, the turnover, the source file. A flag without evidence is an opinion.
Both answers. The headline statistics with the flagged sessions and without them. If the two disagree materially, the conclusion rests on the cleaning rule, and a reader is entitled to know that.
The order. When the rule was fixed relative to the first result, and the data snapshot it was applied to. A result someone else can check needs the file, not a description of the period.
What cleaning is for
Cleaning exists to remove errors of recording. It is not a way to make a distribution resemble the model someone intends to fit to it. In an index record the extreme sessions are not contamination of the sample; they are the part of the sample a risk estimate exists to describe, and on this record deleting 11 of them halves the worst drawdown of 13.7 years. A series cleaned of them describes a milder market than the one that traded, and a rule tested on it has been tested on a market that did not exist.
The work that survives is unglamorous: reconcile the calendar, check each file against its neighbour, test every candidate against series that share its cause, delete only what the evidence convicts, and report the rest flagged. Treating the cleaning rule as a modelling decision, stated in advance and reported both ways, is method rather than a figure to memorise, and it is how this material is taught here.
Frequently asked questions
What is the difference between a bad tick and an outlier?
An outlier is any observation far from the rest. A bad tick is a recorded value that does not describe what traded: a mistyped price, a decimal slip, a stale value carried forward, or two sessions merged into one. The number alone cannot say which it is. In 3,395 sessions of the broad Indian index every one of the 44 sessions beyond three standard deviations was matched by at least 8 of ten sector indices moving the same way, so all of them were real.
Should outliers be removed from price data before a backtest?
Not by default. Deleting the 11 sessions that moved more than 5 per cent, 0.32 per cent of the record, cut the worst drawdown of the broad index from 38.4 to 19.6 per cent. Remove only what you can show is an error, keep the rest flagged, and report results both ways.
Why does a three standard deviation filter fail on market returns?
Because the extremes inflate the standard deviation the filter uses to find them. On this record the whole-sample figure is 1.016 per cent a day while the core left after repeated clipping has 0.799, so the extremes widen the scale by 27 per cent. A window that includes the value being tested is worse: its z-score cannot exceed n minus 1 divided by the square root of n, so a 3 sigma rule over 10 sessions can never flag anything.
Is the median absolute deviation the robust fix?
It fixes the inflation and exposes the other problem. The robust scale is 0.753 per cent against a standard deviation of 1.016, so a robust 3 rule flagged 98 real sessions, 2.9 per cent of the record. A robust method measures departures from a normal core, and daily returns depart from normality as a matter of fact, not error.
What does winsorising do to a return series?
It keeps the session and clamps its size to the rule's boundary, so the count of sessions is unchanged but the tail is capped. It is gentler than deletion and still decides the answer: clamping at 5 per cent took the worst drawdown from 38.4 to 28.6 per cent and excess kurtosis from 17.5 to 4.1.
How can I tell a data error from a real market move?
Look outside the series. A real market move appears in related series that share its cause, arrives with abnormal turnover, and persists. In the planted error experiment a cross check against ten sector indices caught 98 per cent of closes recorded five standard deviations low and flagged only 2 real sessions. Every sector-only move in the record came with at least 2.2 times its usual turnover.
What changed about the official closing price on 3 August 2026?
For stocks with derivatives the close is now the equilibrium price of a closing auction, bounded at 3 per cent either side of the average price of trades between 3:00 and 3:15 pm, which becomes the close if no auction price is found. Before that date, and still for other stocks, the close was the volume weighted average price of the last 30 minutes. SEBI proposed further changes in a consultation paper of 12 September 2026.
Why can a downloaded daily series contain returns that span two sessions?
Because Indian exchanges trade on some weekends and the archive is missing some weekdays. A downloader that requests weekdays only never asks for budget day, muhurat or recovery drill sessions, and one that reads a not found reply as a holiday skips real sessions. In the record as a weekday downloader first collected it, 25 returns spanned more than one session, and one of them became a false three sigma outlier.
Can cleaning make risk look worse rather than better?
Yes, which is the clearest sign that a cleaning rule is a model, not a correction. A rolling robust filter deleted 94 sessions and raised excess kurtosis from 17.5 to 22.2, because it removed ordinary moves in calm windows and kept the crisis extremes. Removing the 12 sessions with a bad print's reversal signature deepened the worst drawdown to 44.0 per cent.
Are bad ticks common in daily Indian index data?
Rare at the level of the daily close, because the close is already an aggregate: a 30 minute average, and since 3 August 2026 an auction price for stocks with derivatives. Once this record was reconciled no recording error remained among its extremes. The defects that did exist were calendar defects: 26 disagreements between files, almost all of them sessions a downloader never fetched.
As at 23 September 2026. Market data runs to 2026-09-18. The closing auction framework described here is under active review: the consultation paper of 12 September 2026 proposes changes to auction timings and to expiry day settlement, with comments invited until 3 October 2026. Confirm the current closing price rules and the session calendar on the exchange's circulars before relying on anything here.
How the figures were produced. Daily closes of the broad fifty share index and of ten sector indices (banking, information technology, consumer staples, pharmaceuticals, automobiles, metals, energy, realty, media and public sector banks) were read from the exchange's daily all index close files: 3,371 weekday files from 2013-01-01 to 2026-09-18 plus the 14 weekend special session files fetched from the same archive, which the shared cache has held since 23 September 2026, with index names stitched across the late 2015 renaming. Each file's close minus its reported points change was checked against the previous file's close, with a tolerance of the larger of 0.06 points and 0.012 per cent; 11 missing weekday closes were recovered as the next file's close minus its reported change, a method that matched the special files' own closes in 14 of 14 cases. Returns are daily log returns; volatility is the population standard deviation times the square root of 252; excess kurtosis is the fourth standardised population moment minus three; the worst drawdown is the largest peak to trough fall of the path the returns imply, with removed sessions contributing no move. Thresholds: fixed at 5 and 10 per cent; 3 standard deviations of the whole record, once and repeated until nothing more is removed; 4 standard deviations of the previous 63 sessions; 3 robust standard deviations, 1.4826 times the median absolute deviation, of the whole record and of the 21 sessions centred on each session; the reversal signature, beyond 3 standard deviations and at least half undone the next session; and the cross check, beyond 4 robust standard deviations with a median sector move under half the broad move, or with fewer than 7 of 10 sectors moving the same way. Winsorisation clamps to each rule's own boundary. Planted errors: each of 3,321 eligible sessions in turn had its close multiplied by the exponential of minus m times 0.01016, for m of 3, 5 and 8, with the next close unchanged; the experiment is exhaustive, so it uses no random numbers, seeds or replications. The regime thirds split those sessions at the one third and two thirds points of the trailing 63 session standard deviation. The closing auction measurement compares the last traded price with the close in the exchange's daily security bhavcopy for equity series stocks with at least one trade, 2026-05-01 to 2026-09-18, dating each file by the session recorded inside it, so the 4 files in that span that were requested for holidays and repeat an earlier session count once or not at all; stocks with derivatives are the 210 underlyings of stock futures in the derivatives bhavcopy of 18 September 2026. Every figure is a gross measurement of a published index or file, not of any tradable outcome.
What could not be verified. The regulator's website could not be reached from this environment, so the circular of 16 January 2026 is cited through the exchange circulars that quote its number and date, and the consultation paper of 12 September 2026 through contemporaneous press reports of its contents rather than the paper itself. The exchange circular that sets out the 30 minute closing method for stocks outside the auction, NSE/CMTR/67774 of 30 April 2025, could not be retrieved, so that description rests on the consultation paper's summary as reported. The purpose of the short Saturday sessions of 11 May 2013 and 22 March 2014 was not established, and nor was the reason the archive lacks twelve weekday files. The list of stocks with derivatives is the one current on 18 September 2026, applied to the whole measurement window.
Bharath Shiksha is an educational publisher and not a SEBI-registered investment adviser or research analyst. Nothing on this page is a recommendation, a forecast or a statement about any future period.
Ready to go deeper than this article?
Bharath Shiksha is a 90-volume curriculum across 6 stages, from chart reading at ₹14,999 through capital raising, or the full bundle at ₹1,49,999. Data handling, from what counts as a session to what counts as an error, is taught as part of method, with the evidence for every decision written down before the result is seen.
Take the free diagnostic →