The reported number is the least informative thing in a performance claim
The short answer
A reported result is a number attached to a search, and the search is the part that is almost never reported. Six questions decide whether the number means anything, and they rank by decisiveness almost exactly opposite to how often they are answered: how many variants were tried, whether the result was chosen after being seen, what costs were charged, what period was used and who chose it, how large the sample was, and what became of the discards. The first is the most informative and is usually missing, not because it was withheld but because it was never written down. To calibrate: in a simulation where no variant has any edge at all, the best of 100 candidates measured over 3 years typically shows a ratio of annual result to annual variability of 1.42. That is what nothing looks like once it has been searched. Most claims are unverifiable rather than false, and the two require entirely different responses.
This page is about the structure of claims, not about anyone who makes them. Every figure here was computed from the method stated beside it, and the simulations carry their parameters and seeds so they can be rerun.
A result is a number attached to a search
A performance figure describes what happened to one rule over one stretch of history under one set of assumptions. That description can be completely accurate and still carry almost no information, because the quantity a reader wants is not the figure. It is the figure relative to what the same process would have produced with nothing in it.
Those two come apart entirely. A ratio of annual result to annual variability of 1.4 over three years is an achievement if the rule was written down first and run once. The identical 1.4 is unremarkable as the survivor of a hundred attempts. Nothing about the number distinguishes them, because the distinguishing information lives in the process, and the process is the part that does not get written up. The reported figure is not wrong. It is the one quantity invariant to what you are trying to learn.
Six questions, ordered by how much they decide
| Question | What the answer decides | How often a claim answers it |
|---|---|---|
| 1. How many variants were tried? | The bar the number has to clear before it is surprising at all | Almost never |
| 2. Was the result chosen after being seen? | Whether this was a test or a description of something already known | Almost never |
| 3. What costs were charged? | The sign of the result, deterministically, at any real turnover | Sometimes, usually as a single figure |
| 4. What period, and who chose it? | Whether the window was fixed in advance or found in the data | Usually stated, rarely justified |
| 5. How large was the sample? | The width of the uncertainty the number is carrying | Usually stated |
| 6. What became of the discards? | Whether the winner was remarkable among its own siblings | Effectively never |
Read the right hand column against the left. Decisiveness and disclosure run in almost opposite directions, and that is not a coincidence. The answers at the top were never recorded; the answers at the bottom fall out of the output of any backtesting run. Disclosure follows what is easy to retrieve, not what is worth knowing.
Question five, what a given sample size can establish, is worked through in a separate piece on detecting an edge, and what a result should be compared against is handled in the piece on choosing the right null. Neither is repeated here. This page is about the questions sitting above both of them.
The trial count is the most informative number, and it is almost never recorded
Testing many candidates and reporting the best one is not a subtle error. It is the ordinary shape of research and it is unavoidable, since nobody finds anything worth reporting on the first attempt. The problem is that the conventional way of judging a result assumes exactly one attempt was made.
The arithmetic is closed form and unforgiving. If a single candidate with no edge clears a conventional threshold one time in twenty, the chance that at least one of several independent candidates clears it is one minus the chance that none does. After 14 candidates that passes one half. Not one in twenty. More likely than not.
| Variants searched | Typical ratio of the winner | Winner exceeded this one search in twenty | Chance at least one clears the single test bar |
|---|---|---|---|
| 1 | 0.00 | 0.95 | 5.0 per cent |
| 5 | 0.65 | 1.34 | 22.6 per cent |
| 14 | 0.96 | 1.55 | 51.2 per cent |
| 20 | 1.05 | 1.62 | 64.2 per cent |
| 100 | 1.42 | 1.90 | 99.4 per cent |
| 500 | 1.73 | 2.14 | 100.0 per cent |
| 2,000 | 1.96 | 2.34 | 100.0 per cent |
The second column is the one to sit with. Every candidate in that population has an edge of exactly zero. The winner of a search over 100 of them typically measures 1.42, and over 2,000 it typically measures 1.96. These are not unlucky outcomes or adversarial constructions. They are the median, and half of all such searches do better.
The third column is the bar a result from such a search must clear to be as surprising as a single untested result, rising from 0.95 for one candidate to 2.34 for 2,000. A reader who does not know the count does not know which row applies, and therefore does not know what bar the figure was up against.
One honest qualification. This assumes the candidates are independent, and real variants are usually related, so twenty settings of one rule are fewer than twenty experiments. That reduces the inflation without removing it, and it cuts both ways: a researcher who exhausts one rule and then moves to a structurally different one has made two genuinely independent attempts and will not count it that way.
Why the missing number is usually an absence, not a concealment
The uncharitable interpretation of the gap is both unfair and analytically wrong. Research does not proceed as a series of labelled experiments. It proceeds as continuous adjustment: a rule is written, it behaves oddly somewhere, a condition is added, a parameter is moved, a filter is introduced because one stretch of the record looked wrong. None of that feels like a new hypothesis while it is happening. It feels like fixing the same one. By the time something is worth showing anyone, the number of distinct attempts is not merely undisclosed; it is unknown to the person who made them, and no amount of good faith afterwards reconstructs it.
That changes how the question should be asked. Demanding a trial count as an accusation produces defensiveness and no information. Asking it as a methodological question produces the answer that is actually available, which is a description of the search: whether a protocol was written first, whether variants were logged, whether the reported result was the first thing tried or the last. That description is obtainable and nearly as useful as the count.
Costs are the one question with a deterministic answer
Everything above is statistical, and a statistical objection can in principle be met with more data. The cost question is not like that. Turnover multiplied by round trip cost is the annual drag, nothing about it is uncertain, and it resolves completely once the two numbers are stated. The table computes that drag at three stated cost levels and expresses it in the same ratio units used throughout, so it can be set directly against a reported figure.
| Turnover | 0.10 per cent per round trip | 0.20 per cent | 0.40 per cent | Ratio consumed at 0.20 per cent |
|---|---|---|---|---|
| 2 round trips a year | 0.2 | 0.4 | 0.8 | 0.03 |
| 12 round trips a year | 1.2 | 2.4 | 4.8 | 0.16 |
| 52 round trips a year | 5.2 | 10.4 | 20.8 | 0.69 |
| 250 round trips a year | 25.0 | 50.0 | 100.0 | 3.33 |
The bottom row decides most intraday claims. At 250 round trips a year and a stated 0.20 per cent round trip cost, the drag is 50 percentage points a year, which in ratio units is 3.33. A result must clear that before costs merely to reach zero after them. Set it against the winner column of the search table and the difficulty is visible: the drag alone exceeds what a search over thousands of empty candidates typically produces.
These cost levels are stated parameters rather than measurements; what an Indian round trip is built from component by component is set out in the piece on the real cost of a trade. The point is not the levels. It is that a claim stating its turnover but not its cost model has omitted the only variable in the discussion that can be settled by multiplication. And a result given as net with a single cost figure is not obviously better than one given as gross with the model shown, since the second can be recomputed under other assumptions and the first cannot be taken apart at all.
The period is a parameter, and the best window inside a flat record is not flat
A reporting period feels like a fact about the data rather than a decision. It is a decision whenever the record could have been cut somewhere else and was cut there. The effect is easy to underestimate, so here it is measured: generate a record with no edge at all, ten years long, and find the best twelve month stretch inside it.
| Length of the full record | Typical best window ratio | Best window ratio, one record in ten |
|---|---|---|
| 3 years | 1.46 | 2.46 |
| 5 years | 1.84 | 2.73 |
| 10 years | 2.26 | 3.06 |
| 20 years | 2.56 | 3.28 |
Inside a ten year record that went nowhere, the typical best year measures 2.26. That figure is produced by nothing whatsoever, and a reader shown only that year has been shown an accurate description of a real period. Note the direction of the table: the longer the underlying record, the better its best window, so a long track record makes selective reporting easier rather than harder.
The structural question is therefore not "what period was used" but "how many periods were available, and who chose among them". Any claim covering a window shorter than the available history carries this problem whether or not anyone intended it to. The defence is cheap and almost never offered: report across several start dates and give the range, so a reader can see whether the conclusion survives moving the window.
Survivorship runs in three directions
Survivorship is usually discussed once, as a property of a data set that dropped its delisted constituents. It is better understood as a filter operating at three separate stages, each pushing the same way and each compounding the last.
Surviving strategies. The search problem above, seen from the other end. Out of every set of variants one is reported and the rest are deleted, so the reported one is the maximum of its set, and the maximum of a set with no edge in it is positive by construction.
Surviving accounts. A record that stops cannot be observed continuing. Simulating 40,000 accounts with no edge whatsoever and then looking only at those that never sustained a decline beyond a stated size produces a positive average from a population whose true average is zero.
| Decline treated as a stopping point | Share that survived | Mean of everybody | Mean of survivors only |
|---|---|---|---|
| 15 standard deviations | 11.9 per cent | -0.05 | +27.85 |
| 25 standard deviations | 51.8 per cent | -0.05 | +14.18 |
| 40 standard deviations | 86.8 per cent | -0.05 | +4.90 |
| 60 standard deviations | 98.8 per cent | -0.05 | +0.61 |
At a stopping point of 25 standard deviations, about 52 per cent of these accounts survive and their mean result is +14.18 against a population mean of -0.05. Nobody had an edge. The filter created the appearance of one, and created more of it the tighter it was.
Surviving publishers. The same arithmetic one level up. A record that does not work stops being shown, and whoever produced it moves on without announcing it, so the visible population is the upper part of a distribution rather than the whole of it. Truncation gives the size of that effect exactly.
| Records remain visible above a ratio of | Share visible | True mean of everyone | Mean of what a reader sees |
|---|---|---|---|
| 0.0 | 50.0 per cent | 0.00 | 0.46 |
| 0.3 | 30.2 per cent | 0.00 | 0.67 |
| 0.6 | 14.9 per cent | 0.00 | 0.90 |
| 1.0 | 4.2 per cent | 0.00 | 1.23 |
Even the mildest filter does damage. If records simply stop being shown once they turn negative, which is the weakest possible version of the effect, the visible average is 0.46 against a true average of zero. Raise the threshold and the gap widens. The three stages then multiply, because a reader encounters a result that was the best of its search, produced by an account still running, published by someone still publishing.
Unverifiable is not false, and the difference is the whole of fair reading
Everything above is a reason to withhold confidence. None of it is a reason to allege anything, and collapsing the two is the commonest error a sceptical reader makes. Three states are distinguishable, and only the third supports any conclusion about the claim itself.
| State | What it means | Correct response | Common error |
|---|---|---|---|
| Verified | An independent record confirms the figures | Grade it on the questions verification does not answer | Reading verification as endorsement of the method |
| Unverifiable | No independent record exists either way | Assign no evidential weight and move on | Reading it as disproved, or as proved by its confidence |
| Contradicted | An available record disagrees with the figures | The only state supporting a conclusion about the claim | Reaching this state from the one above it |
The overwhelming majority of published claims sit in the middle row, which is uninformative in both directions: no independent record exists because none was ever created, the normal condition of a private research process rather than evidence of anything. Keeping the distinction changes what you do next. Reading an unverifiable claim as false puts you in a dispute about somebody's character that you cannot settle and that teaches you nothing. Reading it as unverifiable leaves you holding a useful fact: this claim cannot move your beliefs, so go on looking for evidence that can. The second is a working position. The first is an argument.
What the verification perimeter actually covers, and where it stops
India now has formal infrastructure for verifying performance figures, and its exact boundary is more useful to know than its existence. A framework for a Past Risk and Return Verification Agency was laid down by a circular dated 4 April 2025. An agency was recognised for the role, an exchange consented to act as the data centre under it, a pilot was inaugurated in December 2025, and the framework was operationalised on 4 May 2026 by a circular dated 29 April 2026. Covered entities were given three months to enrol, a deadline of 3 August 2026 extended to 3 September 2026 by a circular dated 3 August 2026. A covered entity that has not enrolled may not communicate certified past performance data to clients or prospective clients, and from 3 May 2028 a covered entity may communicate only verified metrics, with no use of past performance data for any period before the framework was operationalised.
What it verifies is specific. Several dozen risk and return metrics are computed independently from transaction data drawn from exchange and clearing records rather than from figures supplied by the entity. That construction is what makes it verification rather than attestation, and it also means a covered entity cannot present a selected favourable stretch, because the computation runs on the whole record the settlement system holds.
| Situation | Inside the perimeter? | Why |
|---|---|---|
| A registered investment adviser communicating past performance | Yes | A covered category carrying on a regulated activity |
| A registered research analyst communicating past performance | Yes | A covered category carrying on a regulated activity |
| An algorithmic trading offering by a covered participant | Yes | Brought in explicitly by the framework |
| A person not registered in any covered category | No | The perimeter attaches to the registered person, not to the claim |
| A historical simulation of a rule, by anyone | No | Metrics are computed from settled transactions, and a simulation produced none |
| A personal trading record shown as education | No | Not a communication of performance by a covered entity to clients |
The last two rows are the structurally interesting ones, and neither describes anybody evading anything. A backtest cannot enter a transaction based verification system because there are no transactions in it. That is not a gap in the framework but a consequence of what verification means when it is done properly, and it follows that the largest category of published strategy claim by volume, the simulated one, sits outside any transaction based verification, whoever publishes it and however scrupulous they are.
An older layer sits alongside it. An advertisement code for the same registered categories, issued in April 2023 and effective from 1 May 2023, restricts what may appear in their communications, including references to past performance, superlative claims and anything implying an assured outcome, and requires prior approval from the relevant supervisory body. Again the perimeter is defined by the registered person. So the absence of verification on a claim tells you almost nothing about the claim, because the great majority of claims sit structurally outside the system rather than having failed it.
What a well presented result looks like
A checklist of suspicions is half useful without a positive standard beside it, and the standard is achievable. Nothing below requires anything beyond keeping records while working.
| What is stated | What it lets a reader do |
|---|---|
| The number of variants considered | Apply the right row of the search table instead of guessing |
| The protocol, fixed before the data was examined | Tell a test from a description of something already seen |
| Gross and net side by side, cost model shown | Recompute it under different cost assumptions |
| The whole period, including what did not work | See whether the window was found or fixed |
| The sample size and what it could have detected | Know whether a null result would have been visible |
| The alternative it was compared against | Judge whether the comparison carried content |
| The distribution of the discarded variants | See whether the winner stood apart from its siblings |
| The size at which it was run, and capacity | Know whether the result is available at scale |
The seventh row is the most powerful and the rarest. If a search produced two hundred variants and the whole distribution of their results is shown with the reported one marked on it, a reader sees immediately whether the winner stands apart from its siblings or sits in the upper part of a cloud shaped exactly as the search table predicts. Showing the discards answers the first question, the sixth and most of the second in one chart, and costs nothing beyond having kept them.
A result failing this standard is not thereby a bad result. It is an ungraded one. Almost all honest work falls short of it, for the ordinary reason that the records were not kept at the time.
Now run the checklist on your own work
Applied outwards, this checklist mostly produces a long list of results you cannot grade, which is true and not very useful. Applied inwards it produces something you can act on. Take the most recent thing you found that looked promising and answer the six questions in order. How many variants did you try before this one, counting every parameter moved and every filter added? Was the reporting period settled before you saw the results or after? What cost model did you charge, and would the result survive doubling it? Over how many observations, and what could that number have detected? And what became of everything you discarded: do you still have it, and what did its distribution look like?
Most people fail on the first two, and the failure is instructive rather than shameful. The trial count is unrecoverable because it was never recorded. The period was settled after the results were seen because that is when the attractive window became visible. Neither involves dishonesty, and neither can be repaired retrospectively.
That is the finding of this page. The checklist is not primarily a reading tool; it is a specification for a research record, and every item on it can only be satisfied while the work is being done. A trial log costs one line per attempt. A protocol written before the data is examined costs ten minutes. Keeping the discards costs disk space. The whole difference between a gradeable result and an ungradeable one is a set of records that are nearly free at the time and impossible to create afterwards. So when you meet a claim you cannot grade, the useful response is not to decide whether to believe it. It is to notice which record would have settled it, and to make sure that record exists in your own work.
Frequently asked questions
What is the single most useful question to ask about a performance claim?
How many variants were tried before this one was reported. That number sets the bar the result has to clear, and without it the result cannot be graded. A ratio that would be striking from a single pre-specified rule is ordinary as the winner of a search over a few hundred candidates, and nothing about the number itself tells you which situation you are in.
Why is the trial count almost never disclosed?
Because it was usually never recorded. Research proceeds by trying something, adjusting it and trying again, and the adjustments are not experienced as separate experiments at the time. By the point a result is worth reporting, the count cannot be reconstructed from memory. Its absence is far more often a recording failure than a concealment.
Does a missing trial count mean the claim is wrong?
No. It means the claim is unverifiable on that dimension, which is a different and much more common state than being false. An unverifiable claim carries no evidential weight, and that is the correct handling. Treating it as a fabrication puts you in an argument about someone's honesty instead of an argument about evidence.
What does a selected best result look like when nothing has an edge?
In the simulation on this page, the best of one hundred zero edge variants measured over a three year record typically shows a ratio of annual result to annual variability of about 1.4, and one search in twenty produces a winner near 1.9. No variant in that population has any edge by construction. That is the calibration for what is unremarkable.
Why do costs come before sample size in the ordering?
Because costs are deterministic and statistics are not. A stated turnover multiplied by a stated round trip cost gives a drag that can be computed exactly, and at high turnover that drag is large enough to reverse the sign of a result with no uncertainty involved. A sample size question widens a confidence range. A cost question can delete the result.
Is a strategy claim covered by any Indian verification requirement?
The perimeter attaches to a registered person carrying on a regulated activity, not to a claim as such. Under the framework operationalised in May 2026, registered investment advisers, registered research analysts and algorithmic trading providers have past risk and return metrics computed by an independent agency from transaction data drawn from exchange and clearing records. Anyone not registered in those categories sits outside it entirely.
Can a backtest ever be independently verified under that framework?
Not by that mechanism, because it computes metrics from settled transactions and a historical simulation produced none. This is a structural boundary rather than a gap anyone is exploiting: there is nothing in the settlement record for a simulation to be checked against. The largest category of published strategy claim by volume is therefore outside transaction based verification, whoever publishes it.
What does a genuinely well presented result contain?
The number of variants considered, the protocol fixed before the data was examined, gross and net side by side with the cost model stated, the whole period rather than a chosen stretch, the sample size and what it could have detected, the alternative it was compared against, and the distribution of the discards. A reader shown the discards can see whether the winner was remarkable among its own siblings.
What happens when I apply this checklist to my own research?
Most people fail on the first two questions, because the trial count was never kept and the reporting period was settled after the results were seen. That is the intended outcome. The checklist is not primarily a filter for other people's claims. It is a specification for a research record, and the only moment it can be satisfied is while the work is being done.
How these numbers were produced. A candidate with no edge, measured over a 3 year record, produces a sample ratio of annual result to annual variability that is approximately normal with mean zero and standard deviation one over the square root of the number of years, which is 0.5774 here. The search table is the exact distribution of the maximum of that many independent such candidates: the typical winner is that standard deviation multiplied by the inverse normal of one half raised to the power of one over the count, and the one in twenty column raises 0.95 to the same power. Those closed forms were checked against a Monte Carlo of 40,000 searches with a fixed seed and agreed to within 0.003 ratio units at every count checked. The cost table multiplies a stated turnover by a stated round trip cost; the 15 per cent annual variability used to convert it into ratio units is a stated conversion factor, not a measurement of any market. The survivorship figures come from 40,000 simulated accounts of 500 trades each with a per trade edge of exactly zero and a fixed seed, filtered only by whether a peak to trough decline of the stated size ever occurred. The publisher figures are the closed form mean of a normal distribution truncated at the stated threshold. The best window figures come from 1,500 simulated records per length with a fixed seed, taking the largest sum over any contiguous 246 observation window and dividing by the square root of that window. All simulation results are illustrative of the structure of the problem, and are not measurements of or predictions about any market, any strategy or any person's results.
The position is stated as at September 2026. Circulars, enrolment deadlines and the scope of the verification framework are amended from time to time; confirm the current position directly against the regulator's own circulars before relying on anything here, and take advice on your own circumstances.
Ready to go deeper than this article?
Bharath Shiksha is a 90-volume curriculum across 6 stages, from chart reading at ₹14,999 through capital raising, or the full bundle at ₹1,49,999. Reading a claim well and building a record that could survive being read are the same skill, and both are taught here as method rather than as a set of warnings.
Take the free diagnostic →