The reported number is the least informative thing in a performance claim

The short answer

A reported result is a number attached to a search, and the search is the part that is almost never reported. Six questions decide whether the number means anything, and they rank by decisiveness almost exactly opposite to how often they are answered: how many variants were tried, whether the result was chosen after being seen, what costs were charged, what period was used and who chose it, how large the sample was, and what became of the discards. The first is the most informative and is usually missing, not because it was withheld but because it was never written down. To calibrate: in a simulation where no variant has any edge at all, the best of 100 candidates measured over 3 years typically shows a ratio of annual result to annual variability of 1.42. That is what nothing looks like once it has been searched. Most claims are unverifiable rather than false, and the two require entirely different responses.

This page is about the structure of claims, not about anyone who makes them. Every figure here was computed from the method stated beside it, and the simulations carry their parameters and seeds so they can be rerun.

A result is a number attached to a search

A performance figure describes what happened to one rule over one stretch of history under one set of assumptions. That description can be completely accurate and still carry almost no information, because the quantity a reader wants is not the figure. It is the figure relative to what the same process would have produced with nothing in it.

Those two come apart entirely. A ratio of annual result to annual variability of 1.4 over three years is an achievement if the rule was written down first and run once. The identical 1.4 is unremarkable as the survivor of a hundred attempts. Nothing about the number distinguishes them, because the distinguishing information lives in the process, and the process is the part that does not get written up. The reported figure is not wrong. It is the one quantity invariant to what you are trying to learn.

Six questions, ordered by how much they decide

Six questions, ordered by how much each one decides Six stacked bars of decreasing width. The widest at the top is the number of variants tried, which decides the most. Below it in order come whether the result was chosen after being seen, what costs were charged, what period was used, how large the sample was, and what became of the discarded results. Width represents how much the answer changes the reading of the claim. Answer the top one and the rest often stop mattering 1. How many variants were trieddecides everything, it sets the bar the number must clear2. Was it chosen after being seendecides whether this was a test at all, or a description3. What costs were chargeddecides the sign of the result, deterministically4. What period, and who chose itdecides whether the window was found or fixed in advance5. How large was the sampledecides the width of the uncertainty around the number6. What became of the discardsdecides whether the winner was ordinary among its siblings Almost every published claim answers the bottom two and is silent on the top two.
The ordering is not by how often the question is asked. It is by how much the answer moves the reading, which runs almost exactly opposite to how often it is answered.
The structural questions, in order of how much the answer moves the reading
QuestionWhat the answer decidesHow often a claim answers it
1. How many variants were tried?The bar the number has to clear before it is surprising at allAlmost never
2. Was the result chosen after being seen?Whether this was a test or a description of something already knownAlmost never
3. What costs were charged?The sign of the result, deterministically, at any real turnoverSometimes, usually as a single figure
4. What period, and who chose it?Whether the window was fixed in advance or found in the dataUsually stated, rarely justified
5. How large was the sample?The width of the uncertainty the number is carryingUsually stated
6. What became of the discards?Whether the winner was remarkable among its own siblingsEffectively never

Read the right hand column against the left. Decisiveness and disclosure run in almost opposite directions, and that is not a coincidence. The answers at the top were never recorded; the answers at the bottom fall out of the output of any backtesting run. Disclosure follows what is easy to retrieve, not what is worth knowing.

Question five, what a given sample size can establish, is worked through in a separate piece on detecting an edge, and what a result should be compared against is handled in the piece on choosing the right null. Neither is repeated here. This page is about the questions sitting above both of them.

The trial count is the most informative number, and it is almost never recorded

Testing many candidates and reporting the best one is not a subtle error. It is the ordinary shape of research and it is unavoidable, since nobody finds anything worth reporting on the first attempt. The problem is that the conventional way of judging a result assumes exactly one attempt was made.

The arithmetic is closed form and unforgiving. If a single candidate with no edge clears a conventional threshold one time in twenty, the chance that at least one of several independent candidates clears it is one minus the chance that none does. After 14 candidates that passes one half. Not one in twenty. More likely than not.

What a search over candidates with no edge produces. Exact for independent candidates, over a 3 year record.
Variants searchedTypical ratio of the winnerWinner exceeded this one search in twentyChance at least one clears the single test bar
10.000.955.0 per cent
50.651.3422.6 per cent
140.961.5551.2 per cent
201.051.6264.2 per cent
1001.421.9099.4 per cent
5001.732.14100.0 per cent
2,0001.962.34100.0 per cent

The second column is the one to sit with. Every candidate in that population has an edge of exactly zero. The winner of a search over 100 of them typically measures 1.42, and over 2,000 it typically measures 1.96. These are not unlucky outcomes or adversarial constructions. They are the median, and half of all such searches do better.

The third column is the bar a result from such a search must clear to be as surprising as a single untested result, rising from 0.95 for one candidate to 2.34 for 2,000. A reader who does not know the count does not know which row applies, and therefore does not know what bar the figure was up against.

What a search produces when nothing has an edge Two rising curves against a flat horizontal line. As the number of variants searched increases along a logarithmic horizontal axis, the typical result of the best variant rises, even though every variant has no edge by construction. The flat line is the bar a single untested result would have to clear. The curves cross it early and keep climbing. Every variant here has no edge whatsoever. This is what the winner still shows. 0.00.51.01.52.02.5 1101001,000 the bar a single result would have to clear, 0.95 1.051.421.96 typical winner one winner in twenty Number of variants searched Vertical axis: measured ratio of annual result to annual variability over a 3 year record. Exact for independent variants.
The bar a result has to clear is not a property of the result. It is a property of the search that produced it, and the search is the part that does not appear in the claim.

One honest qualification. This assumes the candidates are independent, and real variants are usually related, so twenty settings of one rule are fewer than twenty experiments. That reduces the inflation without removing it, and it cuts both ways: a researcher who exhausts one rule and then moves to a structurally different one has made two genuinely independent attempts and will not count it that way.

Why the missing number is usually an absence, not a concealment

The uncharitable interpretation of the gap is both unfair and analytically wrong. Research does not proceed as a series of labelled experiments. It proceeds as continuous adjustment: a rule is written, it behaves oddly somewhere, a condition is added, a parameter is moved, a filter is introduced because one stretch of the record looked wrong. None of that feels like a new hypothesis while it is happening. It feels like fixing the same one. By the time something is worth showing anyone, the number of distinct attempts is not merely undisclosed; it is unknown to the person who made them, and no amount of good faith afterwards reconstructs it.

That changes how the question should be asked. Demanding a trial count as an accusation produces defensiveness and no information. Asking it as a methodological question produces the answer that is actually available, which is a description of the search: whether a protocol was written first, whether variants were logged, whether the reported result was the first thing tried or the last. That description is obtainable and nearly as useful as the count.

Costs are the one question with a deterministic answer

Everything above is statistical, and a statistical objection can in principle be met with more data. The cost question is not like that. Turnover multiplied by round trip cost is the annual drag, nothing about it is uncertain, and it resolves completely once the two numbers are stated. The table computes that drag at three stated cost levels and expresses it in the same ratio units used throughout, so it can be set directly against a reported figure.

Annual drag in percentage points, at stated round trip costs. The final column converts the middle cost level into ratio units at a stated annual variability of 15 per cent.
Turnover0.10 per cent per round trip0.20 per cent0.40 per centRatio consumed at 0.20 per cent
2 round trips a year0.20.40.80.03
12 round trips a year1.22.44.80.16
52 round trips a year5.210.420.80.69
250 round trips a year25.050.0100.03.33

The bottom row decides most intraday claims. At 250 round trips a year and a stated 0.20 per cent round trip cost, the drag is 50 percentage points a year, which in ratio units is 3.33. A result must clear that before costs merely to reach zero after them. Set it against the winner column of the search table and the difficulty is visible: the drag alone exceeds what a search over thousands of empty candidates typically produces.

These cost levels are stated parameters rather than measurements; what an Indian round trip is built from component by component is set out in the piece on the real cost of a trade. The point is not the levels. It is that a claim stating its turnover but not its cost model has omitted the only variable in the discussion that can be settled by multiplication. And a result given as net with a single cost figure is not obviously better than one given as gross with the model shown, since the second can be recomputed under other assumptions and the first cannot be taken apart at all.

The period is a parameter, and the best window inside a flat record is not flat

A reporting period feels like a fact about the data rather than a decision. It is a decision whenever the record could have been cut somewhere else and was cut there. The effect is easy to underestimate, so here it is measured: generate a record with no edge at all, ten years long, and find the best twelve month stretch inside it.

The best year inside a record that goes nowhere A wandering line representing ten years of results generated with no edge at all, ending close to where it began. One twelve month stretch of it is shaded and drawn in a brighter colour. That stretch, reported on its own, shows a strong measured ratio, even though the full record shows nothing. Ten years generated with no edge at all best 12 months, ratio 2.22 Full record ratio over the whole ten years: +0.07 One simulated path, fixed seed, chosen so the full ten years end where they began. Not a measurement of any market or strategy. Reporting only the shaded stretch would state nothing untrue about the shaded stretch.
A period is a free parameter with hundreds of settings. The best one inside a record that went nowhere is not modest, and choosing it requires no dishonesty at all.
The best twelve month window inside a record generated with no edge. Simulation, 1,500 runs per record length, fixed seed.
Length of the full recordTypical best window ratioBest window ratio, one record in ten
3 years1.462.46
5 years1.842.73
10 years2.263.06
20 years2.563.28

Inside a ten year record that went nowhere, the typical best year measures 2.26. That figure is produced by nothing whatsoever, and a reader shown only that year has been shown an accurate description of a real period. Note the direction of the table: the longer the underlying record, the better its best window, so a long track record makes selective reporting easier rather than harder.

The structural question is therefore not "what period was used" but "how many periods were available, and who chose among them". Any claim covering a window shorter than the available history carries this problem whether or not anyone intended it to. The defence is cheap and almost never offered: report across several start dates and give the range, so a reader can see whether the conclusion survives moving the window.

Survivorship runs in three directions

Survivorship is usually discussed once, as a property of a data set that dropped its delisted constituents. It is better understood as a filter operating at three separate stages, each pushing the same way and each compounding the last.

Survivorship in three directions Three rows, each showing a wide faint box for a whole population and a narrow bright box for the part of it that remains visible, with an arrow between them. The first row is strategies, where the search keeps only the best variant. The second is accounts, where the ruined ones stop reporting. The third is publishers, where a weak record quietly ends. In every row the population average is zero by construction and the visible figure is positive. Three filters, all pointing the same way Strategiesall variants built, mean 0.00typical 1.42the search keeps the bestAccountsall accounts opened, mean -0.05mean +14.18the ruined ones stop reportingPublishersall records produced, mean 0.00mean 0.90a weak record quietly ends Nothing in any row has an edge. Every visible mean is positive anyway.
Simulated and computed on populations with no edge by construction. Each row is in its own units, so the comparison that matters is within a row and not across them. Survivorship is not one bias applied once. It is the same bias applied at three separate stages, each of which compounds the last.

Surviving strategies. The search problem above, seen from the other end. Out of every set of variants one is reported and the rest are deleted, so the reported one is the maximum of its set, and the maximum of a set with no edge in it is positive by construction.

Surviving accounts. A record that stops cannot be observed continuing. Simulating 40,000 accounts with no edge whatsoever and then looking only at those that never sustained a decline beyond a stated size produces a positive average from a population whose true average is zero.

Accounts with no edge, 500 trades each, filtered only by survival. Simulation, 40,000 runs, fixed seed. Results in units of the per trade standard deviation.
Decline treated as a stopping pointShare that survivedMean of everybodyMean of survivors only
15 standard deviations11.9 per cent-0.05+27.85
25 standard deviations51.8 per cent-0.05+14.18
40 standard deviations86.8 per cent-0.05+4.90
60 standard deviations98.8 per cent-0.05+0.61

At a stopping point of 25 standard deviations, about 52 per cent of these accounts survive and their mean result is +14.18 against a population mean of -0.05. Nobody had an edge. The filter created the appearance of one, and created more of it the tighter it was.

Surviving publishers. The same arithmetic one level up. A record that does not work stops being shown, and whoever produced it moves on without announcing it, so the visible population is the upper part of a distribution rather than the whole of it. Truncation gives the size of that effect exactly.

What a visibility threshold does to an observed population whose true mean is zero. Closed form, over a 3 year record.
Records remain visible above a ratio ofShare visibleTrue mean of everyoneMean of what a reader sees
0.050.0 per cent0.000.46
0.330.2 per cent0.000.67
0.614.9 per cent0.000.90
1.04.2 per cent0.001.23

Even the mildest filter does damage. If records simply stop being shown once they turn negative, which is the weakest possible version of the effect, the visible average is 0.46 against a true average of zero. Raise the threshold and the gap widens. The three stages then multiply, because a reader encounters a result that was the best of its search, produced by an account still running, published by someone still publishing.

Unverifiable is not false, and the difference is the whole of fair reading

Everything above is a reason to withhold confidence. None of it is a reason to allege anything, and collapsing the two is the commonest error a sceptical reader makes. Three states are distinguishable, and only the third supports any conclusion about the claim itself.

Three states a claim can be in, and the response each one calls for
StateWhat it meansCorrect responseCommon error
VerifiedAn independent record confirms the figuresGrade it on the questions verification does not answerReading verification as endorsement of the method
UnverifiableNo independent record exists either wayAssign no evidential weight and move onReading it as disproved, or as proved by its confidence
ContradictedAn available record disagrees with the figuresThe only state supporting a conclusion about the claimReaching this state from the one above it

The overwhelming majority of published claims sit in the middle row, which is uninformative in both directions: no independent record exists because none was ever created, the normal condition of a private research process rather than evidence of anything. Keeping the distinction changes what you do next. Reading an unverifiable claim as false puts you in a dispute about somebody's character that you cannot settle and that teaches you nothing. Reading it as unverifiable leaves you holding a useful fact: this claim cannot move your beliefs, so go on looking for evidence that can. The second is a working position. The first is an argument.

What the verification perimeter actually covers, and where it stops

India now has formal infrastructure for verifying performance figures, and its exact boundary is more useful to know than its existence. A framework for a Past Risk and Return Verification Agency was laid down by a circular dated 4 April 2025. An agency was recognised for the role, an exchange consented to act as the data centre under it, a pilot was inaugurated in December 2025, and the framework was operationalised on 4 May 2026 by a circular dated 29 April 2026. Covered entities were given three months to enrol, a deadline of 3 August 2026 extended to 3 September 2026 by a circular dated 3 August 2026. A covered entity that has not enrolled may not communicate certified past performance data to clients or prospective clients, and from 3 May 2028 a covered entity may communicate only verified metrics, with no use of past performance data for any period before the framework was operationalised.

What it verifies is specific. Several dozen risk and return metrics are computed independently from transaction data drawn from exchange and clearing records rather than from figures supplied by the entity. That construction is what makes it verification rather than attestation, and it also means a covered entity cannot present a selected favourable stretch, because the computation runs on the whole record the settlement system holds.

What the verification perimeter reaches, as at September 2026
SituationInside the perimeter?Why
A registered investment adviser communicating past performanceYesA covered category carrying on a regulated activity
A registered research analyst communicating past performanceYesA covered category carrying on a regulated activity
An algorithmic trading offering by a covered participantYesBrought in explicitly by the framework
A person not registered in any covered categoryNoThe perimeter attaches to the registered person, not to the claim
A historical simulation of a rule, by anyoneNoMetrics are computed from settled transactions, and a simulation produced none
A personal trading record shown as educationNoNot a communication of performance by a covered entity to clients

The last two rows are the structurally interesting ones, and neither describes anybody evading anything. A backtest cannot enter a transaction based verification system because there are no transactions in it. That is not a gap in the framework but a consequence of what verification means when it is done properly, and it follows that the largest category of published strategy claim by volume, the simulated one, sits outside any transaction based verification, whoever publishes it and however scrupulous they are.

An older layer sits alongside it. An advertisement code for the same registered categories, issued in April 2023 and effective from 1 May 2023, restricts what may appear in their communications, including references to past performance, superlative claims and anything implying an assured outcome, and requires prior approval from the relevant supervisory body. Again the perimeter is defined by the registered person. So the absence of verification on a claim tells you almost nothing about the claim, because the great majority of claims sit structurally outside the system rather than having failed it.

What a well presented result looks like

A checklist of suspicions is half useful without a positive standard beside it, and the standard is achievable. Nothing below requires anything beyond keeping records while working.

The positive standard, and what each element lets a reader do
What is statedWhat it lets a reader do
The number of variants consideredApply the right row of the search table instead of guessing
The protocol, fixed before the data was examinedTell a test from a description of something already seen
Gross and net side by side, cost model shownRecompute it under different cost assumptions
The whole period, including what did not workSee whether the window was found or fixed
The sample size and what it could have detectedKnow whether a null result would have been visible
The alternative it was compared againstJudge whether the comparison carried content
The distribution of the discarded variantsSee whether the winner stood apart from its siblings
The size at which it was run, and capacityKnow whether the result is available at scale

The seventh row is the most powerful and the rarest. If a search produced two hundred variants and the whole distribution of their results is shown with the reported one marked on it, a reader sees immediately whether the winner stands apart from its siblings or sits in the upper part of a cloud shaped exactly as the search table predicts. Showing the discards answers the first question, the sixth and most of the second in one chart, and costs nothing beyond having kept them.

A result failing this standard is not thereby a bad result. It is an ungraded one. Almost all honest work falls short of it, for the ordinary reason that the records were not kept at the time.

Now run the checklist on your own work

Applied outwards, this checklist mostly produces a long list of results you cannot grade, which is true and not very useful. Applied inwards it produces something you can act on. Take the most recent thing you found that looked promising and answer the six questions in order. How many variants did you try before this one, counting every parameter moved and every filter added? Was the reporting period settled before you saw the results or after? What cost model did you charge, and would the result survive doubling it? Over how many observations, and what could that number have detected? And what became of everything you discarded: do you still have it, and what did its distribution look like?

Most people fail on the first two, and the failure is instructive rather than shameful. The trial count is unrecoverable because it was never recorded. The period was settled after the results were seen because that is when the attractive window became visible. Neither involves dishonesty, and neither can be repaired retrospectively.

That is the finding of this page. The checklist is not primarily a reading tool; it is a specification for a research record, and every item on it can only be satisfied while the work is being done. A trial log costs one line per attempt. A protocol written before the data is examined costs ten minutes. Keeping the discards costs disk space. The whole difference between a gradeable result and an ungradeable one is a set of records that are nearly free at the time and impossible to create afterwards. So when you meet a claim you cannot grade, the useful response is not to decide whether to believe it. It is to notice which record would have settled it, and to make sure that record exists in your own work.

Frequently asked questions

What is the single most useful question to ask about a performance claim?

How many variants were tried before this one was reported. That number sets the bar the result has to clear, and without it the result cannot be graded. A ratio that would be striking from a single pre-specified rule is ordinary as the winner of a search over a few hundred candidates, and nothing about the number itself tells you which situation you are in.

Why is the trial count almost never disclosed?

Because it was usually never recorded. Research proceeds by trying something, adjusting it and trying again, and the adjustments are not experienced as separate experiments at the time. By the point a result is worth reporting, the count cannot be reconstructed from memory. Its absence is far more often a recording failure than a concealment.

Does a missing trial count mean the claim is wrong?

No. It means the claim is unverifiable on that dimension, which is a different and much more common state than being false. An unverifiable claim carries no evidential weight, and that is the correct handling. Treating it as a fabrication puts you in an argument about someone's honesty instead of an argument about evidence.

What does a selected best result look like when nothing has an edge?

In the simulation on this page, the best of one hundred zero edge variants measured over a three year record typically shows a ratio of annual result to annual variability of about 1.4, and one search in twenty produces a winner near 1.9. No variant in that population has any edge by construction. That is the calibration for what is unremarkable.

Why do costs come before sample size in the ordering?

Because costs are deterministic and statistics are not. A stated turnover multiplied by a stated round trip cost gives a drag that can be computed exactly, and at high turnover that drag is large enough to reverse the sign of a result with no uncertainty involved. A sample size question widens a confidence range. A cost question can delete the result.

Is a strategy claim covered by any Indian verification requirement?

The perimeter attaches to a registered person carrying on a regulated activity, not to a claim as such. Under the framework operationalised in May 2026, registered investment advisers, registered research analysts and algorithmic trading providers have past risk and return metrics computed by an independent agency from transaction data drawn from exchange and clearing records. Anyone not registered in those categories sits outside it entirely.

Can a backtest ever be independently verified under that framework?

Not by that mechanism, because it computes metrics from settled transactions and a historical simulation produced none. This is a structural boundary rather than a gap anyone is exploiting: there is nothing in the settlement record for a simulation to be checked against. The largest category of published strategy claim by volume is therefore outside transaction based verification, whoever publishes it.

What does a genuinely well presented result contain?

The number of variants considered, the protocol fixed before the data was examined, gross and net side by side with the cost model stated, the whole period rather than a chosen stretch, the sample size and what it could have detected, the alternative it was compared against, and the distribution of the discards. A reader shown the discards can see whether the winner was remarkable among its own siblings.

What happens when I apply this checklist to my own research?

Most people fail on the first two questions, because the trial count was never kept and the reporting period was settled after the results were seen. That is the intended outcome. The checklist is not primarily a filter for other people's claims. It is a specification for a research record, and the only moment it can be satisfied is while the work is being done.

How these numbers were produced. A candidate with no edge, measured over a 3 year record, produces a sample ratio of annual result to annual variability that is approximately normal with mean zero and standard deviation one over the square root of the number of years, which is 0.5774 here. The search table is the exact distribution of the maximum of that many independent such candidates: the typical winner is that standard deviation multiplied by the inverse normal of one half raised to the power of one over the count, and the one in twenty column raises 0.95 to the same power. Those closed forms were checked against a Monte Carlo of 40,000 searches with a fixed seed and agreed to within 0.003 ratio units at every count checked. The cost table multiplies a stated turnover by a stated round trip cost; the 15 per cent annual variability used to convert it into ratio units is a stated conversion factor, not a measurement of any market. The survivorship figures come from 40,000 simulated accounts of 500 trades each with a per trade edge of exactly zero and a fixed seed, filtered only by whether a peak to trough decline of the stated size ever occurred. The publisher figures are the closed form mean of a normal distribution truncated at the stated threshold. The best window figures come from 1,500 simulated records per length with a fixed seed, taking the largest sum over any contiguous 246 observation window and dividing by the square root of that window. All simulation results are illustrative of the structure of the problem, and are not measurements of or predictions about any market, any strategy or any person's results.

The position is stated as at September 2026. Circulars, enrolment deadlines and the scope of the verification framework are amended from time to time; confirm the current position directly against the regulator's own circulars before relying on anything here, and take advice on your own circumstances.

Related guides

How many trades before you can tell an edge from luck

Read →

Benchmarking a strategy against the right null

Read →

The real cost of an Indian trade

Read →

Ready to go deeper than this article?

Bharath Shiksha is a 90-volume curriculum across 6 stages, from chart reading at ₹14,999 through capital raising, or the full bundle at ₹1,49,999. Reading a claim well and building a record that could survive being read are the same skill, and both are taught here as method rather than as a set of warnings.

Take the free diagnostic →