Educational Reference

Monte Carlo for Traders: Your Equity Curve Is One Draw From a Distribution

The curve at the bottom of your journal is not a property of your strategy. It is one realisation of a random process, and if the same trades had arrived in a different order you would be looking at a different worst drawdown, a different worst losing run, and possibly a different opinion about whether the system is worth trading. Monte Carlo simulation makes that hidden range visible. Instead of one curve you generate thousands and read off the percentiles. This page runs one, on a seeded 240-trade record, and publishes every number it produced.

The finding, stated first. The simulated record used here fell 14.7 percent at its worst. Rebuild the identical trades five thousand times and the median worst drawdown is 14.1 percent, the 95th percentile is 24.6 percent and the deepest is 41.8 percent. Nothing about the strategy changed. Only the order and the draw changed. A trader who had sized for 14.7 percent and met 24.6 percent would have been holding a position about 1.8 times the size the arithmetic permits. All figures illustrative and simulated.

One curve, and the thousand curves it was drawn from

Consider what a completed trading record actually is. You had an edge, or you did not, and either way each trade drew an outcome from some underlying distribution. Some were wins, some losses, some larger than expected, some scratched. Then those outcomes arrived in a particular sequence, and that sequence compounded into the curve you now stare at. The strategy determined the distribution. Chance determined the order and the exact mix.

This matters because almost every risk number a trader quotes is a property of the order, not of the distribution. Maximum drawdown is the deepest fall from any peak, which depends entirely on whether the losses happened to bunch together. The longest losing streak is a pure ordering statistic. Time spent underwater is an ordering statistic. Yet these are precisely the numbers people use to decide position size, to decide whether to abandon a system, and to decide how much psychological punishment they are signing up for. All three are being estimated from a sample of exactly one.

The fix is mechanical rather than clever. Take the trades you have, rebuild the curve many thousands of times with the order or the draw changed, and record the risk statistic each time. What comes back is not a number but a distribution, and a distribution can be interrogated. What was the median worst drawdown? What did the worst one in twenty look like? How far does the tail run beyond that? Those are answerable questions, and the answers change decisions.

There is a companion idea in the validation literature that this page assumes rather than repeats. Our page on walk-forward analysis deals with a different failure: a backtest scored on the same data used to choose its settings. That is a question about whether the edge is real. Monte Carlo asks a question that only arises after you accept the edge is real, which is what the range of outcomes around it looks like. The two are complementary and neither substitutes for the other. Both sit downstream of the wider catalogue in our guide to backtesting integrity, which lists the eight ways a curve can lie before any of this arithmetic starts to mean anything.

One caution before the numbers. Resampling can only rearrange the outcomes it is given. It has no opinion about a market condition absent from your record, it cannot invent a loss larger than any you have suffered, and it will happily produce beautifully precise percentiles from a record contaminated by optimistic fills. Everything below inherits the quality of the trades it is built from.

The trade record this page uses, stated in full

Rather than borrow a live record, this page generates one from a stated distribution, so the assumptions are visible and the exercise reproduces exactly. The generator is deliberately unremarkable: a modest win rate, a payoff ratio better than one to one, and enough untidiness in the individual outcomes that it does not behave like a coin flip with fixed stakes.

The simulation inputs. Every number on this page derives from this single seeded configuration. Illustrative and simulated.
InputSettingWhy it is set this way
Record length240 tradesAbout two years for a trader taking ten positions a month. Long enough for streaks to appear, short enough to be a realistic journal
Win rate41 percentA trend-following profile: wrong more often than right, paid more when right
Losing trades85 percent stop out at 1.0R, 15 percent gap through at 1.6RNot every loss is exactly one R. Slippage and gaps are part of the distribution, not an exception to it
Winning trades20 percent at 1.0R, 70 percent at 2.0R, 10 percent at 4.0RMost winners reach the target, some are cut early, a few run. The average win is 2.00R
ExpectancyPlus 0.177R per tradeA genuine positive edge, stated openly. The point of the page is what happens to a system that works
Position sizing1.0 percent of current equity risked per tradeFixed fractional and compounding, so drawdowns are expressed as a percentage of the equity at the time
Simulations5,000 paths per methodEnough to settle the 95th percentile to a few tenths of a percentage point
SeedFixedThe whole run reproduces identically, so any figure here can be checked or disputed

Drawing once from that generator produced the record this page treats as the observed history. It contains 98 winners and 142 losers, a realised win rate of 40.8 percent, an average win of 2.02R and an average loss of 1.06R, for a total of 47.0R over the 240 trades. Compounded at 1.0 percent risk per trade it finished 55.1 percent above its starting equity, with a worst drawdown of 14.7 percent and a longest losing run of eight.

That is the curve a trader would take to their journal review. It is a good record. It is also the only one they will ever see, and the rest of this page is about the ones they will not.

The same trades, three hundred times over

The first thing to do with a record is to stop treating it as unique. Resample it, plot the results on top of each other, and see where the actual curve sits inside the cloud.

The same system, three hundred times over Each faint line is one resampling of the identical 240 trades. The bright line is the record that actually happened. 0.75x 1.00x 1.25x 1.50x 1.75x 2.00x 2.25x 2.50x 2.75x the curve you actually got trade 1 trade 240 +3% 5th percentile finish +55% median finish +136% 95th percentile finish 4% finished below the start
Three hundred resampled paths built from the same 240 trades, drawn with replacement. The bright line is the record that happened. Every faint line is a history the same system could have produced. Illustrative and simulated.

The spread is the finding. Across the full five thousand resamplings the median path finished 55.1 percent above its start, the 5th percentile finished 2.8 percent above, and the 95th finished 135.9 percent above. The worst single path ended 36.0 percent below where it started, and 3.9 percent of paths finished below their starting equity altogether. That last figure deserves a moment, because the generator has a genuinely positive expectancy of 0.177R per trade. About one run in twenty-five of a system that works loses money over two years, purely because of how the draws fell.

Notice what this does to the ordinary experience of reviewing a losing year. A trader whose account is down after 240 trades will look for a cause, and the assumption that there must be one is nearly universal. Sometimes there is. But on this generator a losing two years is a routine outcome that requires no explanation beyond arithmetic. The corresponding error runs the other way too. The trader who happened to draw the 95th percentile path will conclude that their process is excellent, and no amount of reviewing the trades will reveal that most of the difference was the draw.

The visible spread on this chart comes from a single unchanging system, which makes any comparison narrower than that spread a measurement of noise. That applies to comparing two traders on one year, and to comparing yourself with your own best year.

The drawdown you lived through, and the one to plan for

Maximum drawdown is the number most people size against, and it is the number most badly served by a sample of one. Here is what it looks like as a distribution.

The drawdown you lived through is not the drawdown to plan for Deepest peak-to-trough fall in 5,000 resamplings of the same 240 trades, at 1.0 percent risk per trade. Illustrative and simulated. 0% 5% 10% 15% 20% 25% 30% 35% 40% 45% deepest peak-to-trough fall over the 240 trades observed 14.7% 95th percentile 24.6% 99th 30.8% The gap that ends accounts The record showed 14.7 percent. 45 percent of resamplings of the very same trades were worse, and the 95th percentile is 24.6 percent.
Deepest peak-to-trough fall across 5,000 resamplings of the identical trades, at 1.0 percent risk per trade. The observed record sits near the middle of its own distribution; the coral region is the tail beyond the 95th percentile. Planning for the green line and meeting the coral region is the failure mode this page exists to describe. Illustrative and simulated.

The observed record fell 14.7 percent at its worst. The median across resamplings is 14.1 percent, so the record was almost exactly typical: 44.9 percent of resamplings of its own trades were worse. Nothing about the observed drawdown was unlucky. That is precisely why it is dangerous. A trader whose worst drawdown was unusually mild might at least suspect they had been fortunate. A trader whose worst drawdown was typical has no such warning, and will reasonably treat it as the number to plan around.

The planning numbers are further out. The 75th percentile is 17.7 percent, the 95th is 24.6 percent, the 99th is 30.8 percent, and the deepest of the five thousand reached 41.8 percent. The gap between the observed 14.7 percent and the 95th percentile 24.6 percent is a factor of 1.7, and that factor is the whole practical content of the exercise. It is the amount by which a sizing decision anchored on the record is too aggressive.

The right way to read a distribution like this is to pick your percentile before you look at it. If you can survive the 95th percentile without changing your behaviour, you can trade the system. If it would make you stop, you are not sizing for the system you have, you are sizing for the version of it you happened to be shown. The tail beyond your chosen percentile does not go away because you chose it: in five thousand paths, the runs deeper than 24.6 percent appeared two hundred and fifty times.

There is a rider on all of this that gets lost. These percentiles describe the worst drawdown over a specific horizon, 240 trades. Trade the same system for six years instead of two and you have three times as many chances to hit a bad stretch, so the distribution of the worst drawdown over the whole period shifts deeper. Whatever percentile you plan for is tied to the length of the record you simulated, and lengthening your career is one of the things that makes it worse.

The losing run you have to sit through

Drawdown is what the account does. The losing streak is what you have to do, which is keep placing the next trade after the last several have failed. It is the statistic most responsible for traders abandoning working systems, and it has a distribution of its own.

How long a losing run should you be prepared to sit through? Longest run of consecutive losses in each of 5,000 resamplings of the same 240 trades. Illustrative and simulated. 4 5 6% 6 15% 7 20% 8 18% 9 14% 10 9% 11 7% 12 4% 13 2% 14 15 16 17 or more longest run of consecutive losing trades observed: 8 What the tail says Median longest streak 9. One run in twenty reached 14 or more. The worst of the 5,000 was 25.
Longest run of consecutive losses in each of 5,000 resamplings of the same trades. The observed record's worst run of eight is the single most common outcome, which is exactly why it feels like the natural maximum. One run in twenty produced a streak of fourteen or longer. Illustrative and simulated.

The observed record's worst run was eight losses in a row. Across resamplings the median longest run is nine, the 75th percentile is eleven, the 95th is fourteen and the 99th is seventeen. The worst of the five thousand was twenty-five consecutive losses. At a 41 percent win rate none of this is surprising arithmetic, but it is thoroughly surprising to live through, and the gap between eight and fourteen is the gap between an unpleasant month and an abandoned system.

The practical value of this chart is that it converts a psychological question into a planning question. Before you start, you can state the number of consecutive losses you must be able to take without altering your process, and you can pick that number from the distribution rather than from your imagination. On this system the honest answer is fourteen, not eight, and knowing that in advance is different in kind from discovering it at loss number nine.

Our page on risk of ruin makes the general point that long losing runs are far more common than intuition allows, using the frequency of streaks at a given win rate. What resampling adds is specificity. That page tells you streaks like this happen; this distribution tells you what your record in particular implies, including the tail, and gives you a number you can commit to in writing before the run begins.

What reshuffling cannot tell you, and why that is worth publishing

The obvious way to resample is to shuffle. Take the 240 trades you have, deal them in a new random order, rebuild the curve. It preserves the exact set of outcomes and changes only the sequence, which feels like the cleanest possible way to isolate the effect of luck in the ordering.

Run it five thousand times and something inconvenient turns up. Every single permutation finished at exactly the same value. Not approximately, exactly: the standard deviation across five thousand shuffled final results was of the order of one ten-trillionth of a percentage point, which is floating-point dust rather than variation.

The reason is arithmetic. Under fixed fractional sizing each trade multiplies your equity by one plus a fixed fraction of its R multiple, and multiplication does not care what order the factors come in. The product of the same 240 factors is the same product however you shuffle them. So a permutation test cannot say anything at all about the distribution of final outcomes. It can only speak to path statistics: drawdown, streaks, time underwater. Those it handles well, because they are entirely properties of the ordering.

This is worth stating plainly because it is a common misreading. A trader who shuffles their trades, sees a tight band of final results, and concludes that their end-of-period outcome is highly reliable has learned nothing except that they used a multiplication. The tight band is a mathematical identity, not evidence. To say anything about the spread of final outcomes you have to let the mix of trades vary, which means resampling with replacement.

Two footnotes keep this honest. If you size in fixed rupee amounts rather than as a fraction of equity, the account is additive rather than multiplicative and the order still does not change the total, for the same reason. And if your position size depends on recent results, through a rule that cuts size after a losing run for example, then order does matter to the final figure, because the sizing rule reads the sequence. Under that kind of rule a permutation test regains its meaning, and it becomes one of the better ways to test whether the sizing rule helps at all.

Time to recovery, the number nobody computes

Depth is only half of a drawdown. The other half is duration, and duration is what actually decides whether a trader is still running the system at the end of it. A 20 percent fall recovered in three weeks is an anecdote. The same 20 percent fall that takes fourteen months to reclaim is a career event.

Measuring it needs a definition, and two are useful. The first is time to recovery: the number of trades from the bottom of the deepest drawdown until the account makes a new high. The second is the longest stretch underwater: the greatest number of consecutive trades spent anywhere below a prior peak, which is closer to how the period feels from the inside.

In the observed record, the deepest trough was reclaimed 26 trades later, and the longest continuous stretch below a prior high was 86 trades. Across resamplings the median time to recovery is 25 trades, the 75th percentile is 42, the 95th is 81 and the 99th is 121. The underwater figures are much heavier: a median of 69 trades, a 75th percentile of 100, a 95th of 169 and a 99th of 224. At ten trades a month, the 95th percentile underwater stretch is about seventeen months, and the 99th is most of the two-year record.

There is a censoring problem in these numbers and it should be visible rather than buried. In 27.4 percent of resampled paths the deepest drawdown was never recovered before the record ran out. Those runs are excluded from the recovery percentiles above, so the published figures understate the true wait. That is structural rather than a defect: the deepest drawdown is disproportionately likely to be the unfinished one, because a fall that is still deepening has not yet been capped by a recovery. Any recovery statistic computed on a finite record inherits the bias, and quoting one without the censoring rate beside it is a small dishonesty.

What the duration figures change in practice is the review interval. Traders commonly decide to reassess a system after three or six bad months. On this distribution, three months of underwater is entirely ordinary and carries almost no information about whether the edge has broken. A review triggered at that point will nearly always fire on a healthy system, and the usual response, adjusting something, is how a working process gets dismantled by noise.

The percentiles, side by side

Averages are the wrong summary for any of these quantities, because every one of them is bounded on one side and has a long tail on the other. The mean maximum drawdown across resamplings is 15.1 percent, which sounds mild and is close to useless, because the number that matters is the one you have to survive rather than the one you expect.

Percentiles across 5,000 resamplings of the same 240 trades, at 1.0 percent risk per trade. Reshuffle preserves the exact set of trades; bootstrap draws with replacement from it. Illustrative and simulated, not a track record.
QuantityObservedMedian75th95th99th
Max drawdown, reshuffled14.7%13.9%16.5%21.3%24.7%
Max drawdown, bootstrap14.7%14.1%17.7%24.6%30.8%
Longest losing run, reshuffled89101416
Longest losing run, bootstrap89111417
Final result, reshuffled+55.1%+55.1%+55.1%+55.1%+55.1%
Final result, bootstrap+55.1%+55.1%+83.7%+135.9%+174.8%
Trades to recover the worst fall26254281121
Longest stretch underwater8669100169224

Two rows in that table are the argument in miniature. The reshuffled final result is a constant, for the reason given above, and printing it as five identical cells is the clearest way to show that the method has nothing to say there. And the bootstrap 5th percentile final result, which does not fit in the columns shown, is a gain of 2.8 percent, meaning one run in twenty of a system with a real edge produces two years of work for approximately nothing.

Two ways to resample, and when each is valid

The two methods used throughout this page answer different questions, and choosing between them without noticing is a common error.

Reshuffling permutes the observed trades. The multiset is fixed: exactly the same 98 winners and 142 losers, exactly the same R multiples, in a new order. It answers the question of how much of your path was down to the sequence in which your results arrived. It is the conservative choice, because it refuses to imagine any trade you did not actually take.

Resampling with replacement, the bootstrap, draws 240 trades at random from your record, allowing repeats and omissions. The mix changes: a particular large loss might appear three times or not at all. It answers a wider question, which is what a different but comparable set of trades from the same underlying distribution would have done. That is closer to the question you actually face, because your next 240 trades will certainly not be a permutation of your last 240.

Three ways of asking, three different answers Maximum drawdown at 1.0 percent risk per trade. Bars span the 5th to the 95th percentile. Illustrative and simulated. The one record that happened 14.7% A single number. It is a sample of size one. A. Reshuffled: same trades, new order 9.6% to 21.3% Holds the multiset fixed, so it can only re-time what you already have. B. Resampled with replacement 8.8% to 24.6% Lets the mix of trades vary too, so the range widens. The process that generated them 9.2% to 26.3% Visible only because this is a simulation. Both methods sit inside it. 0% 10% 20% 30% 40% maximum drawdown
The same quantity asked three ways. Bars span the 5th to the 95th percentile, the notch is the median, the dashed whisker reaches the 99th. The bottom row is the true generating process, visible here only because the record is simulated. Both resampling methods sit inside the truth, and the reshuffle sits furthest inside it. Illustrative and simulated.

Because this record was generated rather than collected, the true distribution can be computed as well, by drawing two and a half thousand fresh records from the same generator. That row is a luxury no real trader has, and it is instructive. The true 95th percentile maximum drawdown is 26.3 percent. The bootstrap of the single observed record reports 24.6 percent. Reshuffling that same record reports 21.3 percent.

So both methods understate the truth, and reshuffling understates it more. The reason is that resampling conditions on the sample you happened to get. Your 240 trades are themselves one draw, and the variation between records is a source of spread that neither method can see from the inside. The bootstrap recovers part of it by letting the mix vary. Reshuffling recovers none of it, because it holds the mix fixed by construction. The practical instruction is to run both, treat the reshuffle as a floor rather than an estimate, and remember that even the wider of the two is an underestimate of what the world can produce.

The assumption both methods share

Reshuffling and bootstrapping both assume your trades are independent, meaning the outcome of one carries no information about the next. That assumption is why it is legitimate to move them around. If losses genuinely cluster, moving them around destroys the clustering, and the simulation will report a gentler risk profile than the strategy actually has.

This is not a theoretical worry, and it is measurable. To show the size of the effect, the same generator was rewired so that wins and losses persist, with a lag-one autocorrelation of 0.35 in the win indicator and the identical marginal win rate and identical R multiples. Nothing about the trades changed; only their tendency to arrive in runs.

The clustered process actually produces a median maximum drawdown of 21.3 percent and a 95th percentile of 37.8 percent, against 14.8 percent and 26.3 percent for the independent version. Its 95th percentile longest losing run is twenty-two, against fourteen. But reshuffle the records that process generates and the reshuffled distribution reports a 95th percentile drawdown of 24.0 percent and a 95th percentile losing run of 13.6. The reshuffle understates the real drawdown risk by nearly fourteen percentage points and cuts the streak estimate by more than a third. If you took the reshuffled answer at face value you would size roughly one and a half times larger than the process permits.

The same check run on independent records behaves as it should, which is the reason to trust the comparison rather than the assertion. Reshuffling independent records reports a 95th percentile of 22.9 percent against a true 26.3 percent, a small conditioning gap and nothing like the collapse seen under clustering. That is the positive control: a diagnostic that fires on clustered data and stays quiet on independent data is measuring clustering, not noise.

The diagnostic you can run on your own record is direct. Compute your actual longest losing streak, then compute the distribution of longest streaks across reshuffles of your own trades, and see where your real number sits. On the observed record here the answer is reassuring: the lag-one autocorrelation of the win indicator is 0.038, effectively zero, and the observed streak of eight sits in the thick middle of the reshuffled distribution. On a clustered record the real streak lands out in the tail of its own reshuffled distribution, and that mismatch is the signal. When it appears, resampling individual trades is the wrong tool and you should resample blocks of consecutive trades instead, which preserves the local structure that plain shuffling destroys.

Serial dependence is not exotic in trading. Any strategy whose edge depends on a market condition will lose repeatedly while that condition is absent, which is clustering by another name. Trend-following systems are the standard example, and they are also the systems whose users are most likely to run this kind of simulation.

What this changes about position size

Everything above is decoration unless it changes a decision, and the decision it changes is how large to trade. Suppose a trader sets a policy ceiling: a drawdown of 20 percent is the most they can take without their process breaking down. That is a personal number, and the point is that it should be set first and the size derived from it, rather than the other way round.

The derivation was run by recomputing the whole simulation at every risk level from 0.25 percent to 4.0 percent per trade, rather than by scaling a single answer, because compounded drawdown does not scale linearly with bet size.

Two answers to the question of how much to risk Drawdown against risk per trade, recomputed from scratch at every level. The gap between the curves is the sizing error. Illustrative and simulated. 0% 10% 20% 30% 40% 50% 60% 70% the 20 percent ceiling this trader set 0.5% 1% 1.5% 2% 2.5% 3% risk per trade, percent of equity 0.78% sized off the 95th percentile 1.38% sized off the one record The three curves drawdown measured three ways observed record 95th percentile 99th percentile The failure mode Risking 1.38 percent keeps the one observed path inside the ceiling, yet 44 percent of resamplings breach it anyway. The whole argument in one line Planning for the observed drawdown permits 1.38 percent per trade. Planning for the 95th percentile permits 0.78 percent. The distribution costs you 43 percent of your position size and buys you the right to stay in the game.
Maximum drawdown against risk per trade, with the entire simulation re-run at each level. The gold curve is what the one observed record did; the green and coral curves are the 95th and 99th percentiles across resamplings. The horizontal distance between the two dots is the sizing error. Illustrative and simulated.

Sized off the observed record, the ceiling permits 1.38 percent of equity per trade. Sized off the 95th percentile of the bootstrap, it permits 0.78 percent. Sized off the 99th, 0.62 percent. The naive answer is 1.76 times the prudent one, which means a trader who anchors on their own backtest is carrying nearly double the position the arithmetic supports.

The consequence is quantifiable rather than rhetorical. At 1.38 percent per trade, the observed path does stay within the 20 percent ceiling, exactly as designed. But 44.1 percent of resamplings of that same record breach the ceiling, the 95th percentile drawdown becomes 32.9 percent and the 99th becomes 39.8 percent. The trader has set a rule and, on the evidence of their own trades, has a better than two in five chance of breaking it. That is not a risk policy, it is a hope with a number attached.

On an illustrative account of Rs 5,00,000, the difference is a risk budget of about Rs 6,910 per trade against about Rs 3,920. Both figures are illustrative and simulated. The prudent number is 43 percent smaller, and that reduction is the entire cost of the exercise. What it buys is the ability to meet the 95th percentile outcome without abandoning the system, which is the only circumstance in which a positive expectancy ever gets the chance to pay.

None of this replaces the sizing arithmetic itself, which our page on the position sizing formula covers: the stop distance, the rupee risk, the resulting quantity. Monte Carlo does not change that calculation. It changes the input to it, by replacing the risk-per-trade figure you guessed with one you derived from a drawdown you are actually prepared to meet.

How this differs from risk of ruin

Risk of ruin and Monte Carlo are related and routinely confused, and the distinction is worth being precise about because they answer different questions and mislead in different ways.

Risk of ruin computes a single probability: the chance that the account falls past a threshold from which it does not return, given a win rate, a payoff ratio and a bet size. It is usually evaluated analytically, it treats ruin as an absorbing state, and its answer is one number. Used properly it is a hard constraint. If a sizing choice produces a meaningful probability of ruin, no other consideration matters and the size is wrong. Our risk of ruin calculator runs that computation directly.

Monte Carlo describes the whole distribution of survivable outcomes as well as the fatal one. That difference matters because an account rarely stops trading because it reached zero. It stops because a drawdown arrived that the trader had never imagined, could not sit through, and responded to by cutting size at the bottom or abandoning the system entirely. On this run that drawdown is the 24.6 percent at the 95th percentile, which is nowhere near ruin and would not register in a ruin calculation at all. Risk of ruin is silent about it, correctly, because it is not ruin. The distribution is not.

There is a second difference in what each can carry. The standard risk-of-ruin formula assumes fixed win and loss sizes and independent trades. A simulation can use the actual R multiples out of your journal, gaps and partial exits and all, and can be extended to handle clustering. The two are best used together: risk of ruin as the constraint that rules out sizes outright, and the distribution as the tool that chooses among the sizes that remain.

The boundary of the method

A simulation that produces a precise percentile invites more confidence than it has earned. It is worth being explicit about which questions this arithmetic answers and which it merely appears to.

What resampling a trade record can and cannot establish.
QuestionCan Monte Carlo answer it?
How deep a drawdown should I be prepared for?Yes. This is the question it answers best, as a percentile you choose in advance
How long a losing run is normal for my system?Yes, provided your trades are close to independent
How long might I sit underwater?Yes, with the censoring rate published alongside, because the deepest drawdown is often unfinished
What position size does my drawdown ceiling permit?Yes, by re-running the simulation at each candidate size rather than scaling one answer
Does my strategy have an edge at all?No. Resampling assumes the edge in your record is real. That is a separate test
Will my backtest survive out of sample?No. Resampling reuses the same trades, so it cannot detect fitting to the past
What happens in a condition absent from my record?No. The simulation can only rearrange outcomes it has been given
Could a loss larger than any I have taken occur?No. Your worst historical loss is a hard ceiling on every simulated path
Are my costs and fills modelled correctly?No. Every error in the input record is reproduced faithfully in all five thousand outputs

The last three rows are the ones that bite. A record of 240 trades taken entirely in a rising market contains no bear-market trades, so no resampling of it will ever produce a bear-market drawdown, and the resulting percentiles will be confidently wrong. This is the same failure the pillar guide catalogues as regime myopia, arriving here in a more persuasive costume, because now it has five thousand simulations and a percentile table behind it.

Running one on your own record

The procedure is short and the discipline is in the order of operations rather than in the code.

Express your trades in R first. Convert every trade to a multiple of the amount you risked on it, so that an illustrative trade risking Rs 4,000 and returning Rs 8,000 records as plus 2R regardless of the position size at the time. Without this step you are resampling rupee amounts drawn from different account sizes and different levels of conviction, and the resulting distribution describes nothing coherent.

Decide your drawdown ceiling and your percentile before you run anything. Both are personal and both are opportunities to cheat later. Write down that you are planning for the 95th percentile and that your ceiling is a specific percentage, and save that note. If you choose the percentile after seeing the distribution, you will choose the one that lets you keep the position size you already wanted.

Run both methods and report both. The reshuffle is your floor and the bootstrap is your working estimate. If they disagree substantially on drawdown, that disagreement is itself informative: it means the specific mix of trades in your record is doing a lot of work, which is a reason to distrust a short sample.

Check the independence assumption rather than assuming it. Compare your real longest losing run to the reshuffled distribution of longest runs. If your actual record sits far out in the tail of its own reshuffles, your trades are clustered and you should resample blocks rather than individual trades.

Publish the censoring. If some fraction of your simulated paths never recovered their deepest drawdown within the record length, say so next to the recovery percentiles. A recovery figure quoted without it is systematically optimistic.

Re-run it at every candidate position size. Do not compute one distribution and scale it. Compounding makes the relationship between bet size and drawdown non-linear, and the error runs in the direction that flatters you.

Decide in advance what the answer obliges you to do. This is the step that makes the difference between a decision tool and an ornament. If the 95th percentile drawdown exceeds what you can take, the required response is to reduce size, not to re-run the simulation with a friendlier percentile. Writing that commitment down while the answer is still unknown is the only reliable protection, because by the time the number arrives you will have reasons.

What the distribution is actually for

It would be easy to read this page as an argument that trading records are meaningless. That is not the conclusion. The conclusion is narrower and more useful: a single equity curve answers one question well and several others badly, and knowing which is which turns it from a source of false confidence into a source of information.

What a record does tell you reliably, given enough trades, is the shape of your trade distribution. Your win rate, your average win in R, your average loss, the frequency of the outsized outcomes at both ends. Those are properties of the strategy and they converge as the sample grows. What a record tells you unreliably is everything that depends on the order: the drawdown, the streaks, the time underwater, the final figure. Those are properties of a single draw and they do not converge, because there is only ever one of them.

Monte Carlo is simply the operation that separates the two. It takes the part of your record that is reliable, the distribution of trade outcomes, and uses it to generate the part that is not, thousands of times, so that the range becomes visible. Everything it tells you was already implicit in your trades. It was merely invisible, because you only ever ran the experiment once.

What changes when the range is visible is the quality of your commitments. You can state a drawdown you are prepared to meet and derive a size from it, name the number of consecutive losses that will not cause you to intervene, and set a review interval long enough that it does not fire on ordinary variance. You can also stop reading a good year as evidence of skill and a bad one as evidence of failure, because you now know how wide the band of outcomes is for a fixed process. Those are habits rather than techniques, and they are the substance of what a serious quantitative curriculum spends its time on. If the arithmetic here was the interesting part rather than the tedious part, that is the method we teach.

FAQ

Frequently asked questions

It is the practice of rebuilding your equity curve thousands of times from the same underlying trades, changing only the order in which they arrive or the exact mix that gets drawn, and then reading off the range of outcomes. One backtest gives you a single number for maximum drawdown. A Monte Carlo run gives you a distribution, so you can say what the median drawdown was, what the worst one in twenty looked like, and how far the tail extends beyond that.

A few thousand is enough for medians and quartiles. The run on this page used 5,000 paths. Repeating that run twelve times with different random seeds moved the 95th percentile drawdown by a standard deviation of 0.26 percentage points, which is small enough to quote. If you care about the 99th percentile or beyond you want ten thousand or more, because the tail is estimated from progressively fewer observations. Doubling the count is cheap, so there is no reason to be stingy.

Run both, because they answer different questions. Reshuffling holds your exact set of trades fixed and asks what a different arrival order would have done to the path. Resampling with replacement lets the mix vary too, and asks what a different but comparable set of trades from the same distribution would have done. The second is wider and closer to the question you actually face, because your next 240 trades will not be a permutation of your last 240.

Because compounding is order independent. If each trade multiplies your equity by one plus a fixed fraction of its R multiple, the product of those factors is the same whatever order you multiply them in. Every permutation of the run on this page finished at exactly the same value, identical to thirteen decimal places. Reshuffling changes the path, never the destination, so it can tell you about drawdown and streaks but nothing at all about the spread of final outcomes.

Not the one your backtest showed. In the run on this page the single observed record fell 14.7 percent at its worst, while the 95th percentile across resamplings was 24.6 percent and the 99th was 30.8 percent. The honest planning number is a percentile you choose in advance, and the tail beyond it is the part you must be able to survive without changing anything. Both figures are illustrative and simulated.

Risk of ruin collapses everything into one probability: the chance you fall past a threshold you cannot come back from. Monte Carlo describes the whole distribution of survivable outcomes as well. An account rarely stops trading because it reached zero. It stops because a drawdown arrived that the trader had not imagined and could not sit through, and on this run that drawdown is the 24.6 percent at the 95th percentile rather than anything close to ruin. Risk of ruin is silent about that outcome because it is not ruin. The distribution is not.

Around a hundred is the point at which the exercise starts telling you something, and the more the better. Below about fifty, resampling mostly re-describes the accident of a small sample and the percentiles will move a lot when you add trades. Remember that resampling can only rearrange the outcomes you already have. If your record contains no large loss, no simulation built from it will ever produce one.

Then plain reshuffling understates your risk, sometimes badly. Losses that cluster produce deeper drawdowns than the same losses scattered at random. In the check on this page a deliberately clustered process actually produced a 95th percentile drawdown of 37.8 percent, while reshuffling records drawn from it reported 24.0 percent. The diagnostic is to compare the longest losing streak in your real record against the reshuffled distribution, and to treat a real record that sits in the far tail as a warning.

No. It means the risk arithmetic of the trades you have recorded is now visible instead of hidden. The simulation inherits every flaw in the record it is built from: optimistic fills, costs left out, a sample drawn entirely from one kind of market. It also cannot manufacture a loss larger than any you have experienced. A clean distribution built on a contaminated record is a precise answer to the wrong question.

Partly. A spreadsheet can shuffle a column of R multiples and recompute a running equity and drawdown, and repeating that a few hundred times with a recalculation key is tedious but workable. Anything past that, and certainly the percentile tables, is a short script. The arithmetic itself is simple: the only operations involved are drawing from a list, multiplying, and taking a running maximum.

Method note

How the numbers on this page were produced

Every figure comes from a single deterministic simulation written in numpy and seeded so that it reproduces identically on each run. A 240-trade record was drawn once from the stated generator and is treated throughout as the observed history. Two resampling methods were then applied to it, 5,000 paths each: a permutation of the observed trades, and a draw of 240 trades with replacement from the same multiset. Equity compounds at 1.0 percent of current equity risked per trade. Maximum drawdown is the deepest fall from any running peak; the longest losing run counts consecutive negative trades; time to recovery counts trades from the deepest trough to the next new high, with unrecovered paths excluded and their share reported; the longest stretch underwater counts consecutive trades below a prior peak. The sizing curves were produced by re-running the full simulation at each risk level rather than by scaling a single result. The independence check compares the true distribution of a deliberately autocorrelated generator against what reshuffling reports on records drawn from it, with the same comparison run on independent records as a control.

All results are illustrative and simulated. They are not a track record, they are not a forecast, and they are not an indication of what any strategy would produce in a live account. The purpose of the exercise is to demonstrate the relationship between a single observed path and the distribution it was drawn from, which is a property of arithmetic rather than of any particular market.

Related

Continue reading

Next step

Find your starting stage. Everything else follows from there.

Educational reference only. No buy, sell or hold recommendations. All results shown are illustrative and simulated.