Educational Reference

Statistical Arbitrage: What It Is, What It Is Not, and Why the Name Misleads

The word arbitrage means a profit that is known before you commit, because the same claim is priced two ways at the same instant. Statistical arbitrage keeps the word and discards the property. It is a portfolio of many small, individually unreliable bets whose edge exists only in aggregate, and only for as long as the statistical relationship behind it keeps holding. This page builds one in a seeded simulation, states every parameter, and reports what the arithmetic produces, including the part that does not flatter the idea.

The finding, stated first. In the simulated book on this page the average round trip earns about 50.0 basis points of the capital committed to it before charges. The Indian cash-delivery charge stack computed here takes 42.7 of those, which is 85.4 percent of the gross edge, and leaves 7.3. What remains is real but it is thin enough that a modest rate of relationship breakdown cancels it outright. On these numbers, statistical arbitrage is not a viable retail path, and the reasons are specific rather than vague.

The word is doing damage before the strategy starts

Arbitrage, used strictly, describes something with an unusual property: the profit is an accounting fact rather than a forecast. The same claim exists in two places at two prices at the same instant, you buy the cheap one and sell the dear one simultaneously, and the two positions cancel. Nothing about the future has to cooperate. The trade can fail on execution, on settlement, or on a counterparty, but it cannot fail because the market moved, because you are not exposed to the market moving. A page on what arbitrage trading actually means sets out the ordinary forms of it.

Statistical arbitrage keeps the market-neutral shape and abandons everything that made the shape safe. It does not trade the same claim twice. It trades different claims whose prices have been observed to move in a stable relationship, and it bets that when the relationship stretches it will come back. The bet can be wrong. It is wrong reasonably often. What holds the strategy up is not the certainty of any one trade but the average of a great many of them, and that average only exists if the relationship keeps being a relationship.

This is not a semantic complaint. The word sets the expectation, and the expectation sets the position size. Someone who believes they are running an arbitrage will size for a bounded outcome, will treat a loss as an execution failure rather than a normal draw, and will not build the one piece of machinery the strategy genuinely requires, which is a way of noticing that the relationship has stopped working. The rest of this page is an attempt to replace the word with numbers.

Arbitrage has one outcome. This has a distribution. Both rows use the same horizontal scale: the result of a single round trip, in basis points of the capital committed to it. TRUE ARBITRAGE every outcome sits here and nothing anywhere else on this line STATISTICAL ARBITRAGE 20,000 simulated round trips zero mean +7.3 bp worst 5 percent below −104.3 bp best 5 percent above 62.0 bp −200 −150 −100 −50 0 50 100 150 200 result of one round trip, basis points of committed capital 25.5% chance the trade loses money +7.3 bp mean result 50.7 bp spread of results 0.14 mean divided by spread
Both rows share one horizontal scale. True arbitrage has one outcome, and it is known before the trade is placed. The lower row is 20,000 simulated round trips from the model described further down this page. The average is +7.3 basis points; the spread of results around it is 50.7. The edge is nearly invisible next to the noise it sits in, which is the whole reason the strategy has to be run in bulk. Illustrative and simulated.
The two things the same word is used for. The right-hand column describes the model built on this page. Illustrative and simulated.
PropertyArbitrage, strictlyStatistical arbitrage
What is exploitedThe same claim priced two ways at the same momentA relationship between different claims that has held in the past
When the result is knownBefore you commitOnly after enough trades for the average to mean something
Chance one trade losesNil by construction if both legs settle25.5 percent in this model
What makes it workAn identity: two prices for one thing cannot both be rightA regularity that was true in the sample you measured
Trades requiredOneHundreds a year, across many relationships
Sensitivity to costsHigh, because the gap is small and closes fastExtreme: charges took 85.4 percent of the gross edge here
How it endsSomeone reaches the gap before you doThe relationship stops holding, and you find out late
Where the risk sitsIn execution and settlementIn the assumption itself

Read the middle rows together. A strategy that needs hundreds of trades a year to establish its own average is a strategy where the cost of a trade is not a detail, it is the main term. That relationship between breadth and cost sensitivity is the actual economics of statistical arbitrage, and it is what the rest of this page computes.

Many small bets, one shared assumption

The simplest instance of the general class is a pair: two instruments whose prices are tied together, a spread that oscillates around a mean, and a rule that trades the extremes. The statistical machinery for the pair case, the distinction between correlation and a genuinely stationary spread, the two-step test that establishes it, the hedge ratio and the reversion half-life, is covered in depth in the guide to cointegration and pairs trading on the NSE. This page assumes all of that and asks the question one level up: what happens when you run not one of these but two hundred, and what is the resulting thing actually worth.

The general class is wider than pairs. A relationship can be one instrument against a basket, an instrument against a factor model's prediction of it, or a whole set of instruments tied by a common exposure. The construction differs; the economics do not. In every case you are holding a large number of positions, each with a small expected gain, each individually likely to be swamped by noise, and each resting on the same kind of assumption, which is that a pattern measured in the past will continue.

That last clause is where the whole class lives or dies, and it is worth being precise about what is being assumed. It is not that the relationship is a law. Nobody claims that. It is that the relationship will persist for long enough, after you have found it, to pay for the cost of finding it and the cost of trading it. Those are two separate claims and both are empirical. The first can be checked. The second is almost never checked, because checking it requires the arithmetic on this page.

There is a second reason the strategy is built in bulk, and it is the reason the industry version uses leverage. If the average trade nets a handful of basis points, the return on the capital reserved for the book is small even when everything works. The way that becomes a business is to hold the book at several times the capital, which multiplies a small stable number into a usable one. Leverage multiplies the losses in exactly the same proportion, and it removes the ability to sit through a stretch where the relationships are all wrong at once. Retail rarely has access to that leverage on a cash-market book, which turns out to matter less than it sounds, because the unlevered number computed below is not one that leverage would rescue.

The model, stated in full so it can be argued with

The relationships below are simulated rather than drawn from live instruments, deliberately. A simulation lets every parameter be stated openly, keeps the exercise reproducible, and avoids implying that any particular Indian instrument or pair would have produced any particular outcome. The costs are the exception: those are computed from published Indian charge schedules, because the entire argument turns on whether the charges are large relative to the edge, and an invented charge would settle that question by assumption.

Every parameter in the model. All figures on this page are produced from this single configuration with a fixed seed. Illustrative and simulated.
ElementSettingWhy it is set this way
The relationshipA spread that pulls back toward its mean, half-life 15 trading daysFast enough to trade several times a year, slow enough not to be pure microstructure noise
Spread sizeEquilibrium standard deviation 1.20 percent of one leg's valueA modest dislocation for two related large names. A larger figure would flatter the strategy
EntryStandardised spread beyond 2.0, measured on a 120-day rolling windowThe rolling window is what a real book uses. The threshold is a choice, and choices like it are the subject of a separate study of parameter sensitivity
ExitStandardised spread back inside 0.5, or beyond 4.0 against the position, or 30 trading days elapsedA target, a divergence stop and a time stop. All three fire in the results below
PositionOne leg long and one leg short at equal value, so gross exposure is twice one legAll results are quoted in basis points of that gross exposure, which is the capital the position ties up
ChargesThe Indian cash-delivery stack, 42.7 basis points of gross per round tripStatutory rates from published schedules. Brokerage, depository markup, borrow and slippage are stated assumptions
ScreeningA unit-root test at the 5 percent level, critical value and power both simulated rather than quotedThe critical value came out at minus 2.865, in line with published tables, which is the check that the code is right
The bookN slots, each reserving one Nth of capital, idle between tradesSlot utilisation came out at 25.5 percent, so most of the reserved capital is not working at any moment
DataFully simulated, one fixed seed, 18,971 trades in the base runReproduces identically on every run. No live instrument is implied

Two of those rows carry more weight than the rest. The spread size decides how much there is to capture per trade, and it has been set at the modest end on purpose, because setting it higher is the easiest way to manufacture a favourable answer. The charge row decides how much of that capture survives, and it is the one part of the model that is not a choice.

What one relationship earns, and what the charges take

Run the rule on a relationship that genuinely reverts and the outcome is respectable. Across 18,971 simulated trades the average round trip gained 54.1 basis points of the capital committed to the position before charges, 88.1 percent of trades were positive, the average holding period was 16.2 trading days, and each relationship produced about 4.0 trades a year. Most positions closed at the target; 2.2 percent hit the divergence stop and 16.7 percent ran out of time. That is a genuine, if unspectacular, edge.

Except that not every relationship in a real book is a genuine one, and the arithmetic of how they get in matters. Screen the pairs available from a hundred-name universe and you have 4,950 candidates. Suppose one in twenty is genuinely related, which is generous. The test's power was simulated directly against the model's own alternative and came out at 40.6 percent, so the scan produces roughly 235 passes that are nothing at all against roughly 101 that are real. Seven in ten of the relationships that pass a single scan would be spurious. Prescreening on economic logic, which raises the share of genuinely related candidates from one in twenty to about one in seven before the test runs, cuts that to about 41.1 percent. Confirming every survivor on a second stretch of data the scan never touched cuts it again, to about 7.9 percent, and that last figure is what the book below is built with.

A relationship that is not a relationship is not a disaster on any single trade. Its spread is a random walk, so fading its dislocations is close to a fair game before charges: those trades averaged minus 1.31 basis points gross, with slightly more than half positive. What they do is pay the full round-trip cost for the privilege, and occupy a slot that a real relationship could have used. Blending them in at the computed rate brings the book's average gross result down to 50.0 basis points a trade.

The friction at which the edge is gone, and where the Indian stack lands Net result of one round trip as the cost of that round trip rises. Both horizontal scales below are the same. NET RESULT PER ROUND TRIP, BASIS POINTS 40 20 0 −20 −40 −60 edge gone at 50.0 bp every round trip loses money to the right of this line computed stack 42.7 bp leaves 7.3 bp of the 50.0 bp gross 0 20 40 60 80 100 120 round trip friction, basis points of the capital committed to the pair WHAT THE 42.7 BASIS POINTS ARE MADE OF, ON THE SAME SCALE 42.7 bp Securities transaction tax 20.00 Stock borrow 7.60 Execution slippage 6.00 Brokerage 4.00 Depository debit 2.12 Stamp duty 1.50 Tax on fees 0.83 Exchange charge 0.61 Regulator fee 0.02
The green line is the average result of one round trip as friction rises; it reaches zero at 50.0 basis points. The Indian cash-delivery stack computed for this model is 42.7 basis points, which leaves 7.3. The bar below uses the same horizontal scale, so its length is literally how much of the distance to the breakeven point the charges already cover. Statutory rates are from published charge schedules; brokerage, depository markup, borrow and slippage are stated assumptions. Illustrative.

Now the charges. On an illustrative two-leg position of one lakh rupees a leg, the round trip is four executions, and the Indian cash-delivery stack comes to 854 rupees, illustrative, which is 42.7 basis points of the two lakh at risk. The single largest line is the transaction tax at 20.00 basis points, 47 percent of the whole stack, because it is charged on every purchase and every sale of both legs. Borrow on the short leg adds 7.60, slippage 6.00, brokerage 4.00, the depository debit 2.12, stamp duty 1.50, tax on the fee lines 0.83, and the exchange and regulator charges the remaining 0.63 between them. Statutory 22.97 basis points, commercial and assumed 19.73.

Set that against a gross edge of 50.0 and the position is stark. The strategy breaks even at 50.0 basis points of friction and the stack is 42.7, so 85.4 percent of the edge is gone before the relationship is even asked to perform. What survives is 7.3 basis points a trade. A trader holding positions for months and turning over a handful of times a year can absorb 42.7 basis points without much thought. Here it is almost the entire result. The wider question of how much of a backtest survives contact with real fills and real charges is worked through separately in the study of the gap between a backtest and a live account; this page needs only the level, and the level is decisive.

One consequence deserves stating plainly, because it is counterintuitive. Because the friction is roughly fixed per round trip while the capture depends on how far the spread travels, a slower relationship is worse than a faster one even though it sounds safer. Rerunning the same engine at a half-life of 5 trading days produced 58.4 basis points a trade after charges; at 40 days it produced minus 24.6, which is to say the trade loses money on average. Those three runs use genuine relationships only and hold the spread size fixed while the speed changes, both of which flatter the fast case, so read them as the direction of the effect rather than its size. The comfortable, slow-moving relationship is the one the charge stack kills first.

Where the money actually comes from

Seven basis points a trade, four trades a year, is 0.29 percent a year on the capital a relationship reserves. That is the whole edge of one relationship, and stated on its own it is not a strategy, it is a rounding error with paperwork. The reason statistical arbitrage exists at all is that this number does not have to be earned one relationship at a time.

The same edge, held with five, fifty and two hundred relationships Simulated distribution of one year of book results. The average never moves. Only the width does. Five relationships positive in 75.4% of simulated years worst 5 percent below −0.49% Fifty relationships positive in 97.4% of simulated years worst 5 percent below 0.04% Two hundred relationships positive in 100.0% of simulated years worst 5 percent below 0.17% −1.5% −1.0% −0.5% 0.0% 0.5% 1.0% 1.5% one year of book results, percent of the capital the book reserves 0.29% average year, every case 21 relationships for a 9 in 10 chance of a profitable year 34 relationships for 19 in 20
Each lane is the same per-trade edge, held with a different number of relationships, simulated 20,000 times. The average year is 0.29 percent in all three cases. What changes is the width: with five relationships the year is close to a coin toss, and with two hundred it is not. This lane chart assumes every relationship stays alive for the whole year and that none of them move together, and both assumptions are removed later on the page. Illustrative and simulated.

Look at what changes across the three lanes and what does not. The average year is 0.29 percent in all three cases. Adding relationships does not raise the return at all, and anyone who tells you their edge improves as they add names is describing something other than this. What changes is the width of the distribution. With five relationships the spread of annual outcomes is 0.45 percent against an average of 0.29, so the year is close to a toss-up: 75.4 percent of simulated years finished positive and the worst twentieth finished below minus 0.49 percent. With two hundred, the spread falls to 0.07 and essentially every simulated year finished positive.

This is the intuition behind the standard result that a manager's outcome depends on skill per bet multiplied by the square root of the number of independent bets. The skill per bet here is fixed by the model, and it is small: the average trade is 0.14 of a standard deviation. Multiplying a small number by the square root of a large one is the only lever available, and it is a lever on reliability, not on return. In the simulation it took 21 relationships to reach a nine in ten chance of a positive year and 34 to reach nineteen in twenty.

That reframes what the strategy is. It is not a way of being right about anything in particular. It is a machine for converting a barely detectable per-trade advantage into an outcome stable enough to plan around, and the machine is made of arithmetic rather than insight. Which means the machine has requirements, and the requirements are where retail attempts come apart. Note also what the lane chart assumes: every relationship stays alive for the whole year, and no two of them move together. Both assumptions are removed below, and each removal is expensive.

Independence is the load-bearing assumption, and it is the weakest one

The square-root effect counts independent bets. Two hundred positions are two hundred bets only if their outcomes are unrelated. They are not, and the reason is structural rather than accidental. Spread dislocations are not random events scattered across unrelated names. They widen when liquidity thins, when a factor rotates, when the market takes risk off, and those things happen to the whole book at once. A modest amount of shared behaviour is enough to do serious damage, because the damage compounds with the number of positions.

Two hundred related bets are not two hundred bets How many genuinely independent bets a book of N relationships is worth, once a common factor moves them together. EFFECTIVE INDEPENDENT BETS 1 3 10 30 100 300 5 10 20 50 100 200 400 no common factor common factor 0.05 common factor 0.15 common factor 0.30 relationships in the book two hundred at 0.15 is worth 6.5 SHARE OF SIMULATED YEARS THAT FINISHED POSITIVE, TWO HUNDRED RELATIONSHIPS UNLESS STATED 100.0% independent, 200 89.0% factor 0.05 76.6% factor 0.15 70.0% factor 0.30 75.4% independent, 5
Effective breadth is the count of independent bets a book actually behaves like, both scales logarithmic. At a common factor of 0.15, two hundred relationships behave like 6.5. Note the shape: past about fifty relationships the curves are flat, so adding names buys almost nothing. The chips below give the share of simulated years that finished positive in each case. Illustrative and simulated.

The arithmetic is unforgiving. A book of N positions with an average pairwise correlation of rho behaves like N divided by one plus N minus one times rho. At a correlation of 0.15, two hundred positions behave like 6.5. At 0.30 they behave like 3.3. Even a correlation of 0.05, which most people would describe as negligible, cuts two hundred down to 18.3. And look at the shape of those curves: past about fifty positions they are flat, which means that from there on, adding relationships costs capital and attention and buys almost nothing.

The simulated outcomes match the arithmetic. Two hundred relationships with a common factor of 0.15 finished positive in 76.6 percent of simulated years, against essentially all of them when the relationships were independent, and the worst twentieth of years finished below minus 0.37 percent. That is very close to what five independent relationships produced. Fifty relationships at the same common factor produced 75.2 percent, statistically indistinguishable from the two hundred. Two hundred related bets and fifty related bets are the same book.

This is the mechanism behind the pattern of statistical arbitrage books having long quiet stretches and then all going wrong in the same week. Nothing has to break for that to happen. The positions were never as independent as the position count implied, and the count is the number everyone quotes. It is also the reason a scan that finds relationships by searching a single sector produces the worst kind of book: every relationship in it shares the same driver, so the position count is large and the effective breadth is close to one.

The practical test is short and almost nobody runs it. Take the daily returns of each relationship in the book, compute the average pairwise correlation, and divide your position count by one plus N minus one times that number. Whatever comes out is how many bets you actually have. If that figure is under ten, the year is a toss-up no matter how many lines are on the screen.

The failure mode nobody models: how long it takes to notice

Every treatment of this family of strategies mentions that relationships break. Very few of them ask the next question, which is how long you would take to find out, and what the delay costs. That question has an answer, and the answer is the most useful thing on this page.

The experiment: take a relationship that has reverted reliably for 900 trading days, end it on a known day so that the spread becomes a random walk with a slow drift, and then run a monitor. The monitor is the same unit-root test used to admit the relationship in the first place, recomputed every day on a rolling 250-day window. It declares the relationship dead when the test statistic rises above a threshold.

The first result arrives before the experiment does. Using the nominal five percent critical value as a monitor is useless: run it daily for 500 days on a relationship that never breaks and it declared a healthy relationship dead in every single one of 400 runs. A test built to be wrong five percent of the time on one look is wrong nearly always when you look five hundred times. The threshold had to be recalibrated by simulation so that a healthy relationship triggers in one monitoring period in twenty, which moved it from minus 2.865 to minus 0.951. That recalibration is itself a thing almost no retail implementation does.

The relationship stops working, and the test takes months to say so One simulated relationship. It is stationary until the marked day, then it is not. The monitor is the same test, run daily. 0 the relationship ends here spread trades monitor threshold, tuned to one false alarm in twenty the nominal five percent value, useless as a monitor declared dead, 158 days late test statistic −250 −125 0 +125 +250 +375 trading days, counted from the day the relationship ended 133 days average delay before the test says so 164.6 bp lost while waiting 157.4 bp of that, caused by the delay alone 5.7 years of that relationship's contribution
One simulated relationship that is mean reverting until the marked day and a random walk with drift afterwards. The lower panel is a rolling unit-root test run every day, with its threshold tuned so that a healthy relationship raises a false alarm in one monitoring period in twenty. The shaded band is the delay between the relationship ending and the monitor saying so. The four figures underneath are averages across 400 simulated breakdowns, not this single path. Illustrative and simulated.

With the monitor properly calibrated, the delay is 133 trading days on average and 106 at the median, and the slowest tenth of cases took more than 289 days. During the delay the relationship kept generating entries, because the spread was still crossing the entry threshold, just no longer coming back. An average of 2.8 trades were placed in that window, of which 1.7 lost money, and the cumulative cost was 164.6 basis points of the capital allocated to that relationship.

Now decompose that. If the relationship had been retired on the day it actually broke, the loss would have been 7.3 basis points, which is just the position that happened to be open at that moment. Essentially the entire 164.6 basis points is the price of the delay, not the price of the break. Against the 28.8 basis points that a working relationship contributes in a year, one breakdown costs 5.7 years of that relationship's output.

An obvious response is to monitor the money instead of the statistics, on the theory that a dead relationship will show up in its own profit and loss faster than in a test. It does not. Calibrated to the identical false-alarm budget, so that a healthy relationship trips the boundary in one monitoring period in twenty, the profit-based monitor was slower: 201 days on average, against 133 for the statistical test. The reason is the same in both cases. A per-trade edge that is small relative to its own noise means a run of losses is not evidence of anything for a long time, and a boundary loose enough not to fire on healthy relationships is loose enough to be reached slowly.

The divergence stop does not rescue this either, and the reason is worth being precise about. In the simulation only about one in fifteen of the trades placed after the break hit the stop; most simply drifted to the time limit. More importantly, a stop closes a position and leaves the relationship in the book, so the next dislocation is traded again, and again. The stop is a limit on one trade. The detector is a decision about the relationship. A book can bleed for six months with every individual risk rule behaving exactly as specified.

Put the pieces together and there is a single number that decides whether the strategy survives its own mortality. Each breakdown costs about 164.6 basis points and each live relationship contributes about 28.8 a year, so the book's edge is exactly cancelled if about 17.5 percent of relationships die each year, which is roughly one in six. At a ten percent annual breakdown rate the strategy keeps 12.3 of its 28.8 basis points. Nobody knows the true rate, which is precisely the problem: the strategy's viability rests on a quantity that is not measured, is not usually estimated, and does not appear in any backtest.

Capacity, and why published approaches decay

The last property worth being honest about is that the edge shrinks as the money grows, and it shrinks for a mechanical reason rather than a mysterious one. The charge stack computed above is proportional to trade value, so it does not change with size. Market impact does. Past a certain order size you begin to move the price you are trying to trade against, and the standard description of that effect has impact growing roughly with the square root of your share of a name's daily traded value.

The model leaves 7.28 basis points of headroom after ordinary charges. Take an illustrative name whose daily traded value is around forty crore rupees, take single-name daily volatility of about one and a half percent, and take an impact coefficient of one, which is a mid range choice and is stated here as an assumption rather than a measurement. The headroom is consumed at a participation rate of 0.059 percent of daily traded value. Spread across two hundred relationships, that corresponds to an illustrative book of roughly 9.4 crore rupees. Across fifty, roughly 2.4 crore. Halve the impact coefficient and the capacity roughly quadruples, because the relationship is quadratic; double it and capacity falls to a quarter. The precise numbers depend entirely on that coefficient, which is why the shape matters more than the level: capacity scales with the square of the edge you have left.

Two things follow. The first is that a thin edge has almost no capacity, because squaring a small number produces a very small one. The second is the reason approaches decay after they are described in public. As more capital chases the same relationships, the dislocations are competed away before they widen, which lowers the gross capture. Lowering the gross capture lowers the headroom, and lowering the headroom lowers the capacity quadratically. The decay is not a story about secrets leaking. It is the arithmetic of a crowded trade with a square-root cost function.

It also explains a difference that puzzles people. An institutional book runs the same idea profitably at a size a retail account cannot approach, and does so with a per-trade edge smaller than the one modelled here. It manages that on turnover and on friction, not on insight: far more trades a year across far more relationships, at a fraction of the per-trade cost. Neither of those two levers is available to a retail account, and both of them are the levers that matter.

The three reasons retail attempts fail, and the test for each

The failures are not mysterious and they are not about discipline. Each of the three is a quantity that can be measured before any capital is committed, and each has a specific diagnostic that takes an afternoon.

Each failure mode, the measurement that detects it in advance, and what that measurement produced in this model. Illustrative and simulated.
Failure modeThe diagnostic that catches it before you tradeWhat it produced here
Not enough independent betsCount the relationships, compute their average pairwise return correlation, then divide the count by one plus N minus one times that correlation. The result, not the count, is your breadthFive relationships: positive in 75.4 percent of simulated years. Two hundred at a common factor of 0.15: 76.6 percent, which is the same book
Charges above the per-trade edgeCompute the full round-trip stack in basis points of gross exposure, then compute the mean gross result per trade in the same units. The second must exceed the first by a margin you can defend42.7 basis points of charges against 50.0 of gross edge, so 85.4 percent consumed and 7.3 left
No breakdown detectorBuild the monitor, calibrate its threshold on relationships you know are healthy until its false-alarm rate is one in twenty, then measure its delay on relationships you know are broken. Both numbers, not just the first133 trading days of delay, 164.6 basis points lost during it, which is 5.7 years of that relationship's contribution

The third row is the one that is almost never built, and it is the one that decides the outcome. A trader can get the breadth right and the costs right and still lose steadily, because the book is quietly full of relationships that stopped being relationships some months ago and nothing in the system is designed to notice. The entry signal fires on all of them, every time the dead spread wanders far enough from a mean it no longer has.

What the arithmetic leaves standing

The honest summary is short. Under the model's own assumptions, with every relationship genuine, every relationship independent, and none of them ever breaking, the strategy earns about 0.29 percent a year on the capital it reserves, before any allowance for the capital sitting idle 74.5 percent of the time. Remove the independence assumption and the reliability collapses to what five relationships would have produced. Add a realistic rate of breakdown and the remaining edge is cancelled at about one relationship in six dying each year. At the charge level computed from published Indian schedules, this is not a viable retail path, and saying so is more useful than any amount of technique.

That is a statement about a particular strategy at a particular cost level, not a claim that the underlying statistics are worthless. They are not. Two things on this page transfer to work that has nothing to do with spread trading, and they are the reason the exercise is worth doing even if you never place one of these trades.

The first is the breadth arithmetic. Any time an edge is described as small but reliable, the question is how many genuinely independent instances of it you can hold, and then the harder question of what makes them independent. The answer is almost always fewer than the position count suggests, and the gap between the two is where most disappointment in systematic trading comes from. The division by one plus N minus one times rho takes thirty seconds and it has spoiled a great many plans that looked fine.

The second is the detection question, and it generalises further. Every rule that assumes something will keep being true has a hidden requirement: a way of learning that it has stopped, a measured delay before that learning arrives, and a cost accrued during the delay. Most strategies have no answer to the first part, let alone the other two. Working out those three numbers before committing capital is not a technique, it is a habit, and it is the sort of habit that separates a system from a set of rules that happen to have worked. If the arithmetic on this page was the interesting part rather than the tedious part, that is the method we teach.

FAQ

Frequently asked questions

No, and the word is the problem. Arbitrage in the strict sense means the same claim is priced two ways at the same moment, so buying one and selling the other locks a known amount before you commit. Statistical arbitrage has none of those properties. It trades different claims whose relationship has held in the past, the profit is unknown when the position goes on, and about one trade in four loses money in the model on this page. What it borrows from arbitrage is the market-neutral shape of the position, not the certainty.

Pairs trading is the simplest possible instance of statistical arbitrage: one relationship, two instruments, one spread. Statistical arbitrage is the general class, and the general class has properties a single pair does not. The edge on any one relationship is too small and too unreliable to be worth trading on its own. It only becomes a strategy when many relationships are held at once, which introduces questions about independence, capacity and monitoring that never arise when you are watching a single spread.

Because the per-trade edge is tiny relative to the noise around it. In the model here the average trade nets about seven basis points with a spread of about fifty, so a single trade tells you almost nothing. The reliability of the aggregate improves with the square root of the number of independent bets, which is why the same edge held with five relationships produced a positive year in roughly three simulated years in four, while the same edge held with two hundred produced one in essentially all of them. Breadth does not raise the average. It narrows the distribution around it.

In the model on this page the average round trip earns about fifty basis points of the capital committed before costs, so the edge is exactly gone at fifty basis points of round-trip friction. The Indian cash-delivery stack computed here comes to about forty-three, which leaves roughly seven. That is the whole difficulty in one line: a level of friction that a trader holding for months would barely register consumes most of a statistical arbitrage edge, and a slightly worse fill or a slightly dearer borrow removes the rest.

Far longer than most people assume. In the simulation a relationship stops reverting on a known day, and a rolling unit-root test tuned to raise a false alarm on one healthy relationship in twenty took an average of about 133 trading days to say so, with the slowest tenth taking more than a year. The delay is not a flaw in the test. It is the arithmetic of needing enough post-break observations to distinguish a broken relationship from an unlucky stretch of a working one.

It caps one trade. It does not retire the relationship. In the simulation only about one in fifteen of the trades placed after the break hit the divergence stop, and every trade that closed for any reason left the relationship in the book, so the next dislocation was traded again. The stop is a limit on a single position and the detector is a decision about the relationship. Confusing the two is why a book can bleed for months with every individual risk rule working exactly as designed.

Because a test with a five percent error rate applied to thousands of candidates produces a great many false passes. Screening 4,950 candidate pairs from a hundred-name universe, with one in twenty genuinely related and the test's power simulated at about forty-one percent, gives roughly 235 false passes against roughly 101 real ones. Seven in ten of the pairs that pass would be spurious. Economic prescreening and a second confirmation on untouched data cut that sharply, but the raw scan is closer to a random number generator than to a discovery process.

Because the cost of trading is not fixed. Beyond a certain order size you start moving the price you are trying to trade at, and the effect grows roughly with the square root of your share of the daily traded value. With about seven basis points of headroom left after the ordinary charges, the model runs out of room at a very small share of a typical name's daily value. That is the mechanical reason a technique that works at one size stops working at ten times that size, and one reason published approaches tend to decay.

The computation on this page does not support it. The three requirements pull against each other: enough independent relationships to make the average meaningful, friction well below the per-trade edge, and a monitor that retires a relationship before it has bled for months. At the charges computed here the friction alone takes about eighty-five percent of the gross edge, and the remainder is thin enough that a modest breakdown rate cancels it. That is the honest reading of the numbers, and it is stated here rather than left for a live account to discover.

Two habits. First, whenever a strategy's edge is described as small but reliable, ask how many genuinely independent bets it takes and then ask what makes them independent, because a common factor collapses breadth faster than anything else. Second, for any rule that assumes something will keep being true, work out in advance how you would learn that it had stopped, how long that would take, and what the delay would cost. Most strategies have no answer to the second question at all.

Method note

How the numbers on this page were produced

Every figure comes from one deterministic simulation, seeded so that it reproduces identically on each run. Spreads are simulated first-order autoregressive series with the stated half-life and equilibrium standard deviation; positions are opened and closed by the stated rule on a rolling standardised spread; and results are quoted in basis points of the gross exposure of the two-leg position. The unit-root critical value used throughout was not quoted from a table, it was simulated under a driftless random walk over 20,000 replications, and came out at minus 2.865, which matches published values and is the check that the implementation is correct. The test's power against the model's own alternative was simulated the same way over 8,000 replications. The breakdown experiment uses 400 broken relationships and 400 healthy ones for calibration.

Charges are the exception to the simulation. The statutory lines, the transaction tax on each purchase and sale, the stamp duty on the buy side, the exchange transaction charge, the regulator turnover fee and the tax on the fee lines, are computed from published charge schedules current at the time of writing. Brokerage, the depository markup, the stock borrow fee and execution slippage are commercial rather than statutory, vary by provider and by instrument, and are stated assumptions marked as such in the model table. The market impact coefficient used in the capacity section is likewise a stated assumption; the point of that section is the quadratic shape rather than the level.

All results are illustrative and simulated. They are not a track record, they are not a forecast, and they are not an indication of what any strategy would produce in a live account. Nothing here is a recommendation to trade or invest. The purpose of the exercise is to expose the relationship between per-trade edge, breadth, friction and detection delay, which is a property of the strategy's structure rather than of any particular market.

Related

Continue reading

Next step

Find your starting stage. Everything else follows from there.

Educational reference only. No buy, sell or hold recommendations. All results shown are illustrative and simulated.