Educational Reference
Cup and Handle: Where the Definition Does the Work
Most chart patterns are described in a sentence. The cup and handle is described in a specification: a prior advance of a stated size, a depth inside a stated band, a duration inside a stated range, a base that must be rounded rather than sharp, a handle no deeper than a stated fraction of the cup, and a handle that must drift downward. That precision is usually offered as evidence of rigour. This page treats it as a hypothesis, codes the whole definition, runs it across 9,000,000 generated daily bars, and then removes the clauses one at a time to find out what each of them is actually doing.
The finding, stated first. The full classical definition fired 443 times across 9,000,000 generated daily bars, about one instance per 81 instrument-years. Measured against a control matched for instrument, direction, holding period and target distance, and drawn only from bars the pattern itself did not select, the difference over twenty bars was smaller than a hundredth of a percentage point. Removing the clauses one at a time never shifted it by more than 0.34 of a percentage point, while relaxing all of them at once took the population from 443 to 52,291 on the same data. At 443 instances the test was in any case blind to anything smaller than 1.2%. All results on this page are illustrative and simulated.
The most specified pattern in common use
Set the cup and handle beside its neighbours and the difference is immediate. A double bottom is two lows at a similar level. A triangle is a range that narrows. A head and shoulders is three peaks with the middle one highest. Each is a shape, loosely stated, and each leaves the reader to supply the tolerances. The cup and handle arrives with the tolerances already filled in, and it arrives with more of them than any other pattern in general circulation.
The usual reading of that is favourable. A vague pattern can be seen anywhere, so a pattern with published bounds looks like a discipline: it tells you when the thing in front of you does not qualify, and a rule that can say no is worth more than one that cannot. That reasoning is sound as far as it goes. The wider taxonomy of consolidation shapes, and where a rounded base sits inside it, is laid out in the guide to chart patterns in Indian stocks, and this page goes underneath one entry in it.
But there is a second reading that is at least as plausible and is almost never stated. Every clause in a specification has to have come from somewhere. Nobody derived a depth band from a model of how supply is absorbed. The bands were set by looking at instances that had already worked and describing what they had in common. That is a legitimate way to generate a hypothesis and a disastrous way to confirm one, because the description will fit the examples it was drawn from whether or not anything general is going on.
Two consequences follow, and they pull in the same direction. The first is arithmetic and certain: every additional clause reduces the number of instances that qualify. A definition with eleven conditions attached selects a much smaller population than one with two, and a smaller population means a wider confidence interval around anything you measure on it. The second is subtler and only probable: every additional clause is another number that could have been reverse-engineered from remembered winners, and there is no way to tell from the outside which clauses those are.
Put together, the specification is a certain cost against an uncertain benefit. That framing turns the question into something answerable. Instead of arguing about whether the constraints are wisdom or folklore, code them, run them, and then take them out one at a time and watch what happens to the number of instances and to the measured outcome. If the result barely moves while the population moves by two orders of magnitude, the constraints were not doing the work they are credited with. That experiment is the centre of this page, and everything before it exists to make it interpretable.
Writing the whole thing down is where it gets awkward
The moment you try to code the standard description, it stops being a specification and becomes a list of questions. A prior advance, yes, but measured over how many bars, and from what starting point? A depth band, yes, but measured from which high to which low? A rounded base, but rounded according to what measurement? A handle of a stated relative depth, but relative to which rim, the left or the right, since they are rarely at the same price? A break above the pivot, but by how much, and does an intraday poke count?
Some of those questions have conventional answers and some have none. The prior advance is usually quoted as a percentage without a window; the depth band is usually quoted without saying which rim; the rounding clause is quoted in every description of the pattern and quantified in none of them. What follows is the set of choices used here, stated in full so that every number on this page can be reproduced or disputed.
| Clause | Value used | Where the number came from, and what it costs |
|---|---|---|
| Prior advance | At least 30% in the 120 bars before the left rim | The pattern is taught as a continuation, so it needs something to continue. This clause was the first to fail more often than any other except the depth band, and removing it alone took the population from 443 to 1,268 |
| Cup depth | Between 12% and 33% from the left rim to the low | A band commonly given for the pattern. Descriptions rarely say which rim the depth is measured from, and the two rims sit at different prices in almost every instance |
| Cup duration | Between 35 and 325 bars, rim to rim | About seven weeks to sixty-five weeks, a range commonly given. It is wide enough to exclude very little, and so it proved: relaxing the floor moved the population by about a fifth |
| Rounded, not V-shaped | At least 0.35 of the cup spent in its lowest quarter by depth | Stated in words wherever the pattern is described and, as far as this exercise could establish, never as a number. This threshold was invented here, and the fact that it had to be invented is itself the evidence discussed further down |
| Right rim recovery | Price must return to within 5% of the left rim | Without a tolerance here the cup can never complete, since price almost never returns to exactly the old high. Descriptions say the right side returns to the old high; they do not say how close is close enough |
| Handle duration | Between 5 and 25 bars | About one to five weeks. Removing the bounds took the population from 443 to 1,035, so this clause is doing more work than its brief mention in most descriptions suggests |
| Handle depth, absolute | No more than 12% from the right rim | The only clause that turned out to be entirely inert. Removing it left the identical 443 instances, every one of them, because the relative clause below always binds first |
| Handle depth, relative | No more than a third of the cup depth | The classical rule, and the one that keeps the handle in the upper part of the base. It removes structures where the pullback is really a second decline |
| Handle drift | The handle must not slope upward | Stated in words rather than as a number wherever the pattern is described, so a slope threshold had to be invented here as well. Measured by what it costs once the other clauses are satisfied, this is the most expensive clause in the definition: removing it takes the population to 1,593 |
| The break | A close above the higher rim by 0.10 of the average true range | A bare close above a line is noise at the top of a base. Entry is on the next bar's open, because entering on the bar that generated the signal grants the test information from the future |
| What is a swing low | A low with 5 higher lows on each side | Included for completeness, and it turned out to be inert. See below: changing it changed nothing whatsoever |
One of those entries deserves attention before anything is measured, because it is the sort of thing only coding reveals. The swing-low rule is inert. Changing it from three bars either side to fifteen produced not merely the same count but the identical set of instances, every one of them. The reason is that another clause already requires the cup low to be the lowest point between the two rims, and a point that is the lowest across an entire cup is automatically a swing low under any reasonable definition of one. One clause quietly makes another irrelevant, and nothing in the written description hints at it.
How often it actually happens
Across 9,000,000 daily bars, which is 3,000 independent generated instruments of 3,000 bars each and about 36,000 instrument-years, the full definition produced 443 completed structures. Overlapping detections are normally the largest source of double counting in a study like this, and here there were none at all: no two instances on the same instrument fell inside one holding period of each other, so all 443 are independent observations. That is roughly one instance per 81 instrument-years.
It is worth pausing on what that means for anyone reading the pattern off a screen. A trader following fifty charts would encounter a structure meeting the full classical definition about once every nineteen months. Somebody following two hundred would see two or three a year. That is not an argument against the pattern, but it is an argument against any confident statement about it, because a rule that fires that rarely accumulates evidence very slowly. Ten years of watching fifty charts is about six observations, which is not enough to distinguish a real edge from a run of luck in either direction.
The rejection census says where the population went. The clause that discarded the most structures was the depth band, which threw out 485,505 window evaluations, followed by the prior advance at 420,445 and the requirement that the cup low be the true bottom at 98,401. Structures that declined the right amount but never recovered to the old high accounted for 94,163. The rounding clause, the one that gets the most attention, rejected 14,296. Those are counts of rejected window evaluations rather than distinct chart formations, since several lookback lengths are tried at each candidate low.
That census answers a narrower question than it appears to, and the distinction matters. Being the first clause to fail is not the same as being the clause that costs you instances, because the checks run in an order and an early one absorbs everything a later one would also have caught. Measured the other way, by how much the population grows when a clause is switched off and the rest are left alone, the ranking inverts completely. The depth band, which rejected the most evaluations, costs relatively little: removing it takes the count from 443 to 597. The requirement that the handle drift downward, which rejected 1,231 evaluations and looks trivial in the census, takes it to 1,593. Any statement about which clause is doing the work has to say which of those two things it means, and most do not.
There is an obvious objection to the rarity figure, and it is worth answering with a measurement rather than an argument. The main tape has no unconditional drift by construction, and a pattern that requires a large prior advance followed by a full recovery ought to be commoner on a tape that rises over time. So a third tape was generated, identical in every respect except that the drift states are shifted upward to give it a positive unconditional drift of about 8% a year. The same detector found 603 instances there against 443 on the flat tape. The pattern does become more common when the tape rises, by about a third, and it remains rare in absolute terms. The rarity is mostly a property of the definition rather than of the drift.
The instances that did qualify were larger and slower than the impression a textbook diagram gives. The median cup ran 89 bars from rim to rim, with a quarter under 57 and a quarter over 140. Median depth was 21%. The median handle lasted 11 bars and gave back 3.4% from the right rim. Those are months of chart, not weeks, and the practical implication is that anyone claiming to find several of these a month on a single instrument is using a much looser definition than the one printed in the books.
The clause nobody can define
Every description of the pattern insists that the base must be rounded and not V-shaped. It is presented as the difference between a genuine accumulation and a panic that happened to bounce, and it is repeated more often than any other clause. It is also, as far as this exercise could establish, the only one that never arrives with a number attached. To run the definition at all, one had to be invented.
Here is the one used. Take the cup depth, measured from the higher rim down to the low. Call the lowest quarter of that depth the bottom of the cup. Then measure what share of the bars between the two rims have their low inside that bottom quarter. A perfect parabola spends half its life there. A perfect V spends a quarter. Anything above about 0.35 is therefore closer to a bowl than to a spike, and 0.35 is the threshold this page used.
That paragraph is a reasonable piece of reasoning and it is also an invention of this exercise. No source endorses it, and a different analyst reading the same textbook sentence would be equally entitled to formalise it another way. So a second formalisation was built alongside it, chosen to be as defensible as the first: fit a parabola to the price path across the cup, fit a two-segment V with its vertex at the low, give each three free parameters so the comparison is fair, and ask which one explains the path better. Where the parabola wins, the base is rounded. Where the V wins, it is not.
The two scores were then applied to the same 1,395 cups, which is the population that passes every other clause. If the rounding requirement were pointing at a real property of a chart, the two measurements of it ought to agree closely. They do not. The rank correlation between them is 0.20. Holding the strictness matched, so that each metric calls exactly the same number of cups rounded, they agree on 207 instances, disagree on 470, and among the quarter of cups each considers most rounded the overlap is 49%.
That is a stronger result than it might look, and it does not depend on any outcome measurement. It says that the most repeated clause in the specification does not identify a stable set of objects. Two analysts, both being careful, both formalising the same sentence honestly, would sort the same charts into different bins. Any statistic quoted for rounded bases therefore inherits an unquoted choice made by whoever computed it.
The invented threshold has a second problem, which is that it sits on a cliff. At a cut-off of 0.20 the definition admits 941 cups; at 0.35, 442; at 0.50, 149; at 0.60, 37. Small movements in a number that nobody publishes change the population by a factor of several. What did not change across that sweep was the measured outcome: it stayed within 0.05 of a percentage point of its matched base rate everywhere the sample held above two hundred instances, and only began to wander once the sample fell below a hundred and fifty, which is what a small sample always does. The clause is enormously consequential for how many patterns you find and apparently inconsequential for what follows them.
Measuring what followed, against a base rate
A pattern statistic on its own means very little. If a breakout is followed by a gain sixty percent of the time, the question is what fraction of comparable moments are followed by a gain, because the answer might also be sixty percent. Leaving that comparison out is usually what makes a pattern statistic look impressive.
Putting it in is harder than it looks, and the hard part is not the part anybody writes about. For each detected breakout the control takes a long position exactly as the pattern does, on the same instrument, held for the same 20 bars, with a target the same distance away as a fraction of its own entry price. Two hundred control entries are drawn per event. Everything about the two arms is matched except the reason for entering. One choice remains, and it turns out to be decisive: which bars the control is allowed to be drawn from.
The natural answer, and the one this exercise began with, is a window either side of the event, 250 bars back and 250 bars forward, so that market conditions are comparable. That window is contaminated, and the contamination is specific rather than vague. A cup and handle is defined by a decline, a recovery and a pullback that together span months. A window of 250 bars either side of the entry therefore contains the very price path that caused the detection. The control is being drawn from data the pattern selected, the two arms are no longer independent, and the comparison stops being a comparison.
That is testable rather than arguable. If a symmetric result reflects something real about the pattern, moving the control to bars strictly after the entry should leave it roughly where it was. If it reflects contamination, it should collapse, and restricting the control to bars strictly before the entry should exaggerate it.
It collapses. Drawn only from bars before the entry, the pattern arm comes in 2.16 percentage points below its control, 4.84 standard errors. Symmetric at 250 bars it is 1.09 points and 2.45 standard errors, which is the number this page would have published had the question never been asked. Widen the symmetric window to a thousand bars and it fades to 0.24 points. Draw the control only from bars after the entry and the difference is 0.05% at 0.11 standard errors, which is nothing at all.
The mechanism is easy to see once the progression is in front of you. The bars before a cup-and-handle entry are, by construction, the bars of a large decline and a full recovery. Control entries drawn from them buy into the middle of that recovery, at prices well below the breakout, so they do well and the pattern looks bad beside them. That is not a fact about cups. It is a fact about sampling the control from the same stretch of tape the pattern used to select itself, and it would have been reported here as a finding about the pattern.
So every number on this page uses a control drawn only from bars strictly after the entry. That is not simply a weaker control, which would make a null uninformative. On the tape with a real effect inserted it still finds the effect at 10.8 standard errors, so a null under it is a measurement rather than a loss of power. All 196,500 control draws made across the run had a valid forward window, so nothing quietly fell back to a wider one.
The lesson generalises well past this pattern. Any definition that spans months selects the neighbourhood it sits in, so a time-local control centred on the event is contaminated for all of them, and it will flatter or damn the pattern depending on which way the selected path happened to run. The general mechanics of why an obvious level attracts price and then rejects it are covered in the guide to breakouts. The point here is narrower and about method: a base rate has to be drawn from data the pattern did not choose.
With the repaired control, the two arms on the tape with no memory are the same distribution. The pattern arm returned 0.41% over twenty bars and its matched controls returned 0.41%. The difference between them is smaller than a hundredth of a percentage point, with a ninety-five percent interval running from 0.89 points below the control to 0.91 points above it. The proportion of positive outcomes was 51.0% against 50.6%. There is no part of the distribution where the pattern arm is distinguishable from a long entry taken for no reason.
A null on its own would still be worthless, because a detector that finds nothing might simply be broken. So the identical detector, the identical thresholds and the identical base-rate machinery were run over a second tape, built the same way but with one genuine effect inserted: after price had been well below its recent high and then closed above it, the following twenty bars carried extra upward drift. The effect was written purely in terms of a price channel and a drawdown, and never in terms of a cup, a handle, a rounding score or a depth band, so the detector had to find it on its own.
It did. On that tape the pattern arm averaged 6.38% against 2.05% for matched random entries, a difference of 4.32 percentage points at 10.8 standard errors, on 633 instances. The machinery can see an effect. On the tape where there was nothing to see, it saw nothing.
The experiment: taking the clauses out one at a time
Everything so far is setup. The question the page exists to answer is whether the specification earns its place, and there is a direct way to test that: relax exactly one clause, leave the rest of the definition untouched, and report both the population and the measured outcome. Do it for every clause. Then do it once more with all of them relaxed at once, leaving only the bare skeleton of a decline, a recovery to the old high, a pullback and a break.
Running the ablation on the tape with nothing in it answers one half of the question: do the constraints buy anything when there is nothing to find? Running it on the control tape answers the other half, and it is the more interesting one: when there genuinely is an effect present, do the constraints help you see it, or do they throw away the sample you need in order to see it?
| Clause relaxed | Found, tape A | Difference from base, tape A | Found, tape B | Difference from base, tape B |
|---|---|---|---|---|
| Nothing relaxed, full definition | 443 | 0.00% | 633 | 4.32% |
| The prior advance | 1,268 | 0.18% | 1,422 | 3.45% |
| The cup depth band | 597 | −0.22% | 938 | 4.12% |
| The cup duration floor | 534 | −0.18% | 749 | 4.24% |
| The rounding requirement | 1,395 | 0.10% | 1,846 | 3.36% |
| Handle depth against the cup | 609 | −0.24% | 807 | 4.17% |
| Handle depth in absolute terms | 443 | −0.01% | 633 | 4.34% |
| The handle duration bounds | 1,035 | −0.23% | 1,357 | 3.65% |
| The downward handle drift | 1,593 | 0.34% | 2,185 | 3.64% |
| All of them, bare skeleton | 52,291 | 0.14% | 52,329 | 2.41% |
Read the left half first. On the tape with no memory the difference from the matched base rate never left a band a third of a percentage point wide, no matter which clause was removed. Every single-clause variant landed between −0.24% and 0.34%, and not one of them reached two standard errors. The full definition itself sat at 0.00%, which is zero to two decimal places. Meanwhile the population ran from 443 up to 52,291. The specification changed how many things you find by a factor of more than a hundred and changed what followed them by nothing that could be measured.
One row in that column deserves a caveat rather than a headline. The bare skeleton, on 52,291 instances, produced a difference of 0.14% at 3.70 standard errors, which clears a conventional significance test. It is also 0.14 of a percentage point, found on a sample large enough to detect 0.11%. That gap between statistical detectability and economic meaning is the whole reason sample size has to be quoted alongside significance. An allowance of a few basis points a trade would consume that difference several times over. A sample of fifty thousand will always find something, and the useful question is whether what it finds is large enough to matter.
The right half is the more useful test, and it complicates the story rather than confirming it, which is why it is worth reporting carefully. On the tape where a real effect exists, the strict definition did concentrate that effect: 4.32 percentage points per instance against 2.41 for the skeleton. So the constraints are not pure decoration when there is something genuine to detect. They select events that sit closer to the mechanism.
But look at what the concentration cost. The strict definition reached 10.8 standard errors of confidence on 633 instances. The skeleton reached 54.2 on 52,329. The loose definition, which per instance found a weaker effect, established that the effect exists far more decisively, because confidence scales with the square root of the sample and the sample differed by a factor of about eighty. If your purpose is to find out whether an effect is there at all, the specification actively works against you.
There is a further detail in the same table that is easy to skim past. Of the eight single-clause relaxations, seven enlarged the population and the eighth left it exactly where it was. Not one of them cost any confidence on the control tape. The weakest came in at 10.8 standard errors against the full definition's 10.8, and the loosest single relaxation reached 17.9. If any individual clause were carrying real information about which recoveries matter, removing that clause should have diluted the effect enough to see. Removing them one at a time never did.
What a sample this size supports, and what it does not
The honest way to close a study like this is to state what the evidence can carry, and the limiting quantity is not the argument but the sample.
With 443 instances and the spread of twenty-bar outcomes measured on them, the smallest true difference this test could reliably have found is 1.2%. Anything smaller is invisible here. That is a demanding bar: an edge of half a percentage point per twenty bars, repeated, would be a substantial thing to own, and this sample could not have seen it. The inserted effect on the control tape was several times larger than the threshold, which is why it showed up clearly, and it was made that large deliberately so that a null on the other tape would mean something.
The two most-quoted numbers for the pattern can be reported, with the same caution attached. The measured move, which projects the depth of the cup upward from the breakout, was reached inside the twenty-bar holding period 8% of the time against 5% for matched controls with the same target distance, direction and window. Extended to sixty bars the figures were 23% and 18%. Those two gaps are the only place on this page where the pattern arm comes out ahead of its control, and they are worth reading rather than quoting.
The reason is in the spread, not the average. Outcomes following a detected breakout have a standard deviation of 9.4% against 8.3% for the controls, and the extra width is on both sides: the worst tenth of pattern outcomes runs below −10.2% against −8.9% for the controls, while the best tenth runs above 12.2% against 9.9%. The definition selects moments of higher volatility, which is not surprising given that it demands a large decline and a full recovery. A more volatile entry reaches any distant level more often, in either direction. Reaching an upside target more often is therefore not the same as making more money, and on the mean the two arms are indistinguishable. The deadline is an undeclared parameter as well: quote no time limit at all and the hit rate climbs toward certainty, because a wandering price eventually reaches most levels.
False breaks, counted as a close back below the breakout level within 5 bars, ran at 57% on the flat tape and 47% on the tape with a real effect in it. That is the most robust number on the page, and it is the one a plan has to survive.
| Claim | Does this page support it? |
|---|---|
| The full classical definition is rare | Yes, strongly. 443 instances in 36,000 instrument-years is a count, not an estimate, and it holds on a rising tape too |
| The rounding clause does not identify a stable set of patterns | Yes. Two careful formalisations of the same sentence correlate at 0.20 and disagree about 470 of 1,395 cups. This needs no outcome data at all |
| The constraints do not improve the measured outcome | Yes on the flat tape, across every single-clause relaxation and a hundredfold change in population, using a control the pattern could not have selected |
| A control drawn from bars either side of the event is safe to use | No, and this is the sharpest methodological result here. On this data that control turned a difference of 0.05% into one of −1.09% at 2.45 standard errors, purely by sampling the price path the pattern selected |
| The pattern reaches its measured-move target more often than a control | Yes, but as a volatility effect rather than a directional one. The pattern arm's outcomes are wider on both sides and its mean is the same, so the target and the drawdown both arrive more often |
| The constraints have no value at all | No. On the control tape they concentrated a real effect, roughly doubling it per instance. They also cost so much sample that the looser definition established it far more confidently |
| The pattern carries no edge in Indian equities | No. This tested generated data, not any exchange. It shows what the pattern does when nothing is behind it, which is the baseline a real study needs |
| The pattern carries an edge somewhere | Not shown either. The control tape proves only that the detector would find an effect if one were present in the data given to it |
| A small measured difference here is meaningful | No. Below about 1.2% over twenty bars this sample cannot distinguish an effect from noise. The loosest variant can resolve 0.11% and finds 0.14%, which is detectable and still too small to trade |
| Any of this is a reason to take a position | No. Nothing here is a trade trigger, a recommendation or a forecast, and the outcome spread is wide enough that no single instance is predictable |
The most important limitation is the familiar one. Generated data has no earnings, no policy announcements, no index rebalancing and no order books, so it cannot tell you what happens when a real base forms ahead of a real catalyst. What it can do, and what real data cannot, is provide a case where the true answer is known in advance. When a statistic reproduces on a tape with nothing in it, that statistic is not evidence, and identifying which of the familiar cup-and-handle numbers fall into that category is the contribution here.
A second limitation is that this tests one coded definition out of many possible ones. The ablation is precisely the reason to expect a different detector to find a different population. Anybody who wants to argue that a better formalisation would produce a different answer is making a testable claim, and the way to settle it is to write that formalisation down as code and run it against the same base rate.
What to keep, once the clauses are gone
None of this makes a rounded base useless to look at. It relocates the useful part, and it is a smaller and sturdier thing than the specification.
Keep the structural observation, drop the geometry. Underneath every clause there is one plain statement: price fell a long way, took a long time to come back, and is now testing the level it fell from. That is a question about supply at a specific price, and it can be examined directly by looking at what happened when price was last at that level. None of that requires a depth band or a rounding score.
Treat a long specification as a warning, not a credential. The instinct that a precise rule is a serious rule is exactly backwards when the precision has no source. Ask of any clause where its number came from. If the answer is that it was observed in examples that worked, the clause is a description of a sample and it will shrink your next sample without telling you anything.
Count before you conclude. The single most useful number in this whole exercise is 443, because it sets a ceiling on what anything else can mean. Before quoting any pattern statistic, ask how many instances it rests on and what the smallest detectable difference at that sample size would be. Most published pattern statistics do not report either, and a great many of them would not survive the question.
Budget for the false break first. 57% of breaks came straight back within a week on data where nothing at all was happening. That is not a market pathology to be outsmarted; it is the arithmetic of drawing a boundary at a prior high and asking whether price crossed it. Any plan built on this pattern has to survive that rate as its normal case. The same exercise run on a narrowing range produced a similar rate and a similar conclusion about where the information is not.
Run the ablation yourself, on whatever you use. Take your own rule, remove one condition, and rerun. If nothing changes except the number of trades, you have found a condition that costs you sample size and buys nothing. That single habit, applied to a handful of rules, will retire more of them than any amount of reading. If working that way appeals more than memorising shapes, it is the method we teach.
FAQ
Frequently asked questions
What is a cup and handle pattern?
It is a rounded decline and recovery, the cup, followed by a small pullback near the recovered high, the handle, and then a close above that high. What separates it from every other consolidation is the number of conditions attached to it: a prior advance, a depth range, a duration range, a requirement that the base be rounded rather than sharp, a handle no deeper than about a third of the cup, and a handle that drifts down rather than up. Seventeen separate numbers had to be supplied before any of it could be detected in code, and the sources that describe the pattern supply only some of them.
How common is the cup and handle?
Under the full classical definition, rare. The detector on this page found 443 non-overlapping instances across 9,000,000 generated daily bars, which works out at roughly one instance per 81 instrument-years. A trader watching fifty charts would see one every year or two. Relax every clause and keep only the bare skeleton of a decline, a recovery and a break, and the same data yields 52,291. The frequency you quote is a property of your definition, not of the market.
Does the rounding requirement mean anything?
It is the most repeated clause in the specification, it is stated in words every time, and this exercise could not find a version of it that carried a number. So a number had to be invented. Two reasonable formalisations were built here: how much of the cup is spent in its lowest quarter, and whether a parabola or a two-segment V fits the path better. On 1,395 cups the two scores rank the same structures with a correlation of only 0.20 and disagree about 470 of them at matched strictness. Two people applying the same textbook sentence carefully would select different patterns.
Do the constraints improve the result?
Not on the tape where nothing was there to find. Every clause was removed one at a time and the measured difference from the matched base rate stayed within a fraction of a percentage point of where the full definition left it, while the population moved by up to two orders of magnitude. On the control tape, where a real effect had been inserted, the strict definition did concentrate that effect somewhat, but it also discarded so many instances that the loosest variant measured the same effect far more confidently.
Is the measured move target reached?
The measured move projects the depth of the cup upward from the breakout level, and it is the standard target. In this test it was reached within the twenty-bar holding period in 8% of cases against 5% for matched controls with the same target distance, direction and window. Given sixty bars instead of twenty the figures were 23% and 18%. Those are the only numbers on the page where the pattern arm comes out ahead, and they read as a volatility effect rather than a directional one: outcomes after a detected breakout are wider on both sides than the controls, so a distant level of any kind is reached more often, while the average outcome is the same. The deadline is also an undeclared parameter in every version of the claim.
How often does the breakout fail?
Counting a close back below the breakout level within five bars, 57% of breaks came straight back. That happened on a tape with no news, no order flow and no participants, so it is not evidence of anything predatory. It is what a boundary drawn at a prior high does when price is wandering. Any plan built on this pattern needs to survive that rate as its normal case rather than treat it as the exception.
How should a base rate for a chart pattern be built?
Not from a window centred on the event, which is the obvious choice and the one this exercise started with. A pattern whose definition spans months selects the stretch of chart around it, so control entries drawn from either side of the detection come from the same price path that caused the detection. On this data that single choice turned a difference of 0.05 of a percentage point above the control into one of 1.09 points below it, at 2.45 standard errors, with nothing else changed. Drawing the control only from bars after the entry removes the overlap, and it was checked on a tape containing a known effect to confirm it still detects one.
Why does the sample size matter so much here?
Because it sets a floor on what any conclusion can mean. With 443 instances and the spread of outcomes measured here, the smallest true difference this test could reliably have detected is about 1.2% over twenty bars. Most claims made for chart patterns are smaller than that. A sample this size cannot confirm them and cannot refute them, and saying so is more useful than reporting a number that cannot bear weight.
Why test on simulated data instead of Indian stocks?
Because a generated tape can be built with a known answer inside it and real data cannot. Two tapes were used: one where nothing makes a recovery predictive, so any measured edge has to be an artefact, and one with a real effect deliberately inserted, which proves the detector can see an effect when there is one to see. That pairing separates claims about the cup and handle from claims about the definition of a cup and handle. It cannot tell you what the pattern does on any exchange, and nothing here should be read as saying it can.
Is a more specified pattern a more reliable one?
That is the assumption this page set out to test, and the test does not support it. Specificity has a guaranteed cost, which is sample size, and an unproven benefit. Every extra clause is also an extra opportunity for a threshold to have been set by looking at examples that worked. The way to tell the difference is to remove the clauses one at a time and see whether the result moves, which takes an afternoon and is the exercise the whole page is built around.
So what is worth taking from the pattern?
The structural observation underneath it, which needs none of the clauses: price fell a long way, took a long time to get back, and is now testing the level it fell from. That is a supply question about a specific price, and it can be examined directly. The rounding, the depth band and the handle proportions are the parts that carry thresholds nobody can source, and those are the parts the ablation on this page found nothing for.
Method note
How the numbers on this page were produced
Every figure comes from one deterministic simulation, seeded so that it reproduces identically on each run. Two primary tapes were generated, each of 3,000 independent instruments of 3,000 daily bars, 9,000,000 bars in total. Volatility clusters and drift regimes switch on both, so declines, long recoveries and quiet basing occur naturally and are never inserted by hand; the three drift states are symmetric and equally likely, so the unconditional drift of the tape is zero. The first tape contains no mechanism that could make a recovery predictive. The second adds one effect, defined entirely in terms of a drawdown and a channel breakout and never in terms of cup or handle geometry, and exists to prove the detector can find an effect when one is present. A third tape, identical except for a positive drift of about 8% a year, was generated only to check whether the rarity of the pattern is an artefact of a flat tape.
Detection runs on the open, high, low and close of the generated bars using the thresholds in the definition table. Because no published description says how far back the left rim may be sought, the detector tries a ladder of lookbacks and keeps the shortest cup that satisfies every clause. Signals are taken on the close and positions are opened at the next bar's open, so no result uses information that was not available at the time. Outcomes are measured over 20 bars. Overlapping detections are discarded so that no two measured outcomes share a holding period, and on this data none had to be.
The base rate is the part of the method that needed correcting during the work, so it is stated in full. Each control entry is matched to its event on instrument, direction, holding period and target distance as a fraction of entry price, at two hundred draws per event for the headline results and one hundred for each ablation variant. Control entries are drawn only from bars strictly after the event, within 250 bars of it. An earlier version of this exercise drew them from 250 bars either side, which is the obvious choice and is wrong here, because a window that wide around the entry contains the decline and recovery that caused the detection. The sensitivity of the result to that choice is reported on the page rather than hidden, and the forward-only control was checked on the tape with a known effect to confirm it still detects one. Every one of the 196,500 control draws in the run had a valid forward window. Returns are shown before costs; an illustrative allowance for the charge stack would apply equally to both arms and would not change the difference between them.
All results are illustrative and simulated. They are not a track record, they are not a forecast, and they are not an indication of what any pattern would produce in a live account or on any Indian security or index. The purpose of the exercise is to establish which properties of the cup and handle are consequences of its definition, which is a question about the definition rather than about any particular market.
Related