Educational Reference
Flags and Pennants: The Pattern Is the Impulse, Not the Shape
A flag is a brief, shallow drift against the direction of a sharp move. A pennant is the same drift with boundaries that close toward each other instead of running parallel. Almost every treatment spends its attention on the little shape and treats the move in front of it as scenery. This page treats the move as the subject, writes both halves as code strict enough to run, detects every occurrence across 1,800,000 generated daily bars, and reports what the detector found.
The finding, stated first. Hold the consolidation definition completely fixed and move only the impulse requirement, from one average true range to six. On a tape built with no directional memory the result does not move: 0.08 percentage points per extra unit of impulse, 1.51 standard errors from zero. Run the identical sweep on a tape that does contain a continuation effect and the same measurement climbs from 1.22% to 2.66%, a gradient of 0.49 points per unit at 9.15 standard errors. The test can see a sharper impulse paying off. Where nothing was there to be found, it found no gradient. A second finding was published here and has been withdrawn: this page reported that flags taken with the prevailing trend underperformed by 0.56 percentage points at 3.58 standard errors. Half of that comparison group was drawn from the impulse that defines the pattern. Rebuilt so it cannot be, the figure is 0.24 points at 1.51. All results on this page are illustrative and simulated.
The famous half is the smaller half
Open any description of a flag and count the words. The overwhelming majority go to the rectangle: how many bars it should run, how steeply it should lean, how tightly it should hold, whether the boundaries are parallel or converging and therefore whether the thing is a flag or a pennant. The sharp move that came first is usually granted a sentence. It is called the pole, or the flagpole, and treated as the thing the shape is attached to rather than the thing the shape is about.
That ordering is backwards, and it is possible to say so before running anything. Take away the sharp move and what remains is a few bars of quiet sideways drift. Quiet sideways drift is not rare. It is close to the default state of a price series, and a definition that finds it will find it constantly. Take away the drift instead and what remains is a strong directional thrust, which is genuinely uncommon and genuinely measurable. If either half of this pattern carries information, the prior should be heavily on the half that is scarce.
There is a second reason to suspect the shape, and it only becomes visible when you try to code the thing. Every clause that describes the drift turns out to be expressed as a fraction of the move in front of it. Shallow means shallow relative to the impulse. Brief means brief relative to how fast the impulse travelled. Narrow means narrow compared with the impulse height. The shape has no independent existence in the definition at all: it is a set of ratios whose denominator is the impulse. You cannot even state what a flag is without first stating what a flagpole is, and almost nobody states the second.
So this page reverses the usual emphasis. It defines the impulse with as much care as the drift, detects both, and then runs one experiment that the literature does not: it varies the impulse requirement alone, holding every other clause frozen, and asks what a stricter impulse actually bought. Where flags and pennants sit in the wider family of shapes is laid out in the guide to chart patterns in Indian stocks; the work here goes underneath that map and tests one entry on it.
Writing the impulse down is where it gets difficult
Here is the verbal definition, close to the standard one. A flag is a short, shallow consolidation that slopes against a sharp preceding move, resolving in the direction of that move. Try to run it. What is sharp? Sharp compared with what, measured over how many bars? What is short, and short in bars or in proportion to the move? What is shallow, and shallow measured from where? What counts as sloping against, given that a drift is never exactly counter to anything? And when do parallel boundaries become converging ones?
Seven questions, and the sentence answers none of them. That is not a complaint about sloppy writing. It is the reason two people can look at the same chart and disagree in good faith about whether there is a flag on it, and it is the reason published statistics about flags are difficult to compare with each other. Every one of those statistics belongs partly to whoever chose the missing numbers.
Two choices in the table below deserve to be flagged before the results arrive. The first is that the impulse is measured against the average true range from before the move began. Using the range during the move would let a violent thrust inflate its own yardstick, so that every sharp move scores about the same and the threshold stops discriminating. The second is that the drift is allowed to run for as long as it keeps qualifying, up to a cap, rather than being cut at a fixed length. Both are decisions nobody publishes, and both are measured for their consequences further down.
| Clause | Value used | Why this value, and what it costs |
|---|---|---|
| How sharp the impulse must be | Net move of 3.0 average true ranges over 5 bars | The single most consequential number on this page, and the one no published definition supplies. The range is taken from before the move so the thrust cannot inflate its own yardstick |
| How clean the impulse must be | Net move at least 0.55 of the distance actually travelled | Separates a thrust from a week of churn that happens to end higher. It turned out to be nearly inert: removing it entirely changed the population by under two percent |
| How brief the drift must be | Between 5 and 20 bars | The drift runs as long as it keeps qualifying, up to the cap. Cutting it at a fixed short length instead is measured further down and barely moved anything |
| How shallow the drift must be | Gives back no more than 50% of the impulse | The clause that makes it a pause rather than a reversal. Loosening it from a third to nearly two thirds roughly doubled the population |
| How narrow the drift must be | Whole range under 60% of the impulse height | Stops a wide, violent chop being read as a tidy consolidation. Note that this is a ratio: without an impulse it has no denominator and no meaning |
| That the drift is a pause | Runs on past the impulse by no more than 25% of it | Rejects the case where price simply kept going, which is not a pattern but the absence of one. This clause fired 49,390 times |
| That the drift leans back | With-impulse slope under 0.15 of the impulse height | Allows a flat or counter-sloping drift and rejects one still travelling with the move. Written as a limit rather than a requirement, because insisting on a counter-slope discards half the real cases |
| Flag or pennant | Boundaries converge by 35% or more | Does the whole job of naming, on its own. One cut-off separates the two words, and moving it inside a reasonable range renames one pattern in five |
| What counts as a break | Close beyond the boundary by 0.10 average true ranges | A bare close beyond a line drawn through a narrow drift is noise. The buffer is small enough not to miss real exits |
| Entry and holding | Next bar open, held 20 bars | Entering on the bar that produced the signal would quietly hand the test information from the future, which is the most common silent error in pattern studies |
How often a real one occurs
Across 1,800,000 daily bars, which is 600 independent generated instruments of 3,000 bars each and about 7,200 instrument years, the definition produced 5,501 candidate patterns. After discarding overlaps, so that no two measured outcomes share a holding period on the same instrument, 3,042 independent instances remained. That is roughly 0.42 per instrument per year, or about one every two and a half years on any given chart.
That is a strikingly small number, and it is small for one reason. The impulse clause is doing nearly all the filtering. Relax it from 3.0 average true ranges to one and the population rises to 6,322; tighten it to six and it falls to 308. Nothing else in the definition comes close to that leverage: loosening the shallowness requirement moved the count by about a factor of two, the narrowness requirement by about three, and the cleanliness requirement by almost nothing at all.
The rejection census tells the same story from the other side. The most common reason a qualifying impulse produced no pattern was that no drift of the required kind followed it at all, 70,013 times. Next came price simply carrying on rather than pausing, 49,390 times, then a give-back too deep to count as a pause, 21,221 times, then a drift too wide, 6,279 times, and finally 2,013 cases where the boundaries widened so much that neither name applied. These are counts of rejected evaluations rather than of distinct chart formations, but the ordering is the point: most sharp moves are not followed by anything a strict definition will call a flag.
The instances that did qualify were brisk. The mean drift ran 9.2 bars and the median 7, with the middle half falling between 6 and 11, and when the break came it came almost immediately, on average 1.6 bars after the drift stopped qualifying. Only 14 of 5,501 candidates failed to break at all inside the window allowed. The mean impulse behind a detected pattern was 4.3 average true ranges, comfortably above the 3.0 floor, because a threshold selects the tail of a distribution rather than a point on it. One more measured detail is worth having, because it quietly contradicts the picture in every textbook: the median drift sloped 0.9 percent of the impulse height, which is to say it did not lean against the move at all. It sat flat. The counter-trend drift that gives the pattern its name is the exception rather than the rule.
Measuring what followed, against a base rate
A pattern statistic on its own says almost nothing. If a break is followed by a gain sixty percent of the time, the useful question is what fraction of random moments on the same data are followed by a gain, because the answer may also be sixty percent. Leaving that comparison out is usually what makes a pattern statistic look impressive.
The base rate here is built to be unfair to the pattern in every respect except the one under test. For each detected break, two hundred random entry points are drawn from the same instrument, within 250 bars of the event, taking a position in the same direction and holding it the same 20 bars. The only difference between the two arms is that one entered because a flag broke and the other entered for no reason at all. The canonical trade is measured, which means only breaks that went the same way as the impulse: a flag that broke backwards is a failed flag, not a short signal.
Those control entries are drawn only from bars that fall after each pattern. An earlier version of this page drew them from 250 bars on either side, which sounds more even-handed and is the single worst decision it made. A flag is defined by a thrust of 3.0 average true ranges. Put the control window around the pattern and the backward half of it contains that thrust, so the comparison group is partly made of the move the pattern was selected for. The next section is entirely about what that cost, because it cost this page a published finding and it moved every level on it.
On the tape with no memory the two distributions are close to the same distribution. The pattern arm returned 0.29% over twenty bars and the matched random arm −0.01%, putting the pattern 0.30 percentage points ahead, which is 2.79 standard errors from zero, with a ninety-five percent interval running from 0.09 points to 0.52. The share of positive outcomes was 52.4% against 50.0%. The tails are close: a tenth of pattern outcomes were worse than 6.6 percent down against 7.3 for the controls, and a tenth better than 7.2 up against 7.3.
That number is not zero and this page is not going to pretend it is. Under the older, symmetric control the same arm read 0.15 points behind at −1.37 standard errors. Moving the control window forward moved it 0.45 points, from a negative that was not significant to a positive that is, on a tape that is supposed to contain nothing. Both readings cannot be right, and as it turns out neither is: what the two of them bracket is a comparison group that is contaminated in one direction and a horizon mismatch in the other. That is unpicked in the next section, with a placebo ladder, because a page that only reports the correction it likes is not doing the thing it claims to do.
Measuring the break in whichever direction it went, which is the arm the companion page on triangles reports and therefore the comparable one, gives 5,365 instances and a gap of 0.16 percentage points, 1.89 standard errors. Both arms say the same thing, and both moved the same way when the control window did.
A null result from a detector that finds nothing is worth nothing, because the detector might simply be broken. So the identical detector, the identical base-rate machinery and the identical thresholds were run over a second tape, built the same way but with one genuine effect inserted: whenever the ten-bar return was large relative to its own recent scale, the following thirty bars carried extra drift in that direction. The effect is written purely in return space. It says nothing about a consolidation, nothing about shallow drift, nothing about boundaries and nothing about a break, so the flag detector had to find it unaided.
It did, emphatically. On that tape the pattern arm averaged 1.89% against 0.10% for matched random entries, a difference of 1.79% at 15.51 standard errors, with 62.2% of outcomes positive against 50.5%. That is the power check, and it is the number that gives the forward-only control the right to report a null anywhere else on this page: it is not a comparison group that reports nothing for everything. It is also, of the four control windows tried, the one that finds the planted effect most clearly. Strictness cost no power at all here.
| Variant | Found | Broke with the impulse | Mean, 20 bars | Matched base rate | Target reached | False breaks |
|---|---|---|---|---|---|---|
| Flag | 1,933 | 56% | 0.40% | 0.01% | 21% vs 21% | 40% |
| Pennant | 1,109 | 56% | 0.11% | −0.04% | 15% vs 18% | 32% |
| Both, weakest impulse band | 1,297 | 54% | 0.00% | 0.04% | 37% vs 44% | 42% |
| Both, strongest impulse band | 293 | 60% | 0.04% | 0.01% | 9% vs 11% | 37% |
| Drift shape alone, no impulse | 58,996 | not defined | 0.10% | 0.01% | 27% vs 27% | 46% |
| All, control tape | 3,292 | 61% | 1.89% | 0.10% | 25% vs 20% | 33% |
Two rows deserve a second look. The measured move, which projects the impulse height from the break and is the standard target for this pattern, was reached inside the holding period 19% of the time against 20% for matched random entries; over sixty bars, 43% against 41%. It is reached often enough to be memorable and slightly less often than chance. And the false break rate, counting a close back inside the boundary within 5 bars, was 37% on the tape where nothing was happening and 33% on the tape where something was. Roughly one break in three comes straight back regardless, which is what a boundary drawn through a narrow drift does. Why obvious levels attract and then reject price is the subject of the guide to breakouts.
The comparison group, taken apart
This page published a context finding that is now withdrawn, and it published every level on the page against a comparison group that was contaminated. The withdrawal comes first, then the evidence, then the part that is still unresolved, because leaving the last one out would be the same mistake in a nicer suit.
The finding, as published. Split the instances by whether the impulse ran the same way as the preceding sixty-bar move, the standard trade-with-the-trend refinement. Measured against a comparison group drawn from 250 bars on either side, the agreeing group came in 0.56 percentage points below its own base rate at 3.58 standard errors across 1,513 instances, while the disagreeing group came in 0.22 points above. The page read that as the refinement making things worse, and cited the sibling triangle study as independent agreement.
The finding, corrected. Against a comparison group drawn only from bars after each pattern, the agreeing group reads 0.24 points at 1.51 standard errors and the disagreeing group 0.36 at 2.41. The gap between them is gone, and what is left is the two cells sitting in the same place. The refinement is not supported here and it is not contradicted here. It is untested.
Look at which arm moved. The two pattern arms never budged: 0.29% with the trend and 0.30% against it, which is the same number twice, because moving a control window cannot change what the pattern returned. The two comparison groups moved a great deal: the with-trend control fell from 0.85% to 0.05%, and the against-trend control from 0.08% to −0.06%. The entire context effect was in the control, and specifically in the half of the control window that sat on top of the sixty-bar move being split on.
Three predictions follow from that, and they can be checked separately. If the contamination is in the backward half, then restricting the control to the backward half alone should exaggerate the effect, widening the window should dilute it, and pushing it fully forward should remove it. Backward only: 1.45 points at 9.19 standard errors, more than two and a half times the published figure. A thousand bars either side: 0.04 points at 0.27. Forward only: 0.24 at 1.51. Monotone, in the predicted order, exactly as a control drawn from the pattern's own bars would behave.
The placebo that settles it. A window argument can be waved away as a judgement call, so here is a measurement that cannot. Build an event set that is nothing but the impulse: every bar where the 3.0 average true range thrust fires, entered the next bar in its own direction, with no drift clause, no channel, no break and no flag. Measure that against the surrounding control the page used to use. It comes in 0.21 percentage points below its own base rate at 4.42 standard errors. A comparison group that shows the impulse significantly underperforming itself is not a strict control, it is a circular one. Against the forward-only control the same placebo reads 0.06 points at 1.26, and across six independently generated tapes it averages −0.01 points at −0.17 standard errors, which is zero to as many decimal places as this exercise can produce.
The full ladder, every rung measured against the same forward-only control on the same tape, reads: random bars with a coin-flip direction, 0.04 points; random bars with the direction set by the prior five-bar move, 0.06; the impulse alone, 0.06; the detected flags, 0.30. The first rung is the harness testing itself on entries that carry no pattern whatsoever, and it reads zero, which is the only reason the rest of the ladder means anything.
What is still unresolved, stated plainly. The forward-only control is the right control and it is not a neutral one. Across six independently generated tapes the flag arm reads 0.19 points above it at 3.01 standard errors, where the same six tapes read 0.26 points below the symmetric control at −3.85. Both controls are biased and they are biased in opposite directions. The symmetric one is understood. The forward-only one is not, and three things are known about it. It is not the arithmetic of quoting returns as ratios, because the same excess appears unchanged in log space, 0.31 points against 0.30. It is not a long-short imbalance, because the event set is 48.7% long. And it is at least partly the generator: switch the drift off in all three regime states, so that the log price becomes an exact martingale and nothing measurable from the past can predict anything, and the same six tapes give 0.11 points at 2.44 standard errors against 0.19 at 3.01 with the drift on.
The honest reading of that last pair is that roughly half the residual is the tape and the remaining half is unaccounted for at six tapes of resolution, which is not enough resolution to say more. What it means for everything else on this page is a bound rather than a point: on data built to contain nothing, the flag arm measures somewhere between zero and a fifth of a percentage point over twenty bars, against a control that is itself uncertain at that scale, before any cost. Twelve basis points of round-trip cost sits inside that band. The page's conclusions are stated at that resolution from here on, and the one conclusion that does not depend on it at all is the next section, because a gradient does not care what constant the control adds to every point on it.
The central experiment: only the impulse threshold moves
Everything so far describes one setting of one definition. The experiment this page exists for is narrower and more useful. Freeze the consolidation clauses entirely, change nothing about how shallow, how brief, how narrow or how sloped the drift must be, and move only the number that decides how sharp the move in front of it has to be. If the pattern carries content anywhere, this is where it should show up, because this is the clause that is supposed to separate a real flag from a coincidence.
On the tape with no memory a stricter impulse bought nothing, and the way to see that is to read the line rather than any single point on it. Every point on the coral line now sits at or a little above zero rather than below it, lifted by the residual the previous section could not fully account for. None is further from zero than 0.44 percentage points and none is beyond 3.58 standard errors, while the population falls from 6,322 instances to 308. What matters here is that the line is level: trading one twentieth of the available opportunities produced the same thing, whatever that thing is, as trading all of them. Under the older symmetric control every point on this line sat below zero instead, and it sagged at the strict end, because a bigger impulse contaminates a backward-looking control more. The sag was the control, not the pattern. The level was never the finding on this page. The gradient was.
Nested thresholds share their events, though, so the differences between neighbouring points on that line are not independent of each other. The cleaner version detects once at the loosest setting and splits the resulting population into non-overlapping bands by how strong each impulse actually was, which gives independent cells and a trend that can carry a standard error. On the tape with no memory that trend is 0.08 percentage points per additional unit of impulse, with a standard error of 0.05, or 1.51 standard errors from zero. Flat.
The obvious objection to a flat line is that the measurement might be incapable of producing anything else. That is what the second tape is for, and here it earns its place twice over. Running the identical sweep on the tape that does contain a continuation effect produces a line that climbs the whole way, from 1.22% at the loosest impulse to 2.66% at the strictest. In the disjoint bands the trend is 0.49 percentage points per unit at 9.15 standard errors, and the band-by-band figures climb from 0.37% in the weakest to 2.60% in the strongest, a sevenfold spread across cells that share no events.
That is the whole argument in one comparison. The experiment is demonstrably capable of detecting a payoff to impulse strength, at a sample size and an effect size of the same order as the null case, because it detected one on demand. When the same experiment is run where no such payoff exists, it reports none. The flat line is a measurement, not a failure to measure.
One thing did rise with impulse strength on both tapes, and it is worth dwelling on because it is the statistic most often quoted as proof. The share of breaks that went in the direction of the impulse, the continuation rate, climbed from 53.7% in the weakest band to 60.0% in the strongest, on the tape with no memory. Overall it stood at 56% of 5,487 breaks. A trader running this filter on real data would observe exactly that: demand a sharper pole and a larger majority of your flags resolve the right way. The observation is real, it is reproducible, and it happens where nothing whatsoever is going on.
The cause is geometric and worth understanding because it generalises. The drift is defined to lean against the impulse or to sit flat, which means the boundary on the impulse side of the drift is the one price is closest to and the one it reaches first. The bigger the impulse, the more sharply the definition constrains the drift to hug that side, and the more lopsided the outcome becomes. A continuation rate is therefore a property of how you defined the drift, and any figure quoted for it has to be netted against a geometric baseline before it can be read as evidence about markets.
Flag against pennant
The two names describe one difference: whether the drift's boundaries stay roughly parallel or close toward one another. Under this definition they were separated by a single cut-off, and the population split 1,933 flags to 1,109 pennants. That the difference is a cut-off rather than a category is easy to demonstrate: move it and the names move with it.
On the measures that decide whether the distinction is worth carrying, the two names are indistinguishable. The continuation rate was 56.2% for flags against 55.7% for pennants, a gap of 0.30 standard errors. The twenty-bar outcome differed by 0.29 percentage points at 1.27 standard errors, and both sat on the same side of their own base rates, flags by 0.39 points and pennants by 0.15, a difference between the two names well inside the residual the comparison group carries. Whatever a pennant is, it is not a flag that resolves better.
On two other measures they separated clearly, and the reason is mechanical rather than predictive. Flags reached the measured move 21.4% of the time against 14.7% for pennants, a difference of 4.75 standard errors, and flags produced false breaks 39.9% of the time against 31.6%, 4.66 standard errors apart. Both follow from the same fact. A converging drift has its boundaries closest together at the moment of the break, so the break happens from a tighter position and is less likely to be immediately undone, and the drift itself ran longer, 10.1 bars against 8.7, leaving less of the holding period for a distant target to be reached. Neither difference is about what the market intends to do next.
The naming itself is unstable in the way this kind of tolerance usually is, though less dramatically than one might expect. Holding the detected population completely fixed and moving the convergence cut-off across the middle of its plausible range, 2,407 of 3,042 patterns, 79%, kept the name they started with. The reason it is not worse is instructive: there are only two bins here and one cut-off, whereas the companion page on triangles had three bins whose boundaries were set by a single tolerance doing two unrelated jobs at once, and there only fifty-eight percent of instances held their label. Fewer names, less reshuffling. It remains true that one pattern in five changes identity for no reason connected to the chart.
Worth noting in passing: 791 of the 3,042 detected drifts had boundaries that widened rather than narrowed, and the median convergence across the whole population was 22%. The parallel channel that illustrations show is not the typical case; it is the middle of a continuum with converging drifts on one side and quietly broadening ones on the other.
What the impulse clause is actually worth
If the impulse is the subject, the fair question is what happens when it is removed. That cannot be done by setting its threshold to zero, because every clause describing the drift is a ratio with the impulse height underneath it, and a zero impulse leaves those clauses with nothing to be a fraction of. The workaround is to re-anchor the same three clauses to average true range, so shallow and narrow and brief keep their meaning while the requirement for a preceding thrust disappears entirely.
The bare shape is common, which was the prediction. It appeared 58,996 times against 3,042 for the full definition, about 8.2 times per instrument per year rather than 0.42. Anyone scrolling a chart and seeing quiet drifts everywhere is seeing something real; they are just not seeing flags, because the flag is the impulse.
On the tape with no memory the bare shape sits 0.09% above its matched base rate, inside the same unresolved residual as everything else measured on that tape and a quarter the size of the full definition's 0.30%. Two detectors, one nineteen times more selective than the other, landing in the same small band on a tape that should contain nothing. On the tape that did contain a continuation effect they part company decisively: the bare shape came in 0.40% above its matched base rate at 14.08 standard errors, while the same shape behind a real impulse came in 1.79% above at 15.51. The gap between those two is more than four times anything either detector produced where there was nothing to find, which is the comparison the conclusion rests on.
That is the clearest statement this page can make about where the content sits. When there was something to find, requiring an impulse in front of the shape multiplied the measured excess by about four, at the cost of discarding nineteen occurrences in twenty. The shape contributed something, because a quiet drift often does follow a thrust and the bare detector picked up some of the same events by accident, but it was the minority contribution. The impulse clause is the pattern. The rectangle is how you notice it.
The counterpart of that result is the one already reported: on the tape with no memory, making the impulse requirement stricter did not help either. Presence and severity are different questions, and the answers here differ. Requiring an impulse changes what the detector is looking at. Requiring a bigger one changes only how many instances survive, and on data with nothing in it, it changes nothing else.
How much of this is the definition rather than the data
Every number above rests on ten thresholds, and a result that survives only at one setting is a result about that setting. The consolidation clauses were varied one at a time with the impulse held fixed. Loosening the shallowness requirement from a third of the impulse to nearly two thirds moved the population from 1,674 to 3,288; loosening the narrowness requirement moved it from 1,086 to 3,584; changing the cap on drift length between twelve and thirty bars moved it from 3,093 to 2,997, which is barely at all. Against the forward-only comparison group every one of those settings sits within 0.37 percentage points of zero, essentially the same spread as the 0.35 the symmetric version gave, and the ordering of the clauses by leverage is unchanged. The control window moved all of these levels together by roughly one constant, which is exactly why a sensitivity study built on differences survives the change of control and a claim about the level does not.
The window rule is worth a paragraph on its own, because it is where this page and its companion diverge. The triangle study found that its unstated rule about how far back to look dominated everything else: switching from the shortest qualifying window to a single fixed one cut its population from 8,036 to 711, a collapse of more than nine in ten. The same substitution here does almost nothing. Cutting the drift at the shortest qualifying length rather than letting it run produced 3,107 instances against 3,042, with a mean drift of 5.1 bars against 9.2.
The reason is worth stating because it is the one structural advantage this pattern has. A triangle has no anchor: its start is wherever the analyst decides to begin looking, so the window rule creates the pattern. A flag has an anchor, and it is the impulse. The drift starts on the bar after the thrust ends, and no choice about lookback can move that. Demanding an impulse buys the definition something real, which is a start date that is not a matter of opinion, even though it buys nothing measurable in outcomes on this data.
The two pages once agreed on a substantive finding and it has been withdrawn from both, which is worth spelling out because the agreement was the reason it survived as long as it did. Both reported that taking the pattern with the prevailing trend made things worse. Both were running the same comparison machinery on the same generator at the same seed, so their agreement was never independent of each other. Two measurements sharing the instrument that produced them will agree whether or not the claim is true. Before citing one of these pages as corroborating another, the question to ask is not whether they agree but whether anything about them could have disagreed.
What this supports and what it does not
A simulation can establish some things and not others, and being exact about which is which is the difference between a study and an opinion with numbers attached.
| Claim | Does this page support it? |
|---|---|
| The drift cannot be defined without the impulse | Yes, and structurally rather than statistically. Every clause describing the shape is a ratio whose denominator is the impulse height, which is why the impulse requirement cannot simply be set to zero |
| The impulse clause does the filtering | Yes. It takes the population from 58,996 to 3,042 on identical data, and on the tape containing a real effect it multiplied what was recovered by about four |
| A stricter impulse improves the outcome | Not on data with no memory: the trend was 0.08 points per unit at 1.51 standard errors. On the control tape the same sweep gave 0.49 points per unit at 9.15, so the test had the power and simply found nothing to report |
| A rising continuation rate is evidence of continuation | No. It rose with impulse strength on both tapes, including the one with nothing in it, because the drift is defined to lean against the move. It is a property of the definition |
| Flags and pennants are two different patterns | No. They differ by 0.30 standard errors in continuation and 1.27 in outcome, and one in five changes name when the cut-off moves. The two measures that do separate them follow from the geometry |
| Taking a flag with the prevailing trend does worse than random | No, and this page said otherwise until the comparison group was rebuilt. Against a control drawn only from bars after the pattern the figure is 0.24 points at 1.51 standard errors. Untested here, not refuted |
| A time-local comparison group is a safe control for a chart pattern | No, and that is the transferable result. An event set consisting of nothing but the impulse reads 0.21 points below a surrounding control at 4.42 standard errors, which is the impulse being measured against itself |
| The forward-only comparison group is neutral | Not entirely, and the page reports the residual rather than hiding it. Across six independent tapes the flag arm sits 0.19 points above it at 3.01 standard errors. About half of that is the generator's own drift regimes; the rest is unaccounted for at this resolution |
| Roughly one break in three comes straight back | Yes, under this definition, on both tapes. It is the most robust number on the page and the largest practical cost of trading the shape |
| Flags carry no edge in Indian equities | No. This tested generated data, not any exchange. It shows what the pattern does when nothing is behind it, which is the baseline a real study needs before it can claim anything |
| Any of this is a reason to take a position | No. Nothing here is a trade trigger, a recommendation or a forecast, and the outcome distributions are wide enough that no single instance is predictable |
The largest limitation is the obvious one. Generated data has no earnings, no policy announcements, no index rebalancing and no order books, so it cannot say what happens when a real thrust is driven by a real event and a real pause follows it. What it can do, and what real data cannot, is supply a case where the true answer is known in advance. When a statistic reproduces on a tape with nothing in it, that statistic is not evidence, and sorting the familiar flag numbers into that category is the contribution here.
A second limitation is that this tests one definition. A different set of ten numbers is a different detector, which is exactly why the sensitivity work above exists. Anybody who wants to argue that a better definition would give a different answer is making a testable claim, and the way to settle it is to write that definition as code and run it.
What to do with a flag instead
None of this makes a sharp move followed by a pause uninteresting. It relocates the interesting part, and it changes what is worth measuring.
Measure the impulse, not the rectangle. The one clause on this page that changed a result was the requirement for a real thrust in front of the shape. A thrust is measurable directly, in units of recent range, without drawing a single line. If the reason a chart has your attention is that price just moved hard and then stopped, then measure the move and the stopping, because both measurements are unambiguous and the drawing is not.
Do not read a continuation rate as continuation. The share of flags resolving with the trend rose with impulse strength on a tape containing no trend information at all. That is the single most quotable statistic about this pattern and it is manufactured by the definition. Before believing any version of it, ask what the same number would be on data with nothing in it.
Stop paying attention to which of the two names applies. Flags and pennants were indistinguishable in outcome and in continuation here, and one instance in five swapped names when a cut-off moved. Time spent deciding whether the boundaries are parallel enough is time spent on a filing decision.
Budget for the false break before the target. About one break in three closed straight back inside the boundary, on both tapes, and the measured move was reached slightly less often than for random entries with the same target and window. Any plan built on this shape needs the reversal as its normal case rather than its exception.
Write the definition down before you look for instances. This is the habit that transfers. Forced to say how sharp is sharp, and against what yardstick, and measured over how many bars, you discover how much of your pattern recognition was a decision rather than an observation. Doing that once is worth more than reading ten descriptions, and if that way of working appeals more than the taxonomy it replaces, it is the method we teach.
FAQ
Frequently asked questions
What is the difference between a flag and a pennant?
Only whether the two boundaries of the drift stay roughly parallel or close toward each other. Both sit behind the same thing: a sharp move followed by a brief, shallow pause. On the test run here the two were separated by a single cut-off, and moving that cut-off across a reasonable range left 79% of 3,042 patterns with the name they started with. They also behaved the same way afterwards, differing by 1.27 standard errors in outcome and 0.30 standard errors in how often the break went with the preceding move.
How sharp does the move before a flag have to be?
No published definition supplies a number, which is the problem this page starts from. The detector used here required a net move of 3.0 times the average true range over 5 bars, measured with the range from before the move so the thrust cannot inflate its own yardstick. That single number is the most consequential choice in the whole definition: loosening it to one average true range multiplied the population by roughly two, and tightening it to six divided it by ten.
Does demanding a sharper impulse improve the outcome?
On the tape with no memory in it, no. Sweeping the impulse requirement from one average true range to six, with the consolidation definition held completely fixed, moved the result by 0.08 percentage points per extra unit of impulse, in the unhelpful direction, which is 1.51 standard errors from zero. On the tape that did contain a genuine continuation effect the same sweep raised the excess from 1.22% to 2.66%, a gradient of 0.49 points per unit at 9.15 standard errors. The experiment can see an impulse gradient when one exists. On the tape where none existed it saw none.
Is the little shape worth anything on its own?
Much less than the move in front of it. Re-anchoring the same drift clauses to average true range, so the impulse requirement can be removed entirely, the shape alone was found 58,996 times against 3,042 for the full definition, about 8.2 times per instrument per year rather than 0.42. On the tape carrying a real continuation effect the bare shape came in 0.40% above its matched base rate while the same shape behind an impulse came in 1.79% above. The shape is common, the impulse is rare, and the filtering is done almost entirely by the impulse.
How often do flags break in the direction of the preceding move?
In this test 56% of 5,487 breaks went the same way as the impulse, and the rate climbed with impulse strength, from 53.7% in the weakest band to 60.0% in the strongest. That happened on a tape built with no directional memory whatsoever, so none of it is evidence of continuation. It is what happens when a drift is defined to lean against a move: the boundary on the side of the move sits closer to price and gets reached first. A continuation statistic has to be netted against that before it means anything.
How common are flags and pennants?
Under this definition, rare. Across 1,800,000 daily bars the detector produced 5,501 candidates and 3,042 independent tradable instances, about 0.42 per instrument per year, or roughly one every two and a half years on any given chart. That is a direct consequence of demanding a real impulse. Relax the impulse to one average true range and the count rises to 6,322; remove it altogether and the drift shape alone appears 58,996 times. Any frequency quoted for this pattern is a statement about the quoter's impulse threshold.
Does price reach the measured move target?
Less often than random entries did. The usual target projects the height of the impulse from the breakout. It was reached inside the twenty-bar holding period 19% of the time against 20% for matched random entries taking the same direction, the same distance and the same window. Extended to sixty bars the figures were 43% and 41%. The target gets reached often enough to be memorable and no more often than chance, and quoting it without a deadline makes the hit rate approach certainty for reasons that have nothing to do with the pattern.
What is the false break rate on a flag?
Counting a close back inside the projected boundary within 5 bars, 37% of breaks were false. On the control tape, where a real continuation effect existed, the rate was still 33%. Flags and pennants differed here, at 39.9% against 31.6%, which is one of only two measures on which the two names separated, and it follows from the geometry: a converging boundary is closer to price at the moment of the break, so the break happens sooner and holds more often.
Should I only take a flag in the direction of the larger trend?
This page said yes-it-makes-things-worse and has withdrawn that answer. The original split reported the with-trend group 0.56 percentage points below its matched base rate at 3.58 standard errors. The comparison group for that split was drawn from 250 bars on either side of each pattern, so half of it came from the very sixty-bar move the split was made on. Drawn only from bars after the pattern, the with-trend group reads 0.24 points at 1.51 and the against-trend group 0.36 at 2.41. The refinement is untested by this data rather than refuted by it. The sibling triangle page reported the same reversal and has withdrawn it for the same reason, and since both pages share one generator and one comparison harness, their original agreement was never independent evidence.
Why does it matter where the random comparison entries are drawn from?
Because on this page it decided the sign of the answer. The comparison entries were originally drawn from 250 bars on either side of each pattern, and a flag is defined by a thrust of 3.0 average true ranges, so the backward half of that window was full of the thrust. The proof is a placebo: an event set consisting of nothing but the impulse, with no drift and no break, reads 0.21 percentage points below that surrounding control at 4.42 standard errors, which is the impulse being scored against itself. Drawn only from bars after the pattern, the same placebo reads 0.06 points, and random bars with a coin-flip direction read 0.04, which is how the harness proves it is not biased. The forward-only control is not merely stricter: on the tape with a real effect planted in it, it finds that effect at 15.51 standard errors, more than any other window tried.
Why test this on simulated data rather than Indian stocks?
Because a generated tape can be built with a known answer inside it and real data cannot. Two tapes were used: one with no directional memory, where any measured edge has to be an artefact, and one with a genuine continuation effect written purely in return space, which acts as a control proving the detector can find an effect when one is there. Without that pairing, a null result is indistinguishable from a broken detector. It cannot tell you what flags do on any particular exchange, and nothing here should be read as saying it can.
Method note
How the numbers on this page were produced
Every figure comes from one deterministic simulation, seeded so that it reproduces identically on each run. Two tapes were generated, each of 600 independent instruments of 3,000 daily bars, 1,800,000 bars in total. Volatility clusters and drift regimes switch on both, so sharp thrusts and quiet pauses occur naturally and are never inserted by hand; the three drift states are symmetric and equally likely, so the unconditional drift of the tape is zero. The first tape contains no mechanism that could make direction predictable. The second adds one effect, defined entirely in return space and never in flag terms, and exists to prove the detector can find a continuation effect when one is present. The generator is the one used by the companion page on triangles, at the same seed, which was checked by generating its first 600 instruments with both pieces of code and comparing every value; a deliberately different seed was checked at the same time to confirm the comparison can fail.
Detection runs on the open, high, low and close of the generated bars using the thresholds in the definition table. Signals are taken on the close and positions opened at the next bar's open, so no result uses information that was not available at the time. Outcomes are measured over 20 bars. The base rate draws two hundred entries per event from the same instrument, matched for direction and holding period, from within 250 bars after each pattern and never from before it, because the bars before a flag contain the impulse that defines it. The window is varied four ways and all four are reported, and the harness is checked on entries carrying no pattern at all, where it reads zero within its own noise. Overlapping detections are discarded so that no two measured outcomes share a holding period. Returns are shown before costs; an illustrative allowance of twelve basis points per round trip would apply equally to both arms and would not change the difference between them, though it would sit well above the difference that was measured.
All results are illustrative and simulated. They are not a track record, they are not a forecast, and they are not an indication of what any pattern would produce in a live account or on any Indian security or index. The purpose is to establish which properties of a flag are consequences of its definition, which is a question about the definition rather than about any particular market.
Related