Educational Reference

Flags and Pennants: The Pattern Is the Impulse, Not the Shape

A flag is a brief, shallow drift against the direction of a sharp move. A pennant is the same drift with boundaries that close toward each other instead of running parallel. Almost every treatment spends its attention on the little shape and treats the move in front of it as scenery. This page treats the move as the subject, writes both halves as code strict enough to run, detects every occurrence across 1,800,000 generated daily bars, and reports what the detector found.

The finding, stated first. Hold the consolidation definition completely fixed and move only the impulse requirement, from one average true range to six. On a tape built with no directional memory the result does not move: 0.08 percentage points per extra unit of impulse, 1.51 standard errors from zero. Run the identical sweep on a tape that does contain a continuation effect and the same measurement climbs from 1.22% to 2.66%, a gradient of 0.49 points per unit at 9.15 standard errors. The test can see a sharper impulse paying off. Where nothing was there to be found, it found no gradient. A second finding was published here and has been withdrawn: this page reported that flags taken with the prevailing trend underperformed by 0.56 percentage points at 3.58 standard errors. Half of that comparison group was drawn from the impulse that defines the pattern. Rebuilt so it cannot be, the figure is 0.24 points at 1.51. All results on this page are illustrative and simulated.

The famous half is the smaller half

Open any description of a flag and count the words. The overwhelming majority go to the rectangle: how many bars it should run, how steeply it should lean, how tightly it should hold, whether the boundaries are parallel or converging and therefore whether the thing is a flag or a pennant. The sharp move that came first is usually granted a sentence. It is called the pole, or the flagpole, and treated as the thing the shape is attached to rather than the thing the shape is about.

That ordering is backwards, and it is possible to say so before running anything. Take away the sharp move and what remains is a few bars of quiet sideways drift. Quiet sideways drift is not rare. It is close to the default state of a price series, and a definition that finds it will find it constantly. Take away the drift instead and what remains is a strong directional thrust, which is genuinely uncommon and genuinely measurable. If either half of this pattern carries information, the prior should be heavily on the half that is scarce.

There is a second reason to suspect the shape, and it only becomes visible when you try to code the thing. Every clause that describes the drift turns out to be expressed as a fraction of the move in front of it. Shallow means shallow relative to the impulse. Brief means brief relative to how fast the impulse travelled. Narrow means narrow compared with the impulse height. The shape has no independent existence in the definition at all: it is a set of ratios whose denominator is the impulse. You cannot even state what a flag is without first stating what a flagpole is, and almost nobody states the second.

So this page reverses the usual emphasis. It defines the impulse with as much care as the drift, detects both, and then runs one experiment that the literature does not: it varies the impulse requirement alone, holding every other clause frozen, and asks what a stricter impulse actually bought. Where flags and pennants sit in the wider family of shapes is laid out in the guide to chart patterns in Indian stocks; the work here goes underneath that map and tests one entry on it.

Writing the impulse down is where it gets difficult

Here is the verbal definition, close to the standard one. A flag is a short, shallow consolidation that slopes against a sharp preceding move, resolving in the direction of that move. Try to run it. What is sharp? Sharp compared with what, measured over how many bars? What is short, and short in bars or in proportion to the move? What is shallow, and shallow measured from where? What counts as sloping against, given that a drift is never exactly counter to anything? And when do parallel boundaries become converging ones?

Seven questions, and the sentence answers none of them. That is not a complaint about sloppy writing. It is the reason two people can look at the same chart and disagree in good faith about whether there is a flag on it, and it is the reason published statistics about flags are difficult to compare with each other. Every one of those statistics belongs partly to whoever chose the missing numbers.

Two choices in the table below deserve to be flagged before the results arrive. The first is that the impulse is measured against the average true range from before the move began. Using the range during the move would let a violent thrust inflate its own yardstick, so that every sharp move scores about the same and the threshold stops discriminating. The second is that the drift is allowed to run for as long as it keeps qualifying, up to a cap, rather than being cut at a fixed length. Both are decisions nobody publishes, and both are measured for their consequences further down.

The coded definition. Every threshold is a choice, and every choice is stated. Illustrative simulated data throughout.
ClauseValue usedWhy this value, and what it costs
How sharp the impulse must beNet move of 3.0 average true ranges over 5 barsThe single most consequential number on this page, and the one no published definition supplies. The range is taken from before the move so the thrust cannot inflate its own yardstick
How clean the impulse must beNet move at least 0.55 of the distance actually travelledSeparates a thrust from a week of churn that happens to end higher. It turned out to be nearly inert: removing it entirely changed the population by under two percent
How brief the drift must beBetween 5 and 20 barsThe drift runs as long as it keeps qualifying, up to the cap. Cutting it at a fixed short length instead is measured further down and barely moved anything
How shallow the drift must beGives back no more than 50% of the impulseThe clause that makes it a pause rather than a reversal. Loosening it from a third to nearly two thirds roughly doubled the population
How narrow the drift must beWhole range under 60% of the impulse heightStops a wide, violent chop being read as a tidy consolidation. Note that this is a ratio: without an impulse it has no denominator and no meaning
That the drift is a pauseRuns on past the impulse by no more than 25% of itRejects the case where price simply kept going, which is not a pattern but the absence of one. This clause fired 49,390 times
That the drift leans backWith-impulse slope under 0.15 of the impulse heightAllows a flat or counter-sloping drift and rejects one still travelling with the move. Written as a limit rather than a requirement, because insisting on a counter-slope discards half the real cases
Flag or pennantBoundaries converge by 35% or moreDoes the whole job of naming, on its own. One cut-off separates the two words, and moving it inside a reasonable range renames one pattern in five
What counts as a breakClose beyond the boundary by 0.10 average true rangesA bare close beyond a line drawn through a narrow drift is noise. The buffer is small enough not to miss real exits
Entry and holdingNext bar open, held 20 barsEntering on the bar that produced the signal would quietly hand the test information from the future, which is the most common silent error in pattern studies
The impulse is the pattern. The drift is the smaller half. One detected instance on generated daily bars. Every clause below had to be given a number before this could be found at all. the impulse: 5 bars, 3.9 ATR the drift: 11 bars close beyond the boundary entry the next bar open WHAT WAS CHECKED Impulse, in ATR 3.9 Impulse efficiency 1.00 Bars in the drift 11 Deepest give-back 9% Drift slope −0.04 Boundary convergence 0% Named Flag The shaded block on the left is the whole of what this page tests. The block on the right is what the pattern is named after. Every clause describing the drift is expressed as a share of the impulse: how much of it was given back, how wide the drift is against it, how far past it price ran. The little shape cannot be defined without the move in front of it, which is the first result on this page.
One detected instance, drawn from the generated daily bars the detector actually ran on and chosen from the middle of the detected population rather than for its outcome: its impulse of 3.9 average true ranges sits beside a population median of 4.0. The gold block is the impulse and the green block is the drift, with the fitted boundaries carried forward to the break. Every clause describing the green block is a ratio whose denominator sits in the gold one. Illustrative simulated data.

How often a real one occurs

Across 1,800,000 daily bars, which is 600 independent generated instruments of 3,000 bars each and about 7,200 instrument years, the definition produced 5,501 candidate patterns. After discarding overlaps, so that no two measured outcomes share a holding period on the same instrument, 3,042 independent instances remained. That is roughly 0.42 per instrument per year, or about one every two and a half years on any given chart.

That is a strikingly small number, and it is small for one reason. The impulse clause is doing nearly all the filtering. Relax it from 3.0 average true ranges to one and the population rises to 6,322; tighten it to six and it falls to 308. Nothing else in the definition comes close to that leverage: loosening the shallowness requirement moved the count by about a factor of two, the narrowness requirement by about three, and the cleanliness requirement by almost nothing at all.

The rejection census tells the same story from the other side. The most common reason a qualifying impulse produced no pattern was that no drift of the required kind followed it at all, 70,013 times. Next came price simply carrying on rather than pausing, 49,390 times, then a give-back too deep to count as a pause, 21,221 times, then a drift too wide, 6,279 times, and finally 2,013 cases where the boundaries widened so much that neither name applied. These are counts of rejected evaluations rather than of distinct chart formations, but the ordering is the point: most sharp moves are not followed by anything a strict definition will call a flag.

The instances that did qualify were brisk. The mean drift ran 9.2 bars and the median 7, with the middle half falling between 6 and 11, and when the break came it came almost immediately, on average 1.6 bars after the drift stopped qualifying. Only 14 of 5,501 candidates failed to break at all inside the window allowed. The mean impulse behind a detected pattern was 4.3 average true ranges, comfortably above the 3.0 floor, because a threshold selects the tail of a distribution rather than a point on it. One more measured detail is worth having, because it quietly contradicts the picture in every textbook: the median drift sloped 0.9 percent of the impulse height, which is to say it did not lean against the move at all. It sat flat. The counter-trend drift that gives the pattern its name is the exception rather than the rule.

A flag, a pennant, and the tape they were found on All three panels are generated daily bars. The lower panel is one unbroken stretch with every detected instance marked. Flag: boundaries stay roughly parallel impulse 4.3 ATR, drift 11 bars boundaries widened 3% Pennant: boundaries converge impulse 3.8 ATR, drift 11 bars boundaries narrowed 51% One instrument, 800 consecutive bars, every detected instance marked 5 instances in this stretch. Shaded green where the impulse was upward, coral where it was downward.
A detected flag and a detected pennant, both from generated daily bars, with the impulse shaded gold and the fitted drift boundaries carried to the break. Both were chosen from the interquartile range of impulse strength and both broke without immediately reversing, so the drawn direction matches the label; neither was selected on its return. The lower panel is one unbroken stretch of the same tape with every detected instance marked. Read the sparseness of the lower panel as the result it is: a strict impulse requirement is why the pattern is rare. Illustrative simulated data.

Measuring what followed, against a base rate

A pattern statistic on its own says almost nothing. If a break is followed by a gain sixty percent of the time, the useful question is what fraction of random moments on the same data are followed by a gain, because the answer may also be sixty percent. Leaving that comparison out is usually what makes a pattern statistic look impressive.

The base rate here is built to be unfair to the pattern in every respect except the one under test. For each detected break, two hundred random entry points are drawn from the same instrument, within 250 bars of the event, taking a position in the same direction and holding it the same 20 bars. The only difference between the two arms is that one entered because a flag broke and the other entered for no reason at all. The canonical trade is measured, which means only breaks that went the same way as the impulse: a flag that broke backwards is a failed flag, not a short signal.

Those control entries are drawn only from bars that fall after each pattern. An earlier version of this page drew them from 250 bars on either side, which sounds more even-handed and is the single worst decision it made. A flag is defined by a thrust of 3.0 average true ranges. Put the control window around the pattern and the backward half of it contains that thrust, so the comparison group is partly made of the move the pattern was selected for. The next section is entirely about what that cost, because it cost this page a published finding and it moved every level on it.

What followed the break, against the base rate Twenty bars after entry, in the direction of the impulse. Controls are random entries on the same tape, drawn only from bars after the pattern. flag and pennant breaks matched random entries −15% −10% −5% 0% 5% 10% 15% Tape with no memory 3,042 patterns against 608,400 matched random entries difference +0.30 points, t = 2.79 −15% −10% −5% 0% 5% 10% 15% The same detector on a tape that does contain a continuation effect 3,292 patterns against 658,400 matched random entries difference +1.79 points, t = 15.51 The lower panel is the control. It proves the detector and the measurement can see a continuation effect when one is present, which is what makes the upper panel worth reading. Where the base rate is drawn from decides the answer The with-the-trend result, measured four times. Nothing changes but the bars the control entries may come from. −2.0% −1.0% 0.0% +1.0% +2.0% backward only 250 either side 1,000 either side forward only +0.69 t 6.00 +1.24 t 10.76 +1.69 t 14.64 +1.79 t 15.51 −1.45 t −9.19 −0.56 t −3.58 +0.04 t 0.27 +0.24 t 1.51 the withdrawn with-the-trend claim, on the tape with no memory a known effect, measured the same four ways
Both arms of the test, on two different tapes, with the comparison group drawn only from bars after each pattern. In the upper panel nothing in the data makes direction predictable and the two distributions sit almost on top of each other. In the lower panel a genuine continuation effect was inserted, and the detector finds it at a size no reading of the upper panel approaches. The bottom panel slides the control window from behind each pattern to in front of it. The withdrawn with-the-trend claim walks from large and negative to nothing while a known effect does not move at all. Illustrative simulated data.

On the tape with no memory the two distributions are close to the same distribution. The pattern arm returned 0.29% over twenty bars and the matched random arm −0.01%, putting the pattern 0.30 percentage points ahead, which is 2.79 standard errors from zero, with a ninety-five percent interval running from 0.09 points to 0.52. The share of positive outcomes was 52.4% against 50.0%. The tails are close: a tenth of pattern outcomes were worse than 6.6 percent down against 7.3 for the controls, and a tenth better than 7.2 up against 7.3.

That number is not zero and this page is not going to pretend it is. Under the older, symmetric control the same arm read 0.15 points behind at −1.37 standard errors. Moving the control window forward moved it 0.45 points, from a negative that was not significant to a positive that is, on a tape that is supposed to contain nothing. Both readings cannot be right, and as it turns out neither is: what the two of them bracket is a comparison group that is contaminated in one direction and a horizon mismatch in the other. That is unpicked in the next section, with a placebo ladder, because a page that only reports the correction it likes is not doing the thing it claims to do.

Measuring the break in whichever direction it went, which is the arm the companion page on triangles reports and therefore the comparable one, gives 5,365 instances and a gap of 0.16 percentage points, 1.89 standard errors. Both arms say the same thing, and both moved the same way when the control window did.

A null result from a detector that finds nothing is worth nothing, because the detector might simply be broken. So the identical detector, the identical base-rate machinery and the identical thresholds were run over a second tape, built the same way but with one genuine effect inserted: whenever the ten-bar return was large relative to its own recent scale, the following thirty bars carried extra drift in that direction. The effect is written purely in return space. It says nothing about a consolidation, nothing about shallow drift, nothing about boundaries and nothing about a break, so the flag detector had to find it unaided.

It did, emphatically. On that tape the pattern arm averaged 1.89% against 0.10% for matched random entries, a difference of 1.79% at 15.51 standard errors, with 62.2% of outcomes positive against 50.5%. That is the power check, and it is the number that gives the forward-only control the right to report a null anywhere else on this page: it is not a comparison group that reports nothing for everything. It is also, of the four control windows tried, the one that finds the planted effect most clearly. Strictness cost no power at all here.

Results by variant, each against its own direction-matched, holding-matched, time-local base rate. Both tapes are generated. All figures illustrative and simulated, not a track record and not a forecast.
VariantFoundBroke with the impulseMean, 20 barsMatched base rateTarget reachedFalse breaks
Flag1,93356%0.40%0.01%21% vs 21%40%
Pennant1,10956%0.11%−0.04%15% vs 18%32%
Both, weakest impulse band1,29754%0.00%0.04%37% vs 44%42%
Both, strongest impulse band29360%0.04%0.01%9% vs 11%37%
Drift shape alone, no impulse58,996not defined0.10%0.01%27% vs 27%46%
All, control tape3,29261%1.89%0.10%25% vs 20%33%

Two rows deserve a second look. The measured move, which projects the impulse height from the break and is the standard target for this pattern, was reached inside the holding period 19% of the time against 20% for matched random entries; over sixty bars, 43% against 41%. It is reached often enough to be memorable and slightly less often than chance. And the false break rate, counting a close back inside the boundary within 5 bars, was 37% on the tape where nothing was happening and 33% on the tape where something was. Roughly one break in three comes straight back regardless, which is what a boundary drawn through a narrow drift does. Why obvious levels attract and then reject price is the subject of the guide to breakouts.

The comparison group, taken apart

This page published a context finding that is now withdrawn, and it published every level on the page against a comparison group that was contaminated. The withdrawal comes first, then the evidence, then the part that is still unresolved, because leaving the last one out would be the same mistake in a nicer suit.

The finding, as published. Split the instances by whether the impulse ran the same way as the preceding sixty-bar move, the standard trade-with-the-trend refinement. Measured against a comparison group drawn from 250 bars on either side, the agreeing group came in 0.56 percentage points below its own base rate at 3.58 standard errors across 1,513 instances, while the disagreeing group came in 0.22 points above. The page read that as the refinement making things worse, and cited the sibling triangle study as independent agreement.

The finding, corrected. Against a comparison group drawn only from bars after each pattern, the agreeing group reads 0.24 points at 1.51 standard errors and the disagreeing group 0.36 at 2.41. The gap between them is gone, and what is left is the two cells sitting in the same place. The refinement is not supported here and it is not contradicted here. It is untested.

Look at which arm moved. The two pattern arms never budged: 0.29% with the trend and 0.30% against it, which is the same number twice, because moving a control window cannot change what the pattern returned. The two comparison groups moved a great deal: the with-trend control fell from 0.85% to 0.05%, and the against-trend control from 0.08% to −0.06%. The entire context effect was in the control, and specifically in the half of the control window that sat on top of the sixty-bar move being split on.

Three predictions follow from that, and they can be checked separately. If the contamination is in the backward half, then restricting the control to the backward half alone should exaggerate the effect, widening the window should dilute it, and pushing it fully forward should remove it. Backward only: 1.45 points at 9.19 standard errors, more than two and a half times the published figure. A thousand bars either side: 0.04 points at 0.27. Forward only: 0.24 at 1.51. Monotone, in the predicted order, exactly as a control drawn from the pattern's own bars would behave.

The placebo that settles it. A window argument can be waved away as a judgement call, so here is a measurement that cannot. Build an event set that is nothing but the impulse: every bar where the 3.0 average true range thrust fires, entered the next bar in its own direction, with no drift clause, no channel, no break and no flag. Measure that against the surrounding control the page used to use. It comes in 0.21 percentage points below its own base rate at 4.42 standard errors. A comparison group that shows the impulse significantly underperforming itself is not a strict control, it is a circular one. Against the forward-only control the same placebo reads 0.06 points at 1.26, and across six independently generated tapes it averages −0.01 points at −0.17 standard errors, which is zero to as many decimal places as this exercise can produce.

The full ladder, every rung measured against the same forward-only control on the same tape, reads: random bars with a coin-flip direction, 0.04 points; random bars with the direction set by the prior five-bar move, 0.06; the impulse alone, 0.06; the detected flags, 0.30. The first rung is the harness testing itself on entries that carry no pattern whatsoever, and it reads zero, which is the only reason the rest of the ladder means anything.

What is still unresolved, stated plainly. The forward-only control is the right control and it is not a neutral one. Across six independently generated tapes the flag arm reads 0.19 points above it at 3.01 standard errors, where the same six tapes read 0.26 points below the symmetric control at −3.85. Both controls are biased and they are biased in opposite directions. The symmetric one is understood. The forward-only one is not, and three things are known about it. It is not the arithmetic of quoting returns as ratios, because the same excess appears unchanged in log space, 0.31 points against 0.30. It is not a long-short imbalance, because the event set is 48.7% long. And it is at least partly the generator: switch the drift off in all three regime states, so that the log price becomes an exact martingale and nothing measurable from the past can predict anything, and the same six tapes give 0.11 points at 2.44 standard errors against 0.19 at 3.01 with the drift on.

The honest reading of that last pair is that roughly half the residual is the tape and the remaining half is unaccounted for at six tapes of resolution, which is not enough resolution to say more. What it means for everything else on this page is a bound rather than a point: on data built to contain nothing, the flag arm measures somewhere between zero and a fifth of a percentage point over twenty bars, against a control that is itself uncertain at that scale, before any cost. Twelve basis points of round-trip cost sits inside that band. The page's conclusions are stated at that resolution from here on, and the one conclusion that does not depend on it at all is the next section, because a gradient does not care what constant the control adds to every point on it.

The central experiment: only the impulse threshold moves

Everything so far describes one setting of one definition. The experiment this page exists for is narrower and more useful. Freeze the consolidation clauses entirely, change nothing about how shallow, how brief, how narrow or how sloped the drift must be, and move only the number that decides how sharp the move in front of it has to be. If the pattern carries content anywhere, this is where it should show up, because this is the clause that is supposed to separate a real flag from a coincidence.

The central experiment: only the impulse threshold moves The consolidation definition is held completely fixed. The one number that changes is how sharp the move in front of it must be. tape WITH a continuation effect tape with no memory excess return over the matched base rate −1.0% −0.5% 0.0% +0.5% +1.0% +1.5% +2.0% +2.5% +3.0% +3.5% 1.0 1.5 2.0 2.5 3.0 3.5 4.0 5.0 6.0 how sharp the impulse must be, in ATR over 5 bars patterns found, tape with no memory 6,322 5,935 5,089 4,055 3,042 2,196 1,535 723 308 1 to 6 ATR +1.2% to +2.7% +0.2% to 0.0% The reading On the tape that contains an effect, a sharper impulse raises the measured excess from +1.2% to +2.7%. With no memory the same sweep does nothing: +0.08 points per extra ATR across disjoint bands, 1.51 standard errors from zero.
The same generated tapes throughout; the only thing that changes between points is the impulse threshold. Bars show how many patterns survive each setting on the tape with no memory. The green line is the control: a sharper impulse pays on a tape that has something to pay. The coral line is the same measurement where there is nothing, and whatever height it sits at, it does not tilt. Illustrative simulated data.

On the tape with no memory a stricter impulse bought nothing, and the way to see that is to read the line rather than any single point on it. Every point on the coral line now sits at or a little above zero rather than below it, lifted by the residual the previous section could not fully account for. None is further from zero than 0.44 percentage points and none is beyond 3.58 standard errors, while the population falls from 6,322 instances to 308. What matters here is that the line is level: trading one twentieth of the available opportunities produced the same thing, whatever that thing is, as trading all of them. Under the older symmetric control every point on this line sat below zero instead, and it sagged at the strict end, because a bigger impulse contaminates a backward-looking control more. The sag was the control, not the pattern. The level was never the finding on this page. The gradient was.

Nested thresholds share their events, though, so the differences between neighbouring points on that line are not independent of each other. The cleaner version detects once at the loosest setting and splits the resulting population into non-overlapping bands by how strong each impulse actually was, which gives independent cells and a trend that can carry a standard error. On the tape with no memory that trend is 0.08 percentage points per additional unit of impulse, with a standard error of 0.05, or 1.51 standard errors from zero. Flat.

The obvious objection to a flat line is that the measurement might be incapable of producing anything else. That is what the second tape is for, and here it earns its place twice over. Running the identical sweep on the tape that does contain a continuation effect produces a line that climbs the whole way, from 1.22% at the loosest impulse to 2.66% at the strictest. In the disjoint bands the trend is 0.49 percentage points per unit at 9.15 standard errors, and the band-by-band figures climb from 0.37% in the weakest to 2.60% in the strongest, a sevenfold spread across cells that share no events.

That is the whole argument in one comparison. The experiment is demonstrably capable of detecting a payoff to impulse strength, at a sample size and an effect size of the same order as the null case, because it detected one on demand. When the same experiment is run where no such payoff exists, it reports none. The flat line is a measurement, not a failure to measure.

One thing did rise with impulse strength on both tapes, and it is worth dwelling on because it is the statistic most often quoted as proof. The share of breaks that went in the direction of the impulse, the continuation rate, climbed from 53.7% in the weakest band to 60.0% in the strongest, on the tape with no memory. Overall it stood at 56% of 5,487 breaks. A trader running this filter on real data would observe exactly that: demand a sharper pole and a larger majority of your flags resolve the right way. The observation is real, it is reproducible, and it happens where nothing whatsoever is going on.

The cause is geometric and worth understanding because it generalises. The drift is defined to lean against the impulse or to sit flat, which means the boundary on the impulse side of the drift is the one price is closest to and the one it reaches first. The bigger the impulse, the more sharply the definition constrains the drift to hug that side, and the more lopsided the outcome becomes. A continuation rate is therefore a property of how you defined the drift, and any figure quoted for it has to be netted against a geometric baseline before it can be read as evidence about markets.

Flag against pennant

The two names describe one difference: whether the drift's boundaries stay roughly parallel or close toward one another. Under this definition they were separated by a single cut-off, and the population split 1,933 flags to 1,109 pennants. That the difference is a cut-off rather than a category is easy to demonstrate: move it and the names move with it.

Flag against pennant, and the tolerance that separates them Two names for one drift. The upper panel compares them measure by measure; the lower panel moves the cut-off that assigns the name. 0% 25% 50% 75% Broke with the impulse difference 0.4 points, 0.30 standard errors 56.2% 55.7% Reached the measured move difference 6.7 points, 4.75 standard errors 21.4% 14.7% Came straight back inside difference 8.3 points, 4.66 standard errors 39.9% 31.6% flag, 1,933 instances pennant, 1,109 instances The same 3,042 patterns, renamed by moving one cut-off converged 15% 1,281 1,761 converged 25% 1,611 1,431 converged 35% 1,933 1,109 converged 45% 2,246 796 converged 55% 2,532 510 How many kept the same name Across the middle three cut-offs only 2,407 of 3,042 patterns, 79%, held one label. Nothing about the patterns changed; the cut-off moved.
The upper panel compares the two names measure by measure, with the difference in standard errors printed beside each. The lower panel takes one fixed population of detected patterns and relabels it five times by moving only the convergence cut-off. Two of the four measures separate the names; the two that matter for anybody trading them do not. Illustrative simulated data.

On the measures that decide whether the distinction is worth carrying, the two names are indistinguishable. The continuation rate was 56.2% for flags against 55.7% for pennants, a gap of 0.30 standard errors. The twenty-bar outcome differed by 0.29 percentage points at 1.27 standard errors, and both sat on the same side of their own base rates, flags by 0.39 points and pennants by 0.15, a difference between the two names well inside the residual the comparison group carries. Whatever a pennant is, it is not a flag that resolves better.

On two other measures they separated clearly, and the reason is mechanical rather than predictive. Flags reached the measured move 21.4% of the time against 14.7% for pennants, a difference of 4.75 standard errors, and flags produced false breaks 39.9% of the time against 31.6%, 4.66 standard errors apart. Both follow from the same fact. A converging drift has its boundaries closest together at the moment of the break, so the break happens from a tighter position and is less likely to be immediately undone, and the drift itself ran longer, 10.1 bars against 8.7, leaving less of the holding period for a distant target to be reached. Neither difference is about what the market intends to do next.

The naming itself is unstable in the way this kind of tolerance usually is, though less dramatically than one might expect. Holding the detected population completely fixed and moving the convergence cut-off across the middle of its plausible range, 2,407 of 3,042 patterns, 79%, kept the name they started with. The reason it is not worse is instructive: there are only two bins here and one cut-off, whereas the companion page on triangles had three bins whose boundaries were set by a single tolerance doing two unrelated jobs at once, and there only fifty-eight percent of instances held their label. Fewer names, less reshuffling. It remains true that one pattern in five changes identity for no reason connected to the chart.

Worth noting in passing: 791 of the 3,042 detected drifts had boundaries that widened rather than narrowed, and the median convergence across the whole population was 22%. The parallel channel that illustrations show is not the typical case; it is the middle of a continuum with converging drifts on one side and quietly broadening ones on the other.

What the impulse clause is actually worth

If the impulse is the subject, the fair question is what happens when it is removed. That cannot be done by setting its threshold to zero, because every clause describing the drift is a ratio with the impulse height underneath it, and a zero impulse leaves those clauses with nothing to be a fraction of. The workaround is to re-anchor the same three clauses to average true range, so shallow and narrow and brief keep their meaning while the requirement for a preceding thrust disappears entirely.

The shape on its own, and the shape after an impulse The same drift clauses re-anchored to average true range so the impulse requirement can be switched off entirely. How often it fires, per instrument per year Drift shape alone, no impulse required 8.19 After an impulse of 3.0 ATR 0.42 58,996 occurrences against 3,042. Strictness is the whole filter. Excess over the matched base rate 0 Shape alone +0.40% t +14.1 +0.09% t +3.3 Shape after an impulse +1.79% t +15.5 +0.30% t +2.8 tape with an effect tape with no memory What the impulse clause is worth On the tape with no memory both detectors sit just above their base rates, by +0.09% and +0.30%, a residual the placebo section takes apart. On the tape that does contain a continuation effect the gap is an order larger: +0.40% for the bare shape against +1.79% behind an impulse. Requiring the move in front turns a shape that fires 8.2 times a year into one that fires 0.42, and it is the only clause on this page that changed a result.
The same drift clauses, re-expressed against average true range so the impulse requirement can be switched off. Left, how often each fires. Right, what each measured on both tapes. Where there is nothing, both sit in the same small band; where there is something, the bare shape recovers a fraction of what the shape behind an impulse recovers. Illustrative simulated data.

The bare shape is common, which was the prediction. It appeared 58,996 times against 3,042 for the full definition, about 8.2 times per instrument per year rather than 0.42. Anyone scrolling a chart and seeing quiet drifts everywhere is seeing something real; they are just not seeing flags, because the flag is the impulse.

On the tape with no memory the bare shape sits 0.09% above its matched base rate, inside the same unresolved residual as everything else measured on that tape and a quarter the size of the full definition's 0.30%. Two detectors, one nineteen times more selective than the other, landing in the same small band on a tape that should contain nothing. On the tape that did contain a continuation effect they part company decisively: the bare shape came in 0.40% above its matched base rate at 14.08 standard errors, while the same shape behind a real impulse came in 1.79% above at 15.51. The gap between those two is more than four times anything either detector produced where there was nothing to find, which is the comparison the conclusion rests on.

That is the clearest statement this page can make about where the content sits. When there was something to find, requiring an impulse in front of the shape multiplied the measured excess by about four, at the cost of discarding nineteen occurrences in twenty. The shape contributed something, because a quiet drift often does follow a thrust and the bare detector picked up some of the same events by accident, but it was the minority contribution. The impulse clause is the pattern. The rectangle is how you notice it.

The counterpart of that result is the one already reported: on the tape with no memory, making the impulse requirement stricter did not help either. Presence and severity are different questions, and the answers here differ. Requiring an impulse changes what the detector is looking at. Requiring a bigger one changes only how many instances survive, and on data with nothing in it, it changes nothing else.

How much of this is the definition rather than the data

Every number above rests on ten thresholds, and a result that survives only at one setting is a result about that setting. The consolidation clauses were varied one at a time with the impulse held fixed. Loosening the shallowness requirement from a third of the impulse to nearly two thirds moved the population from 1,674 to 3,288; loosening the narrowness requirement moved it from 1,086 to 3,584; changing the cap on drift length between twelve and thirty bars moved it from 3,093 to 2,997, which is barely at all. Against the forward-only comparison group every one of those settings sits within 0.37 percentage points of zero, essentially the same spread as the 0.35 the symmetric version gave, and the ordering of the clauses by leverage is unchanged. The control window moved all of these levels together by roughly one constant, which is exactly why a sensitivity study built on differences survives the change of control and a claim about the level does not.

The window rule is worth a paragraph on its own, because it is where this page and its companion diverge. The triangle study found that its unstated rule about how far back to look dominated everything else: switching from the shortest qualifying window to a single fixed one cut its population from 8,036 to 711, a collapse of more than nine in ten. The same substitution here does almost nothing. Cutting the drift at the shortest qualifying length rather than letting it run produced 3,107 instances against 3,042, with a mean drift of 5.1 bars against 9.2.

The reason is worth stating because it is the one structural advantage this pattern has. A triangle has no anchor: its start is wherever the analyst decides to begin looking, so the window rule creates the pattern. A flag has an anchor, and it is the impulse. The drift starts on the bar after the thrust ends, and no choice about lookback can move that. Demanding an impulse buys the definition something real, which is a start date that is not a matter of opinion, even though it buys nothing measurable in outcomes on this data.

The two pages once agreed on a substantive finding and it has been withdrawn from both, which is worth spelling out because the agreement was the reason it survived as long as it did. Both reported that taking the pattern with the prevailing trend made things worse. Both were running the same comparison machinery on the same generator at the same seed, so their agreement was never independent of each other. Two measurements sharing the instrument that produced them will agree whether or not the claim is true. Before citing one of these pages as corroborating another, the question to ask is not whether they agree but whether anything about them could have disagreed.

What this supports and what it does not

A simulation can establish some things and not others, and being exact about which is which is the difference between a study and an opinion with numbers attached.

The boundary of what this exercise establishes. Illustrative simulated data.
ClaimDoes this page support it?
The drift cannot be defined without the impulseYes, and structurally rather than statistically. Every clause describing the shape is a ratio whose denominator is the impulse height, which is why the impulse requirement cannot simply be set to zero
The impulse clause does the filteringYes. It takes the population from 58,996 to 3,042 on identical data, and on the tape containing a real effect it multiplied what was recovered by about four
A stricter impulse improves the outcomeNot on data with no memory: the trend was 0.08 points per unit at 1.51 standard errors. On the control tape the same sweep gave 0.49 points per unit at 9.15, so the test had the power and simply found nothing to report
A rising continuation rate is evidence of continuationNo. It rose with impulse strength on both tapes, including the one with nothing in it, because the drift is defined to lean against the move. It is a property of the definition
Flags and pennants are two different patternsNo. They differ by 0.30 standard errors in continuation and 1.27 in outcome, and one in five changes name when the cut-off moves. The two measures that do separate them follow from the geometry
Taking a flag with the prevailing trend does worse than randomNo, and this page said otherwise until the comparison group was rebuilt. Against a control drawn only from bars after the pattern the figure is 0.24 points at 1.51 standard errors. Untested here, not refuted
A time-local comparison group is a safe control for a chart patternNo, and that is the transferable result. An event set consisting of nothing but the impulse reads 0.21 points below a surrounding control at 4.42 standard errors, which is the impulse being measured against itself
The forward-only comparison group is neutralNot entirely, and the page reports the residual rather than hiding it. Across six independent tapes the flag arm sits 0.19 points above it at 3.01 standard errors. About half of that is the generator's own drift regimes; the rest is unaccounted for at this resolution
Roughly one break in three comes straight backYes, under this definition, on both tapes. It is the most robust number on the page and the largest practical cost of trading the shape
Flags carry no edge in Indian equitiesNo. This tested generated data, not any exchange. It shows what the pattern does when nothing is behind it, which is the baseline a real study needs before it can claim anything
Any of this is a reason to take a positionNo. Nothing here is a trade trigger, a recommendation or a forecast, and the outcome distributions are wide enough that no single instance is predictable

The largest limitation is the obvious one. Generated data has no earnings, no policy announcements, no index rebalancing and no order books, so it cannot say what happens when a real thrust is driven by a real event and a real pause follows it. What it can do, and what real data cannot, is supply a case where the true answer is known in advance. When a statistic reproduces on a tape with nothing in it, that statistic is not evidence, and sorting the familiar flag numbers into that category is the contribution here.

A second limitation is that this tests one definition. A different set of ten numbers is a different detector, which is exactly why the sensitivity work above exists. Anybody who wants to argue that a better definition would give a different answer is making a testable claim, and the way to settle it is to write that definition as code and run it.

What to do with a flag instead

None of this makes a sharp move followed by a pause uninteresting. It relocates the interesting part, and it changes what is worth measuring.

Measure the impulse, not the rectangle. The one clause on this page that changed a result was the requirement for a real thrust in front of the shape. A thrust is measurable directly, in units of recent range, without drawing a single line. If the reason a chart has your attention is that price just moved hard and then stopped, then measure the move and the stopping, because both measurements are unambiguous and the drawing is not.

Do not read a continuation rate as continuation. The share of flags resolving with the trend rose with impulse strength on a tape containing no trend information at all. That is the single most quotable statistic about this pattern and it is manufactured by the definition. Before believing any version of it, ask what the same number would be on data with nothing in it.

Stop paying attention to which of the two names applies. Flags and pennants were indistinguishable in outcome and in continuation here, and one instance in five swapped names when a cut-off moved. Time spent deciding whether the boundaries are parallel enough is time spent on a filing decision.

Budget for the false break before the target. About one break in three closed straight back inside the boundary, on both tapes, and the measured move was reached slightly less often than for random entries with the same target and window. Any plan built on this shape needs the reversal as its normal case rather than its exception.

Write the definition down before you look for instances. This is the habit that transfers. Forced to say how sharp is sharp, and against what yardstick, and measured over how many bars, you discover how much of your pattern recognition was a decision rather than an observation. Doing that once is worth more than reading ten descriptions, and if that way of working appeals more than the taxonomy it replaces, it is the method we teach.

FAQ

Frequently asked questions

Only whether the two boundaries of the drift stay roughly parallel or close toward each other. Both sit behind the same thing: a sharp move followed by a brief, shallow pause. On the test run here the two were separated by a single cut-off, and moving that cut-off across a reasonable range left 79% of 3,042 patterns with the name they started with. They also behaved the same way afterwards, differing by 1.27 standard errors in outcome and 0.30 standard errors in how often the break went with the preceding move.

No published definition supplies a number, which is the problem this page starts from. The detector used here required a net move of 3.0 times the average true range over 5 bars, measured with the range from before the move so the thrust cannot inflate its own yardstick. That single number is the most consequential choice in the whole definition: loosening it to one average true range multiplied the population by roughly two, and tightening it to six divided it by ten.

On the tape with no memory in it, no. Sweeping the impulse requirement from one average true range to six, with the consolidation definition held completely fixed, moved the result by 0.08 percentage points per extra unit of impulse, in the unhelpful direction, which is 1.51 standard errors from zero. On the tape that did contain a genuine continuation effect the same sweep raised the excess from 1.22% to 2.66%, a gradient of 0.49 points per unit at 9.15 standard errors. The experiment can see an impulse gradient when one exists. On the tape where none existed it saw none.

Much less than the move in front of it. Re-anchoring the same drift clauses to average true range, so the impulse requirement can be removed entirely, the shape alone was found 58,996 times against 3,042 for the full definition, about 8.2 times per instrument per year rather than 0.42. On the tape carrying a real continuation effect the bare shape came in 0.40% above its matched base rate while the same shape behind an impulse came in 1.79% above. The shape is common, the impulse is rare, and the filtering is done almost entirely by the impulse.

In this test 56% of 5,487 breaks went the same way as the impulse, and the rate climbed with impulse strength, from 53.7% in the weakest band to 60.0% in the strongest. That happened on a tape built with no directional memory whatsoever, so none of it is evidence of continuation. It is what happens when a drift is defined to lean against a move: the boundary on the side of the move sits closer to price and gets reached first. A continuation statistic has to be netted against that before it means anything.

Under this definition, rare. Across 1,800,000 daily bars the detector produced 5,501 candidates and 3,042 independent tradable instances, about 0.42 per instrument per year, or roughly one every two and a half years on any given chart. That is a direct consequence of demanding a real impulse. Relax the impulse to one average true range and the count rises to 6,322; remove it altogether and the drift shape alone appears 58,996 times. Any frequency quoted for this pattern is a statement about the quoter's impulse threshold.

Less often than random entries did. The usual target projects the height of the impulse from the breakout. It was reached inside the twenty-bar holding period 19% of the time against 20% for matched random entries taking the same direction, the same distance and the same window. Extended to sixty bars the figures were 43% and 41%. The target gets reached often enough to be memorable and no more often than chance, and quoting it without a deadline makes the hit rate approach certainty for reasons that have nothing to do with the pattern.

Counting a close back inside the projected boundary within 5 bars, 37% of breaks were false. On the control tape, where a real continuation effect existed, the rate was still 33%. Flags and pennants differed here, at 39.9% against 31.6%, which is one of only two measures on which the two names separated, and it follows from the geometry: a converging boundary is closer to price at the moment of the break, so the break happens sooner and holds more often.

This page said yes-it-makes-things-worse and has withdrawn that answer. The original split reported the with-trend group 0.56 percentage points below its matched base rate at 3.58 standard errors. The comparison group for that split was drawn from 250 bars on either side of each pattern, so half of it came from the very sixty-bar move the split was made on. Drawn only from bars after the pattern, the with-trend group reads 0.24 points at 1.51 and the against-trend group 0.36 at 2.41. The refinement is untested by this data rather than refuted by it. The sibling triangle page reported the same reversal and has withdrawn it for the same reason, and since both pages share one generator and one comparison harness, their original agreement was never independent evidence.

Because on this page it decided the sign of the answer. The comparison entries were originally drawn from 250 bars on either side of each pattern, and a flag is defined by a thrust of 3.0 average true ranges, so the backward half of that window was full of the thrust. The proof is a placebo: an event set consisting of nothing but the impulse, with no drift and no break, reads 0.21 percentage points below that surrounding control at 4.42 standard errors, which is the impulse being scored against itself. Drawn only from bars after the pattern, the same placebo reads 0.06 points, and random bars with a coin-flip direction read 0.04, which is how the harness proves it is not biased. The forward-only control is not merely stricter: on the tape with a real effect planted in it, it finds that effect at 15.51 standard errors, more than any other window tried.

Because a generated tape can be built with a known answer inside it and real data cannot. Two tapes were used: one with no directional memory, where any measured edge has to be an artefact, and one with a genuine continuation effect written purely in return space, which acts as a control proving the detector can find an effect when one is there. Without that pairing, a null result is indistinguishable from a broken detector. It cannot tell you what flags do on any particular exchange, and nothing here should be read as saying it can.

Method note

How the numbers on this page were produced

Every figure comes from one deterministic simulation, seeded so that it reproduces identically on each run. Two tapes were generated, each of 600 independent instruments of 3,000 daily bars, 1,800,000 bars in total. Volatility clusters and drift regimes switch on both, so sharp thrusts and quiet pauses occur naturally and are never inserted by hand; the three drift states are symmetric and equally likely, so the unconditional drift of the tape is zero. The first tape contains no mechanism that could make direction predictable. The second adds one effect, defined entirely in return space and never in flag terms, and exists to prove the detector can find a continuation effect when one is present. The generator is the one used by the companion page on triangles, at the same seed, which was checked by generating its first 600 instruments with both pieces of code and comparing every value; a deliberately different seed was checked at the same time to confirm the comparison can fail.

Detection runs on the open, high, low and close of the generated bars using the thresholds in the definition table. Signals are taken on the close and positions opened at the next bar's open, so no result uses information that was not available at the time. Outcomes are measured over 20 bars. The base rate draws two hundred entries per event from the same instrument, matched for direction and holding period, from within 250 bars after each pattern and never from before it, because the bars before a flag contain the impulse that defines it. The window is varied four ways and all four are reported, and the harness is checked on entries carrying no pattern at all, where it reads zero within its own noise. Overlapping detections are discarded so that no two measured outcomes share a holding period. Returns are shown before costs; an illustrative allowance of twelve basis points per round trip would apply equally to both arms and would not change the difference between them, though it would sit well above the difference that was measured.

All results are illustrative and simulated. They are not a track record, they are not a forecast, and they are not an indication of what any pattern would produce in a live account or on any Indian security or index. The purpose is to establish which properties of a flag are consequences of its definition, which is a question about the definition rather than about any particular market.

Related

Continue reading

Next step

Find your starting stage. Everything else follows from there.

Educational reference only. No buy, sell or hold recommendations. All results shown are illustrative and simulated.