Educational Reference

Shooting Star and Inverted Hammer: One Bar, Two Names, Two Stories

These two patterns are the same bar. A small body at the bottom of the range, a long upper shadow, little or nothing below. Put a ruler on the candle and there is no measurement that separates them, because there is nothing to separate. The only difference is what came before it, so the name is assigned entirely by prior context. That makes this pair the cleanest test available of the claim that context, rather than shape, carries whatever information a candle contains. This page does not argue the point. It writes one definition as code, finds every instance across 1,500,000 generated daily bars, sorts them by the trend that preceded them, and measures whether the two piles behave differently.

The finding, stated first. Across 4,252 bars that qualified after an advance and 3,921 that qualified after a decline, the two groups differed by 0.12 percentage points over the next ten bars, with a ninety-five percent interval running from 0.20 points below zero to 0.44 above. The same measurement applied to a tape with a real effect planted in it separated the two groups by 2.42 points, so the instrument was not blind. All results on this page are illustrative and simulated.

Two names for a bar that only has one shape

Most candle patterns are at least arguably distinct objects. An engulfing bar and a doji look different from each other, and whatever you think of what they mean, you can tell them apart with your eyes shut. The upper-shadow family is not like that. A shooting star and an inverted hammer are pixel for pixel the same candle, and every published description of them agrees on this. The difference is not in the drawing, it is in the caption.

The lower-shadow half of this family is covered in full in our guide to the hammer candlestick, which owns the hammer and its identical twin the hanging man, along with the anatomy conventions, the confirmation discipline, the volume reading and the position-sizing arithmetic that a long shadow forces. This page is the mirror image and it is deliberately a different kind of page: not a description of what the upper-shadow bar records, but a measurement of whether the two names it can receive describe anything different. The wider taxonomy that both pages sit inside is set out in the guide to candlestick charts.

Here is why this particular pair is worth the trouble of measuring. In almost every other pattern argument, shape and context are tangled together. If a breakout works better in a trend, you cannot easily tell whether the trend improved the breakout or the trend was doing all the work by itself. The upper-shadow family gives you a rare clean experiment, because the shape is held exactly constant by construction. Two piles of bars, identical in every measurable respect, sorted only by what preceded them. If the two piles behave the same, the naming convention is describing nothing at all. If they behave differently, the difference has to be the context, because there is nothing else left.

That is a real question with a real answer, and it is answerable without any argument about whether candles work. It also has a second, less comfortable form. Suppose the two piles do behave the same. There is still the question of whether the bar itself does anything, pooled across every context, compared with an ordinary bar picked at random nearby. Most published discussion skips straight past that question, because it is busy distinguishing the two names. Both questions are answered below, and one of them produced a small result that complicates the tidy story.

Writing down the one definition

There is only one bar, so there is only one definition to code. That is worth saying plainly, because the literature presents two patterns and the instinct is to write two detectors. Two detectors would be two copies of the same five clauses with a different label at the end, and writing them separately is exactly the sort of thing that hides the point.

The verbal definition is short: a small real body at the lower end of the range, an upper shadow at least twice the height of that body, and little or no lower shadow. Every clause in that sentence stops the code from compiling until a number is supplied, and none of the numbers are in the sentence. How small is small? Small relative to what? What counts as the lower end of the range, and measured how, from the body's top or its bottom? How little is little? And what stops a bar so tiny that its shadow is noise from qualifying on a technicality?

Five thresholds answer those questions, and all five are choices.

The single coded definition. There is one bar, so there is one definition. Every threshold is a choice, and every choice is stated. Illustrative simulated data throughout.
ClauseValue usedWhat it does, and what it costs
Upper shadowAt least 2.0 times the real bodyThe convention in common use. On its own it passes 19.2% of bars. Sweeping it from 1.0 to 4.0 moved the population from 25,842 to 11,065, which is measured further down
Body positionTop of the body inside the lower 0.40 of the rangeThe clause that actually does the work, and the one almost never stated numerically. It passes only 7.8% of bars on its own, the strictest of the five
Lower shadowNo more than 0.25 of the upper shadowWhat makes the bar one-sided rather than merely volatile. Relaxing it to 1.00 took the population from 23,436 to 32,372 on identical data
Body floorReal body at least 0.05 of the rangeBelow this the bar is a gravestone doji, a different object with its own family. Excluding it keeps this page from quietly annexing the doji
Bar sizeRange at least 0.60 of the 20-bar average true rangeStops a shrunken bar qualifying on ratios alone. The median detection ran 0.92 of average true range, so this floor binds on the small tail rather than the middle
ColourNot usedDeliberately absent. A sixth threshold on a body that is small by construction would have bought very little and cost one more arbitrary number

Two things emerged from writing that table that were not visible from the prose version. The first is that the body-position clause is doing far more work than the famous one. On its own, the requirement that the top of the body sit inside the lower forty percent of the range admits 7.8% of all bars, while the twice-the-body shadow rule admits 19.2%. The clause everybody quotes is the loose one.

The second is worse, and it only shows up when you sweep. Setting the shadow multiple to 1.0 and setting it to 1.5 return exactly the same population, 25,842 detections both times. The reason is arithmetic rather than coincidence: if the top of the body sits inside the lower forty percent of the range, then the upper shadow is at least sixty percent of the range and the body is at most forty percent, so the shadow is already at least one and a half bodies. The most-quoted clause in the definition is inert below 1.5, because a clause nobody states has already enforced it. That is the sort of hidden coupling you only find by being forced to write the thing down.

One bar, one definition A detected bar from the generated daily tape, with every clause of the definition measured on its four prices. The measurements high body top body low low upper shadow real body lower shadow Upper shadow, in bodies 3.50 Body top, share of range 0.297 Lower shadow, share of upper 0.137 Range, in ATR20 0.93 The five clauses, and what each one costs Share of the 298,100 sampled bars that passes the clause on its own Upper shadow at least 2.0 times the real body 19.2% Top of the body inside the lower 0.40 of the range 7.8% Lower shadow no more than 0.25 of the upper shadow 15.8% Real body at least 0.05 of the range 95.8% Range at least 0.60 of the 20-bar average true range 74.4% All five together 1.92% of bars, 23,436 non-overlapping detections across 1,500,000 daily bars about 3.9 a year on any one chart The clause every guide quotes is the loosest of the five. The one that does the work is rarely stated as a number.
One detected bar from the generated tape, with each clause measured on its four prices. The bars on the right show how much of the tape each clause admits on its own, which is the honest way to see which threshold is carrying the definition. The clause every guide quotes is the loosest of the five. Illustrative simulated data.

Applied to 1,500,000 daily bars, which is 500 independent generated instruments of 3,000 bars each and about 6,000 instrument-years, the definition produced 28,764 qualifying bars. Discarding overlaps, so that no two measured outcomes share a holding period on the same instrument, leaves 23,436 independent detections. That is 3.9 per instrument per year, or roughly one every three months on any given chart. The typical detection had an upper shadow worth 3.5 bodies, a body top sitting 0.29 of the way up the range, and a lower shadow worth 0.12 of the upper one. The bars in the far tail were far more emphatic than the convention requires: a tenth of them had shadows longer than 9.4 bodies.

Two thirds of them do not get a name at all

Now the classification. The convention says a shooting star appears after an advance and an inverted hammer after a decline, which is a sentence with no numbers in it. Coding it requires deciding how much of a move counts, over how many bars, and measured against what.

The test used here reads only the closes, and it stops on the bar before the candidate, so that the candidate's own high and close can never contribute to the classification that names it. Take the change in the close over the previous ten bars, divide it by the recent daily volatility scaled to a ten-bar horizon, and you get a standardised prior move. A reading at or above 1.00 is an advance and the bar becomes a shooting star; a reading at or below the negative of that is a decline and the identical bar becomes an inverted hammer. Anything in between is neither.

The immediate consequence is one the convention never mentions. Of 23,436 detections, 4,252 followed an advance and 3,921 followed a decline. The remaining 15,263, almost two thirds of the population, followed neither. They came after a market that was going sideways, or drifting mildly, or doing something that a ten-bar test cannot call in either direction. No name in the convention applies to them, and yet they are the same bar and they are the bulk of the population.

In practice nobody leaves them unnamed. A chart with an upper-shadow bar in the middle of a range gets called a shooting star if the reader is bearish and an inverted hammer if the reader is bullish, and neither reading has any basis in a rule that was written down in advance. That is not a criticism of anybody's eyesight. It is the predictable result of a classification scheme whose only input is a trend test that was never specified.

The same bar, twice, with two different names Two detected bars from the generated tape, chosen because their geometry is almost identical. Only what came before them differs. Shooting star prior 10 bars advanced. Standardised prior move +2.97, and the whole drawn window +30.9%. the 10 bars the trend test reads the bar the same bar, rescaled shadow 3.50 bodies body top 0.297 lower 0.137 range 0.93 of ATR20 Inverted hammer prior 10 bars declined. Standardised prior move −1.48, and the whole drawn window −11.7%. the 10 bars the trend test reads the bar the same bar, rescaled shadow 3.43 bodies body top 0.301 lower 0.140 range 0.93 of ATR20 Nothing measured on the bar itself separates the two panels. The prior ten bars do all of the naming.
Two detections from the generated tape, chosen because their geometry is almost identical: their body positions agree to four thousandths, their lower-shadow fractions to three, their range to two hundredths of an average true range, and their shadow multiples to within a tenth of a body. The shaded band is the ten bars the trend test reads, and the bar itself is excluded from it. Every measurement taken on the bar is almost the same in both panels; only the prior ten bars differ. The pair was chosen on that geometry and on the prior move being legible in the drawn window. No outcome entered the choice. Illustrative simulated data.

The threshold has a second consequence, which is that it is unstable. Hold the 23,436 detections completely fixed, change nothing about any bar, and simply move the trend threshold across the three settings either side of the one used here. Only 5,780 of the 8,173 bars that received a name at the middle setting kept that name across all three. The rest either swapped from one name to the other or fell out of the naming scheme entirely. The market did nothing. A number that no textbook prints moved slightly.

Building a comparison that is not circular

A pattern statistic on its own means nothing. If a shooting star is followed by a fall sixty percent of the time, the question is what fraction of ordinary moments on the same data are followed by a fall, because the answer might also be sixty percent. That much is standard. What is not standard, and what turns out to decide the answer on a page like this one, is where the comparison entries are drawn from.

The natural choice is to draw random entries from the bars surrounding each detection, so the market conditions are comparable. For a test conditioned on prior trend, that choice is circular, and the circularity is not subtle once it is pointed at. The bars immediately before a shooting star are the advance that made it a shooting star. Comparing the bar against random entries taken from inside that advance is comparing it against the thing being used to classify it.

So every control entry on this page is drawn from bars strictly after the detection, between twenty-two and two hundred and fifty bars later, far enough ahead that neither the control's own trend test nor its holding period can reach back and touch the event or the ten bars being measured. Each control is also matched on the same prior context, takes a position in the same direction, holds for the same ten bars, and is required not to satisfy the shape definition itself. Two hundred controls are drawn for each detection, and the comparison is paired, so each detection is measured against its own controls rather than against a single pooled average.

That repair is necessary and it is not sufficient. A forward-only control is still not neutral, because a control drawn from a different stretch of tape is not matched on whatever slow-moving condition the detection happens to be sitting inside. The way to find out how big that problem is, rather than to argue about it, is to run the identical machinery on bars that are not the pattern. Take the same number of bars in each context class, require them to fail the shape definition, and push them through the same control apparatus. Those placebo bars carry nothing by construction. Anything the machinery reports for them is the machinery.

It reports a great deal. Under the forward-only control the shooting stars came out 0.60 percentage points ahead of their controls, and the placebo bars came out 0.32 points ahead of theirs. Most of the apparent excess was never about the bar. Every headline number on this page is therefore the difference of two differences: the pattern's excess over its control, minus the placebo's excess over its own.

Where you draw the control decides what you find The identical detections, measured four times against control entries taken from four different windows. Percentage points over 10 bars. Shooting stars: the raw excess, and the same measurement on placebo bars 0 0.2 0.5 0.8 only bars BEFORE inside the move that named it 250 bars either side the usual choice 1000 bars either side wider, still straddling only bars AFTER nothing before the bar the bar placebo bars the difference The placebo bars are not the pattern. They move with it anyway, so most of the raw excess is the harness. And the comparison the page actually rests on, under all four t −0.67 t −0.18 t −0.30 t +0.73 0 The gap between the two names, with its 95 percent interval. Every window says the same thing: no difference. The raw number is a property of the control window. The comparison the page rests on is not.
The pattern arm, the placebo arm and the difference between them, measured four times against four different control windows. Both arms swing by around four tenths of a percentage point as the window moves, and they swing together, which is what it looks like when a measurement is partly reporting its own apparatus. The lower panel is the comparison the page actually rests on, the gap between the two names, which stays inside one standard error of zero under every one of the four. Illustrative simulated data.

The progression in that figure is the reason the repair was worth making. Draw the controls only from bars before the detection and the shooting star's raw excess is 0.53 points. Draw them from two hundred and fifty bars either side and it falls to 0.21. Widen to a thousand bars either side and it is 0.22. Draw them only from after and it is 0.60. Four defensible-sounding choices, a range of nearly three to one, and no way to tell from the number itself which one you are reading. Once the placebo arm is netted out, the gap between the two names sits between 0.11 points below zero and 0.12 above it under all four, and never reaches one standard error.

The central experiment: do the two names differ?

With the comparison built, the question can finally be put properly. Take the 4,252 bars that qualified after an advance and the 3,921 that qualified after a decline. Both groups are measured the same way, as a long position taken at the next bar's open and held ten bars, so the two sit on one scale and can be subtracted. The bearish reading of a shooting star is the exact negative of its long figure, so reporting it separately would be reporting the same number twice with the sign changed, and the page says so rather than doing it.

Does the prior trend change what the bar does next? Excess over a forward-only, context-matched control, minus the same measurement on placebo bars, in percentage points over 10 bars. Tape A: nothing in the data makes anything predictable 0 −0.9 −0.5 0.5 0.9 Shooting star 4,252 bars +0.27% t +2.42 Inverted hammer 3,921 bars +0.15% t +1.31 No name applies 15,263 bars +0.09% t +1.59 Gap between the two names +0.12%, interval −0.20% to +0.44%, t +0.73 Tape B: a real effect was planted, and the same test finds it 0 −2.6 −1.3 1.3 2.6 Shooting star 4,605 bars −0.91% t −8.20 Inverted hammer 3,978 bars +1.51% t +11.00 No name applies 13,581 bars +0.05% t +0.76 Gap between the two names −2.42%, interval −2.77% to −2.08%, t −13.71 Where nothing is happening the two names are indistinguishable. Where something is happening the identical test separates them.
The comparison the whole page exists to make, with ninety-five percent intervals. On the tape with nothing in it the three intervals overlap each other heavily and sit within a third of a percentage point of zero, so the gap between the two names crosses zero comfortably even though the shooting-star arm on its own is marginally above it. On the tape with a planted effect, the identical test pulls the two names 2.42 points apart and the unnamed class stays where it was. The lower panel is the calibration: without it, a null in the upper panel could just mean the instrument is broken. Illustrative simulated data.

The shooting stars finished 0.27 percentage points ahead of their matched placebo comparison over ten bars. The inverted hammers finished 0.15 ahead of theirs. The gap between them is 0.12 percentage points, with a standard error of 0.16 and a ninety-five percent interval running from 0.20 below zero to 0.44 above. That is not a small effect that failed to reach significance. It is an interval tight enough to rule out any difference larger than about four tenths of a percentage point in either direction, on more than eight thousand named instances.

Read in the direction the names imply, which means selling after the shooting star and buying after the inverted hammer, the combined exercise returned −0.07 percentage points, with an interval from 0.23 below to 0.09 above. It is on the wrong side of zero and it is indistinguishable from zero, which are two different statements and both of them are true.

One detail in that pair of numbers should be flagged rather than buried, because it is the opposite of what the reader is expecting. Both arms lean the same way, upward, and the shooting-star arm is far enough above zero on its own to clear two standard errors. A bar that is supposed to be the bearish half of the family is the one with the mildly positive drift after it. What that is, and why it should not be treated as a finding, is the subject of the next section.

A null result is worth nothing unless the instrument can be shown to detect something. So the identical detector, the identical control machinery, the identical placebo arm and the identical thresholds were run over a second tape, generated the same way but with one real effect deliberately inserted. The effect is stated purely as two distances, how far the session high stood above the open and above the close in average true ranges, with the sign taken from the twenty-bar change in the close. It never mentions a real body, a lower shadow, or any ratio between the parts of a bar, which are the only things the detector measures. It fires on roughly one bar in six across the whole tape, far broader than the definition, so the detector has to find it by intersection rather than by being handed it.

It found it. On that tape the two names separated by 2.42 percentage points, more than thirteen standard errors, and the folklore reading returned 1.19 points at a similar margin. The instrument works. On the tape where there was nothing to find, it found nothing.

Both context classes on both tapes, each against its own forward-only, context-matched control, and each with the placebo arm subtracted. All figures are ten-bar outcomes in percentage points, illustrative and simulated, and are not a track record or a forecast.
GroupDetectionsPattern over controlPlacebo over controlDifferenceStandard errort
Shooting star, tape A4,2520.600.320.270.112.42
Inverted hammer, tape A3,921−0.29−0.450.150.121.31
No name applies, tape A15,2630.04−0.050.090.051.59
All detections, tape A23,4360.08−0.050.130.052.94
Shooting star, control tape4,605−0.640.27−0.910.11−8.20
Inverted hammer, control tape3,9781.29−0.221.510.1411.00

The row worth pausing on is the third one. The bars that received no name at all behaved essentially the same as the two named groups, 0.09 points against 0.27 and 0.15. Whatever the population of upper-shadow bars is doing on this tape, it is doing it uniformly, and the trend that assigns the name is not sorting it into anything.

Does the bar carry anything at all?

The second question is the one the naming argument distracts from. Forget both labels. Pool every detection, named or not, and ask whether an upper-shadow bar is followed by anything different from what follows an ordinary bar.

It is, slightly, and the honest thing is to publish it rather than to round it to zero. Pooled across all 23,436 detections the bar was followed by 0.13 percentage points more than the matched placebo comparison over ten bars, with an interval from 0.04 to 0.22 and a t statistic of 2.94. On the strength of the number alone that is a result.

Four things about it are worth stating before anybody does anything with it. It is 0.13 of a percentage point over ten bars, which is roughly the cost of the round trip that would be needed to collect it. It points upward, which is the wrong direction for the shooting star's bearish reading and the right direction only for the half of the population called an inverted hammer. It is present in all three context classes including the one with no name, so it is not a property of either label. And this page ran several dozen comparisons, so a single t statistic near three, in a set that size, is not a finding anyone should defend.

What followed the bar, against what followed a matched control Ten-bar outcomes for shooting stars and for the control entries drawn to match them. Both tapes, same axes. Tape A: no memory in the data −12% −6% 0% 6% 12% 10-bar outcome the bar matched control mean +0.35% control −0.24% positive 52% vs 47% Tape B: a planted effect the detector had to find on its own −12% −6% 0% 6% 12% 10-bar outcome the bar matched control mean −2.54% control −1.90% positive 30% vs 34% Read in the direction the names imply, selling the shooting star and buying the inverted hammer: −0.07% points on tape A (t −0.83), against +1.19% points on tape B (t +13.63).
The full ten-bar outcome distribution for shooting stars against the entries drawn to match them, on both tapes. On both tapes the two curves overlap almost completely, because a ten-bar outcome is mostly noise. The whole argument lives in a difference between two means that is a small fraction of the spread around them, which is why the means are printed rather than left to the eye. Note that the tape A means already differ before the placebo arm is subtracted, which is the harness effect the rest of the page removes. Illustrative simulated data.

The obvious explanation for a small positive residual is arithmetic rather than behavioural. The definition selects bars with long shadows, long shadows come with high volatility, volatility is persistent, and the expected simple return over a fixed window rises with variance even when the underlying drift is exactly zero. That is a real effect and it needs no market in it. It was checked rather than assumed. Measured in log returns, where the variance term disappears, the pooled figure is 0.12 points against 0.13 in simple returns, and the realised volatility of the ten bars following a detection is only 1.05 times that of a placebo bar, which accounts for about 0.010 of a point. The convexity explanation is not the explanation. Something small and real is there in the generated data, it is smaller than the friction required to reach it, and this page cannot tell you what it is.

The classical repair is confirmation, and the standard is a close beyond the bar rather than a touch, which for the bearish reading means the next bar closing below the shooting star's low. That was tested too, with entry moved to the open after the confirming bar. Only 1,580 of 4,252 shooting stars were confirmed on that definition, about 37%, and the confirmed subset finished 0.18 points from its placebo comparison with a t statistic of 1.02. For inverted hammers the bullish equivalent, a close above the bar's high, occurred just 465 times out of 3,921 and produced 0.03 standard errors of nothing. Waiting for the close remains good discipline for the reason the parent page gives, which is that it fixes an entry and therefore fixes what is being risked. It did not turn the bar into information here.

How much of this is the definition rather than the data

Every number above sits on five shape thresholds and one trend threshold. A finding that survives only at one setting is a finding about that setting, so the honest report is what happens when they move.

Move one threshold and the population moves with it The same generated tape throughout. Only the definition changes, and the classification changes with it. How many bars the definition finds by the upper shadow to body multiple 25,842 1.0 25,842 1.5 23,436 2.0 19,097 2.5 15,611 3.0 11,065 4.0 upper shadow, in real bodies 1.0 and 1.5 return the identical population. The gap between the two names shooting star minus inverted hammer, 95pc interval 0 0.5 1.0 1.0 1.5 2.0 2.5 3.0 4.0 upper shadow, in real bodies Where the gap opens it opens the wrong way for both names. The other tolerance: how much of a prior move earns a name at all Same detected bars every time. The number above each column is the share that receives any name. 0.50 64% 0.75 48% 1.00 35% 1.25 25% 1.50 17% standardised prior move required before a name is assigned shooting star inverted hammer no name applies Of 8,173 named bars, 5,780 keep one name across the middle three. The bar does not change. The label does.
The same generated tape throughout; only the definition changes. Top left, how many bars each shadow multiple admits. Top right, the gap between the two names at each setting with its interval. Below, what happens to the three classes as the trend threshold moves. At the two strictest shadow settings a gap does open, and it opens in the direction opposite to what both names claim. Illustrative simulated data.

The shadow multiple is the clause the brief for this exercise singled out, and it behaves in two distinct ways. Below 1.5 it does nothing whatsoever, for the arithmetic reason given earlier. Above 2.0 it starts cutting hard: 23,436 detections at 2.0, 15,611 at 3.0, 11,065 at 4.0. So the population is a tolerance setting, exactly as the triangle exercise found for its own classification, and anybody quoting a frequency for this bar without publishing their multiple is quoting their own settings.

What happens to the result as the multiple tightens is the uncomfortable part. At 1.0, 1.5, 2.0 and 2.5 the gap between the two names is indistinguishable from zero. At 3.0 it reaches 0.47 points at 2.39 standard errors, and at 4.0 it reaches 0.74 at 3.07. The same thing happens when the body-position clause is tightened: at 0.30 rather than 0.40 the gap is 0.63 points at 3.03 standard errors. It survives being measured in log returns, so it is not the variance artefact either, coming in at 0.74 points at the strictest shadow setting.

Before anyone reaches for that, note its direction. The gap is positive, which means the bars called shooting stars did better over the following ten bars than the bars called inverted hammers. That is the reverse of both names. And it appears on a tape built with no directional memory in it, where by construction no market mechanism exists to produce it. A statistic that shows up at the strict end of a threshold sweep, on data where nothing is happening, pointing the wrong way, in a set of several dozen comparisons, is a description of the definition and of multiple testing. It is not a discovery about markets. It is, however, precisely the kind of number that would be reported as one if it had turned up on real charts at the setting somebody happened to pick.

The trend threshold moves the classification even more freely. At the loosest setting tested, 64% of detections receive one of the two names; at the strictest, 17% do. The shooting star population runs from 7,735 down to 2,047 without a single bar changing shape. Across the middle three settings, 5,780 of 8,173 named bars keep their name and the rest do not. The other thresholds tell the same story in miniature: the lower-shadow cap moves the population between 10,536 and 32,372, the size floor between 6,672 and 33,077, and the holding period from five bars to twenty leaves the gap between the names within 0.35 points of zero throughout.

What survives all of it is the absence of a difference at any setting a reasonable person would have chosen in advance, and the presence of one only where the population has been cut to a third and the direction has flipped against both names. The population of upper-shadow bars is extremely sensitive to the definition. The finding that the two names do not sort that population is not.

What this supports and what it does not

A simulation can establish some things and cannot establish others, and being precise about which is which is the difference between a study and an opinion with numbers attached to it.

The boundary of what this exercise establishes, and what it leaves entirely open. Illustrative simulated data.
ClaimDoes this page support it?
The two names describe one bar, distinguished only by prior contextYes, and it is definitional rather than empirical. One detector found both; the label was attached afterwards from a trend test
The two names sort the population into groups that behave differentlyNo. The gap was 0.12 points over ten bars with an interval of about four tenths of a point either side, on 8,173 named instances
The naming convention covers the bars it is applied toNo. 15,263 of 23,436 detections followed neither an advance nor a decline, so no name in the convention applies to about two thirds of them
Which name a bar receives is stableNo. Only 5,780 of 8,173 named bars kept one name across the three trend settings either side of the one used
The bar carries something over an ordinary bar, pooled across contextsMarginally, and not usefully. 0.13 of a point over ten bars, upward, present in the unnamed class too, and smaller than the cost of collecting it
Confirmation by a close beyond the bar rescues the signalNo, not on this data. It also fired rarely, on 37% of shooting stars and far less often for the bullish case
Where the control entries come from changes the answerYes, and it is the largest single effect measured here. One arm moved from 0.21 to 0.60 points on the window choice alone, and placebo bars moved with it
These bars carry no information in Indian equitiesNo. This tested generated data, not any exchange. It shows what the pattern does when nothing is behind it, which is the baseline a real study needs before it can claim anything
Any of this is a reason to take a positionNo. Nothing here is a trade trigger, a recommendation or a forecast, and the outcome distributions are wide enough that no single instance is predictable

The largest limitation is the obvious one. Generated data has no earnings, no policy announcements, no index rebalancing and no order book, so it cannot tell you what happens when a real rally runs into real overhead supply. What it can do, and what real data cannot, is provide a case where the true answer is known in advance. When a statistic reproduces on a tape with nothing in it, that statistic is not evidence, and separating the numbers that fall into that category from the ones that do not is the whole contribution here.

A second limitation is that this tests one definition and one trend rule. A different set of six numbers is a different detector, and the sensitivity study above is precisely the reason to expect a different detector to find a different population. Anybody who wants to argue that a better definition would produce a different answer is making a testable claim, and the way to settle it is to write that definition down as code and run it against a control that was not drawn from the move being tested.

What to do with an upper shadow instead

None of this makes a long upper shadow useless to look at. It makes the two names useless, which is a much smaller loss than it sounds, and it relocates the part that was worth having.

Read the fact, not the label. The bar records one thing that is not in dispute: price traded well above where the session settled. Somebody was paying up higher and stopped. That is a fact about a single period of trading, it is checkable against the four prices, and it needs no name attached to it. The moment you decide whether to call it a star or a hammer, you have added an unstated trend threshold to a fact that did not need one.

Treat the trend test as the claim it is. If the reason a bar looks bearish to you is that it followed a rally, then the rally is your evidence and the bar is decoration. Say it that way and the position becomes arguable, because someone can ask how much of a rally, over how many bars, and whether that has ever been measured. Say it the other way and the trend hides inside a candle name where nobody has to defend it.

Write your definition before you go looking. This is the habit that transfers to everything else on a chart. Being forced to say how long a shadow must be, and how far up the range a body may sit, and how much prior move earns a name, is how you discover how much of your pattern recognition was a decision rather than an observation. Doing it once for this bar is worth more than reading ten descriptions of it, and the same exercise on a breakout or a support level produces the same uncomfortable clarity.

Know where your comparison group came from. The single largest number measured on this page was not produced by the market. It was produced by moving the control window, and it moved a headline figure by a factor of nearly three. Any pattern statistic you are shown, anywhere, was measured against something. Asking what that something was, and whether it was drawn from the same stretch of price action that defined the pattern, will dispose of a great deal of published material very quickly.

Keep the doji distinction, and keep it for the right reason. Push the body of this bar down to nothing and it becomes a gravestone doji, which our guide to the doji candlestick covers as part of a family with its own limiting cases. The reason to keep the two apart is not that they mean different things. It is that a definition which quietly absorbs its neighbours cannot be tested, because you can no longer say what it excluded.

What is left after all of that is smaller than the usual telling and considerably more durable. A single bar is a fact about one session. Turning facts into decisions requires a rule written down in advance, a comparison group that was not drawn from the evidence, and a willingness to publish the runs that went nowhere. If that way of working is more interesting to you than the taxonomy it replaces, it is the method we teach.

FAQ

Frequently asked questions

Nothing that can be measured on the bar. Both are a small real body sitting at the bottom of the range, a long upper shadow, and little or no lower shadow. The name is decided entirely by what happened before: after an advance the bar is called a shooting star and read as bearish, after a decline the identical bar is called an inverted hammer and read as tentatively bullish. Because the shape is common to both, the pair is the cleanest available test of whether context or shape is carrying the information.

Not on the data tested here, and the interval is tight enough to be worth quoting. Across 4,252 detections on a generated tape with no directional memory in it, the shooting star and the inverted hammer differed by 0.12 percentage points over the following ten bars, with a 95 percent interval running from 0.20 points below to 0.44 above. Read in the direction the names imply, selling one and buying the other, the whole exercise returned 0.07 percentage points on the wrong side of its own control. Nothing on this page is a reason to take a position.

There is no published answer, which turns out to matter. The usual convention is at least twice the real body, and that is what this page coded. Sweeping the multiple from 1.0 to 4.0 moved the population from 25,842 detections to 11,065 on identical data. More awkwardly, 1.0 and 1.5 returned exactly the same population, because the separate requirement that the body sit inside the lower 40 percent of the range already forces the shadow to be at least one and a half bodies. The clause every guide states does nothing at all until it passes the clause almost none of them state.

The body. A shooting star keeps a small but visible real body, so the open and the close finished apart. A gravestone doji has essentially no body at all: open and close land at the same price at the bottom of a long upper shadow. The detector on this page draws the line at a body worth 5 percent of the range and excludes anything smaller, because the doji family is a different subject with its own limiting cases. That 5 percent is a choice, not a fact, and it is stated with every other threshold.

The definition coded here ignores it entirely, and that is deliberate rather than casual. The claim the bar makes is that price traded well above where it settled, which is a statement about the shadow and about where the body sits in the range. Whether the close finished a little above or a little below the open is a second-order detail on a body that is small by construction. Adding colour as a clause would have added a sixth threshold to a definition that already needed five.

Under the definition used here, about 3.9 times per instrument per year, or roughly one every three months on any given chart. That came from 28,764 raw detections across 1,500,000 daily bars, reduced to 23,436 once overlapping holding periods were discarded. The frequency is a property of the thresholds rather than of the market: loosening the lower-shadow cap alone took the population above 32,000, and tightening the body-position clause took it below 14,000.

Classical practice asks for a close beyond the bar rather than a touch, which for the bearish reading means the next bar closing below the shooting star's low. That was tested. Only 1,580 of 4,252 shooting stars were confirmed that way, and the confirmed subset did not separate from its control either. Confirmation is still worth insisting on for a different reason, which is that it fixes an entry price and therefore fixes what is being risked, but on this data it did not turn the bar into information.

Because a generated tape can be built with a known answer inside it and real data cannot. Two tapes were used. One has no mechanism that could make any bar shape predictive, so anything the detector reports there is an artefact of the definition or of the measurement. The second has a real effect deliberately planted, expressed in terms the detector never looks at, and it exists to prove the instrument can see an effect when one is present. Without that pair, a null result is indistinguishable from a broken detector.

Because for anything conditioned on a prior trend it decides the answer. The obvious control is random entries from the bars surrounding each detection, but the bars before a detection are the very advance or decline that assigned the name, so that control is partly the thing being tested. Moving the control to bars strictly after the detection changed the measured excess for shooting stars from 0.21 percentage points to 0.60. Running the identical machinery on bars that are not the pattern showed most of that swing was the harness rather than the bar, which is why every headline number here is a difference of two differences.

The one thing the bar records that is not in dispute: price traded well above where the session settled, so somebody who was paying up higher stopped paying up. That is a fact about a single period, it is worth noticing at a level that already mattered, and it is nowhere near a decision. What this page argues against is not looking at the bar. It is the belief that attaching one of two names to it, on the strength of an undocumented trend threshold, converts it into a forecast.

Method note

How the numbers on this page were produced

Every figure comes from one deterministic simulation, seeded so that it reproduces identically on each run. Two tapes were generated, each of 500 independent instruments of 3,000 daily bars, 1,500,000 bars in total. Volatility clusters and drift regimes switch on both, so quiet and violent stretches occur naturally and no bar shape is ever inserted by hand; the three drift states are symmetric and equally likely, so the unconditional drift is zero. Upper and lower shadows are drawn from the same distribution, so the generator has no preference for one end of a bar over the other. The first tape contains no mechanism that could make any bar shape predictive. The second adds one effect, stated as two distances measured in average true ranges, how far the session high stood above the open and above the close, with the sign taken from the twenty-bar change in the close. It refers to no real body, no lower shadow and no ratio between the parts of a bar, and it fires on about one bar in six, so the detector inherits it only by intersection. Its size is deliberately large, because a control that cannot be seen is not a control.

Detection runs on the open, high, low and close of the generated bars using the six thresholds in the definition table. The prior-context test reads closes only and stops on the bar before the candidate. Positions are opened at the next bar's open, so no result uses information that was not available at the time, and outcomes are measured over ten bars as simple returns, which is what a position actually makes. Overlapping detections are discarded so that no two measured outcomes share a holding period on one instrument. Control entries are drawn from the same instrument, from bars between twenty-two and two hundred and fifty bars after the detection, matched on the prior-context label, required to fail the shape definition, and entered and held the same way, at two hundred draws per detection with the comparison paired. A placebo arm of the same size, matched on context and required to fail the shape definition, is pushed through the identical machinery, and every headline figure is the pattern's excess minus the placebo's. Returns are shown before costs; an illustrative allowance for a full round trip would apply equally to both arms and would not change the difference between them, though it would sit above most of the differences measured.

All results are illustrative and simulated. They are not a track record, they are not a forecast, and they are not an indication of what any pattern would produce in a live account or on any Indian security or index. The purpose of the exercise is to establish which properties of this bar are consequences of its definition and of the way it is measured, which is a question about definitions and measurement rather than about any particular market.

Related

Continue reading

Next step

Find your starting stage. Everything else follows from there.

Educational reference only. No buy, sell or hold recommendations. All results shown are illustrative and simulated.