Educational Reference
Morning Star and Evening Star: What Three Bars Can and Cannot Tell You
A morning star is a long down bar, then a small-bodied bar that gaps away below it, then an up bar that closes well back into the first. The evening star is the same machine at the top of an advance. They are taught as among the most dependable reversal signals in the candlestick canon. This page does not argue with that. It writes the definition as code, twice, detects every occurrence across 1,800,000 generated daily bars, and reports what the detector found.
The finding, stated first. The classical definition requires gaps, and how strictly you enforce that one clause changes the population by orders of magnitude. Enforced as a chartist would read it, with the middle bar clear of both neighbours, the pattern occurred 11 times in 7,200 instrument years, which is far too few to test and is therefore the first thing this page reports. Enforced between the real bodies, as the classical texts write it, 1,176 times. With the gap clause dropped, which is how most published detectors code it, 8,490 times. Same three bars, same tape, a factor of 772 between the extremes. All results on this page are illustrative and simulated.
A clause written for a different market
Most candlestick patterns are defined by the shapes of their bars. The star patterns are unusual because their definition names something that happens between bars. The middle bar is not merely small and not merely low: it is required to gap away from the bar before it, and the third bar is classically required to gap back away from the middle one. Two gaps, one on each side of the star, and the word star is itself a description of that isolation. A bar that has jumped clear of its neighbours hangs on its own.
That clause was not invented carelessly. It was written for markets that opened once a day, settled once a day, and had long unpriced hours between sessions in which news accumulated and was released in a single jump at the next open. In that world an isolated bar is a genuinely distinct event, because there was a real discontinuity in the auction. The information arrived while nobody could trade, and the gap is the record of it.
Now consider what happens to the same clause in a market that trades continuously through the session and reopens the next morning after a comparatively short and well-lit interval. Gaps still occur, but they are smaller relative to the day's own range, and a body that clears the previous body entirely becomes an uncommon event rather than a routine one. Nothing about the definition changed. The market it describes did. That is a testable proposition rather than a rhetorical one, and it is the proposition this page tests.
There is a second reason to look hard at this clause, and it is more damning than the first. Open several published descriptions of the morning star and you will find the gap named in the prose and absent from the diagram. Open several published detectors and you will find the gap absent from the code as well, replaced by a weaker ordering condition that merely asks the middle bar to sit lower than the first. Nobody announces the substitution. The result is that two people can run what they both call a morning star scan over the same data and disagree about the count by a factor of seven, without either of them having done anything wrong. Where the star sits in the wider family of candle formations is mapped in the guide to reading candlestick charts; the work here goes underneath that map and tests one entry on it.
Writing the definition down, twice
Here is the verbal definition, close to the standard one. A morning star is a long down bar, followed by a small-bodied bar that gaps below it, followed by a strong up bar that closes well into the body of the first. Try to run it. How long is long, and long compared with what? How small is small? Gaps by how much, measured between which prices? How strong is strong? How far is well into, and is the reference point the first bar's midpoint, its open, or somewhere else? And what happens on a chart where the bodies overlap by a hair, which is the overwhelmingly common case?
Six questions, and the sentence answers none of them. That is not a complaint about sloppy writing. It is the reason two analysts can look at the same chart and disagree in good faith, and it is the reason published statistics about star patterns are difficult to compare with one another. Every one of those statistics belongs partly to whoever chose the missing numbers.
Two decisions in the table below deserve to be flagged before the results arrive. The first is that every threshold is expressed against the average true range measured before the pattern begins, never during it. Using the range inside the pattern would let a violent three-bar sequence inflate its own yardstick, so that every star scores about the same and the thresholds stop discriminating. The second is that the gap clause is written as a single number in those same units, which allows it to be moved continuously from a demanding true gap, through zero, into tolerated overlap, and finally switched off. That is what makes the central experiment on this page possible at all.
| Clause | Enforced as written | The common relaxation | What the clause is doing |
|---|---|---|---|
| Bar 1 is long | Real body at least 1.0 times the average true range measured before the pattern | Establishes that there was a decisive session to be reversed. Loosening it to 0.6 raised the count to 3,783, tightening it to 1.6 cut it to 194 | |
| The star is small | Real body no more than 0.35 of bar 1's body | What makes the middle bar a pause rather than a continuation. Nearly inert: moving it between 0.20 and 0.60 changed the population from 997 to 1,228 | |
| Bar 3 is strong | Real body at least 0.50 of bar 1's body, in the opposite direction | Stops a trivial bar being read as a reversal. Tightening it to 0.8 cut the count from 1,176 to 629 | |
| The star is displaced | The star's body midpoint sits beyond bar 1's body midpoint, away from the eventual direction | The minimum that makes the word star mean anything once the gap clause is removed. It is the only positional requirement the relaxed reading keeps | |
| The gap | A true gap on both sides: the star's body clear of bar 1's body, and bar 3's body clear of the star's | No gap required. Bodies may overlap freely | The clause this page exists to measure. It cut the surviving population from 12,526 evaluations to 1,348, a factor of 9.3. Nothing else in the definition comes close |
| Bar 3 closes back in | Closes at least 0.50 of the way back through bar 1's body | The other clause the guides leave loose. Swept separately further down, from 0.30 to 1.10, which moved the count from 1,301 to 330 | |
| Entry and holding | Next bar's open, held 20 bars | Entering on the bar that completed the pattern would quietly hand the test information from the future, which is the most common silent error in pattern studies | |
The two-bar members of this family, where one body swallows or hollows out the one before it, are worked through separately in the deep dive on the bullish engulfing pattern, which covers the engulfing, the piercing line and the harami as a single spectrum of how far the second bar claws back. The star is the three-bar extension of that idea, and the extra bar buys it one thing the two-bar patterns do not have: an explicit pause in the middle, with an explicit clause about how that pause is positioned. That clause is the subject here. The middle bar itself, when its body is small enough to be near zero, is a doji, and a star whose middle bar is a doji is sometimes given its own name, which is a naming decision rather than a measurement.
Three readings, three populations
Across 1,800,000 daily bars, which is 600 independent generated instruments of 3,000 bars each and about 7,200 instrument years, the detector evaluated 3,396,000 candidate positions, counting each bar once as a possible morning star and once as a possible evening star. The funnel narrows fast. 81,692 positions had a long enough first bar. 39,273 of those had a small enough second. 12,685 had a strong enough third, and 12,526 had the star displaced to the correct side.
Then the gap clause runs, and it is the only place in the funnel where an order of magnitude disappears. Of the 12,526 positions that had the right three shapes in the right order, 1,348 had a true body gap on both sides. After the penetration clause and after discarding overlaps so that no two measured outcomes share a holding period on the same instrument, 1,176 independent instances remained under the strict definition and 8,490 under the relaxed one. That is 0.16 per instrument per year against 1.18, or roughly one every six years on a given chart against about one a year.
The third reading is the one a chartist would recognise. When a trader says there is a gap on the chart, they do not mean the bodies cleared: they mean there is a visible hole, with the whole of one bar sitting clear of the whole of the next, highs and lows included. Coded that way, requiring the star's entire range to sit clear of bar 1's range and bar 3's range to sit clear of the star's, the detector found 11 instances in 7,200 instrument years. Eleven. That is not a small effect size, it is an absence of data, and no honest statistic can be computed from it. The measured difference from its base rate was −1.44% with a standard error of 2.45 percentage points, which is a way of saying the number means nothing.
The distribution behind that is worth stating directly, because it is the mechanism. Across the 8,490 relaxed detections, the median gap on the tighter of the two sides was −0.24 of an average true range, which is to say the bodies overlapped. Only 12.3% of them cleared on both sides. On a continuously traded daily tape whose opens gap around the previous close with a standard deviation of about three tenths of the day's volatility, a body that jumps clear of its neighbour is the exception. The classical clause is not describing the common case. It is describing the case that used to be common.
Where the comparison group comes from
A pattern statistic on its own says almost nothing. If a star is followed by a gain fifty five percent of the time, the useful question is what fraction of random moments on the same data are followed by a gain, because the answer may also be fifty five percent. Leaving that comparison out is usually what makes a pattern statistic look impressive. So every number on this page is a difference between two arms: entries taken because a star completed, and entries taken for no reason at all, in the same direction, on the same instrument, held for the same 20 bars.
The construction of that second arm turns out to matter more than anything else on this page, and it is the one decision that pattern studies almost never state. The natural choice is to draw the random entries from a window either side of the event, so that the market conditions are comparable. For a reversal pattern that window is not neutral. A morning star is selected partly by the price path immediately before it: the long down bar is required, and if you additionally condition on there having been a genuine decline to reverse, which the pattern's own logic demands, then the preceding stretch has been chosen twice over. Random long entries drawn from inside a decline lose money. The comparison group is therefore depressed by the very price path that selected the pattern, and the pattern is flattered by an amount nobody measured.
That is a claim, so it is measured rather than asserted. Take the subset that had something to reverse and hold everything else fixed: the same 3,066 detections, the same outcomes, the same holding period, changing only the window the comparison entries come from.
The progression is monotonic and it is the whole argument for the choice made here. Against entries drawn only from bars before the event, the cell reads 0.21% at 1.57 standard errors. Against a symmetric window of 250 bars either side, −0.09% at −0.65. Widen that symmetric window to a thousand bars either side, which dilutes the contamination without removing it, and it becomes −0.28% at −2.11. Draw only from bars after the event and it is −0.33% at −2.45. Same detections, same returns, four answers spanning half a percentage point, and the ordering is exactly what circularity predicts.
Every number on this page therefore uses the forward-only comparison group. Two checks were needed before that could be trusted. The first is that a control which reports zero for everything is not a strict instrument, it is a broken one, so the forward-only window had to be shown capable of finding something. Run on the tape with a real effect planted in it, it reports 0.83% at 9.38 standard errors. It works. The second is that a forward window has an asymmetry of its own, since the pattern sits at the very start of it while the comparison entries are spread across it. Narrowing the window to the next sixty bars, widening it to five hundred, and lagging it so that it begins only after the holding period has finished all give the same answer for the conditioned cell: −0.28%, −0.36% and −0.34% against −0.32% for the window used throughout.
What followed, on two tapes
With the comparison group settled, the result is quickly told. On a tape built with no directional memory, the relaxed definition returned −0.04% over twenty bars and the matched random arm −0.03%, leaving the pattern −0.01% away from its base rate, which is −0.17 standard errors from zero. The share of positive outcomes was 49.2% against 49.8%. The tails match as well, which matters more than the averages: a tenth of pattern outcomes were worse than 8.4 percent down against 7.4 for the controls, and a tenth better than 8.6 up against 7.3.
The strict definition, which is the one the classical texts actually specify, has 1,176 instances rather than 8,490 and correspondingly wider intervals. It came in 0.26% above its base rate at 1.16 standard errors, with a ninety five percent interval running from 0.18 points below zero to 0.71 above. That interval is wide enough to contain outcomes a trader would care about in both directions, which is the honest summary of what a thousand-odd instances can establish.
A null result from a detector that finds nothing is worth nothing, because the detector might simply be broken. So the identical detector, the identical comparison machinery and the identical thresholds were run over a second tape, built the same way but with one genuine effect inserted: whenever a single bar's return is large relative to its own recent scale, the following twenty five bars carry extra drift in the opposite direction. That is a one-bar overreaction effect, written purely in return space. It says nothing about a small middle bar, nothing about a gap, nothing about a third bar closing back into the first and nothing about three bars at all, so the star detector had to find it unaided.
It did find it. On that tape the pattern arm averaged 0.85% against 0.02% for matched random entries, a difference of 0.83% at 9.34 standard errors, with 54.2% of outcomes positive against 50.2%. The strict definition on the same tape recovered 0.50% at 1.98 standard errors from its smaller sample. The machinery can see a reversal effect. On the tape where there was none, it saw none.
| Variant | Found | Per instrument year | Mean, 20 bars | Matched base rate | Difference | Standard error |
|---|---|---|---|---|---|---|
| Gap visible on the chart | 11 | 0.002 | −0.89% | 0.55% | −1.44% | 2.45 points |
| Gap between the bodies, as written | 1,176 | 0.16 | 0.25% | −0.02% | 0.26% | 0.23 points |
| Gap clause dropped | 8,490 | 1.18 | −0.04% | −0.03% | −0.01% | 0.08 points |
| Morning stars only, gap dropped | 4,150 | n/a | 0.20% | 0.18% | 0.02% | 0.12 points |
| Evening stars only, gap dropped | 4,340 | n/a | −0.27% | −0.21% | −0.06% | 0.11 points |
| With something to reverse | 3,066 | n/a | −0.38% | −0.06% | −0.32% | 0.13 points |
| Nothing to reverse | 5,424 | n/a | 0.15% | −0.01% | 0.15% | 0.10 points |
| All, control tape | 8,637 | n/a | 0.85% | 0.02% | 0.83% | 0.09 points |
One row in that table deserves attention because it is the kind of comparison that gets quoted without a base rate. Morning stars averaged 0.20% over twenty bars and evening stars −0.27%, a gap of 0.47 percentage points at 2.91 standard errors. Read raw, that says morning stars are the better pattern by a comfortable margin. Read against their own direction-matched base rates, the two come in at 0.02% and −0.06%, and the difference disappears entirely. A long position and a short position on the same series do not have symmetric arithmetic, and the base rate is what removes that. Almost every published claim that one direction of a pattern works better than the other is this arithmetic, uncorrected.
The central experiment: only the gap clause moves
Everything so far describes two settings of one definition. The experiment this page exists for is narrower. Freeze every other clause, change nothing about how long bar one must be, how small the star must be, how strong bar three must be or how far it must close back, and move only the number that decides how much of a gap is required. If the classical clause carries content, this is where it must show up, because this is the clause that is supposed to separate a real star from a coincidence.
On the tape with no memory, a stricter gap bought nothing. The population fell from 8,490 with the clause removed, to 1,176 at a true body gap, to 313 at a gap of 0.15 average true ranges, to 72 at 0.30, to 10 at 0.50, and to none at all at 0.75. Across that whole collapse the difference from the matched base rate stayed inside half a percentage point of zero at every setting with enough instances to measure, and never reached two and a half standard errors at any of them.
Nested thresholds share their events, though, so the differences between neighbouring points on that line are not independent of each other. The cleaner version detects once with the clause switched off and splits the resulting population into non-overlapping bands by how much gap or overlap each instance actually had, which gives independent cells and a trend that can carry a standard error. On the tape with no memory that trend is 0.77% per additional average true range of gap, with a standard error of 0.36 points, or 2.12 standard errors from zero.
That reading sits close enough to two standard errors to be worth taking seriously, and the control tape is what settles it. Running the identical band split on the tape that does contain a real reversal effect gives a trend of −0.55% per average true range at −1.36 standard errors, which runs the other way. If a bigger gap genuinely selected for a stronger reversal, it would show up at least as clearly where a reversal effect actually exists. It does not show up there at all. The reading on the null tape is one marginal result among the many comparisons this page runs, and the honest description of the gap clause is that it is a population filter rather than an outcome filter: it divides the count by 9.3 and leaves the measurement where it was, on both tapes.
That is the whole argument in one comparison. The experiment is demonstrably capable of detecting a real reversal effect, at a sample size and an effect size of the same order as the null case, because it detected one on demand at 9.34 standard errors. When the same experiment is run where no such effect exists, it reports none. The flat line is a measurement, not a failure to measure.
The other clause the guides leave loose
The gap is not the only place a published definition goes vague. The third bar is required to close well back into the first, and well back is doing as much unexamined work as the gap did. The classical answer is past the midpoint of bar one's body, which is a specific claim: half is meaningful and a third is not. Some guides ask for more, some accept any close inside the body, and a few say nothing at all. It is a second continuous dial and it can be swept the same way.
On the tape with no memory, depth is inert. Split the population into disjoint bands by how far bar three actually closed and the trend is −0.07% per unit of depth with a standard error of 0.26 points, which is −0.26 standard errors from zero. Demanding a close past the midpoint rather than merely inside the body costs 1,301 instances down to 1,176 and changes nothing measurable. Demanding a close all the way past bar one's open, which is the strongest version anybody asks for, leaves 330 and still changes nothing.
The control tape says something more interesting than nothing. There the trend is −0.71% per unit of depth at −2.46 standard errors, and the sign is negative: on a tape where a genuine reversal effect exists, demanding a deeper third bar made the result worse. The reason is mechanical rather than mysterious, and it generalises well beyond this pattern. A deeper close means more of the reversal has already occurred before the pattern is complete and before you can act on it. What you are measuring after entry is whatever is left, and waiting for more confirmation means waiting through more of the move you were hoping to capture. Every confirmation requirement has that cost, and it is invisible in any test that does not have a tape with a known answer to compare against.
A reversal pattern needs something to reverse
The single condition the star's own logic demands is that there be a prior move to turn. A morning star in the middle of a range is three bars in a range. Conditioning on that is also the place where the comparison group had to be repaired, which is why this section comes after that argument rather than before it.
Conditioning is expensive in sample, so the cells are reported with their counts. Requiring a genuine adverse move of at least one average true range across the twenty bars before the pattern left 3,066 of 8,490 instances, or 36.1% of them. Nearly two thirds of the things this definition calls stars had no meaningful move in front of them to reverse, which is worth pausing on: the shape fires overwhelmingly where its own premise does not hold.
The conditioned subset came in −0.32% from its forward-only base rate at −2.39 standard errors, while the 5,424 instances with nothing to reverse came in 0.15% above theirs at 1.53. The difference between the two cells is −0.47 percentage points with a standard error of 0.17, which is −2.83 standard errors. Taken at face value, requiring the one condition the pattern's logic demands made the measurement slightly worse rather than better, on a tape where nothing is happening.
Two cautions before anyone carries that away as a finding. The first is that this page runs a large number of comparisons, and a reading near two or three standard errors among many is not the same thing as a reading near three standard errors from a single pre-registered test. The second is the strict definition, where the same split leaves only 401 and 775 instances and the difference between the cells is −0.79 standard errors, which is nothing at all. Two definitions of the same pattern, two different answers to the same context question, and the smaller sample cannot adjudicate.
What can be said with more confidence is the counterpart on the control tape, because there the direction is predicted in advance. On the tape carrying a real reversal effect, conditioning on a genuine adverse move helped: 1.12% against 0.68%, a difference of 0.45% points at 2.43 standard errors. That is the shape a real context effect makes. The null tape does not make that shape. It makes the opposite one, weakly, in a place where the honest reading is that conditioning bought nothing and may have cost something.
The statistic that looks most like evidence
One number on this page does separate the pattern from its base rate, cleanly and repeatably, and it is the number most often quoted as proof that star patterns work. Project the height of the three bars from the entry, which is the standard target for this family, and ask how often price covered that distance in the trade direction inside the holding period. It was reached 47.6% of the time against 42.4% for matched random entries taking the same direction, the same distance and the same window. Over sixty bars, 66.4% against 63.3%. The strict definition shows the same thing, 47.8% against 41.3%.
A five point gap on a proportion with thousands of observations is not noise. It is also not what it looks like. The first candidate explanation, that stars occur when volatility is high and a fixed distance is easier to cover then, can be tested by restricting the comparison entries to bars whose twenty-bar average range is comparable to the event's. That changed the base rate from 42.4% to 42.0%, which is to say it changed nothing, and it also ruled the explanation out: the twenty-bar average range at a star is 0.0212 of price against 0.0208 at ordinary bars. By that measure a star is not a high-volatility moment.
The sharper hypothesis is that a star is a local range burst that a twenty-bar average cannot see, and that the target is made of the burst. The three bars of a detected star span 0.0494 of price against 0.0318 for three ordinary bars, which is 1.55 times as wide. The target is that same span. Short-horizon volatility clusters, so the bars immediately after a burst cover ground faster than average and a distance scaled to the burst is easier to reach. Restrict the comparison entries to moments whose own trailing three-bar span is comparable and the base rate rises from 42.4% to 46.0%, closing roughly three quarters of the gap, with the residual explained by the matched pool averaging 0.0454 against the event's 0.0494 because bars that wide are not always available nearby.
The mean outcome, meanwhile, does not move at all: −0.04% for the pattern against −0.06% for the range-matched comparison group. So the hit rate is telling you how wide the pattern is, and the pattern is wide by definition, because a long first bar and a strong third bar are two of its six clauses. Any target expressed as a multiple of the pattern's own size inherits this. It is worth carrying to every measured-move statistic you meet, in this family and outside it: a hit rate against a self-scaled target is partly a statement about the size of the thing you measured, and it has to be netted against a size-matched baseline before it can be read as evidence about what happens next.
How much of this is the definition rather than the data
Every number above rests on seven thresholds, and a result that survives only at one setting is a result about that setting. The clauses were varied one at a time. Loosening the requirement on bar one from 0.6 to 1.6 average true ranges moved the population from 3,783 to 194, which is the largest single lever after the gap clause and a reminder that a threshold nobody states can be as consequential as a clause nobody enforces. The size requirement on the star barely mattered, moving the count from 997 to 1,228 across a threefold change in the threshold. The strength requirement on bar three moved it from 1,338 to 629.
The difference between those clauses and the gap clause is not the size of the lever. It is that nobody disagrees about whether the other clauses exist. Everyone requires a long first bar; they differ on how long. The gap clause is different in kind, because it is either present in the definition or silently absent from it, and the two readings are both published under the same name. That is what makes it worth a whole page, and it is why the sibling study of triangle patterns, which runs on the first 250 instruments of this same generated tape at the same seed, ran into a structurally similar problem from a different direction: there the unstated choice was where the pattern is allowed to begin, and it moved that population by more than a factor of ten.
One methodological point deserves stating plainly, because it applies to this page as much as to anything it criticises. A study that reports thirty comparisons will contain readings near two standard errors whether or not anything is happening, and this page reports more than thirty. The two readings here that sit in that zone, the gap trend on the null tape and the context cell, are both reported with their control-tape counterparts precisely so that a reader can see whether they behave like real effects. They do not. The results this page would defend are the ones that are either very large, such as the population multiples, or confirmed on the tape with a known answer, such as the depth penalty and the absence of an outcome effect.
What this supports and what it does not
A simulation can establish some things and not others, and being exact about which is which is the difference between a study and an opinion with numbers attached.
| Claim | Does this page support it? |
|---|---|
| The gap clause decides how many stars exist | Yes, and by a wide margin. Enforced between the bodies it cut the surviving population by a factor of 9.3; enforced as a visible hole on the chart it left 11 instances in 7,200 instrument years |
| The strict version is rare enough to be untestable | The chart-gap reading, yes: 11 instances cannot support any conclusion and the honest response is to say so. The body-gap reading, no: 1,176 instances is a usable sample, though the interval it produces is wide |
| A stricter gap improves the outcome | Not on data with no memory, where the population fell three orders of magnitude with no measurable change. Not on the control tape either, where the same band trend runs the other way at −1.36 standard errors. It is a population filter, not an outcome filter |
| A deeper third bar is a better third bar | No. Flat on the null tape at −0.26 standard errors, and actively worse on the tape containing a real reversal effect at −2.46, because a deeper close means more of the move has already happened |
| The target is reached more often than chance | Yes as a raw number, 47.6% against 42.4%, and no as evidence. A star spans 1.55 times a typical three-bar range and the target is that span; matching the comparison group on recent span closes three quarters of the gap while the mean outcome does not move |
| Morning stars work better than evening stars | No. The 0.47 point raw gap at 2.91 standard errors is the arithmetic of long against short and vanishes against direction-matched base rates |
| Where the comparison group is drawn from matters | Yes, and it is the most transferable result here. The same conditioned cell reads 0.21% against a backward window and −0.33% against a forward one, and the unconditioned population barely moves, which is what keeps the problem hidden |
| Star patterns carry no edge in Indian equities | No. This tested generated data, not any exchange. It shows what the pattern does when nothing is behind it, which is the baseline a real study needs before it can claim anything |
| Any of this is a reason to take a position | No. Nothing here is a trade trigger, a recommendation or a forecast, and the outcome distributions are wide enough that no single instance is predictable |
The largest limitation is the obvious one. Generated data has no earnings, no policy announcements, no index rebalancing and no order book, so it cannot say what happens when a real gap is caused by real news and a real reversal follows it. That matters more for this pattern than for most, because the gap clause is precisely the clause that would carry event information on a real tape. What generated data can do, and what real data cannot, is supply a case where the true answer is known before the measurement starts. When a statistic reproduces on a tape with nothing in it, that statistic is not evidence, and sorting the familiar star numbers into that category is the contribution here.
A second limitation is that this tests one gap regime. The generated opens jump around the previous close with a standard deviation of about three tenths of the day's own volatility, which is a plausible daily tape and is not any particular market. A market with wider overnight jumps would find more strict stars and a market with narrower ones fewer, which is exactly the point about the clause: its population is a property of the gap regime, and the gap regime is a property of the market structure rather than of the pattern.
What to do with a star instead
None of this makes three bars uninteresting. It relocates the interesting part, and it changes what is worth measuring.
Decide which definition you are using, and say so. The word star covers two populations that differ by a factor of 7.2 and a third that is a hundred times rarer again. Any frequency, any hit rate and any win rate quoted for this pattern is a statement about the quoter's gap clause before it is a statement about markets. Ask which one was used. The answer is almost never in the article.
Treat a gap as an event, not as a shape clause. On a real tape the reason a body jumps clear of the one before it is usually that something was announced. That is information, and it is information about the news rather than about the geometry. The clause is worth keeping for what it selects, not for what it draws.
Net every hit rate against a size-matched baseline. The measured-move statistic on this page beat its plain base rate by five points and beat a size-matched one by barely one, and the mean outcome did not move at all. Any target expressed as a multiple of the pattern's own dimensions has this problem, and it is the single easiest way to manufacture a convincing number by accident.
Be suspicious of confirmation that costs you the move. Demanding a deeper third bar made things worse on the only tape where a real reversal existed to be captured. More confirmation is not free; it is paid for out of the part of the move that happens while you wait, and the payment is invisible unless you measure it against something.
Ask where the comparison group came from. This is the most transferable habit on the page. Any claim of the form this pattern beats normal conditions rests on a definition of normal, and if normal was measured over the stretch of chart that produced the pattern, the comparison is circular. It will usually flatter reversal patterns and penalise continuation ones, and it never announces itself.
Write the definition down before you look for instances. Forced to say how long is long, and against what yardstick, and whether the bodies must actually clear, you discover how much of your pattern recognition was a decision rather than an observation. Doing that once is worth more than reading ten descriptions, and if that way of working appeals more than the taxonomy it replaces, it is the method we teach.
FAQ
Frequently asked questions
What is a morning star pattern?
Three bars read as one event. A long down bar, then a small-bodied bar that sits below it and is classically required to gap away from it, then an up bar that closes well back into the first bar's body. The evening star is the mirror image at the top of an advance. The definition is unusual among candlestick patterns because it names a gap explicitly, and that clause turns out to decide almost everything about how often the pattern can be found.
Does the morning star really need a gap?
The classical definition says yes, on both sides of the middle bar, and most published detectors quietly say no. That disagreement is not cosmetic. Coded on the same 1,800,000 generated daily bars, the strict reading found 1,176 instances and the relaxed reading 8,490, a factor of 7.2. Read the gap the way a chartist reads one, as a visible hole with the whole of the middle bar clear of its neighbours, and the count falls to 11. Any statistic quoted for this pattern belongs to whichever of those three definitions the quoter used.
How often does a morning star or evening star actually occur?
Under the relaxed definition, about 1.18 times per instrument per year on this generated tape. Under the strict definition with a true gap on both sides, about 0.16 times, or roughly one every six years on a given chart. With the gap read as a visible hole on the chart, 11 instances across 7,200 instrument years, which is too few to support any conclusion at all. The frequency is a property of the definition long before it is a property of the market.
Did the pattern beat a base rate in this test?
No. On a tape built with no directional memory the relaxed definition returned 0.04 percent below zero over twenty bars against 0.03 percent below zero for matched random entries, a difference of 0.01 percentage points at 0.17 standard errors. The strict definition came in 0.26 points above its base rate at 1.16 standard errors, with a ninety-five percent interval running from 0.18 points below zero to 0.71 above. Neither is distinguishable from nothing. All results are illustrative and simulated.
Why does it matter where the comparison group is drawn from?
Because a reversal pattern is selected using the price path immediately before it, so a comparison group drawn from that same stretch has been chosen by the pattern. Measured against random entries taken only from bars before the event, stars that followed a genuine adverse move came in 0.21 points above their comparison group. Widen the window to both sides and it becomes 0.09 below. Draw only from bars after the event and it is 0.33 below. Same detections, same outcomes, four different answers, and only the last one is free of the circularity.
Does a deeper third bar make the pattern better?
Not on data with nothing in it, and it actively hurt on the tape that did contain a real reversal effect. Splitting the population into disjoint bands by how far bar three closed back into bar one, the trend on the null tape was 0.07 percentage points per unit of depth at 0.26 standard errors, which is flat. On the control tape the trend ran the other way, 0.71 points per unit at 2.46 standard errors, because a deeper close means more of the reversal has already happened before you can act on it.
Should I only take a star when there is a trend to reverse?
It is the one condition the pattern's own logic demands, and it did not help here. Conditioning on a genuine adverse move of at least one average true range over the twenty bars before the pattern left 3,066 of 8,490 instances, and that subset came in 0.32 percentage points below its forward-only base rate at 2.39 standard errors, against 0.15 points above for the rest. On the control tape the same conditioning helped, which is what it should do when a real effect is present. The conclusion is about this tape and this definition, not about any exchange.
The target gets hit more often than chance. Is that not evidence?
It is the most convincing-looking number on the page and it is a measurement artefact. Projecting the height of the three bars from the entry, the target was reached inside twenty bars 47.6 percent of the time against 42.4 percent for matched random entries. But a star spans 1.55 times a typical three-bar range, and short-horizon volatility clusters. Draw the comparison entries from moments with a comparable recent span and the base rate rises to 46.0 percent, while the mean outcome stays unchanged. The hit rate measures how wide the pattern is, not what follows it.
Is a morning star more reliable than an evening star?
It looked that way and it was an artefact of arithmetic. Morning stars averaged 0.20 percent over twenty bars and evening stars 0.27 percent below zero, a gap of 0.47 percentage points at 2.91 standard errors, which is exactly the kind of number that gets quoted. Against their own direction-matched base rates the two came in at 0.02 points above and 0.06 points below, and the difference disappears. A long position and a short position on the same series do not have symmetric arithmetic, and a base rate is what removes that.
Why test this on generated data rather than Indian stocks?
Because a generated tape can be built with a known answer inside it and real data cannot. Two tapes were used here: one with no directional memory, where any measured edge has to be an artefact, and one carrying a genuine one-bar overreaction effect written purely in return space, which acts as a control proving the detector can find an effect when one is present. It found it at 9.34 standard errors. Without that pairing a null result is indistinguishable from a broken detector. Nothing here says what stars do on any particular exchange.
Method note
How the numbers on this page were produced
Every figure comes from one deterministic simulation, seeded so that it reproduces identically on each run. Two tapes were generated, each of 600 independent instruments of 3,000 daily bars, 1,800,000 bars in total. Volatility clusters and drift regimes switch on both, so long bars, small bars and quiet pauses occur naturally and are never inserted by hand; the three drift states are symmetric and equally likely, so the unconditional drift of the tape is zero. Opens jump around the previous close with a standard deviation of three tenths of the day's volatility, which is what gives the tape its gaps. The first tape contains no mechanism that could make direction predictable. The second adds one effect, defined entirely in return space and never in star terms, and exists to prove the detector can find a reversal effect when one is present. The generator is the one used by the companion pages on triangles and on flags, at the same seed, which was checked by generating its first 250 instruments with both pieces of code and comparing every value; a deliberately different seed was checked at the same time to confirm the comparison can fail.
Detection runs on the open, high, low and close of the generated bars using the thresholds in the definition table. Patterns are recognised on the third bar and positions opened at the next bar's open, so no result uses information that was not available at the time. Outcomes are measured over 20 bars. The base rate takes 250 bars strictly after each event on the same instrument, matched for direction and holding period, at two hundred draws per event; the reasons for excluding bars before the event, and the effect of not excluding them, are set out in the body of the page. Overlapping detections are discarded so that no two measured outcomes share a holding period. The base rate is a Monte Carlo draw rather than a closed form, so a cell recomputed in a separate pass moves in the second decimal place of a percentage point; where the same quantity appears twice on this page from two passes, that is why. Returns are shown before costs; an illustrative allowance for the round-trip charge stack would apply equally to both arms and would not change the difference between them, though it would sit well above the differences that were measured.
All results are illustrative and simulated. They are not a track record, they are not a forecast, and they are not an indication of what any pattern would produce in a live account or on any Indian security or index. The purpose is to establish which properties of a star pattern are consequences of its definition, which is a question about the definition rather than about any particular market.
Related