Somebody publishes a chart every year showing that a particular month, day or hour behaves differently, and the chart is accurate about the past. Almost none of these survive contact with the future, and the reason is a statistical mechanism that produces convincing patterns from data containing nothing at all.
Take any series of random numbers, long enough, and search it for regularities. You will find them: months that outperform, days of the week that are negative, hours that trend. The finding is guaranteed rather than surprising, because a long enough random series contains every pattern somewhere, and searching is what surfaces it.
This is the fact that most calendar analysis is built on top of without acknowledging. A chart showing that a particular month has averaged a positive return over fifteen years is a true statement about those fifteen years. Whether it is evidence of anything depends entirely on how many other months, and other rules, were examined before that one was published.
The distinction matters because both look identical when presented. A genuine effect and a coincidence produce the same chart, the same average, and the same confident caption. Nothing in the presentation separates them, and the thing that would separate them is information about the search rather than about the data.
Any long enough series contains every pattern somewhere. Finding one proves you searched, not that the pattern is real.
The formal name is multiple testing, and the arithmetic is unforgiving. If you test one rule and adopt a threshold that a coincidence would clear only one time in twenty, you have a reasonable filter. If you test a hundred rules against the same threshold, you should expect around five to clear it purely by chance, and those five will look exactly like discoveries.
Calendar analysis is close to the worst case for this, because the space of rules is enormous and free to search. Twelve months, seven weekdays, thirty-one dates, twenty-four hours, and every combination of them, applied to every asset and every window, is tens of thousands of testable rules. Some of them will look extraordinary, and the extraordinary ones are the ones that get published.
The correction is conceptually simple and rarely applied: the threshold has to be raised in proportion to how many rules were examined. A finding that would be impressive after one test is unremarkable after ten thousand, and almost nobody publishing a calendar chart states how many they looked at before choosing that one.
The well-known calendar claims have been examined repeatedly by people with no stake in the answer, and the results follow a consistent shape. Most disappear entirely when tested on data from after they were published. Some survive in a weakened form that is smaller than transaction costs. A small number persist and turn out to have a structural cause rather than a calendrical one.
The last category is the interesting one because it explains why the whole field keeps producing candidates. When an effect is real, it is real because something is happening at that time rather than because of the date itself: institutions rebalancing on a schedule, tax years ending, reporting periods closing, or liquidity thinning during holidays. The calendar is a proxy for the cause rather than the cause.
That reframing is the useful one. A claimed calendar effect with no proposed mechanism is a pattern somebody found by searching. A claimed effect with a mechanism that can be stated and checked is a hypothesis about market structure that happens to have a date attached, and it can be evaluated on the mechanism rather than on the chart.
There is a genuine periodic effect in this market and it is not calendrical. Attention arrives in waves: coverage increases, new participants enter, activity and volatility rise together, and then attention decays. That cycle is observable in search interest, in new account creation, and in the composition of flow.
It gets mistaken for seasonality because waves have a rhythm and rhythms invite calendar explanations. But the driver is a feedback loop between price movement and coverage rather than a date, which means it does not repeat on a schedule and cannot be anticipated by looking at a month. It can only be observed while it is happening.
What it does produce is a reliable structural observation: during high-attention periods, the composition of participants shifts toward those trading a narrative rather than a spread, leverage rises, and the market becomes more prone to sharp moves in both directions. That is a statement about fragility that follows from the mechanism, and it requires no view on direction or timing.
One calendar effect in this market does hold up, and it holds up because the mechanism is obvious. Liquidity has a weekly shape: it is thinner at weekends and during the overnight hours of the largest participating regions, because the people and firms providing it are not working. That is not a pattern found by searching, it is a consequence of human schedules.
It is measurable directly rather than inferred from returns. Depth in the order book, spreads, and the cost of executing a given size all vary across the week in a way that repeats and that can be observed on any instrument without any statistical machinery at all. The effect is on execution cost rather than on direction.
That distinction is what makes it usable while monthly claims are not. Knowing that a particular hour costs more to trade in is a fact about your costs that you control by choosing when to trade. Knowing that a particular month has historically been positive is a claim about direction that you cannot act on without taking the risk the claim describes.
The weekly pattern is real because the mechanism is people's working hours. The monthly patterns are claims about direction with no mechanism at all.
A pattern that is genuinely exploitable stops being exploitable once enough people know about it, because acting on it moves the price to the point where it no longer pays. This is not a theoretical objection, it is the documented fate of most published market anomalies across every asset class.
The studies that track this find a consistent shape: effects shrink substantially after publication, and the shrinkage is larger for effects that were easier to trade. Some of that is arbitrage removing a real effect, and some is that the effect was never there and the original finding was a coincidence that regression to the mean corrected.
Either explanation leads to the same practical conclusion. A calendar effect you read about in an article is either false or already competed away, and in both cases acting on it is a poor use of capital. The ones worth anything are the ones nobody has published, which is a set you can only join by doing your own work.
The procedure is short and it settles most claims. Split your data into two periods before looking at anything. Develop or evaluate the rule using only the first, then apply it unchanged to the second and record the result. If it works on the first and not the second, which is the usual outcome, the rule was fitted to noise.
Then check whether it survives costs. A calendar effect worth a fraction of a percent per occurrence is worth nothing if executing it costs more, and many published effects are of exactly that size. Subtract a realistic round-trip cost drawn from your own records rather than an assumed one, and a substantial share of candidates disappear at this step.
Finally, count. Before believing a rule, ask how many rules you considered before arriving at it. If the answer is many, your threshold for believing this one has to be correspondingly higher, and the honest version of that adjustment eliminates almost everything.
Occasionally a rule survives all three checks, and the correct response is more restrained than it feels. A rule that works out of sample, survives costs, and was not selected from a large search is a candidate rather than a discovery, and the next step is to ask what mechanism would produce it.
If a mechanism can be stated, the rule becomes considerably more credible, because a testable cause is evidence in a way that a pattern is not. If no mechanism can be stated, the rule remains a coincidence that happened to survive two tests, which is exactly what a coincidence in a long series does some of the time.
And whatever the answer, size accordingly. A statistical edge measured on limited data has an uncertainty around it that is usually larger than the edge, and a position sized as though the estimate were exact is a position sized on a number that has error bars nobody drew.
None of this means the calendar is irrelevant. It means the useful applications are about cost and risk rather than about direction, and there are several.
Knowing when liquidity thins lets you avoid executing then, which is a saving you capture with certainty rather than a bet you might win. Knowing when scheduled announcements occur lets you avoid holding a position through an event you have no view on, or size for the volatility it will produce. And knowing the weekly shape of the market lets you plan around it rather than being surprised by it.
All three are the same kind of knowledge: descriptive facts about when the market behaves differently, used to manage execution and exposure. That is a smaller claim than predicting returns from a month, it is supportable, and it is worth more in practice than any of the patterns that get charted.
Because any sufficiently long series contains every pattern somewhere, and searching is what surfaces it. A chart of a month's historical average is a true statement about the past. Whether it means anything depends on how many other rules were examined before that one was chosen, which is never stated.
Testing many rules against the same threshold. If a threshold would be cleared by chance one time in twenty, testing a hundred rules produces about five apparent discoveries from nothing. Calendar analysis has tens of thousands of testable rules, which makes it close to the worst case.
A few, and the ones that do turn out to have a structural cause rather than a calendrical one: scheduled rebalancing, tax years ending, or liquidity thinning during holidays. The date is a proxy for the mechanism. A claimed effect with no proposed mechanism is a pattern somebody found by searching.
Yes, in liquidity rather than direction. Depth is thinner at weekends and during the overnight hours of the largest participating regions, because the people providing it are not working. It is measurable directly on any instrument and it affects your execution cost, which you control by choosing when to trade.
Two reasons that lead to the same conclusion. Acting on a genuinely exploitable pattern moves the price until it no longer pays, and many published effects were coincidences that regression to the mean corrected. Studies tracking this find effects shrink substantially after publication, more so for the easily tradeable ones.
Split the data into two periods before looking at anything. Evaluate the rule on the first only, then apply it unchanged to the second. Then subtract a realistic round-trip cost from your own records. Then count how many rules you considered before arriving at this one, and raise your threshold accordingly.
It is a real periodic effect and it is not calendrical. Coverage increases, participants enter, activity and volatility rise, and attention then decays. The driver is a feedback loop between price and coverage rather than a date, so it does not repeat on a schedule and can only be observed while it is happening.
No, but the useful applications are about cost and risk rather than direction. Knowing when liquidity thins lets you avoid executing then, which is a certain saving. Knowing when scheduled announcements occur lets you size for the volatility. Both are supportable in a way that predicting a month's return is not.