Skip to content
Candle Prediction

Method

How confidence is measured, and what would prove it wrong.

Candle Prediction makes two different claims, and they are easy to confuse. The first is about how often it is right. The second is about whether its confidence means anything. Only the second one is load-bearing, and it is the one this page is about.

The claim that is not being made

Directional accuracy on the reserved split is 50.6% at one candle, 50.2% at five, and 51.0% at ten. The project's own acceptance target before the run was 53%. It was not met, and the table below is the run that missed it.

Every one of those confidence intervals contains the baseline. Anyone selling you a chart-reading model that reliably calls direction is selling you something. We are not going to join them by rounding these numbers up.

The final gate

729 non-overlapping windows across 26 symbols, on a slice of history reserved from the start and opened once. Non-overlapping windows. Every interval is a cluster-bootstrap 95% CI. The plan's acceptance target was 53% directional accuracy, and this run did not meet it.

horizondirectiondriftECE80% coverage
1 bar0.506[0.474, 0.538]0.5110.057[0.031, 0.087]0.813[0.790, 0.837]
5 bars0.502[0.464, 0.533]0.4720.029[0.012, 0.067]0.820[0.795, 0.843]
10 bars0.510[0.473, 0.542]0.4880.018[0.004, 0.056]0.849[0.821, 0.874]

Source: report_final_gate_20260811_0551.md. Every interval is a cluster-bootstrap 95% CI, clustered so that windows from the same symbol on the same day cannot pretend to be independent evidence.

The claim that is being made

Expected calibration error runs 1.8% to 5.7% across the three horizons, against a target of 5%. The 80% range contained the outcome 81.3%84.9% of the time, against a promise of 80%.

That is the whole product. A model can sit a point or two above chance on direction and still be genuinely useful, so long as it tells you honestly which of its calls to trust — and refuses to dress up the ones it cannot.

25%25%30%30%35%35%40%40%45%45%50%50%chance · 33.3%
stated probability →  ↑ observed frequencyn = 729 windows · 1-candle horizonECE 5.7%7 of 8 bins inside the corridor

Stated probability against observed frequency at one candle. The corridor is where an honest model's observations should land 90% of the time at each bucket's sample size; the diagonal is where a perfect one would sit. With three possible outcomes the floor is 33.3%, which is the dashed line — not 50%, and the difference matters when the probabilities on offer are in the high thirties.

The three tiers

The engine's raw score is a calibrated probability that a card is right on at least two of its three horizons. The base rate for that event is 34.2%, so the raw score clusters near 34 by construction — which meant 96% of forecasts once displayed as "low". The gauge now shows a percentile of the engine's own output, so every band is populated and "strong" means something checkable: better supported than 90% of the forecasts this engine makes.

tierbandsharelanded rightvs base
WEAKbottom 25%25%30.9%−3.3
TYPICAL25th–90th65%34.5%+0.4
STRONGtop 10%10%40.2%+6.0

n = 7,884 validation predictions. Base rate 34.2%.

There were four tiers at one point. The fourth — the 75th to 90th percentile — realised 35.6% against a 34.2% base, a 1.4-point lift sitting well inside the noise. It was cut. Printing a tier that does not separate is exactly the overclaiming this product refuses everywhere else.

How the evaluation is kept honest

What would prove this wrong

If ECE climbs above 5% on new data, the calibration claim fails and the confidence score stops meaning what it says. If 80% coverage drifts outside roughly 76–84%, the range is the wrong width. If the STRONG tier stops separating from the base rate, the tiers are noise and should be removed — as the fourth one was.

None of those would be hidden. The measured record is the product.

What is deliberately missing

There is no directional call on stocks. Three separate approaches were tested and none beat a coin flip, so none shipped. There is no candlestick pattern edge either: eleven patterns over 1.7 million bars, and every apparent effect dissolved once the last bar's direction, size and close-position were controlled for. A placebo with the shape deleted reproduced the effect. The app shows the pattern counts and makes no claim about them.

There is no news or sentiment feed, because a sentiment label cannot be validated the way a forecast can, and printing one would be the same overclaiming in a different coat.


Candle Prediction estimates probabilities from chart images. It is not financial advice, not a recommendation, and not a signal service. Measured directional accuracy is 51–55%. Never risk money you can’t afford to lose.

← Back to Candle Prediction