How to read an AI confidence score without mistaking it for proof

An answer with an 80% confidence label feels more precise than an answer that says probably. But precise presentation does not establish what the number measures. Before using it, identify the claim, the deadline and the evidence that could settle it. This guide shows how to read ECHO's confidence label and turn a broad scenario into a question you can check.
What the confidence label means in ECHO
Ask ECHO generates a scenario from your event or question. Its response can include possible consequences, a time horizon, projected market moves and an AI confidence value. The current analysis asks the model for a subjective confidence estimate. It does not calculate that score from a published record of comparable forecasts and their resolved outcomes.
Read the percentage as part of the generated analysis. It does not mean that the proposed event has been observed, that the model independently verified every premise, or that a trade has that chance of making money. The projected moves and timings also remain scenario estimates. Their specificity is a reason to inspect the assumptions, not evidence that the estimates have been measured.
Decide which claim could actually be resolved
A score beside 'the company benefits' leaves too much unstated. Does benefit mean more orders, higher revenue, a larger profit margin or a rising share price? Those outcomes can diverge. Choose one before evaluating a prediction, and separate it from the event that motivated the question.
Metaculus's question-writing guidance offers a useful discipline: define the terms, the evidence source and the conditions for resolving the question. You can apply that discipline to your own notes without assuming that ECHO operates a forecasting tournament or automatically tracks resolutions.
- Claim: specify one observable outcome, not a general direction of benefit.
- Horizon: give the last qualifying date and the relevant time zone.
- Evidence: name the record that would establish the outcome.
- Ambiguity: decide how delays, missing records and partial completion will be treated.
An example: the permit is not the opening
Imagine a fictional company, Rowan Transit, receiving permission to build a station. Suppose the permit is the only document you have checked. Permission is the supported fact. Construction finishing, passenger services beginning and the company earning additional profit are separate possible outcomes. None follows automatically from the permit.
For a personal forecasting exercise, define one question: will Rowan's dated operating bulletin confirm that the first passenger service departed from the station by June 30, 2030? This is an invented example, not a real project or ECHO forecast. A planned timetable alone would not confirm that a train departed. Set the departure deadline, time zone and acceptable confirmation before evaluating the outcome.
Now imagine 30 fictional forecasts of clearly defined station openings, each assigned 70%. At review time, 20 are resolved: 14 opened by their deadlines and six did not. Ten remain pending. Among resolved cases, 14 divided by 20 is 70%. Dividing by all 30 would incorrectly count unresolved cases as failures.
| Outcome | Count | Treatment |
|---|---|---|
| Opened by the defined deadline | 14 | Resolved yes |
| Missed the defined deadline | 6 | Resolved no |
| Outcome not yet resolved | 10 | Keep pending |
A matching percentage is a starting point for evaluation
Calibration asks whether stated probabilities correspond to observed frequencies across comparable predictions. Guo and colleagues studied this problem for neural-network classification. Their research helps explain why a model's confidence and its empirical reliability are different properties. It did not test ECHO or establish how an ECHO scenario will perform.
The fictional 70% result is too limited to establish a reliable forecasting system. The sample is small, pending cases might differ from resolved ones, and the result says nothing about other score ranges or horizons. Keep the original forecasts, resolution rules and pending cases visible. Do not quietly remove difficult outcomes or redefine success after seeing what happened.
Use the answer to find the next piece of evidence
When reading an ECHO response, mark which statements come from a source you have opened and which are proposed mechanisms. A permit may be verified while funding, construction capacity and a service date remain unknown. A high confidence label does not fill those gaps. If a projected price move is the headline, inspect its horizon and the assumptions connecting the event to that move.
A focused prompt is more useful than asking whether the company is a winner: 'Illustrative: Rowan Transit has a station permit. What evidence would support passenger service by June 2030, and what could delay it?' Compare the answer with documents you can actually obtain.
Keep the score beside the claim it describes, not beside an unrelated decision. When new evidence arrives, preserve what you previously believed and why you changed it. The useful result is a clearer research question and a visible evidence gap, even when the confidence number offers little additional help.
Questions about a source or explanation? Read our editorial approach and report a correction. AI scenarios remain possibilities to investigate.
Explore an event with Ask ECHO