Say which part of the answer is uncertain.
“The new service probably improves access.” A reader still needs to know what “probably” covers. Did people say it was easier to use? Did more people complete the process? Are you unsure whether the result will hold elsewhere? Adding a cautious word does not explain the weakness in the evidence.
Split the statement into what you know and what remains open. For example: “People we interviewed found the new form easier to follow. We do not yet know whether more applicants finish it.” The reader can see both the finding and the limit, without having to interpret a general confidence label.
When answers can be checked, you can also compare confidence with outcomes. SimpleQA uses short factual questions to examine this relationship.[1] If many comparable answers assigned 70% confidence are correct about 70% of the time, that level of confidence is well calibrated. One correct answer does not establish that track record.
Tell the reader what is uncertain and which evidence would help.
Explain why you are unsure.
Different gaps call for different work. Missing records call for more evidence. An unreliable measure calls for checking how it was collected. Several plausible explanations call for a study that can distinguish them. Evidence from another country calls for checking whether the local conditions are similar. Putting all four problems under “low confidence” hides what to do next.
AI introduces another question: would a different prompt, source selection, or model version change the answer? Repeated answers can reveal inconsistency, but a model can also repeat the same mistake. Research has found that models can assess some of their own uncertainty, while struggling to keep that assessment well calibrated on new tasks.[2] Treat the model’s confidence as something to test.
Attach the limitation to the relevant claim. An official notice may establish when a policy began, while leaving open how consistently it was implemented. Interviews may suggest that access improved, without telling you how many people benefited. Confidence in the date should not make the implementation or impact sound equally well established.
Compare judgments about similar questions.
A good record on one kind of question does not establish reliability on every kind. Checking a date in an official document differs from predicting demand for a new product. If you combine them, the easy factual checks may hide poor forecasts.
The set of comparable judgments is sometimes called a reference class. Choose comparisons that matter for your work: similar questions, sources, methods, settings, and prediction periods. Start with enough cases to learn something, then look for meaningful differences. Very small groups give unstable estimates; very broad groups can hide where confidence fails.
For a forecast, record the date and decide in advance what outcome will count as correct, when it will be checked, and by whom. If a broad conclusion cannot be judged that way, explain the strength and limits of its evidence in words. You do not need to turn every interpretation into a yes-or-no prediction.
Use percentages when you can check what they mean.
A number such as 73% can look more considered than “likely” even when both express the same untested intuition. Use a percentage when the event is clear and you can explain the basis for the estimate. To assess calibration, you also need a record of comparable judgments and their outcomes.
Calibration is one part of a useful forecast. If half the events occur, giving every event a 50% chance can match the overall frequency while helping little with individual cases. Forecasting research also considers sharpness: how concentrated a forecast is, assessed alongside calibration.[3] For yes-or-no events, confident estimates are useful when they distinguish cases and hold up against outcomes. The hypothetical chart below shows estimates that are more confident than their results justify.
Does confidence match the results?
Hypothetical answers grouped by stated confidence
A scoring rule helps compare repeated forecasts. For a binary event, the Brier score averages the squared gap between the predicted probability and the outcome, coded as 1 or 0. An 80% prediction scores 0.04 if the event happens and 0.64 if it does not; lower is better. It belongs to a family of proper scoring rules that reward honest probability estimates in expectation.[4] Logarithmic scoring penalizes confident mistakes especially strongly. Review scores alongside the chart and results for important groups, so one average does not hide a recurring problem.
A precise-looking percentage still needs a basis and a way to be checked.
Give words such as “likely” a clear meaning.
Words can be more useful than a percentage when describing a body of evidence. “Supported by two independent records, but not yet checked in practice” tells the reader why you believe a claim and what is missing. “High confidence” on its own does much less.
Models can learn to express confidence in words that correspond reasonably well to observed probabilities in the tasks studied.[5] Readers may still interpret those words differently. A multinational study of IPCC probability statements found that people’s interpretations often differed from the intended meanings.[6] A familiar word is not necessarily a shared definition.
If you use repeated labels such as “supported,” “provisional,” or “unresolved,” briefly define them. Say whether they describe evidence quality or the chance of an event. If “likely” stands for a numerical range in your method, show that range where the term appears. Keep the supporting explanation close enough for the reader to judge it.
Separate how likely something is from what to do about it.
A 60% chance of a small inconvenience and a 60% chance of irreversible harm call for different responses. Confidence alone cannot determine the action. Consider the cost of being wrong, the cost of waiting, whether the action can be reversed, and what safeguards are available.
Write the estimate and recommendation separately. In a hypothetical project, “We estimate a 65–75% chance the pattern will recur” describes a belief. “Try a limited pilot because the downside is contained and results will arrive within two weeks” explains an action. Someone can then question either the estimate or the reasoning behind the pilot.
A weak signal may justify another interview or a small test without justifying a full rollout. Missing support may justify pausing an expensive expansion without proving that it could never work. Ask which feasible observation would most affect the choice. That gives the next research step a reason to exist.
Make “I don’t know” useful.
When evidence is insufficient, say why. A missing record calls for retrieval; an ambiguous question calls for clarification; conflicting findings call for comparison; a judgment outside the method’s scope calls for another approach or a qualified reviewer. These are different next steps, even if none allows a confident answer yet.
One research approach compares the meanings of several AI answers to detect inconsistent, potentially invented claims.[7] Another, developed for robot planning, uses conformal prediction—a method calibrated on examples—to decide when to ask for help, with statistical guarantees under its stated assumptions.[8] These methods address particular problems. Repeated answers can still share a false belief, and a guarantee does not automatically carry over to a different population or task.
Check both how often the system answers and how often those answers are wrong. Answering fewer questions is useful only if the reduction in errors justifies the extra referrals. Inspect which cases are being deferred, especially by language and user group. Where coverage is poor, improve evidence or human support rather than force uncertain answers. Give deferred cases a clear route, a responsible person, and an expected response time.
A useful “I don’t know” explains what is missing and what to check next.
Use probabilities for clearly defined events and evidence descriptions for claims that do not have a simple yes-or-no outcome. In either case, explain the basis and the next check.
Say what would change the judgment.
An estimate reflects what is known at a particular time. Name the evidence that would strengthen it, weaken it, or reverse it. A new record, a result from a different group, or a failed prediction may matter more than another general review of the topic.
Keep a dated record of important estimates, their evidence, the comparable cases used, and who will check the outcome. When results arrive, compare them with the original estimate before revising the explanation. Retain the earlier version so you can see whether confidence was justified and what changed.
Finish with what the evidence supports now, what remains uncertain, and what action makes sense while that uncertainty remains. Name the next useful check and when the judgment should be revisited. Readers can then act with a clear view of the limits and recognize when to change course.

