Do not review an AI answer as one block of prose. Break it into the statements that would change what you do next. A suggested keyboard shortcut is low risk; a claim about a security setting, legal obligation, health choice, or payment is not. The consequence of being wrong should determine how much checking you do.

Highlight each material claim and classify it. Is it a fact that should have a source, a calculation you can reproduce, an inference from several facts, or a recommendation that depends on your priorities? This prevents a persuasive tone from hiding the jump between evidence and advice. If an answer cites a source, open it and check that it supports the precise claim, not merely the topic.

Prefer primary documentation, standards, product support pages, and original research for factual questions. Check publication or update dates where a product, policy, or price can change. Then test the scope: a statement about one edition, country, device, or account tier may not apply to yours. Record the URL and a short note for the claims you plan to reuse.

Keep uncertainty visible. “I could not verify this” is a better operational result than a false yes. For a high-impact decision, seek an independent second source or a qualified reviewer. If a model proposes an action involving credentials, files, money, or external messages, do not let the model’s confidence substitute for authorization.

This check will not make every answer complete. It does create a reliable stopping point: act only on the portions you can support, and turn the rest into a narrower question with better evidence.

Begin with the decision, not the answer

Before checking sources, write down what the answer would change. Are you choosing a keyboard shortcut, configuring a privacy setting, approving a purchase, or advising another person? The more costly or hard to reverse the consequence, the more evidence you need. This is not a rule that every ordinary question needs a research report. It is a way to avoid spending equal effort on a typo and a claim that could expose an account, money, or someone else’s work.

Make the next action concrete. “I need to understand cloud security” is too broad to check well. “I need to decide whether this account setting permits a contractor to share a folder outside the team” gives you a product, setting, audience, and consequence. A specific decision lets you notice when an answer has supplied generic background instead of the fact you actually need.

If the decision can wait, keep the uncertainty open rather than rushing to turn a plausible answer into a commitment. A good review often ends with a smaller question: which plan applies to this account, what date did the policy change, or who owns approval for this action? Narrow questions are easier to verify and easier for another person to audit.

Turn prose into claim cards

Models often produce fluent paragraphs that mix several kinds of statements. Separate the material ones into small claim cards. A card can contain the claim, its type, the consequence if wrong, the source needed, and your review status. For example: “This setting is available on our plan” is a product fact; “turning it on will reduce our risk” is an inference; “we should enable it today” is a recommendation. They require different checks.

Facts should point to evidence. Calculations should show their inputs and arithmetic. Inferences should state the facts they rely on and the alternative interpretation that might matter. Recommendations should name the priority they optimize—cost, privacy, speed, accessibility, or reliability—because reasonable people can choose differently even when they agree on the facts. This separation is a practical defense against confident language: a polished sentence cannot make an unsupported leap disappear.

Keep cards short enough to verify. If a claim includes a date, geography, product edition, and exception, split it. You will often find that only one part is documented. That is a valuable result. It tells you exactly what to ask next rather than encouraging you to accept the whole sentence because one fragment was true.

Prefer evidence that can answer the exact question

A highlighted document with source annotations on a desk

For product behavior, start with the official documentation or support page for the exact product and plan. For a standard, use the issuing body. For a research finding, locate the original paper or dataset where possible. For a rule that affects legal, medical, tax, or employment choices, find the competent authority and obtain qualified advice when the consequence calls for it. Secondary explainers can be useful orientation, but they should not be the final support for a high-consequence factual claim when a primary source exists.

Open the source; do not rely on the model’s citation label or search-result snippet. Check its publication or update date, author or publisher, and scope. A correct page about a consumer plan may not apply to an enterprise account. A U.S. policy page may not apply in another country. A support article may describe a feature that is disabled by an administrator. Write a short note beside the link explaining the exact sentence or section that supports your claim.

When sources disagree, do not average them into a confident answer. Find the reason for the disagreement: different dates, definitions, jurisdictions, versions, or incentives. The conflict may show that the original question was underspecified. Record it and decide which authority governs your particular situation. A disagreement you can describe is safer than an apparent consensus created by skipping the details.

Check scope, time, and boundary conditions

Many wrong answers are not fabricated; they are applied outside their boundary. Ask who, where, when, and under what configuration the statement holds. A workflow may work only for an administrator. A price may exclude tax or apply only to new customers. A security control may protect stored data but not a shared link. A tutorial may be accurate for the prior version of an app. Add these qualifiers to the claim card rather than treating them as footnotes.

Look for negative cases too. If an answer says a setting “prevents sharing,” search the official wording for exceptions such as existing links, external guests, mobile clients, or inherited permissions. If an answer proposes an automation, ask what it does on missing input, a denied permission, or a timeout. The purpose is not to prove every answer wrong. It is to understand whether the answer has enough boundary information to guide the action you intend.

This is particularly important when an answer touches tools. A model may correctly describe an API while overlooking that your credential lacks a required scope or that a call changes a live record. Confirm the current environment, account, and authorization separately. Knowledge of a feature is not authorization to use it.

Keep an evidence trail that another person can use

Save the URL, access date, claim, and a one-sentence result for information you will reuse. For a decision record, include the alternatives considered, the remaining uncertainty, and who made the final choice. This takes less time than reconstructing a web search after a question arises, and it makes handoffs much better. A colleague should be able to see what was verified without reading an entire chat transcript.

Use proportionate review. A low-risk tip may need only a quick check in current documentation. A purchase, public statement, policy change, or permission change may deserve a second independent source, a test in a non-production environment, or review from the responsible owner. NIST’s risk-management material is useful here because it frames trustworthy use as context and consequences, not as a universal confidence score.

OWASP’s LLM guidance is another reminder that an answer can be persuasive while being incomplete or unsafe. Do not allow a model’s explanation to substitute for permissions, validation, or a responsible reviewer when it proposes an action involving credentials, files, money, or external communication. The evidence check and the authorization check are different controls; use both.

Know when to stop and escalate

Stop when the evidence supports the narrow action, the scope matches, and the remaining uncertainty is acceptable for the consequence. Do not keep researching merely to remove every ambiguity. Conversely, escalate when authoritative evidence is unavailable, sources conflict on a material point, or the decision is outside your authority. A clear “not verified” is a successful output if it prevents an unsafe action.

The goal is not to make AI answers useless. It is to treat them as efficient starting points: a way to generate questions, candidates, and drafts that still pass through evidence and accountability before they affect real work.

A five-minute review pass

For an ordinary, low-risk answer, a fast pass is often enough: state the decision, underline the two or three claims that change it, open the best primary source, check the date and scope, and record what remains unknown. If the answer asks you to change a setting or send something outside your control, add an authorization check and a reversible test. This small routine is faster than treating every output as either unquestionable or useless.

Use the same routine when you write a prompt for someone else. Asking for sources, dates, assumptions, and uncertainty will not guarantee a correct output, but it produces a reviewable starting point. The final standard remains the same: the claim must be supported by evidence appropriate to the decision.

When you cannot find the evidence quickly, preserve the exact wording of the open question and the searches already attempted. That saves the next reviewer from repeating work and stops a provisional answer from gaining authority simply because it was written down first.

If new evidence changes a prior conclusion, update the claim record and communicate the change to people who acted on it. Evidence review is not only a gate before action; it is a maintenance habit when facts, policies, and products evolve.

That final communication should state what changed, why it matters, and whether anyone needs to reverse or repeat an action. A correction with an explicit operational consequence is easier to use than a quiet edit to an old answer.

Over time, these records reveal which questions repeatedly lack evidence. That is a useful signal to improve documentation, obtain a better source, or stop treating a fragile assumption as standard practice.

Limits

No checklist can create evidence where none exists, and primary sources can still be incomplete or change. The review effort should remain proportionate to the risk. For specialized high-stakes questions, use the qualified professional or authority responsible for the decision.