
70 / 78
Check the claims
Verify AI output: extract claims, rank by risk, check primary sources
Treat model output as unverified. Extract discrete claims, rank them by cost-if-wrong, and confirm the risky ones against primary sources.
The short answer
Treat model output as unverified draft text. Language models generate plausible continuations, not retrieved facts, so the verification step is yours: extract discrete claims, rank them by cost-if-wrong, and confirm the high-risk ones against primary sources you open yourself. Five minutes covers a typical answer you will send, sign, or pay for.
Why AI answers need checking
Chat tools write the likeliest next words. Usually that lands close to the truth. Sometimes it produces a sentence that sounds certain and is simply false. People call that a "hallucination," and it can happen with any tool.
OpenAI's own help pages say it plainly: "ChatGPT can make mistakes. Check important information." That line applies to Claude, Gemini, and every other chat tool too.
Floor vote
How often do you check an AI answer before you use it?
One vote per device. Change it any time.
Loading votes…
Step 1: Ask it to list its own claims
After you get an answer, send this:
List every factual claim in your answer as a numbered list:
names, numbers, dates, prices, laws, and quotes. Next to
each one, say how sure you are and where I could check it.
Now you have a checklist instead of a wall of text.
Step 2: Sort by "what if it's wrong?"
Go down the list and put each claim in one of three piles:
- Would cost money or cause trouble. Prices, deadlines, tax rules, legal steps, medical or safety advice. Check every one.
- Would be embarrassing. A wrong name, a wrong date, a quote from someone famous. Check before you publish or send.
- Background. General explanations you are only using to understand a topic. Skim, don't chase.
The five-minute check
- 01ListEvery claim, numbered
- 02SortCost if it's wrong
- 03OpenThe real source
- 04ClickEvery link
- 05CutWhat you can't confirm
Step 3: Check it against something you can open
- Go to the source. The government website for a rule. The company's own page for a price. The actual contract for what it says.
- Click every link. AI tools sometimes cite pages that don't exist, or that say something different. If the page doesn't say it, the claim doesn't count.
- Search the exact words. Put a quote or a specific phrase in quotation marks on Google. A real quote usually shows up with a named source.
- Check the date. Prices, laws, and app menus change. Look for when the page was last updated.
- Get a second source for numbers. Two places that don't copy each other.
Red flags that mean "check this first"
- Very exact numbers with no source.
- A quote from a famous person.
- A law or rule cited by section number.
- "Experts say" with no names.
- A confident answer about something that happened in the last few weeks.
Using a second AI to help
You can paste the answer into a different chat tool and ask, "What in this is likely to be wrong?" It sometimes catches mistakes. It can also be wrong in the same way. Treat it as a helper that points at what to check, not as the final judge.
The five-minute routine
- Ask for the claim list (1 minute).
- Sort into three piles (1 minute).
- Open sources for the risky pile (3 minutes).
If you can't confirm something that matters, take it out or ask a person who knows. "I couldn't confirm this" is a perfectly good reason to leave a line out.
Under the hood
Hallucinations come from the generation objective: the model predicts likely tokens, and a fluent false sentence can be highly likely. Retrieval and web search reduce the rate by grounding the answer in fetched text, but they add a new failure: a real citation attached to a claim the page does not support.
That is why the check is claim-level, not answer-level:
- Extraction turns prose into testable units.
- Risk ranking spends verification effort where errors are expensive.
- Citation checks confirm the cited page exists and actually states the claim.
- Recency checks catch stale training data on prices, rules, and UI paths.
A second model is a weak signal. Different models can share the same training gaps, so agreement between them is not proof.
Field check
Field check
Three questions. Honest answers. No score sent anywhere but this page.
01 / 03
Have you asked an AI to list the claims in its own answer?
Short close
List the claims, check the risky ones against a source you open, and click every link. If you can't confirm it, don't use it.
Keep these
The working rules
Model output is plausible text, not retrieved truth. Then: Claim extraction; Risk ranking; Primary-source verification; Citation and date checks.