Blog

70 / 78

Article·Cross-vertical···3 min

EverydayA story, same factsTechnicalThe deeper cut

Check the claims

Verify AI output: extract claims, rank by risk, check primary sources

Treat model output as unverified. Extract discrete claims, rank them by cost-if-wrong, and confirm the risky ones against primary sources.

The short answer

Treat model output as unverified draft text. Language models generate plausible continuations, not retrieved facts, so the verification step is yours: extract discrete claims, rank them by cost-if-wrong, and confirm the high-risk ones against primary sources you open yourself. Five minutes covers a typical answer you will send, sign, or pay for.

Why AI answers need checking

Chat tools write the likeliest next words. Usually that lands close to the truth. Sometimes it produces a sentence that sounds certain and is simply false. People call that a "hallucination," and it can happen with any tool.

OpenAI's own help pages say it plainly: "ChatGPT can make mistakes. Check important information." That line applies to Claude, Gemini, and every other chat tool too.

Floor vote

How often do you check an AI answer before you use it?

One vote per device. Change it any time.

Loading votes…

Step 1: Ask it to list its own claims

After you get an answer, send this:

List every factual claim in your answer as a numbered list:
names, numbers, dates, prices, laws, and quotes. Next to
each one, say how sure you are and where I could check it.

Now you have a checklist instead of a wall of text.

Step 2: Sort by "what if it's wrong?"

Go down the list and put each claim in one of three piles:

  • Would cost money or cause trouble. Prices, deadlines, tax rules, legal steps, medical or safety advice. Check every one.
  • Would be embarrassing. A wrong name, a wrong date, a quote from someone famous. Check before you publish or send.
  • Background. General explanations you are only using to understand a topic. Skim, don't chase.

The five-minute check

  1. 01ListEvery claim, numbered
  2. 02SortCost if it's wrong
  3. 03OpenThe real source
  4. 04ClickEvery link
  5. 05CutWhat you can't confirm

Step 3: Check it against something you can open

  • Go to the source. The government website for a rule. The company's own page for a price. The actual contract for what it says.
  • Click every link. AI tools sometimes cite pages that don't exist, or that say something different. If the page doesn't say it, the claim doesn't count.
  • Search the exact words. Put a quote or a specific phrase in quotation marks on Google. A real quote usually shows up with a named source.
  • Check the date. Prices, laws, and app menus change. Look for when the page was last updated.
  • Get a second source for numbers. Two places that don't copy each other.

Red flags that mean "check this first"

  • Very exact numbers with no source.
  • A quote from a famous person.
  • A law or rule cited by section number.
  • "Experts say" with no names.
  • A confident answer about something that happened in the last few weeks.

Using a second AI to help

You can paste the answer into a different chat tool and ask, "What in this is likely to be wrong?" It sometimes catches mistakes. It can also be wrong in the same way. Treat it as a helper that points at what to check, not as the final judge.

The five-minute routine

  1. Ask for the claim list (1 minute).
  2. Sort into three piles (1 minute).
  3. Open sources for the risky pile (3 minutes).

If you can't confirm something that matters, take it out or ask a person who knows. "I couldn't confirm this" is a perfectly good reason to leave a line out.

Under the hood

Hallucinations come from the generation objective: the model predicts likely tokens, and a fluent false sentence can be highly likely. Retrieval and web search reduce the rate by grounding the answer in fetched text, but they add a new failure: a real citation attached to a claim the page does not support.

That is why the check is claim-level, not answer-level:

  • Extraction turns prose into testable units.
  • Risk ranking spends verification effort where errors are expensive.
  • Citation checks confirm the cited page exists and actually states the claim.
  • Recency checks catch stale training data on prices, rules, and UI paths.

A second model is a weak signal. Different models can share the same training gaps, so agreement between them is not proof.

Field check

Field check

Three questions. Honest answers. No score sent anywhere but this page.

01 / 03

Have you asked an AI to list the claims in its own answer?

Short close

List the claims, check the risky ones against a source you open, and click every link. If you can't confirm it, don't use it.

Keep these

The working rules

Model output is plausible text, not retrieved truth. Then: Claim extraction; Risk ranking; Primary-source verification; Citation and date checks.

Mark

Saved on this device. No account.

Pass along

Comments

Comments

A short note on the job in this piece. Name optional. Saved on this page, not a profile.

Loading comments…