Once an answer passes the usual check for made-up facts, the mistakes left are the quiet ones: a real-looking source, a number in the right range, a summary that reads faithful, reasoning that flows and is still wrong. You cannot catch those by verifying a citation, because nothing looks off. You catch them by pushing back on the answer until it shows you where it is soft.
Why the mistakes got harder to see
Ethan Mollick, who writes One Useful Thing, puts it plainly: "no matter how good the AI is, it will still make errors and mistakes and still give you confident answers where it is wrong." The trouble is that the confident-but-wrong answer now looks exactly like the confident-and-right one. In one study he ran with colleagues on the "jagged frontier" of AI ability, people given a task that sat just outside what AI does well actually did worse with it than without, because the wrong answer arrived just as polished as a good one.
There is a second trap: AI wants to please you. A Stanford study of eleven chatbots found they endorsed the user's actions far more often than a person would, and backed harmful choices almost half the time. Lead author Myra Cheng put it bluntly: "By default, AI advice does not tell people that they're wrong nor give them tough love." So the answer that agrees with you is the one to trust least.
The quiet errors to watch for
- It agrees with your framing. Ask "was I right to do X" and it leans towards yes. The agreement is a feature of the model, not evidence you were right.
- Confidence hiding a guess. The tone is just as sure whether it knows or is pattern-matching. Sureness tells you nothing about accuracy.
- Right-shaped, wrong answer. The reasoning flows and the number looks plausible, but a step in the middle is off, or the task sat on the wrong side of what AI can do.
- Hidden assumptions. The answer quietly assumes something you never said, and the whole conclusion rests on it.
- A summary that drifts from the source. You gave it a document, and its summary is faithful in feel but adds or softens something the source never said.
The method: interrogate the answer, do not skim it
Mollick's advice is to engage the AI in a back and forth: "Don't just ask for a response, push the AI and question it." That is the whole shift. Instead of scanning the answer and nodding, you make it turn on its own work. A handful of probes do the job.
- Make it argue against itself. Ask for the strongest honest case that its answer is wrong. A model that cannot find a real objection to a shaky claim tells you something too.
- Check it did not just agree with you. Ask where it leaned towards what you wanted rather than what is true, and what the answer would be without your framing. This is the fix for the model's habit of telling you what you want to hear.
- Surface the assumptions. Ask what it assumed that you did not give it, and what happens to the answer if each assumption is false.
- Ask for confidence, and what would change it. Make it separate what it knows from what it is guessing, and name the one fact that would flip each claim.
- Show the one step that matters. Have it lay out the single calculation or reasoning step the conclusion hangs on, in full, so you can re-check that one thing yourself.
- Keep it in the source. If it used a document you gave it, make it redo the answer using only that source and mark anything the source does not cover.
The paste-in that runs all five
Paste this straight after an answer you are about to rely on. It borrows from Nafiul Hasan of Prompt Architects, who calls it red-teaming your own prompt: attack the answer before you trust it.
You are checking your own last answer before I rely on it. Do not defend it, interrogate it. The risk is not an obvious mistake, it is a confident answer that looks completely fine and is quietly wrong, so treat "it reads well" as no evidence that it is right. Be blunt with me. WHAT THIS IS FOR (one line from me, so you know how hard to push) [For example: I am about to send this to my boss, act on it with money, publish it, or I am just curious. The higher the stakes, the harder you dig.] WORK THROUGH THESE IN ORDER 1. ARGUE AGAINST YOURSELF. Build the strongest honest case that your answer is wrong or incomplete. Use only real considerations, no strawmen. If the case against it is weak, say so plainly. 2. CHECK WHETHER YOU JUST AGREED WITH ME. Did any part lean towards what I seemed to want rather than what is true? Point to where you went along with my framing, and say what the answer would be without it. 3. SURFACE YOUR ASSUMPTIONS. List every assumption the answer depends on that I did not give you and you did not verify. For each one, say what happens to the conclusion if it turns out false. 4. RATE YOUR CONFIDENCE, HONESTLY. For each main claim, tell me how sure you are and the single fact that would change your mind. Flag anything you are pattern-matching rather than actually confident about, and mark any claim you could not back without checking a source. Your confident tone is not a signal, so do not let it stand in for evidence. 5. SHOW THE ONE STEP THAT MATTERS. Lay out, in full, the single calculation or reasoning step the whole conclusion hangs on, so I can re-check it myself. If it is a number, show the working. 6. STAY IN THE SOURCE. If your answer used a document or link I gave you, redo it using only that source. For anything the source does not actually cover, write NOT IN SOURCE instead of filling the gap. THEN TELL ME PLAINLY - The one or two things most likely to be wrong here, ranked. - The single fact I should go and check myself before I rely on this. Do not reassure me and do not soften this to keep me happy. If any part is shaky, uncertain, or a guess, say so in plain words. I would rather know now than later.
The honest catch
Interrogating an answer shows you where it is soft, but it cannot make a wrong answer true, and a model can red-team itself and still miss its own error. So for anything that matters, the last step is always yours: confirm one real fact against a real source with your own eyes. The probing does not replace that, it tells you exactly where to point it. When the answer is a pile of specific facts and citations rather than reasoning, start with the faster made-up-answer check instead.
Try it on one answer now
Take the last AI answer you actually used for something, a plan, a summary, a recommendation, and paste the prompt under it. Watch what it flags when you tell it plainly not to reassure you. If you want a tool that argues the other side of a decision from the start, get AI to pressure-test your thinking does exactly that.
A few quick questions
How is this different from checking for made-up facts?
Checking for made-up facts catches the obvious ones: a fake citation, an invented name, a number you can look up. A subtle error passes that check, because the source looks real and the number looks reasonable. You catch it by interrogating the whole answer, not by spot-checking one claim.
Why does AI agree with me so much?
Models lean towards telling you what you want to hear. A Stanford study of eleven chatbots found they endorsed the user's actions far more often than a person would. So if you ask whether you were right to do something, treat a yes with suspicion and ask it to argue the other side.
Does interrogating the answer make it correct?
No. Probing an answer surfaces the weak spots, but it cannot make a wrong answer true. For anything that matters, you still confirm one real fact against a real source yourself. The probing tells you where to look.