A language model predicts what comes next in a sentence. After "studies show that", a percentage is extremely likely to follow — and the model writes that percentage not because it knows it, but because a number fits there. Invented statistics aren't a malfunction; they're a natural consequence of how the tool works. That's why better prompting alone never fully solves it.
The most expensive kind of error
A typo annoys a reader; a fabricated figure costs credibility. It's also the error most likely to slip through, precisely because a number lends the sentence authority. And a number is the most quotable part of a text — so the mistake doesn't stay with you, it travels to whoever cites you.
The fix isn't in the model, it's in the check
Instead of asking the model not to invent, it's far more reliable to check what it produced against sources. The approach is simple: every numeric claim in the draft — percentages, amounts, dates, measures — is parsed out and looked up in the sources you have, namely research findings and the documents you uploaded. Claims with a match pass; claims without one get flagged.
The check itself has to be honest
There's a subtlety here: the check must not silently correct. Deleting an unmatched number, or swapping in a different one, hides the problem and adds a second layer of invention. The right move is to put the claim in front of the writer: "this figure has no match in the sources — fix it or remove it." The decision stays with a person, because a person is the one who knows which source is good enough.
Does a draft get weaker without numbers?
Counter-intuitively, no. When an unsourced number comes out, what replaces it is usually more concrete: an example, a description of a process, a customer situation. You actually have those and your competitor doesn't — whereas an invented percentage is the same everywhere. Sourced numbers stay, and now they're genuinely strong, because something stands behind them.