Skip to content

Home / Research / Limits

Why a fluent answer can still be wrong

Fluency is the feeling that a sentence fits. Generated text is built to produce that feeling. The score inside a language model favours continuations that look like competent prose. Competent prose can describe a real document, a muddled one, or a document that does not exist. The feeling does not sort those cases. This page explains the gap. It does not offer a technique that removes error, and it does not suggest that detecting error is a way to gain in a market.

Readers meet the problem in two places at once: technical explanations, and commentary that borrows a technical tone to talk about firms, prices, or policy. The tone travels. The evidence does not travel with it unless someone carries it.

Fluency is a surface

Grammar, connective tissue, and a calm adverb are cheap for a model. They are the regularities training rewarded. A wrong date can sit inside a clean sentence. A swapped name can sit inside a paragraph that otherwise tracks a real story. The surface is not a signal of the interior. If anything, a polished surface can delay the check, because nothing “sounds off.”

Confidence markers work the same way. “It is clear,” “studies show,” “the consensus is” are patterns. They do not query a study. When those phrases appear without a named paper, treat them as style. Style is allowed in a draft. It is not allowed to stand in for the paper.

Fluency
How easily the sentence reads. Independent of whether it matches a source.
Hallucination
A common name for a generated detail that is not supported. The name does not measure how often it happens.
Hedge
Wording that sounds careful. It can be copied without any actual limit being checked.

The gap between sentence and source

A source has a method: who counted, over what period, with what definition. A sentence can skip the method and keep the adjective. “Unemployment fell” without the survey, the month, and the revision policy is not yet a claim you can use. The model is good at the adjective. The survey note is good at the method. They are not substitutes.

Contradiction is another form of the gap. Two paragraphs in one answer can cite incompatible dates because each paragraph was a locally likely continuation. Local likelihood does not enforce a single timeline. If you need one timeline, build it from the documents, not from the adjacency of two smooth blocks.

fluent sentence source line, often empty
The top line can be complete while the line that would justify it is blank.

Omission is a kind of error

A wrong digit is easy to imagine. A missing condition is easier to miss. “The index rose” may hide that the rise is in one component, or that the series is not seasonally adjusted, or that the office flagged the month as provisional. The sentence is not false in every sense. It is incomplete in the sense that matters for reading. Incompleteness is still a failure of the answer if the answer presented itself as a briefing.

Generated lists do this often. They include the familiar items and drop the exception that the real document spends a paragraph on. The list looks orderly. Order is not coverage. Ask what a careful summary would be forced to mention, then see whether the answer mentioned it.

A check that does not pretend to be certainty

The practical check is small. Find one proper name, one date, and one number. Look them up in a source that is not the model. If any of the three fails, do not repair the answer by asking for a more confident rewrite. Go to the source for the rest of the paragraph as well. A single failure is evidence about the paragraph’s status, not a puzzle to smooth over.

This check does not make the remaining text true. It only refuses to let fluency be the test. There is no method on this page that turns a model into a reliable narrator of markets. Markets move, definitions differ, and losses are possible for anyone who acts. The educational point is prior to action: do not let the surface of a sentence decide for you.

  • Separate how a sentence sounds from what it cites.
  • Test one name, one date, and one number.
  • Treat a missing condition as an error of completeness.

Close

A fluent answer can be wrong because fluency is the objective the model was shaped to meet, and truth is not that objective. The source line is a different kind of object. If it is empty, the answer is still only prose.

Back to the research table