Skip to content

Home / Research / Data

Training data is not a neutral mirror

When a system answers in a smooth voice, it is easy to treat the voice as a reflection of “what is online” or “what people think.” The training data behind that voice is not a mirror. It is a cut: some texts were collected, some were filtered, some were never in reach. This page describes the cut. It does not rank datasets, vendors, or any asset linked to them in the press.

The question is practical for a reader. If a paragraph about markets or technology cites “the data,” the next sentence should say which collection, which dates, and which exclusions. Without those, “the data” is a gesture.

A corpus is a decision

Someone, or some pipeline, chose sources. Public web pages, books under a licence, code repositories, forums, scientific PDFs: each source has a gate. A crawler that skips pages behind a login does not “see the whole web.” A filter that drops short lines or duplicated blocks changes what remains. The remainder is the corpus. Calling it neutral confuses “large” with “complete.”

Omissions are part of the definition. If a language, a region, or a decade is thin in the collection, the model’s continuations will be thin there too. That is not a moral score of the places that are missing. It is a coverage fact. Coverage is something a documentation note can state, and many notes state it only partly.

Corpus
The body of text used in training or in a stated slice of training.
Filter
A rule that drops or keeps items before or during training.
Omission
What the collection never contained. Silence in the output can come from here.

Dates are not optional

Text has a time. A page from 2019 does not describe a rule published in 2025. A model whose training stopped on a given month cannot have that later page in its parameters. Products sometimes add browsing or a fresh index. That is a second data path, with its own dates. Mixing the two paths in one sentence — “the model knows the latest figures” — hides which path was used.

Economic series are a sharp example, and only as an example of dating. An official statistic is revised. A sentence trained on an early release can sound current and still name an old vintage. The right habit is to read the vintage on the statistical office’s own table. The model is not that table.

cutoff later text
Documents to the right of the cutoff are outside the parameters unless another system fetches them.

Language and repetition

If most of the corpus is in one language, continuations in another language rest on a thinner pattern. Repetition matters as well. A claim that is copied across many pages can become easy to regenerate even when it was weakly sourced the first time. Frequency in the corpus is not a vote, and it is not a check. It is frequency.

This is one reason a fluent answer can agree with a popular mistake. The mistake was available, often, in the cut. Availability is not confirmation. A reader who wants the status of a claim still has to leave the continuation and open a source that states its method.

Evaluation data is a second cut

Teams often hold out a set of examples to measure the model. That set is also a selection: tasks someone wrote down, in a format someone chose. A high score on that set means the model matched those examples under those rules. It does not mean the model is generally right, and it does not mean a sentence about a company or a price is sound.

Published benchmarks are useful as descriptions of a test. They become misleading when a headline treats the score as a property of the world. The score is a property of the test. Ask what the test asked. Then stop. This catalog does not convert scores into rankings of products or into any financial conclusion.

  • Name the collection, or say that it was not named.
  • Name the cutoff, and any later tool separately.
  • Treat frequency as frequency, not as proof.

Close

Training data is a cut in time, language, and source. A mirror would show what was there without a gate. There is always a gate. Reading the gate is part of reading the answer.

Back to the research table