If AI Can Create New Knowledge, Why Does It Still Need Human Data?
Recent AI systems have begun producing results that look less like polished imitation and more like genuine discovery. In mathematics, specialised systems have found new constructions or solved difficult problems by combining learned search, synthetic examples and formal evaluation. Generative models can also produce images, programs and arguments that did not occur in their training data in precisely the same form.
This makes concern about limited human-generated training data sound almost contradictory. If a model can create novelty, why can it not manufacture the examples it needs and become increasingly self-sustaining?
The answer is not that synthetic data is inherently inferior, or that only humans can produce useful information. A learning system needs some source of correction that is independent of its current approximation. That correction may come from people, but it may also come from a proof checker, a compiler, a game, an experiment, a sensor or a physical process that refuses to behave as predicted.
The important divide is therefore not human data against machine data. It is between a loop that remains open to information it does not control and one that gradually trains on its own assumptions.
The Scarcity Is Not a Shortage of Words
The world is not about to stop producing text, images, audio or video. Generative systems can already manufacture more content than anyone could consume. Raw quantity, however, is a poor measure of how much information a dataset contains. A billion variations on familiar summaries do not replace a smaller number of original observations, experiments, arguments and experiences.
Training material also performs several different jobs. Large general corpora teach models how language, concepts and cultural references tend to relate. Synthetic examples can provide concentrated practice, fill a known gap or expose a model to deliberately difficult cases. Preference data records what evaluators consider useful, safe or persuasive. New observations update the system when the world, its language or its institutions change.
Generated data can serve some of these purposes exceptionally well. It can be balanced, labelled and targeted more precisely than material collected from the open internet. It can also be produced in domains where privacy, cost or scarcity makes real examples difficult to obtain.
What it cannot guarantee merely by existing is independence. A model generating another million examples from its current distribution may increase the number of rows in a dataset without adding much that the system did not already believe.
Synthetic Learning Works When Something Else Sets the Terms
Some of the strongest AI systems already rely heavily on machine-generated experience. AlphaZero learned chess, shogi and Go through self-play rather than by imitating databases of human games. Yet the system did not decide for itself whether a move was legal or whether a game had been won. The rules fixed the available actions, the board recorded their consequences and the result supplied a reward independent of how convincing the strategy appeared.
AlphaGeometry used a different form of synthetic learning. Its neural component was trained from scratch on millions of machine-generated geometry theorems and proofs rather than on translated human demonstrations. A symbolic deduction engine then constrained the search and extended valid reasoning. The examples were synthetic, but the geometry was not whatever the model wished it to be.
FunSearch made the relationship even clearer. A language model proposed programs, while an evaluator executed and scored them against a defined problem. Programs that failed to run, produced invalid outputs or performed badly were discarded. The system found new mathematical constructions because candidate generation was connected to a test that could reject fluent nonsense.
These examples demonstrate that AI can generate useful training experience and sometimes exceed known human results. They do not demonstrate improvement through unrestricted self-approval. Each system works inside a structure in which failure can arrive from outside the generator: a lost game, an invalid proof or an inferior score.
The synthetic loop is productive because it is not epistemically closed.
Model Collapse Is a Failure of the Pipeline
The copy-of-a-copy metaphor describes a genuine risk, but it is easy to apply too broadly. Research on recursive training with model-generated data found that models can progressively lose information from the original distribution. Low-probability cases disappear first; later generations become narrower, less varied and increasingly unlike the data from which the process began.
Collapse may therefore begin as sameness rather than immediate nonsense. Common patterns become smoother and more dominant, while rare phrasing, minority cases and unusual combinations are produced less often. Once those cases are scarce in generated output, the next model receives even less evidence that they exist.
This does not mean that every dataset containing synthetic examples is doomed. A separate study on accumulating real and synthetic data found that collapse was avoided in its tested settings when original material remained available and generated examples were added rather than repeatedly replacing the previous dataset. The result depends on the mixture, selection procedure and training workflow.
The safer conclusion is narrower than “synthetic data causes collapse.” Synthetic material can expand coverage, target weaknesses and create difficult exercises. The danger arises when a model’s approximation increasingly substitutes for the distribution it was meant to represent, especially when the original evidence becomes inaccessible or underweighted.
A Judge Is Useful Without Being an Oracle
One response is to make the loop more selective. Generate many candidates, ask another model to rank or correct them, retain the strongest and train on the filtered set. This can work well, especially when human evaluation is expensive and the judge is stronger than the generator on the relevant task.
A learned evaluator does not necessarily provide independent evidence. Research on large language models used as judges found substantial agreement with human preferences while also identifying position, verbosity and self-enhancement biases. A judge may reward polished structure, familiar reasoning or answers resembling its own output even when those qualities are imperfect proxies for correctness.
Formal validators occupy a different position. A proof checker applies explicit logical rules. A compiler can reject invalid code, and execution can reveal whether a program returns the required result. A game produces a winner that cannot be negotiated through rhetoric.
Even these checks are only as strong as the specification. A program can pass the supplied tests and fail on an omitted case. A simulation may execute perfectly while modelling the wrong world. An optimiser can become extremely effective at satisfying a measurable target that captures only part of what its designers wanted.
Validation is strongest when the generator cannot redefine success after seeing the result. It remains incomplete whenever the test itself leaves important parts of the problem outside the frame.
Human Data Is Neither Sacred nor Replaceable as One Category
Human-generated material is not automatically true or well grounded. People repeat rumours, imitate fashionable styles, preserve institutional errors and write summaries of summaries. The internet contained recursive distortion, advertising, propaganda and confident misunderstanding long before generative AI entered the chain.
The distinctive value of human data lies partly in its variety and causal independence. Different people encounter different languages, bodies, workplaces, environments and institutions. They record events because something happened, not because a model sampled a likely continuation. A witness, experimenter or patient can add information that was absent from the previous dataset.
Humans also remain part of the evaluation process where no complete objective function exists. There is no proof checker for whether a novel is moving, whether an explanation is fair, whether a medical trade-off is acceptable or which social compromise distributes risk justly. Models can approximate patterns in those judgments, but the patterns do not abolish the values and conflicts from which they were learned.
This does not mean that every future training example must be written by a person. It means that “human data” bundles together several things: linguistic material, fresh observation, preference, cultural change and judgment about what matters. Some can be generated or automated more readily than others.
Novelty Becomes Knowledge Through Resistance
The claim that AI creates new knowledge also requires a distinction between novelty and warranted result. A previously unseen image is novel. An original hypothesis or molecule may be promising. Neither becomes knowledge merely because it was not copied from the training set.
In mathematics, a new construction becomes valuable when its properties can be demonstrated and its improvement measured. FunSearch mattered because the programs it generated produced verifiable results on defined problems, not simply because their code had never appeared before.
Empirical science imposes a different test. A model may infer a promising material, drug candidate or biological structure from existing evidence, and that can be a genuine contribution to discovery. The claim acquires empirical weight only when synthesis, observation or experiment confirms that the world behaves as predicted.
This is the distinction developed in the Journal’s essay on artificial intelligence as a discovery engine. AI can widen the search through possible solutions without making proof, interpretation or experiment obsolete. It changes where scientists begin and which candidates they can afford to consider.
The crucial resource is therefore not content alone but resistance. A closed model can elaborate, recombine and derive consequences from what it already contains; mathematics shows how productive such operations can be. An empirical system cannot discover that a virus mutated, a bridge failed or a population behaved unexpectedly without receiving some signal that was not already implicit in its current model.
Reality does not care how fluent the prediction was. The molecule cannot be synthesised, the machine overheats or the measurement refuses to match. That friction is not an obstacle surrounding discovery. It is part of what distinguishes discovery from elaboration.
A Faster Route to the Outside
AI can still make the discovery loop far more autonomous. A model can generate hypotheses, rank candidates and design tests. Automated equipment can run experiments, sensors can measure the results and the data can return directly to the system that proposes the next attempt.
Humans may increasingly operate at the level of goals, safety, interpretation and the decision about which questions deserve resources rather than manually inspecting every candidate. The process could generate much of its own training experience and proceed at a pace no human laboratory team could match.
It would not be self-improving because it had escaped external data. It would be self-improving because it had built a faster route towards data produced by proof, execution and the physical world.
This connects to the broader question of whether AI merely opens more paths through existing research or begins changing the map of scientific attention. Systems that make certain questions easy to generate and evaluate will influence which problems appear tractable, fundable and important.
So can AI train itself? In bounded domains with reliable rules and evaluators, often to a remarkable degree. In open-ended domains, it can generate abundant practice and useful criticism, but plausibility cannot establish truth and learned preference cannot settle every value conflict.
The future may contain far more synthetic training data than human-authored material. That is not inherently a problem. The problem begins when generated data is no longer connected to fresh observation, preserved diversity, independent evaluation or consequences the model cannot explain away.
The decisive question is not who wrote the next example. It is whether something outside the model still gets to answer back.
Comments
Post a Comment