AI and Scientific Discovery: More Paths, or a New Map?
Artificial intelligence is no longer only speeding up isolated scientific tasks. In some fields, it is beginning to change how candidate solutions are generated, which experiments are attempted and where researchers direct their attention. AI systems can propose hypotheses, search molecular and mathematical spaces, plan experiments, analyse results and generate synthetic examples. They are not replacing science as a whole, but they are redistributing effort within parts of it.
A 2025 perspective in npj Artificial Intelligence surveys the use of large language models across the scientific process, from literature review and hypothesis formation to experimental design and analysis. It also makes an important qualification: AI’s impact on fundamental science, defined as the discovery of new principles or laws, remains limited.
It is tempting to describe the emerging change as straightforward acceleration: faster discovery, cheaper experimentation, more progress. That is part of the story, but the more revealing question is what kind of science AI makes easier—and which parts remain difficult because they have not yet been turned into something a machine can represent, search, score or test.
More Paths Through an Existing Field
In an earlier essay on the low-hanging-fruit theory of innovation, I argued that progress depends not only on harvesting opportunities within an existing field but also on creating new fields in which different opportunities become possible. AI sharpens that distinction rather than resolving it. In many scientific domains, it makes far more of the existing possibility space reachable.
AlphaFold’s performance in CASP14 changed what researchers could expect from computational protein-structure prediction. Its structures were substantially more accurate than those produced by competing methods, and the system achieved near-experimental accuracy in a majority of cases. That did not make experimental structural biology unnecessary. It changed the cost and availability of a difficult intermediate result, allowing many investigations to begin with a useful predicted structure rather than requiring every structure to be determined from the beginning.
In materials science, the A-Lab autonomous laboratory combined computational screening, information extracted from scientific literature, machine learning, active learning and robotics. It generated synthesis recipes, performed experiments, analysed the products through X-ray diffraction and used the results to choose further attempts. The system did not abolish chemical difficulty. It reduced the cost of moving through a sufficiently well-defined set of candidate materials.
The pattern is important. AI can generate more candidates, compare more alternatives and eliminate some weak directions before scarce laboratory time is spent on them. Through the low-hanging-fruit metaphor, it lowers the ladder and makes more branches reachable. A field can be explored with greater intensity without the system necessarily having created the field in which that exploration occurs.
The Map Is Already a Scientific Claim
Current AI systems are particularly effective when a domain can be expressed as structured search: possible protein sequences and conformations, candidate molecules and materials, mathematical statements and proof steps, experimental parameters, computer programs, simulated environments or large bodies of scientific literature. Within such spaces, machines can compare possibilities at a scale no individual researcher could manage.
That process can still be creative. A search system may produce combinations no person had considered, reveal regularities that suggest a new abstraction or find an answer that changes how researchers understand the problem. Creativity does not require that every step occur outside an existing representation.
Its power nevertheless remains shaped by that representation. Before a search can begin, someone—or some earlier system—must decide what counts as a candidate, which variables describe it, which data are available, what constitutes improvement and which features can safely be ignored. These decisions are not neutral preparation. They embody a provisional theory of what the problem is.
This is where the map metaphor becomes more than decoration. AI may discover routes through a map that humans did not see. It can travel faster, compare more destinations and identify neglected regions. The map still determines which territory is visible to the search. If an important property cannot yet be measured, encoded or simulated, it may remain outside even an extraordinarily powerful exploration system.
The danger is not that AI cannot surprise us. It is that surprise within a representation may be mistaken for escape from the representation.
Discovery Needs Something That Can Say No
Generating a candidate is only one part of scientific discovery. Something must also be capable of rejecting it. Candidate generation attracts attention because it produces visible novelty: a structure, molecule, theorem, experiment or paper that did not previously exist. Novelty is not yet knowledge.
AlphaFold’s predictions became scientifically useful because they could be compared with experimentally determined structures and because the system produced confidence estimates that helped researchers judge when a prediction was likely to be reliable. Mathematical systems become more useful when a proof checker can establish whether each step follows. A proposed material becomes empirically significant when it can be produced and characterised.
A-Lab demonstrates both the power and the fragility of this validation loop. A published author correction reported that manual reanalysis confirmed 36 of the system’s 40 claimed successes, while four were inconclusive. It also clarified that materials described as novel were new to the prediction platform, not necessarily new to scientific knowledge.
The correction does not make the autonomous laboratory unimportant. It shows why the evaluator cannot be treated as an administrative detail added after the search. The method used to recognise success determines what the system believes it has discovered.
Different validators provide different kinds of resistance. A proof checker applies formal rules. A compiler can expose a program that does not run. A physical experiment allows matter to refuse the prediction. A sensor can produce measurements that contradict the model. Human experts can recognise that a technically successful result is trivial, unsafe, irrelevant or based on a misleading premise.
A second AI model can also evaluate an answer, but its judgement is not automatically independent. It may share the generator’s training material, assumptions and preference for fluent or conventional outputs. An automated loop can become highly efficient at satisfying its own evaluator without moving closer to the world.
Synthetic data makes the same boundary visible. Generated patients, driving scenarios, molecular configurations or simulated environments can supply rare cases and make repeated testing cheaper. They can be extremely useful, but they do not remove constraint; they relocate it into the simulation’s assumptions. An internally rich synthetic world can remain externally wrong.
As explored in the Journal’s essay on whether something outside the model still gets to answer back, machine-generated possibilities are most valuable when they remain connected to an independent source of correction. A system may generate its own candidates. It should not be the only authority deciding whether they correspond to reality.
When Exploration Begins to Redraw the Map
The boundary between exploring a field and expanding it should not be drawn too sharply. A sufficiently productive search can expose anomalies, regularities and repeated failures that force the underlying representation to change. Exploration can reveal that the map is incomplete.
AlphaFold did more than accelerate an existing manual procedure. It changed what biologists could reasonably assume was computationally available and made predicted structures accessible as inputs across a much wider range of research. Autonomous laboratories could produce a similar shift if they make closed-loop experimentation sufficiently cheap and routine. Regions previously neglected because each attempt required too much labour may become practical to explore.
The most ambitious systems attempt to automate much of the research cycle itself. The AI Scientist generates research ideas, searches literature, writes and executes code, analyses results, produces manuscripts and applies an automated reviewer. Its authors demonstrated the system in machine-learning research, where the object of study, experimental apparatus and many evaluation criteria already exist inside computers.
The achievement is real, but narrower than the phrase “automated scientist” suggests. One operating mode began from human-provided code templates, while a more open-ended mode worked with less scaffolding. When generated manuscripts were submitted to a workshop, researchers manually filtered promising outputs at several stages before choosing three papers. One passed review. The workshop’s acceptance rate was 70%, considerably above that of the main conference, and the authors explicitly concluded that the system did not yet meet top-tier publication standards consistently.
The result matters because it demonstrates that activities once treated as inseparable parts of scientific authorship can be assembled into an automated pipeline. It also reveals how much that pipeline depends on the surrounding research environment. Machine-learning experiments are unusually compatible with automation: the equipment is code, results arrive quickly and many objectives are numerical. The system is exploring a field whose map, instruments and feedback loops already exist inside machines.
Tractability Can Begin to Look Like Importance
Cheap exploration does not merely change what researchers can do. It may change what institutions reward. A programme that produces more candidates, faster experiments and clearer performance measures generates visible evidence of productivity. Work that requires a new instrument, an unfamiliar measurement, a difficult dataset or a theory without immediate computational tests can appear comparatively slow.
This risk is no longer purely speculative. A 2026 Nature analysis of 41.3 million papers found that researchers using AI tools were associated with higher output, more citations and faster progression into leadership roles. At the collective level, however, the volume of scientific topics studied narrowed by 4.63%, and engagement among researchers declined by 22%. AI-augmented work clustered around areas already rich in usable data.
The analysis is observational. It cannot establish that AI alone caused every difference, and researchers who adopt AI may already work in fields with different funding, publication or collaboration patterns. The tension it identifies nevertheless matches the institutional mechanism suggested by the article: tools can expand individual productivity while directing collective effort towards questions that are easiest to formalise and evaluate.
A data-rich field can produce more papers, prototypes and fundable milestones. Those visible successes attract resources, which improve its tools and datasets, making it still more compatible with automation. Other questions may remain difficult not because they matter less, but because the map has not yet been built.
The relevant policy problem is not that scientific institutions will consciously abandon important questions. It is that tractability can begin to resemble importance. Efficiency gains may broaden a field by freeing researchers to attempt previously impractical work, or they may intensify activity around programmes whose objectives are already legible. Which outcome dominates will depend on funding structures, disciplinary norms and whether institutions reward the creation of new measurements and concepts as seriously as the rapid production of searchable results.
More Paths, Better Maps and the Territory Beyond Them
The map metaphor requires three elements. There is the map: the representation that defines what can be searched. There are the paths: the hypotheses, candidates and experiments the system can explore. And there is the encounter with the territory: the proof, observation or experiment capable of showing that the representation was incomplete.
AI is already powerful at multiplying paths. It can also improve maps by identifying contradictions, unexplored regions and recurring failures. In some settings, it can participate in almost the entire loop: propose a candidate, design the test, operate the equipment, interpret the result and choose the next experiment.
The scientific value of that loop depends on where independent information enters it. A system that generates thousands of plausible answers and evaluates them through assumptions shared across the same models may move rapidly without learning much about the world. A system that proposes one imperfect answer and encounters a measurement capable of proving it wrong may have performed more meaningful science.
AI changes the speed, scale and texture of scientific exploration. It lowers the cost of generating hypotheses, prioritising experiments, automating measurements and comparing possibilities beyond unaided human attention. In some fields, this will be transformative, and productive exploration may reveal patterns that alter the original theory so deeply that the distinction between using a map and redrawing it becomes impossible to maintain.
It does not free science from the need for new concepts, instruments, measurements and institutions. Nor does it decide which questions are scientifically or morally important. A model may optimise a search after an objective has been specified; satisfying that objective efficiently does not make it the right objective.
The language of acceleration can therefore mislead. Speed is not direction. A faster search within a known space can be enormously valuable, but it can also make the existing representation feel more complete than it is.
AI may give science more paths through the present map. Some of those paths will expose errors in the map and help redraw it. Whether the technology helps researchers recognise the next field of discovery depends on something less spectacular than generation.
It depends on whether the world is still allowed to say no.
Comments
Post a Comment