Around 2,000 years ago, the Greek biographer Plutarch wrote about a 30-oared ship that was said to have carried the hero Theseus and a group of young Athenians home from Crete. To preserve the vessel, its Athenian caretakers removed the old wood as it decayed and replaced it with stronger timber. They continued, plank by plank, until the repairs produced a philosophical problem: If none of the original material remained, was it still the Ship of Theseus?

The ship became the namesake of a thought experiment, applied to everything from bands that replace their members to classic cars rebuilt with new components. I have found a new application in the growing role of artificial intelligence in knowledge production: In this version, science is the ship, its human-driven practices are the planks, and AI provides the replacement material. Science’s silhouette stays recognizable even as the activity inside it becomes increasingly automated.
Much of the anxiety surrounding an AI takeover assumes a cinematic event in which an autonomous machine emerges overnight and outperforms every living scientist, maybe in the form of a Nobel-worthy discovery. This picture fits our science fiction imagination. It also distracts from the likelier scenario, in which AI displaces features of the scientific process incrementally. Each step that gets transferred to AI appears to offer a meaningful gain in efficiency. No committee has to vote to hand over science to the robots; a sequence of quiet software updates can get us there just as fast.
The first plank sits in education, where future scientists are trained. In a large field experiment, high school mathematics students with access to a standard generative AI interface (based on GPT-4) performed better on practice problems than students who didn’t use it; however, they scored 17 percent worse on the subsequent exam. Students who used a more specialized AI math tutor for the practice problems largely avoided the decline.
The finding captures a basic distinction between completing an assignment and learning how to do it well. Training depends on working through difficulty: Students read challenging texts, make mistakes, and revise their reasoning. This is the process by which we learn how to be discerning, how to reason, and make good inferences. When AI removes these challenges, students lose one of the most critical steps to the process.
Peer review offers a second example. It is one of the most important activities in science, yet is driven by unpaid labor and dubious motivations: informal professional obligation, shame, and deadline panic. These conditions make the use of an AI tool tempting. Researchers examining peer review text submitted to several major artificial-intelligence conferences found that between 6.5 and 16.9 percent could have been substantially modified by large language models, or LLMs. Figures were higher among reviews completed near the deadline. (This research itself was published on a pre-print server and is not yet peer reviewed.)
Each step that gets transferred to AI appears to offer a meaningful gain in efficiency. No committee has to vote to hand over science to the robots.
The role of AI is even moving into manuscript preparation. By tracking sudden increases in words favored by LLMs, an audit showed that at least 13.5 percent of abstracts published in the PubMed database in 2024 had been processed using an LLM.
Neither of these findings can tell us the role that AI played in the overall decisions, but they do show LLM use spreading across all sides of scholarly publishing. But the journal workflow — preparation, submission, review — remains structurally the same, even as portions of it are increasingly carried out by the machines.
In many areas of research, the replacement of human-driven planks with computer code makes the changes harder to catch. But even if contributions made via AI are not technically false, their usage can be a problem. A 2025 investigation based on interviews with working scientists warned that code generated by AI models can quietly influence statistical parameters in ways that the user may not recognize. And so, AI may produce a result that appears authoritative, even when addressing the wrong question. Notably, such a result would be reproducible in the standard sense and therefore pass many of the formal tests for appropriateness.
A portion of a mosaic from the 4th century A.D. depicts Theseus boarding the ship that would carry him home to Athens.
Visual: © KHM-Museumsverband
A reasonable objection to my prediction is that human minds will still oversee the integration of the steps in the scientific process. But the algorithms are coming for the integration. A March report described an “AI Scientist” that generates ideas, searches the literature, writes code, runs computational experiments, interprets the results, generates visuals, prepares a manuscript, and reviews its own output. A manuscript written by this algorithm passed the first round of peer review at a machine-learning workshop.
Another project, named Robin, brings a related approach to experimental biology. Human researchers perform the physical experiments, but Robin proposes hypotheses, analyzes the results, and produces figures. Both the AI Scientist and Robin represent architectural milestones, organizing the individual planks in the process into an integrated, alternative research pipeline.
Why do I consider these legitimate threats to the scientific order, rather than just technological gimmicks? As with most challenges in a profession with many participants and multiple interests, we can point to incentives. Scientists are adopting AI because the tools can be useful for the sorts of outputs that we are evaluated on and rewarded for. A recent study surveying 41.3 million research papers found that authors using AI had higher publication and citation rates. Authors who use it also became project leaders earlier. At the collective scale, however, adoption of AI was associated with a 4.63 percent contraction in the range of topics studied and a 22 percent decline in engagement with others across the scientific community. Those findings cannot establish that AI caused these changes. But the pattern captures a collective-action problem: A strategy that may reward an individual scholar can produce a narrower and less social enterprise when repeated across millions of careers.
Maybe now is the time to draw a line in the sand and organize our efforts toward halting the intrusion of these artificial agents into our science. But a blanket rejection would be irresponsible and based on a fantasy. By now, it should be clear that AI can and will do things — including make discoveries — that we are incapable of, if only because we can’t work as fast and tirelessly as a machine. Furthermore, AI is already too prevalent for a full reversal. The more appropriate approach involves realistic appraisals of the situation, such as in the Leiden Declaration, whereby the community of mathematicians laid out the challenges of navigating their field in an age of AI. Several other fields should follow suit.
I believe that science derives its authority from something more than mastering the pipeline from data generation to publication.
So far gone are we that it may soon be difficult to argue why humans are necessary for any step in the process. One can respond with the self-soothing notion that we need humans to retain enough expertise to inspect the work and challenge the outputs. But this suggestion provides no concrete guidance for why, where, and how we’ll need humans when the recognizable activities of science can be done faster, and almost as well (or better, in some cases) by the machines.
I believe that science derives its authority from something more than mastering the pipeline from data generation to publication. Rather, it comes from interactions within a community of people trained to interrogate evidence and answer for the conclusions they draw from it. And it is time to formalize what this means, and which sorts of human judgment cannot (or should not) be offloaded to AI.
If we do not act, then a full replacement — completed plank by plank — may happen before we know it. By then, the greatest knowledge creation tool in human history will be just another task automated by machines, with the human activity that made AI possible relegated to memory and history.