As the biotechnology industry pours billions of dollars into artificial intelligence–driven drug discovery and development, a foundational problem remains unsolved. A 2024 survey of biomedical researchers found that nearly three in four believe their field is gripped by a reproducibility crisis — the widespread inability of researchers to replicate study findings even when similar methods are used. At the core of this crisis in the life sciences is the quality and authenticity of data that are being reused by millions of researchers, including those building the latest models for AI-enabled biology (often referred to as AIxBio) applications. This is creating a significant risk for AIxBio’s future success.
This is not a new concern. When Walter Goad, an early pioneer in theoretical biology and biophysics at Los Alamos National Laboratory, proposed the first centralized DNA sequence repository in 1979, he recommended collecting complete information on the DNA’s origin, preserving supporting evidence, and assigning each database entry with a validation level based on the available evidence. In effect, the proposal treated the evidentiary record behind a sequence as an integral part of the sequence record itself. It was obvious then as now that digital sequences map not just to a species or an individual but often to actual samples that researchers need access to in order to replicate another researcher’s results.
Yet, as genomic data expanded rapidly over subsequent decades, public repositories increasingly prioritized the volume and accessibility of the sequence data, while much of the detailed provenance needed to evaluate, reproduce, or validate those data remained inconsistently captured or obscured.
This growth in genomic data has fueled important discoveries, such as Nobel Prize–winning work on CRISPR gene editing, microRNAs, mRNA vaccines, and many others. It is also, however, exposing a structural weakness that has been compounding for decades: DNA sequence data are preserved far more reliably than the metadata and methods that describe the origins and production of these data. Metadata — the contextual information describing how, when, where, and from what biological material the data were generated — is needed for researchers to verify, reproduce, or trust the underlying genomics data itself. But in many cases these metadata are missing or contain errors that complicate reuse of the primary data. These lapses in curating metadata tied to genomics data have been noted as having real-world impacts on a diverse range of fields, including the responsiveness of public health agencies, biomedical genomics research in biopharma, and wildlife conservation biology efforts.
In 2021, the U.S. Food and Drug Administration led an analysis of more than 550,000 pathogen genomes and found that approximately a quarter were missing at least one piece of metadata essential for public health surveillance. Imagine borrowing a textbook from a library and not knowing what edition it is. Information in the textbook could be incorrect or out-of-date, but without metadata on the edition or printing date, one could wrongly assume all the information was accurate.
A technically sound analysis can faithfully reproduce a prior finding and still be wrong if it is built on contaminated, mislabeled, or otherwise compromised reference data.
In addition, our organization — ATCC, a biological resource center that collects, stores, and distributes cells, microorganisms, and genetic materials and their associated data sets — conducted a review of genomes labeled as coming from a major reference cell culture collection. We found that approximately 40 percent of records did not list the sequencing technology or bioinformatics methods used to produce the genome, and over 99 percent lacked any description or source information about the originating biological sample. The research also found that in public databases, many DNA sequences from reference microbial strains had likely been generated from derivative isolates, potentially with an unknown chain of custody, rather than the certified source material itself. These gaps in the chain of custody of cell lines create a black box about the strain’s origin, natural history, handling conditions, and the provenance of the data representing those materials.
There’s also an important distinction between scientific reproducibility and correctness. A technically sound analysis can faithfully reproduce a prior finding and still be wrong if it is built on contaminated, mislabeled, or otherwise compromised reference data. Subsequent studies that reuse the same flawed inputs can arrive at the same incorrect conclusion through an entirely valid process — a form of reproducible wrongness that spreads as data are reused.
Walter Goad, an early pioneer in theoretical biology and biophysics at Los Alamos National Laboratory, proposed the first centralized DNA sequence repository.
Visual: Science Source
Instances such as these are not uncommon, and the net effect often delays scientific discovery and wastes resources attempting to reproduce prior results. Thus, under the prevalent “archive first” model that many public genomics repositories embrace, sequences are often preserved more reliably than the evidence needed to interpret, reproduce, or validate them. This approach leaves a growing trust gap as datasets scale up and reuse of the data downstream in AI models expands.
Models are only as reliable as the data they are trained on. For example, the strength of AlphaFold, which is used to predict the 3D structures of proteins, was largely possible because the underlying data came from the Protein Data Bank, a gold-standard resource built over decades on rigorous, expert-curated, independently validated structural data. In contrast, as biological AI scales up on incomplete, inconsistently documented genomics data and rarely verified records, the weakness will grow faster, threatening to slow the pace of scientific discovery and raise the risk of failed drug development programs.
Many newer genomic foundation models have not been built with this level of rigor. For example, developers of Evo 2 and the Nucleotide Transformer — two prominent AI models focused on predicting or generating the DNA sequence of genes or entire genomes from scratch — had to build additional layers of manual curation and filtering on top of public sequence archives before the data were usable for training, precisely because the raw public data were not reliable enough on their own. In effect, these model development efforts expose the underlying inconsistencies in biological training data. Most researchers do not have the budget or time to validate the existing third-party datasets they rely on, and there is little incentive to publish corrections to third-party data when errors are found.
Fortunately, none of these issues are irremediable, and solutions already exist. This isn’t a problem that requires new science to solve. Community standards capable of capturing sample origin, handling history, and computational methods have been developed via projects such as MIxS, Darwin Core, BioCompute, the Functional Annotation of ANimal Genomes, and ENCODE. Enforcement is the missing element. In the major repositories on which the research world relies, critical fields that record provenance remain optional, inconsistently filled in, and weakly checked.
The genomics community has long championed open data. The next step is to champion trustworthy data sharing.
To make meaningful progress, however, enforcement must be managed in ways that reduce friction for people submitting data while materially improving interoperability, traceability, and reuse. Critically, physical reference samples and the metadata that describe their origin, handling, and transformation should be treated as shared scientific infrastructure that is proactively preserved.
Culture collections and biorepositories can uniquely bridge the physical‑to‑digital divide by creating verifiable “digital twins,” anchored to authenticated source materials, while repositories, journals, funders, and tool builders can drive adoption of these practices by establishing clearer minimum deposition requirements, providing better tools for submission, and offering visible incentives for strong provenance.
The genomics community has long championed open data. The next step is to champion trustworthy data sharing so that the future of digital biology and AI‑enabled biology is built on traceable, defensible, and reusable genomic records, rather than on an assumption of trust.
Jonathan Jacobs, Ph.D., is senior director of bioinformatics, and Patrick Boyle, Ph.D., is interim chief scientific officer at ATCC (American Type Culture Collection).