One way to think about this is that, when the system comes across a location where a set of chemically related amino acids will all work (like leucine/isoleucine/valine or serine/threonine), it will use one that’s consistent with the watermark when possible. Another way to look at the process, suggested by one of the people involved in developing the system, is that it searches through the space occupied by functional proteins for the subset that happens to have a sufficient number of watermark amino acids.
As a result, the watermark is randomly distributed across the entire length of the protein, and detecting one isn’t a simple yes-or-no question. You have to scan the whole sequence, knowing the key, and measure how often the amino acids suggested by SynthIDBio actually appear in the final sequence. Google has also developed the software needed to do this.
It’s alive!
The question, then, is whether watermarked proteins are functional. The team used the system to design proteins that physically interact with key natural proteins previously targeted with AI designs. And the watermarked versions worked just fine, binding the intended targets. This isn’t as rigorous a test as finding a catalyst, but it suggests that there’s no reason to expect serious problems in more complicated design tasks.
So as long as a protein is long enough, the system can detect a watermark. How might that be useful? Again, it comes down to biosecurity. When someone orders DNA sequences, the people who make the DNA normally screen the sequence for its ability to encode portions of viruses, toxic proteins, and other similar threats. Right now, however, when they see a protein that doesn’t look similar to anything we already know about—something that’s potentially true for any AI-designed proteins—they can’t assess its threat.

