What Is Semantic Design?
Imagine artificial intelligence (AI) reading a genome like a book and filling in the missing parts so that they make sense. This is the basis of semantic design, a method described in an article published in Nature. The Evo model, developed by scientists, can generate new DNA sequences based on context from prokaryotic genomes. Evo was trained on data from OpenGenome, a collection of more than 80,000 bacterial and archaeal genomes plus 2 million phage and plasmid sequences, totaling 300 billion nucleotides.
The Evo 1.5 model, an expanded version of the original Evo 1, processes sequences up to 8,192 tokens long. In tests, it was able to complete conserved genes such as rpoS from Escherichia coli with 85% accuracy using only 30% of the input sequence. For operons such as trp or modABC, it generated sequences with more than 80% similarity to natural proteins, even when antisense strands were used for guidance.
Generating Toxin-Antitoxin Systems
Semantic design proved effective in creating toxin-antitoxin systems, which bacteria use to defend themselves against phages (viruses that infect and replicate in bacteria). For type II systems, Evo generated the toxin EvoRelE1, with 71% similarity to RelE, which caused a 70% reduction in bacterial growth. This toxin was then used as a prompt to generate the antitoxins EvoAT1 through EvoAT4, which restored growth to 70-100%. EvoAT2 neutralized three natural toxins: RelE, MazF, and YoeB, despite having only 21-27% similarity to natural proteins.
For type III systems, Evo created the toxin EvoT1, which reduced bacterial survival to 33%, without significant similarity to known toxins. The antitoxin EvoAT6, an RNA sequence with 78% similarity to ToxI from Bacillus multifaciens, neutralized ToxN with an 88% success rate. Structures predicted by AlphaFold 3 showed similarities in secondary motifs, even though the sequences were different.
New Anti-CRISPR Proteins
Evo also designed anti-CRISPR proteins (Acr) that block CRISPR-Cas systems. Using prompts from Acr operons, it generated candidates such as EvoAcr1 through EvoAcr5. These protected bacteria against SpCas9 with success rates of 74-101% in survival assays and phage infection tests. EvoAcr1 and EvoAcr2 had no significant similarity to known proteins, yet they functioned robustly. EvoAcr3 had 25% similarity to the sigma-70 factor, EvoAcr4 had 58% similarity to AcrIIA2, and EvoAcr5 had 31% similarity to AcrIIA4.
These proteins were composed of fragments from 15-31 natural proteins, similar to de novo designs from RFdiffusion or BindCraft. The 17% success rate in tests exceeded expectations without relying on structural assumptions.
SynGenome: An Artificial DNA Database
Evo generated SynGenome, a database containing more than 120 billion bases of synthetic DNA derived from 1.7 million prompts from UniProt. Each prompt produced sequences annotated according to Gene Ontology and InterPro. SynGenome preserves natural patterns, such as ORF lengths and Pfam domain frequencies (a correlation of 0.78 with OpenGenome).
The association network in SynGenome revealed links such as DUF2871 with cytochrome c and DUF2797 with rhomboid domains, supporting hypotheses about unknown functions. The database contains chimeric proteins, such as domain fusions, and is available at evodesign.org/syngenome for further research.
This approach opens the door to designing genes beyond natural evolution, with applications in biotechnology.



