
A key cellular enzyme can accurately read an eight-letter genetic alphabet, doubling the four-letter code used by all known life, researchers at the University of California San Diego have demonstrated in two studies published this year.
The claim is specific and the methodology is detailed: the team combined biochemical experiments with high-resolution cryo-electron microscopy to capture structural images of RNA polymerase from Escherichia coli bacteria as it recognised and incorporated synthetic base pairs. The imaging, the researchers say, resolved structures at scales smaller than the width of a single atom.
What the enzyme actually does with an eight-letter genetic alphabet
RNA polymerase is the enzyme that reads DNA and produces RNA, the first step in gene expression. The question the UC San Diego team set out to answer was whether existing cellular machinery could handle synthetic genetic letters without breaking down or making errors. Their images showed that the enzyme identifies synthetic DNA letters using many of the same biochemical and structural signals it relies on to recognise natural base pairs.
That result matters because it suggests cells would not necessarily need wholesale re-engineering to process expanded genetic information. The researchers say it helps explain why the enzyme can accurately copy information written in an expanded code. The key word is “accurately”: cryo-electron microscopy allowed them to observe the process in molecular detail rather than infer it from downstream error rates alone.
A related paper, published in Proceedings of the National Academy of Sciences on 12 August 2026, extended the finding further. The same team found that RNA polymerase can also recognise another synthetic base pair even though it lacks the hydrogen bonds that normally help hold DNA base pairs together. Hydrogen bonding is generally considered essential to DNA’s structure and stability, so the enzyme’s ability to work without it in this context is the kind of result that invites scrutiny of the underlying mechanism.
Two papers, one research group, and what they stop short of claiming
The Nature Communications study, titled ‘Structural Basis of Transcription of the Hachimoji Eight-Letter Alphabet by E. coli RNA Polymerase’, was led by Dong Wang, professor at the UC San Diego Skaggs School of Pharmacy and Pharmaceutical Sciences, and published on 2 September 2026 in Nature Communications. The PNAS study, also led by Wang, appeared three weeks earlier.
The papers do not claim that a working eight-letter organism has been built, or that the expanded alphabet has been deployed in any therapeutic context. What they establish, the researchers say, is a structural and mechanistic foundation: a molecular-level account of how RNA polymerase handles non-natural DNA letters. That is a narrower, more defensible claim, and the distinction is worth preserving.
The potential applications the team points to are described as future possibilities, not near-term products. These include new diagnostic tools, therapeutics, and engineered biological systems with capabilities that do not occur naturally. Earlier research, separately from these two studies, has already used expanded genetic alphabets to create synthetic DNA molecules capable of recognising liver cancer cells, suggesting the direction of travel is real even if the destination remains some distance away.
The broader goal, as the researchers frame it, is expanding the language of DNA beyond what nature settled on. Whether RNA polymerase from E. coli is a reliable proxy for how human cellular machinery would respond to synthetic letters is a question these papers do not answer. Wang’s group has, however, produced two papers in the space of three weeks that give that question a more precise experimental footing than it had before.



