These notes are intended to illustrate to curious folks outside of the field some of the excitement on the frontier of RNA research.
Introduction: The Central Dogma
Most people learn the central dogma of biology in a high school class. The problem is, it is normally written on the blackboard as a simple map: DNA > RNA > protein. Not only is this unidirectional version of the central dogma an oversimplification given what we know now about RNA’s role in the cell, it’s an oversimplification of what was meant even at the time the theory was proposed.
The original idea started with: “Once information has got into a protein it can’t get out again.” Notice that this does not state anything DNA or RNA –– the earliest available sketch of the central dogma, from Crick’s 1956 notes, show clearly that he already suspected RNA could actually transfer information back to DNA or to itself:
Even though that possibility was on the table, the reality of RNA’s role remained long obscured. For decades, textbooks described RNA as simply a messenger between DNA and protein. As sequencing and spatial tools improved, that view collapsed. DNA does code for which protein to make, RNA copies that code and delivers it to the ribosome (our molecular protein printing machines), and the output is the encoded protein. However, the RNA that copies and delivers the code, messenger RNA (mRNA), is only a subset of RNAs. Not only do other RNAs help regulate which DNA has its code copied, the ribosome, that molecular printing machine, is itself functionally a complex RNA. These miracle molecules can act as gradients to direct intracellular traffic, perform complex enzymatic functions like proteins, and can regulate the expression of genes.
One particularly useful way to think about the reality of the genome is a hardware/software split. DNA is analogous to stable hardware storage, and RNA is analogous to the software layer that executes conditional logic, routes information, and configures the system.
(It is also fun to note that the original version of the central dogma, while shown to be continuously true in nature, has been engineered backwards by a group from Stanford who were in fact able to achieve Protein>DNA.)
The “Dark Genome” and the Expansion of RNA Biology
It should now be clear that RNA does more than just carry information for which proteins the ribosome should produce, but the reality is even more extreme than that knowledge implies. Only a very small percentage of the genome, less than 2%, produces proteins. You can account for another 24% of the genome with introns that space coding regions, usually leaving people to sift the remaining ~75% into the category of the “Dark Genome.” Though not qualifying as part of the “unknown” Dark Genome, introns (themselves non-coding) serve a myriad of critical functions, containing regulatory elements like enhancers and silencers, as well as enabling alternative splicing. I suspect it is not unlikely they have as many undiscovered functions as discovered ones and should be in the Dark Genome category, but this is controversial.
Regardless, because early assays filtered for open reading frames and high-abundance transcripts, the Dark Genome regions were originally labeled “junk.” As sequencing improved and became more routine, these supposedly inert regions were revealed to be widely transcribed. That is, cells were producing large numbers of RNAs with no coding potential. This forced a revision of what counts as genomic information and how that information is executed. It pushed the field toward a more complex understanding of the central dogma, in which RNA helps shape the regulatory environment. The aforementioned “junk” regions became known as non-coding RNAs (ncRNAs).
This category further expanded with the introduction of long non-coding RNA (lncRNA), a term that reflects arbitrary tech dev constraints as opposed to any real biological logic. Funnily enough, a transcript longer than 200 nucleotides without a coding region was placed into this group simply because purification columns at the time let through anything less than 200 nucleotides, so what remained was anything larger.
The term lncRNA, like the broader term ncRNA, aggregates molecules by the fact they do not produce protein rather than by their function. As tools improve, these categories will almost certainly fracture into groups by structure, localization, or interaction type rather than by the absence of a protein-coding region.
Several classes of ncRNA and lncRNA are well-characterized, but there are still thousands of unknown non-coding transcripts across the genome. For those, ncRNA and lncRNA remain catchall for many molecules whose functions span architectural scaffolding, spatial organization, transcriptional control, epigenetic modification, and more. Below are a few of the ncRNAs and lncRNAs that are well-characterized to help illustrate how functional the molecules are, as well as some examples of how they perform intricate regulatory control functions.
Well-Characterized ncRNAs and lncRNAs
ncRNAs
tRNA
These RNAs are responsible for adding the right amino acid via the ribosome by carrying both the amino acid and the matching code complementary to the mRNA codon. The first ncRNA characterized was in fact an alanine tRNA found in baker’s yeast.
siRNA
Short double-stranded RNAs that act as guides for Argonaute proteins, leading them to degrade transcripts. They have provided a new paradigm for medicine by allowing us to silence harmful genes. Alnylam, now a $60B company, is at the center of the development of siRNA for medicine.
miRNA
Small ~22 nt RNAs that help tune gene expression across the transcriptome. Each miRNA can influence dozens to hundreds of targets, contributing to developmental patterning and stress responses.
lncRNAs
Xist
If you have ever wondered why human women don’t produce twice as much of their X chromosome proteins as men, Xist is why. Xist controls X chromosome inactivation, and it was discovered partially because of calico cats. Their coat patterns are actually reflective of the random inactivation of either the maternal or paternal X chromosome. Each patch corresponds to a set of cells that made the same inactivation “choice.” This cannot happen in an XY chromosome pattern, so almost all calico cats are female. The only male calico cats have XXY patterns. Efforts to understand X inactivation are what led to the identification of Xist as responsible for the silencing.
(Even more exciting is the magnitude of effect. Only tens to a few hundred Xist molecules per inactive X are enough to shut down the entire chromosome in humans. For context, the human X chromosome is about 155 million base pairs long, and Xist is only about ~17,000 nucleotides. That means 1 XIST nucleotide per ~180 bp of X-chromosome DNA.)
HOTAIR
HOTAIR is transcribed from the HOXC locus and regulates the distal HOXD locus. It guides chromatin regulators to specific regions, showing that lncRNAs can act in trans (at a distance) and influence development and chromatin patterning through targeted recruitment. It also plays a part in epigenetic differentiation of skin between different body parts, and is remarkably “shuttled” from chromosome 12 to chromosome 2 by a protein called Suz-Twelve. This is one of the most well-characterized lncRNAs because it is highly expressed in metastatic breast cancer.
(It was named HOTAIR over STAR [Suz-Twelve Associated RNA]. Howard Chang said to John Rinn at the time, “Well, if it all turns out to be hot air, we’ll have named it correctly.”)
TERC
TERC is the RNA component of telomerase, an enzyme many of you will recognize as central to aging. It serves as both the structural scaffold of the enzyme and the template used for extending telomeres, repetitive regions that protect chromosome ends as our cells divide. Telomere shortening limits the number of times most cells can divide, which is a critical function for most forms of life, and telomerase counteracts this. LncRNA is therefore part of the machinery that governs cell lifespan.
rRNA
Ribosomal RNA is responsible for the catalytic functions of the ribosome. It mediates peptide-bond formation, positions tRNAs, and ensures translational accuracy. Protein components of the ribosome exist mainly for stability, but this was not always obvious. Before the 1980s, the ribosome was instead assumed to be a protein enzyme with RNA scaffolding. The picture changed after Tom Cech and Sidney Altman demonstrated that RNA itself could catalyze reactions, for which they won the Nobel prize. That work opened the door to viewing the ribosome as a ribozyme, which was later confirmed by Harry Noller, Ada Yonath, Venki Ramakrishnan, and Tom Steitz. It was understanding this and learning about the ribosome’s SRL mechanism (a central method of enforcing translational fidelity) that drew my personal interest in ncRNA. I am not alone in this –– their discovery that RNA could act as a ribozyme motivated the whole modern field of ncRNA research.
If you want to learn more about the ribosome, tRNA, and the overall RNA>Protein translation process, check out this video from Oxford University:
How RNA Executes Regulatory Control
A multitude of properties explain how RNA exerts influence without translation. I have listed some of these below.
Structure
RNA folds into shapes that recruit proteins, other RNAs, and DNA, like proteins do. Moreover, RNA’s structure can be pliable, better enabling them to have multiple functions or perform functions of high complexity, catalytic and architectural. RNA allows the cell to build structures without requiring persistent protein and with likely lower energy cost than an all-protein complex assembly. Some lncRNAs act as durable scaffolds, while others exist transiently.
This flexibility helps explain why differentiation, stress responses, and other rapid adjustments often depend on RNA. RNA structure prediction is hard, and structure is context-dependent. It seems that RNA structure preempts sequence, as there is evidence that structure is preserved between species where sequence isn’t.
Timing
The transience and persistence of different kinds of RNA mean that RNA half-life is a regulatory mechanism in and of itself. RNA also uses processing latency and orchestrates kinetic timing competition. The ability to use timing as a mechanism is crucial for many processes, especially those that occur during development.
Spatial Organization
Some RNAs can act as a map. Mitch Guttman’s group has shown that RNAs can shape higher-order structure, recruit regulatory complexes to specific nuclear regions, and organize interactions in trans. Essentially, they have shown that small RNAs can operate as a local chemical gradient map.
Epigenetic Regulation
There is increasing evidence that lncRNAs can influence the epigenome, including site-specific DNA methylation and chromatin accessibility. In plants, RNA-directed DNA methylation is fairly accepted. In mammals, the research is ongoing, but at least several lncRNAs interact with DNA methyltransferases or demethylating enzymes, guiding them to specific loci. John Rinn’s group has shown that the lncRNA FIRRE, for instance, has strong effects on chromatin accessibility, and achieves these effects at unexpectedly high speeds.
Therapeutic Potential
You may be wondering at this point why we don’t already take advantage of lncRNA’s multiple regulatory functions to treat diseases of dysregulated genes like we do with siRNA. This is exactly why I co-founded LincSwitch Therapeutics with John Rinn and Marvin Caruthers (Amgen co-founder and inventor of phosphoramidite synthesis).
Other companies pursuing this space or adjacent to it include HAYA Therapeutics, CAMP4, Flamingo Therapeutics, Amaroq Therapeutics, and a handful of others. I will likely write something more in depth about the RNA therapeutics landscape at a later date.
Technology Development
Our understanding of RNA has expanded only as quickly as the technologies capable of measuring it. Each methodological advance has revealed layers of complexity that earlier techniques could not resolve. Long-read sequencing, including the newer approaches that grew out of John Rinn’s cap-seq work, has been especially important because it recovers full transcript rather than just poly-A-enriched fragments. Spatial transcriptomics has helped clarify where RNAs operate inside the cell. Single-cell multiomics has created a much richer picture by allowing us to see RNA levels, chromatin accessibility, and protein expression all at once. More precise perturbation techniques, like CRISPR interference and activation systems, antisense oligonucleotides, and base editing, have made it easier to attribute function.
Most standard sequencing workflows still rely on poly-A enrichment capture methods, which are mostly blind to non-polyadenylated non-coding RNAs. Most workflows also still rely on short-read, fragment-based methods that cannot recover isoforms because fragmentation and short-read assembly collapse alternative splicing. Moreover, post-processing typically filters out low-abundance transcripts to decrease noise, which undoubtedly leads to occasional signal loss. The result is a heavily filtered view of the transcriptome.
If the goal is to build accurate models of regulatory behavior, the data feeding those models ought to reflect the full range of RNAs that operate in the cell. The problem is similar to training a language model on mostly nouns. Such a model could not reliably recover the grammar that details relativistic relationships and would find it impossible to accurately gauge meaning. We are essentially producing large, clean datasets at the cost of excluding entire categories of transcripts, including many that appear to control structure, timing, and spatial organization.
As biological modeling becomes more ambitious, the need for better data is obvious. I suspect a virtual cell model, or anything approaching it, will require datasets that include more of the Dark Genome and are isoform-aware. Without this information, the underlying regulatory architecture will be only partially visible, and the models built on it will reflect those omissions. The problem is, this will be incredibly expensive. I think it is still worth pursuing. You need to measure a system if you want to model it.
Concluding Thoughts
Non-coding RNA is the next frontier in understanding genomic regulation and decoding life. The terms ncRNA and lncRNA will likely be replaced by mechanistic categories as more functions are revealed. What emerges from this will be an understanding of the genome in which RNA is essentially the operating software.
Complete understanding of the genome as an RNA system, where RNA is essentially the operating software, will eventually allow us, to some extent, to program life. Hundreds of RNA-based drugs are already in development, with tens already approved. Further basic research on the Dark Genome will be critical for transforming medicine from guesswork to models that can predict and alter the flow of disease.
If this was interesting to you, you should consider working on it –– we will need many people to care about how to model healthy systems before we can truly make programmable medicines.
If you’re interested in reading more about this academically, here are some people whose work I recommend reading:
Tom Cech & Sidney Altman
Harry Noller
Elizabeth Blackburn, Carol Greider, & Jack Szostak
Joan Steitz
Thomas Steitz, Venkatraman Ramakrishnan & Ada Yonath
Phillip Sharp & Richard Roberts
Andrew Fire & Craig Mello
Jeannie Lee
Carolyn Brown & Hunt Willard
Jennifer Doudna
John Rinn
Howard Chang
Mitch Guttman
Eric Miska
Robert Sebra
Laura Landweber
Uttiya Basu
Gideon Dreyfuss
Bo Wang
Aviv Regev
Patrick Hsu
Brian Hie
Hani Goodarzi
Sam Sternberg
Feng Zhang
David Liu



