Every gene needs a signal saying “start here.” Scientists have decoded one of the most important — a DNA element called the initiator — using machine learning.
Researchers at the University of California, San Diego, led by James T. Kadonaga with graduate researcher Torrey Rhyne-Carrigg, published the work in Genes & Development.
What an initiator is
Reading a gene begins at a precise point. The machinery transcribing DNA into RNA must be positioned exactly, because starting a few bases off produces a different molecule — potentially a non-functional one.
The initiator is the DNA element marking that start point. It sits at the core promoter, the short stretch immediately around where transcription begins, and it helps recruit and position the transcription machinery.
It has been known about for decades. What has been missing is a reliable description of what it looks like — the sequence pattern distinguishing a real initiator from a similar-looking stretch of DNA that is not one.
Why the pattern resisted description
Some regulatory DNA elements have strong, recognisable consensus sequences that can be searched for reliably. The initiator does not.
Its pattern is degenerate — loosely defined, tolerating substantial variation, with no single sequence that must be present. Traditional approaches produce a consensus so permissive it matches sequences throughout the genome, generating far more false positives than real findings.
That is precisely the sort of pattern-recognition problem where statistical models outperform rules written by hand: the signal exists and is distributed across positions in a way no simple description captures.
The approach
The team used high-throughput sequencing to measure gene activity across roughly 500,000 DNA sequence variants, then trained a machine-learning model on the results.
The design deserves attention, because it is what makes the model trustworthy. Rather than training on existing annotations — which would teach the model to reproduce prior assumptions — they generated the training data experimentally, systematically varying sequences and measuring what each one did.
That produces a model learning the relationship between sequence and function from direct measurement. Half a million variants is also enough to capture a degenerate pattern that a smaller set would miss.
The result
They pinned down the characteristic pattern and found the initiator in about 60% of human genes.
“These AI models were found to provide, for the first time, strong predictions of the presence or absence of the initiator in human genes,” Kadonaga said.
The 60% figure is itself informative. It establishes the initiator as a majority feature rather than universal, meaning a substantial minority of genes start through different arrangements — consistent with the growing recognition that core promoters are more diverse than early models assumed.
Why gene starts matter for disease
Genes must switch on in the right place at the right time for the body to develop and stay healthy, and when that control fails, disease including cancer can follow.
This connects to a broader shift in how genetic disease is understood. Most disease-associated genetic variation identified by large studies lies outside protein-coding regions — in regulatory DNA controlling when and how much a gene is expressed rather than what protein it encodes.
Interpreting those variants has been the field’s central difficulty. A change in a coding region can be assessed against the protein it alters; a change in regulatory DNA requires knowing what that DNA does, which frequently nobody does.
The practical use
A model predicting where initiators sit, and how sequence changes affect them, is a tool for exactly that interpretation problem.
Given a variant of unknown significance in a promoter region, such a model can estimate whether it disrupts the initiator and by how much — converting an uninterpretable finding into a testable hypothesis. Multiplied across the many variants of unknown significance genetic testing generates, that is a meaningful capability.
What it does not do
Predicting an initiator’s presence is not the same as predicting gene expression. The initiator is one element among several in the core promoter, and core promoters are themselves only part of a regulatory system including distant enhancers, chromatin structure and cell-type-specific factors.
Why this style of AI work differs from the headline kind
It is worth distinguishing what happened here from the general-purpose models that dominate discussion of AI in science.
This is a narrow model trained on data the researchers generated specifically to answer one question. It does not reason, does not generalise beyond its task, and would be useless for anything other than recognising initiator sequences. Its value comes precisely from that narrowness — the training data matches the question exactly, and the output can be checked experimentally.
That pattern — generate purpose-built measurements at scale, train a focused model, validate against biology — has quietly become one of the more productive uses of machine learning in molecular biology. It works because the bottleneck was never reasoning; it was extracting a signal from a pattern too degenerate and too high-dimensional for a human to specify by hand.
The result is a tool rather than a discovery engine, which is a less exciting description and a more accurate one.
A model of one component is a component of a model, and the work is fundamental science rather than an immediate clinical tool — the kind that underpins future diagnostics rather than delivering them. Research news.