Google Deepmind’s New Hypothesis Generation Machine: AlphaGenome Atlas Maps Hidden Regulatory Code of the Human Genome

Google Deepmind is making public their extensive AlphaGenome Atlas database. A collection of precomputed predictions made using the AlphaGenome Atlas model that predicts the result of mutating genes in humans and providing the results freely that anyone can query.
Most of the human genome does not encode proteins, about 98% of human genetic variation sits in non-coding DNA, the regulatory regions that control when and where genes are switched on rather than the protein-coding parts. When a patient’s genome is sequenced, they typically carry 4–5 million variants, and almost all of the non-coding ones are “variants of uncertain significance.” Existing tools each fail in a different way: annotation frameworks like ENCODE/VEP lack basepair resolution, conservation-based scores (GPN-Star) miss functional variants that aren’t conserved, and CADD aggregates features but is only as good as its inputs. AlphaGenome itself (their earlier model) can predict a variant’s regulatory effect, but only on-the-fly, one variant at a time, which is too slow for genome-scale work.
AlphaGenome Atlas, from Google DeepMind and collaborators, takes a brute-force computational approach: by mutating essentially every possible base in the human genome and modelling what changes can be expected at the molecular level. Thereby ensuring others do not have to re-run the same query again.
The authors did not merely build another pathogenicity predictor. They computationally evaluated essentially every possible SNV in the human genome (~9 billion) against thousands of molecular phenotypes, generating roughly 27,000 experiment-specific predictions per variant. That transforms AlphaGenome from a model you query one variant at a time into something closer to a reference map of predicted regulatory consequences. Although the disclaimer mentions that outputs can’t be used for any clinical purposes.

THE ATLAS IN NUMBERS
| ~9 billion possible human SNVs | 100M+ observed indels | ~27,000 predictions per variant | 2,601 discovered motif clusters | 253 billion motif instances |
Genome-wide motif map. They ran TF-MoDISco on the ISM scores to discover ~2,601 regulatory motifs de novo, hand-annotated them to 94 main TFs and 122 zinc fingers, and mapped 253 billion motif instances across the genome, per modality and per cell type. This lets a high-scoring variant be linked directly to the specific transcription-factor motif it breaks. The team subsequently mapped motif instances throughout the genome using FiNeMO.
Experimental ChIP-nexus validation showed that predicted motif sites could recover genuine, cell-type-specific binding, supporting the idea that these are more than simple sequence-matching annotations.
The underlying AlphaGenome model considers a 1 Mb genomic context and predicts multiple regulatory readouts, including chromatin accessibility, transcription-factor binding, histone marks, RNA abundance, transcription initiation,polyadenylation, chromatin contacts and splicing.
In silico saturation mutagenesis of the whole genome. They precomputed AlphaGenome predictions for all ~9 billion possible SNVs in hg38 plus >100 million observed indels from gnomAD, UK Biobank and All of Us. Each variant gets ~27,000 scalar predictions spanning chromatin accessibility, TF binding, histone marks, transcription initiation, RNA expression, splicing, polyadenylation and 3D contact maps, across hundreds of cell types.
AVI (AlphaGenome Variant Impact) score. A small neural net that folds those precomputed scores (reduced to 10 modality features) together with AlphaMissense, protein-termination flags, two conservation scores and indel flags — only 18 features versus CADD’s 150+ — into a single PHRED-scaled prioritization score. It’s trained on population allele frequency as a proxy for negative selection (rare = probably impactful, common = probably benign). Critically, they compute SHAP attributions for every score, so you can see why a variant scored high: splicing, promoter activity, protein disruption, etc.

From prediction to a single variant-impact score
Rather than treating AVI as another black-box pathogenicity score, the Deepmind researchers used SHAP attribution (SHapley Additive exPlanations) to show which molecular signals contributed to a variant’s score. A high score can therefore be traced back to predicted changes in splicing, transcription, accessibility, TF binding or other mechanisms.
Variant → Molecular Effect → Regulatory Mechanism → Candidate Biological Consequence (Hypothesis)
A real rare-disease example
The most striking demonstration came from an epileptic encephalopathy case. AVI ranked a previously unresolved DNM1 intronic variant highly because AlphaGenome predicted that it creates a cryptic splice acceptor in a brain-specific transcript, producing a 13-amino-acid exon extension. AVI put the variant at number one among all the proband’s de novo variants, not just “highly.”
The splicing prediction was experimentally supported: saturation mutagenesis of the surrounding intron identified 12 variants producing exon extensions, with AlphaGenome achieving an AUPRC of 0.943, comparable to SpliceAI and Pangolin. Conservation-based approaches largely missed these variants. The evidence was enough to recommend a Likely Pathogenic classification for a child whose case had been unsolved for five years despite an epilepsy panel, exome, genome and RNA-seq.
Why it matters
The real advance is not simply a better variant score. AlphaGenome Atlas turns sequence-to-function prediction into a precomputed perturbation map of the regulatory genome.
That could make millions of previously uninterpretable non-coding variants experimentally tractable—while providing a molecular hypothesis for why a particular base may matter.
Important caveat: Atlas remains a predictive research resource. Its motif maps and variant scores do not establish causality or replace experimental or clinical evidence. However, researchers can independently choose any of the predictions and validate or invalidate them.
Reachout to Researchers with Our Extensive Marketing Network
Modern scientific marketing partner built for life science brands



