Skip to main content
Lab Grimoire
TW EN
Coffee
LLM Internals

Nine billion DNA letter swaps, pre-drawn as a map: what AlphaGenome Atlas ranks

DeepMind’s AlphaGenome Atlas precomputes molecular effects for about nine billion possible single-letter DNA changes and ranks them with AVI scores. A research traffic map, not a clinical diagnosis.

Author
CY
Published
Updated
Nine billion DNA letter swaps, pre-drawn as a map: what AlphaGenome Atlas ranks

One letter change is too many to test one by one

The human genome holds roughly 3 billion DNA letters. At each site the letter can swap to one of three alternatives, so about 9 billion single-letter edits are possible.

Picture a megacity road network. Change one alley and traffic lights, junctions, and transit links may all shift. No lab can close every alley and remeasure the city. In September 2026 Google DeepMind released AlphaGenome Atlas: a roughly 1-petabyte map of predicted molecular effects for those edits, plus a rankable AVI score. It is more like printing the map and legend first than asking every traveler to repave the street.

Looking up a map beats re-simulating traffic each time

Only about 2% of the genome is coding sequence, closer to structural pillars on a blueprint. The other 98% behaves more like a control panel: when genes turn on, how strongly, and how RNA is spliced. Many trait-linked variants sit on that panel.

Coding DNA is the blueprint; non-coding DNA is the control panel

AlphaGenome reads up to 1 Mb of DNA and predicts expression, splicing, chromatin accessibility, and other tracks at base resolution. Its 2026 Nature paper reports matching or beating the strongest external models on 25 of 26 variant-effect evaluations. For working scientists the bottleneck is often scale: too many candidates, too little budget for on-the-fly inference. Nature News notes about 9,000 researchers already used the API; pushing that to the whole genome remains impractical for most groups.

Atlas answers that gap by precomputing on hg38 so people can look values up. DeepMind describes a resource more than 30 times larger than the AlphaFold structure database, and it also covers on the order of 100 million short insertions and deletions seen in biobanks. You open a traffic map instead of rerunning the simulator every time.

Model, score, and early real-world cases

Raw tracks still overwhelm. The team therefore trained the AVI (AlphaGenome Variant Impact) score, folding multimodal predictions with features such as AlphaMissense into one rank, then using SHAP to break the score back into splicing, accessibility, conservation, and related legend items. The technical report puts roughly 27,000 experiment-specific scalars on an average variant; AVI training uses allele-frequency 0.1% as a proxy label. High score means “inspect first,” not “court verdict.”

AVI score acts like a map legend for molecular impact

Collaborators have already used the map to shrink the haystack. In a GREGoR-linked rare-disease effort, AVI helped surface a DNM1-related candidate; the model pointed to a misplaced splice junction and an abnormally extended protein, later aligned with experimental screens. At the University of Exeter, whole-genome data from more than 54,000 UK Biobank participants yielded about 22% more non-coding associations when rare variants were grouped by predicted molecular effect; focusing on the top 1% most impactful non-coding variants flagged 19 regions for body-mass index. It is like painting congestion hotspots red before sending a repair crew.

From variant haystack to ranked hypothesis and lab check

The path is candidate list, AVI ranking, mechanistic hypothesis, then lab checks. Phenotype and family context still sit with you. The map shortens the route; it does not walk it for you.

A map is not a diagnosis

Do not read a high AVI as proof of pathogenicity. The boundary is sharper than the headline.

Take this home: genome-wide single-letter edits can be pre-drawn as a research map, and ranked molecular legends can speed hypotheses. Do not take home a personal genetic report or a prescribing basis. DeepMind states the resource is not clinically validated; AVI’s frequency labels are proxies; training data still miss cell types and some RNA classes; ultra-long-range and trans effects remain incomplete. Commercial terms differ from academic ones, and identifiable sequences need consent before they leave the clinic.

Next time a headline says AI “understands” the genome, ask whether you were handed traffic ranks or a final verdict. Then ask whether anyone closed that alley in a real experiment.

Frequently Asked Questions

How is AlphaGenome Atlas different from the AlphaGenome model?

AlphaGenome is the sequence model that predicts many molecular tracks from DNA. Atlas is the precomputed map of variant effects across the human reference genome, plus AVI scores and a browser, so you look things up instead of re-running full inference for every candidate.

Does a high AVI score mean a variant causes disease?

No. AVI compresses multimodal molecular predictions and some protein or conservation features into a rankable score, trained with allele-frequency proxies. It helps prioritize what to inspect first. Diagnosis still needs phenotype, segregation, and experiments.

How can researchers use it today, and is it clinical?

Non-commercial research access is offered via the portal and API; commercial use is a separate cloud path. DeepMind states the resource is not validated or approved for clinical use and is not a substitute for professional care. Identifiable clinical sequences still need consent and data agreements before external queries.

Found this useful?

Follow for new AI × biomedical research notes:

Or buy me a coffee to keep new content coming.

☕ Buy Me a Coffee