Appearance
2.9 — Genomics, Genetic Testing and CRISPR
The Human Genome Project ran from 1990 to 2003, cost about three billion dollars, and involved thousands of people across six countries. It produced one composite human genome sequence, and it was declared essentially complete in 2003 — though about 8 percent of the genome, mostly highly repetitive regions, stayed unreadable until 2022, when the Telomere-to-Telomere consortium finished it using long-read sequencing.
Today a complete human genome can be sequenced in about a day for a few hundred dollars. That is a cost reduction of roughly ten million fold in twenty years — faster than the improvement in computer chips over the same period, which is the usual benchmark for rapid technological change (Volume I, Chapter 1.5).
This page covers what that capability actually delivers, and — just as importantly — what it does not.
How sequencing works
Sanger sequencing, developed by Frederick Sanger in 1977, works by sabotage. You copy the DNA, but you mix in a small proportion of modified bases that lack the 3′ attachment point needed to add the next one — so whenever one is incorporated, that copy stops there. Each of the four modified bases carries a different fluorescent label. The result is a mixture of fragments of every possible length, each ending in a labelled base. Separate them by size, read the colours in order, and you have the sequence.
It reads about 500 to 1,000 bases per run, accurately. It is still used today for checking a single gene or confirming a single variant.
Next-generation sequencing does the same job massively in parallel. The DNA is chopped into millions of short fragments, each is amplified into a cluster on a flow cell, and the sequencing is imaged one base at a time across all clusters simultaneously. Each read is short — 100 to 300 bases — but there are hundreds of millions of them, and a computer assembles them by finding overlaps. The bottleneck moved from chemistry to computation, which is why genomics is now as much a software discipline as a laboratory one.
Long-read sequencing pushes single molecules through a tiny pore and reads the changing electrical current as each base passes. Reads run to tens or hundreds of thousands of bases. Accuracy per base is lower, but long reads span repetitive regions that short reads cannot resolve — which is exactly what finished the last 8 percent of the genome in 2022.
What gets sequenced, and when
Three levels, and choosing between them is a real clinical decision.
A targeted panel reads a specific set of genes known to matter for a particular question — a cancer panel, an inherited-deafness panel, a cardiomyopathy panel. Cheapest, fastest, deepest coverage, and the least ambiguity, because you only get results about genes you were asking about.
Whole exome sequencing reads only the protein-coding parts — about 1 to 2 percent of the genome, but where the large majority of known disease-causing mutations sit. Costs a fraction of a whole genome.
Whole genome sequencing reads everything, including regulatory regions and structural changes an exome misses.
The diagnostic yield is worth being honest about. In a child with unexplained developmental delay or a suspected rare genetic disease, exome or genome sequencing finds a causative answer in roughly 30 to 40 percent of cases. That is a genuine transformation compared with the diagnostic odyssey families used to face — years of tests with no answer — and it also means the majority still go undiagnosed.
Variants of uncertain significance: the real problem
When you sequence a genome you find around 4 to 5 million places where it differs from the reference. Almost all are harmless. The task is deciding which ones matter, and that is much harder than sequencing.
Variants are classified on a five-point scale: pathogenic, likely pathogenic, uncertain significance, likely benign, benign.
The uncertain category is large and it is a genuine clinical burden. A variant nobody has seen before, in a gene that matters, changing an amino acid — is it the cause of the patient's illness or an irrelevant quirk? Often there is no way to tell yet.
This causes real harm when it is mishandled. A woman told she has a "BRCA variant of uncertain significance" may request preventive surgery she does not need. Reclassification happens constantly as more data accumulates — a variant called uncertain in 2018 may be reclassified as benign in 2024 — which is why laboratories increasingly recontact patients, and why an uncertain result should be treated as no result rather than a weak positive.
And the databases are skewed. Most sequence data comes from people of European ancestry, so a variant common and harmless in a West African or South Asian population is more likely to appear novel and therefore suspicious. This produces measurably worse genetic care for people of non-European ancestry, and it is a data problem rather than a biological one.
Direct-to-consumer testing: what it does and does not tell you
The consumer tests that send you a tube to spit in mostly do not sequence anything. They use a genotyping array, which checks a few hundred thousand pre-selected positions known to vary between people. That is under 0.02 percent of your genome.
What they do reasonably well:
Ancestry estimation, by comparing your pattern of variants with reference populations. The continent-level inferences are solid. The fine-grained percentages are estimates against a particular reference panel, which is why the same sample sent to two companies gives different answers, and why your reported ancestry can change when a company updates its panel.
Relative finding. This works very well, because shared long stretches of identical DNA are unambiguous evidence of a recent common ancestor. It works well enough that it has reunited families, revealed unexpected parentage in large numbers of cases, and — through public genealogy databases — identified criminal suspects from distant relatives' DNA, most famously the Golden State Killer in 2018. People uploading their data for genealogy did not anticipate that use, and it is a genuine and unresolved privacy question, because you cannot consent on behalf of your relatives whose DNA you partly share.
What they do poorly, and where the harm is:
Health risk estimates from these arrays are frequently wrong in both directions. The tests check a small selected set of variants. A test that checks three specific BRCA variants — the ones common in Ashkenazi Jewish populations — will miss the thousands of other pathogenic BRCA variants entirely. A reassuring result is therefore close to meaningless, and people have declined clinical testing on the strength of one.
In the other direction, raw data files from these services are often run through third-party interpretation tools, and a large study found that around 40 percent of variants flagged as pathogenic in such raw data were false positives when checked in a clinical laboratory. The array was never designed for clinical-grade calling of rare variants.
Polygenic risk scores add up the small effects of thousands of variants to estimate risk of common conditions like heart disease or type 2 diabetes. They are scientifically real and improving. They are also weakly predictive at the individual level, poorly transferable between ancestries, and almost never change what you should do — which for heart disease is the same advice regardless of the score. Chapter 18.1.
The honest summary: these tests are entertaining, genuinely good for ancestry and relatives, and should not be used to make a medical decision without confirmation in a clinical laboratory.
Gene therapy
The idea is straightforward: if a disease is caused by a missing or broken gene, deliver a working copy.
Delivery is the hard part, because DNA does not cross membranes and is destroyed in blood.
Viral vectors are the main solution — a virus stripped of its ability to replicate, carrying the therapeutic gene instead. Adeno-associated virus (AAV) is the workhorse: it does not usually integrate into the genome, provokes relatively little immune response, and different natural variants prefer different tissues, so the vector can be chosen to target liver, muscle, retina or nervous system. Its limitation is cargo size — about 4.7 kilobases, which is too small for large genes like dystrophin.
Lentiviral vectors integrate permanently and are used to modify cells outside the body, which are then returned.
Two approaches:
In vivo — inject the vector into the patient. Luxturna for an inherited retinal dystrophy is injected under the retina and restores useful vision. Zolgensma for spinal muscular atrophy type 1 is a single intravenous infusion given to infants; untreated, most of these children died or needed permanent ventilation before two, and treated infants are sitting, standing and in some cases walking. It costs about 2.1 million dollars, making it among the most expensive medicines ever priced.
Ex vivo — remove the patient's cells, modify them in the laboratory, return them. Used for blood disorders and for CAR-T cancer therapy (Chapter 19.7).
The field's history includes a serious failure and it should be stated. In 1999 Jesse Gelsinger, an 18-year-old in a gene therapy trial, died of a massive immune reaction to the viral vector. In the early 2000s, several children treated successfully for severe combined immunodeficiency developed leukaemia because the vector had integrated next to a growth-promoting gene. Both events halted the field for years, and both drove the safety engineering — safer vectors, better immune screening — that made the current generation of approved therapies possible.
CRISPR
Where it came from. CRISPR is a bacterial immune system. Bacteria store short pieces of DNA from viruses that have attacked them, in an array of Clustered Regularly Interspaced Short Palindromic Repeats — which is what the acronym stands for. They transcribe those stored pieces into guide RNAs, and a Cas protein uses each guide to find and cut matching viral DNA on a future attack. It is an adaptive immune system in a bacterium, with a genetic memory of past infections.
Jennifer Doudna and Emmanuelle Charpentier showed in 2012 that the system could be reprogrammed by simply supplying a guide RNA of your choosing, and that it would then cut any chosen DNA sequence. They shared the 2020 Nobel Prize in Chemistry.
Why it changed everything. Targeted DNA cutting existed before — zinc finger nucleases and TALENs did the same job — but retargeting them meant protein engineering, which took months and specialist skill. Retargeting CRISPR means ordering a different twenty-base RNA, which costs a few dollars and arrives in days. The capability was not new; the accessibility was, and that is what turned it from a technique into a revolution.
What happens after the cut. Cas9 makes a double-strand break, and the cell repairs it using the systems from Chapter 2.5.
Non-homologous end joining is fast and sloppy and usually loses or gains a few bases — which causes a frameshift and destroys the gene. So the easy application is knocking a gene out.
Homology-directed repair can be exploited by supplying a template, letting you write in a specific new sequence. This is far less efficient, and it barely works in cells that are not dividing — which includes neurons, muscle and liver cells. Precisely correcting a mutation is much harder than breaking a gene, and popular coverage routinely blurs this.
Newer tools address that. Base editors chemically convert one base to another without cutting both strands at all — C to T, or A to G — which covers a large share of known pathogenic point mutations. Prime editors carry a template and write short specified sequences in directly. Both are more precise and both are in clinical development.
Off-target effects are the main safety concern: the guide can partly match other places in the genome and cause cuts there. Guides are now designed computationally to minimise this and the whole genome is checked afterwards for unintended edits.
What has actually been achieved
Casgevy, approved in the UK and US in late 2023 for sickle cell disease and transfusion-dependent beta-thalassaemia, is the first approved CRISPR therapy. Note carefully what it does: it does not correct the sickle mutation. It disables a regulator called BCL11A in the patient's own blood stem cells, which reactivates fetal haemoglobin — a workaround using a gene knockout, which is the easy operation, rather than a correction. That choice is a direct consequence of the mechanism above.
In vivo editing has been demonstrated. A 2021 trial injected CRISPR components into patients with transthyretin amyloidosis and reduced the abnormal protein by over 80 percent from a single dose — the first demonstration that editing can be done inside the body rather than in a dish.
Agriculture and research are where the technology has had its largest quiet effect: disease-resistant crops, and the ability to knock out any gene in any model organism in weeks rather than years.
The line that was crossed
In November 2018, He Jiankui announced that he had used CRISPR on human embryos and that twin girls had been born with edited genomes. He had disabled the CCR5 gene, which HIV uses to enter cells, aiming to make the children HIV-resistant.
The scientific community's condemnation was close to universal, and it was not squeamishness. The specific objections:
There was no medical need. The father was HIV-positive, but transmission from father to child is preventable by sperm washing, which is routine and safe. Editing an embryo solved a problem that already had a solution.
Embryo editing changes every cell including the germline, so the changes pass to all descendants, who cannot consent.
The edits were not the intended ones. The actual sequences produced were not the known protective CCR5 variant but novel deletions with unknown effects, and at least one child was a mosaic, with different edits in different cells.
And CCR5 does other things. Loss of it raises susceptibility to West Nile virus and appears to worsen influenza outcomes.
He was sentenced to three years in prison in China. The episode produced a broad international consensus that heritable human genome editing should not proceed at present, and that consensus has held. Editing body cells to treat a sick patient is a different matter entirely and is proceeding normally — the distinction is between changing one person with their consent, and changing every one of their descendants without it.
What the next page fixes
Part 2 has covered the code from the molecule to the clinic. Chapter 2.10 closes it with a question everyone has asked and few have been answered properly: why are fingerprints unique, how did we establish that, and how does DNA fingerprinting actually work — including the specific arithmetic behind the "one in a billion" figures quoted in courtrooms, and where that arithmetic has gone badly wrong.