Skip to content

25.6 — From Target to Molecule

A company has decided which protein to block. Now it needs a physical substance that blocks it, that reaches the right place in a human body, that stays there long enough to matter, that does not poison anything else, and that can be manufactured by the ton.

That is five separate problems, and a molecule that solves four of them is worthless. This chapter is how the industry works through all five, and it is the stage where the word "chemistry" stops being a school subject and becomes a project plan.

First you need a way to see whether anything is happening

Before any molecule is tested, somebody has to build an assay. An assay is a laboratory test that produces a number telling you how much your target's activity changed.

Say the target is an enzyme that adds a phosphate group to another protein, driving a cancer cell to divide (Chapter 1.8). The assay puts the enzyme, its substrate and a detection chemistry into a tiny well, and produces a light signal proportional to the reaction. Add a compound that blocks the enzyme, and the light drops.

The assay decides the quality of everything that follows, and two properties are argued about constantly. It must be robust, meaning the difference between full activity and no activity is far larger than the random variation between wells. And it must measure the thing you care about rather than something adjacent — a compound that quenches the light signal looks exactly like a compound that blocks the enzyme, and such false hits waste months.

The result of any assay is usually reported as an IC50 or an EC50. IC50 is the concentration of the compound that reduces the target's activity by half; EC50 is the concentration that produces half the maximum effect. A lower number means a more potent compound, and the units are typically nanomolar or micromolar. Those two abbreviations appear on nearly every discovery slide you will ever see.

High-throughput screening: testing everything you own

A large pharmaceutical company keeps a physical library of compounds, often one to two million of them, stored as tiny amounts in plates. High-throughput screening runs the assay against all of them.

The scale is genuinely industrial. Plates hold 384 or 1,536 wells, robots move them, liquid handlers dispense volumes measured in billionths of a litre, readers photograph the result, and a system stores tens of millions of measurements with the plate, well, compound identifier and run conditions attached.

Out of a million compounds you might get a few thousand that appear active. Then the work begins, because most of them are lying.

Some are interfering with the detection chemistry rather than the target. Some are compounds that stick to any protein — the industry calls these frequent hitters. Some form microscopic aggregates that trap the enzyme non-specifically. Some come from a well where the compound had degraded. So hits are re-tested, tested at several concentrations to give a proper dose–response curve, tested in a second assay of a different design, and confirmed with fresh material of verified purity. What survives is called a confirmed hit, and a screen that started with a million compounds may deliver a few dozen.

Two other routes to a starting point are common enough to name.

Fragment-based discovery screens very small molecules that bind weakly, then grows or joins them. It covers chemical space far more efficiently, because small pieces can be combined in many ways.

Structure-based design starts from the three-dimensional shape of the target — determined experimentally or, increasingly, predicted computationally as described in Chapter 25.4 — and designs a molecule to fit the pocket. This is where the AI work of that chapter actually lands in practice.

Hit to lead, then lead optimisation: the loop that consumes the years

A hit blocks the target in a dish. A drug candidate must do far more, and getting from one to the other is a repeating cycle: design a change, make it, test it, learn, design again. A medicinal chemistry team may run that loop for three to five years on one programme.

What they are optimising is a set of properties that fight each other, which is the central difficulty of the whole discipline.

Potency — the compound must work at a low concentration, because the dose has to be swallowable.

Selectivity — it must hit the intended target far more strongly than related ones. Human kinases number in the hundreds and look alike; a compound hitting fifty of them will have effects everywhere.

Absorption — a tablet must dissolve, survive stomach acid, and cross the gut wall into the blood (Chapter 22.2).

Metabolic stability — the liver's enzymes will attack it, and a compound destroyed in twenty minutes cannot be a once-a-day tablet.

Distribution — it has to reach the tissue where the disease is. A brain drug must cross the blood–brain barrier (Chapter 11.10); a drug for an infection in bone must reach bone.

And a clean safety profile — no blocking of the heart's hERG potassium channel, which lengthens the electrical recovery of heart cells and can cause a lethal rhythm (Chapter 18.5); no damage to DNA; no unexpected toxicity in liver cells.

The frustration of the work is that these trade against each other. Adding a chemical group that improves potency often makes the molecule greasier, which worsens solubility and increases how strongly it binds to blood proteins, which reduces the free drug available to act. Chemists talk about a molecule "getting fat" during optimisation, and fighting that is a real discipline.

A rough guide has been used for decades as a sanity check on oral drugs — Lipinski's rule of five, published in 1997 after looking at what successful oral drugs had in common. Molecular weight under 500 daltons; not too greasy, measured by a partition value under 5; no more than five hydrogen-bond donors; no more than ten acceptors. It is a guideline, violated by plenty of real drugs, and its value is as a warning that you are leaving well-explored territory.

The vocabulary of the stages is worth having straight, because project plans are written in it. A hit is any confirmed active compound. A lead is a hit series with acceptable potency and a workable profile, chosen for full optimisation. A candidate is the single molecule the company commits to development — the point at which the money changes scale, since everything in Chapter 25.7 onwards costs far more than everything before it.

The ADME and safety work that runs alongside

ADME stands for absorption, distribution, metabolism and excretion — what the body does to the drug, taught physiologically in Chapter 22.2. In discovery it is measured on every promising compound, and the tests are quick versions of the real thing.

Solubility and permeability are measured in artificial systems and in a layer of intestinal-type cells grown across a membrane. Metabolic stability is measured by adding the compound to liver microsomes — the fragment of liver cell containing the drug-destroying enzymes — and watching how fast it disappears. Protein binding is measured because only the unbound fraction is active. And drug–drug interaction potential is measured by checking whether the compound blocks or induces the liver enzymes that clear other medicines, which is the mechanism behind the dangerous combinations in Chapter 22.13.

The compound is also profiled against a standard panel of unrelated receptors, ion channels and enzymes, to find out what else it hits before an animal or a person does.

This whole set of data is why the industry stopped losing programmes to pharmacokinetics. In the early 1990s a large share of clinical failures were compounds that simply did not reach useful blood levels. Moving these measurements early, into discovery, largely removed that failure mode — which is worth remembering, because it is a genuine example of a process change fixing a systemic problem, and it is the pattern a services company should be looking for.

When the drug is a biologic, the route is completely different

Everything above assumes chemistry. If the product is an antibody, none of it applies, because you are not designing a molecule — you are finding an immune system that has already designed one, or engineering one that already exists.

Three discovery routes dominate.

Immunise an animal. A mouse is exposed to the target protein, its immune system generates antibodies, and the antibody-producing cells are captured and made immortal so they keep producing one identical antibody. This is the classical method from Chapter 25.2.

Display libraries. Enormous collections of antibody fragments are displayed on the surface of viruses or yeast cells, poured over the immobilised target, and whatever sticks is kept and amplified. Repeat, and binders emerge from a library of billions.

Transgenic animals. Mice engineered to carry human antibody genes produce human antibodies directly, which removes the humanisation step.

Then comes the part with no equivalent in small molecules: developability assessment. The candidate antibody must not clump at high concentration, must be stable at the temperatures it will see, must not be chewed up during purification, must survive freezing and thawing, and must be produced at high enough yield by the cells that will make it. A beautiful binder that aggregates in the vial is not a product.

And the last discovery step is building the factory cell. The gene for the antibody is inserted into a cell line — most commonly Chinese hamster ovary cells — thousands of resulting clones are screened for yield and product quality, one is chosen, and a frozen bank of that clone is created and stored in multiple locations. Every vial the product will ever be made from descends from that bank. Chapter 25.11 shows what is then built around it.

What this stage produces, and what it means for your systems

At the end of discovery a company has: one candidate molecule; a synthetic route to make it; a data package on potency, selectivity, ADME and early safety; a patent application filed; and an internal committee decision to spend the next several hundred million dollars.

The data behind that decision lives in systems your company may well build or run. Compound registration and inventory, electronic laboratory notebooks, screening data management, structure–activity databases, and the analysis environment where chemists compare thousands of compounds across a dozen properties at once.

Two things about this software are worth knowing before you touch it.

Most discovery systems are not GxP-regulated, because no regulatory decision and no patient depends directly on them. This makes discovery the least regulated software environment in the industry, and it is where a services company can move fastest.

But intellectual property makes them extremely sensitive anyway. The date and content of an experiment can decide who owns a patent worth billions, so electronic laboratory notebooks carry strict rules about signing, witnessing and immutability that look very much like the audit trail requirements of Chapter 25.20 — arriving from patent law rather than from drug law. Confusing "not GxP" with "not important" is a mistake that ends engagements.

Next: Chapter 25.7, the work that stands between a chosen molecule and the first human being — animal safety studies, the rules that govern them, and how the first dose in a person is calculated.