What Combinatorial Chemistry can Teach us about the AI revolution in Mathematics and Informatics

Michael T. M. Emmerich, University of Jyväskylä, Finland, September 2026

Chemistry has been an application field of mine for many years, including work on multi-objective optimisation and artificial intelligence in drug design [LMEW2023]. This is one reason why the present discussion about AI in mathematics and informatics feels familiar to me. Drug discovery has already spent decades learning what happens when machines become much better at searching enormous combinatorial spaces.

A small example leads surprisingly far. How many saturated, acyclic hydrocarbons can be built from six carbon atoms? The answer is five.

The title picture shows these five alkanes as carbon trees: n-hexane, 2-methylpentane, 3-methylpentane, 2,3-dimethylbutane and 2,2-dimethylbutane; branching carbons are shown in dark blue. This was exactly the kind of enumeration problem that first sparked my interest in both organic chemistry and graph theory when I was still a schoolboy in Coesfeld, Germany. I learned everything I could about organic chemistry from an introductory textbook my mother had used during her training as a medical technical assistant and a dedicated chemistry teacher in the local gymnasium.

Looking back, this is also part of why I care about the educational side of the current AI debate. Simple examples, worked through by hand, can create an intuition and curiosity that stay with us for a lifetime. In my case, one such example connected chemistry and mathematics long before I knew that this connection had played a role in the history of graph theory itself.

1. Chemistry helped to create graph theory

In the nineteenth century chemists began drawing molecules as atoms connected by bonds. Arthur Cayley noticed that a saturated hydrocarbon without rings has a carbon skeleton that is a tree. In 1874 he turned the chemical question “how many isomers exist?” into a problem of enumerating trees with restricted vertex degrees [Cayley1874].

For an alkane with {n} carbon atoms, every carbon has valence four and the carbon skeleton has {n-1} edges. Hence

\displaystyle {h=4n-2(n-1)=2n+2,}

which gives the familiar formula CnH2n+2. Counting alkanes is therefore counting non-isomorphic trees of maximum degree four. There are six non-isomorphic trees on six vertices; one has a vertex of degree five and is chemically impossible. The remaining five are the hexanes above.

James Joseph Sylvester then borrowed the chemists’ diagrams for algebra. In Nature in 1878 he wrote of a “graph precisely identical with a Kekuléan diagram or chemicograph” [Sylvester1878]. The word survived. So did the deeper idea: a chemical structure is a combinatorial object that can be counted and manipulated mathematically.

The numbers grow quickly. For {n=1,2,\ldots,14} carbon atoms the alkane counts begin

1, 1, 1, 2, 3, 5, 9, 18, 35, 75, 159, 355, 802, 1858.

Pólya later put such enumeration on a general footing [Polya1937]. Even for these very simple molecules the number grows asymptotically roughly like {2.8155^n n^{-5/2}} [OEIS].

2. Sometimes the important step is to leave the search space

The most famous six-carbon hydrocarbon (the orange molecule in the title picture) is not enumerated by this procedure: benzene, C6H6. No enumeration of saturated carbon trees, however fast, could ever have produced benzene. It lies outside that search space.

In 1865 Kekulé proposed the six-membered carbon ring with alternating bonds. Twenty-five years later he told the famous story of dozing by a fireplace in his hometown Ghent and seeing a snake seize its own tail. Whether taken literally or not, the image came to a mind that had spent years thinking about valence and carbon chains. Pasteur’s phrase fits well: chance favours the prepared mind.

For me, this is the interesting point. The decisive step was not to enumerate trees faster. It was to change the space of possibilities. That is one thing human intuition can do.

3. The chemical ocean

Once we allow rings, multiple bonds and elements such as nitrogen, oxygen, sulfur and halogens, the space becomes enormous. Reymond and co-workers enumerated 166.4 billion small organic molecules with up to 17 heavy atoms under their rules [RDBR2012]. A widely quoted estimate puts the number of possible drug-like molecules around {10^{60}} or more [BMG1996].

An ocean of {10^{60}} possible drug-like molecules remains an ocean, no matter how fast the compute is.

This is where my own work connects to the story. In multi-objective drug design, one is not usually looking for the molecule that maximises a single number. Potency, selectivity, toxicity, solubility, synthesizability and other properties have to be balanced. AI can generate and rank candidates, and it can help navigate trade-offs, but the central problem remains one of orientation: which region of chemical space is worth exploring, which objectives matter, and which model should we trust? [LMEW2023]

Chemistry has used combinatorial libraries, high-throughput screening, virtual screening, structure-based design and now machine learning for decades. These methods have enabled important discoveries. Yet the overall productivity of pharmaceutical R&D has not risen in proportion to our computational power. Scannell and co-authors famously described a long decline in new drugs per inflation-adjusted research dollar as Eroom’s law [SBBW2012].

There are many reasons for this decline. But one possibility worth thinking about is that something hard to formalise can be weakened when search becomes increasingly automated: chemical intuition and tacit knowledge. Lombardino and Lowe argued already in 2004 for restoring more of the medicinal chemist’s creative role [LL2004]. Hann and Keserű recently made a related case in the age of AI [HK2025].

We may end up fishing with better and better tools while gradually understanding less well where to fish, and why.

4. The same ocean in mathematics and informatics

The analogy is not exact, but it is close enough to be useful. Proofs, algorithms and programs also have finite combinatorial representations, and candidate spaces grow explosively. A Boolean formula with {n} variables has {2^n} truth assignments; a travelling-salesperson tour through {n} cities is one of {(n-1)!/2}. A more practical example from my own work is home health care routing and scheduling: caregivers have to be assigned and routed to patients while respecting time windows, qualifications, continuity of care and other constraints. The resulting optimisation problems are NP-hard and have very direct consequences for patients, caregivers and health-care costs. In such applications compute costs can explode quite fast [ABE2025].

The P versus NP question asks whether the gap between checking and finding is fundamental. A problem is in NP if a proposed solution can be checked in polynomial time, and in P if a solution can be found in polynomial time. Cook’s 1971 paper was tellingly titled “The complexity of theorem-proving procedures” [Cook1971]. If, as almost everybody believes but nobody has proved,

\displaystyle {\mathrm{P} \neq \mathrm{NP},}

then no algorithm, however clever and however much compute it is given, solves all these combinatorial oceans efficiently in the worst case. A neural network, a language model or an AI agent is, in the end, an algorithm. Thousands of GPU cores can move the boundary of what we can solve, but they do not make combinatorial explosion disappear.

AI can find excellent solutions for particular instances, learn useful structure, suggest heuristics and occasionally produce a brilliant idea. But if P ≠ NP, it cannot guarantee optimal answers to NP-hard problems in polynomial time. The machines, too, have to fish. What helps is structure: knowing that a particular family of instances has a shape that can be exploited. In chemistry we might call the analogous ability chemical intuition.

And still, in polynomial time we can do impressive things.

5. What I take from this

If AI increasingly generates proofs, algorithms and code, an important educational question is what happens when fewer people develop the expertise needed to understand why something works, where a genuinely new idea might be found, or when the machine is heading in the wrong direction.

Chemistry does not prove that this will happen in mathematics or computer science. But it gives us a useful experiment to look at. Better search tools do not automatically give better orientation. And decisive ideas sometimes change the search space instead of searching the existing one faster.

That seems to me a good argument for continuing to give students a solid foundation and the opportunity to develop their own intuition, while also teaching them how to use AI effectively. Draw the five hexanes by hand before asking a program to enumerate the 1858 tetradecanes. Learn to prove and to program before delegating everything to a machine.

The two are not in conflict. The second relies on the first.

Code

The enumeration behind the figure is essentially one line of Python with networkx:

alkanes = [t for t in nx.nonisomorphic_trees(n) if max(d for _, d in t.degree()) <= 4]

The full script also identifies the main chain, names the five isomers and draws the figure.

Please cite this note as: M. T. M. Emmerich, “What Combinatorial Chemistry Can Teach Us About the AI Revolution in Mathematics and Informatics”, Mathematical Playground / MODA News, emmerix.net, September 2026.

References

[ABE2025] S. Atta, V. Basto-Fernandes, M. T. M. Emmerich. A Concise Review of the Home Health Care Routing and Scheduling Problem. Operations Research Perspectives 15:100347, 2025. doi:10.1016/j.orp.2025.100347

[BMG1996] R. S. Bohacek, C. McMartin, W. C. Guida. The art and practice of structure-based drug design: a molecular modeling perspective. Medicinal Research Reviews 16:3–50, 1996.

[Cayley1874] A. Cayley. On the mathematical theory of isomers. Philosophical Magazine 47:444–447, 1874.

[Cook1971] S. A. Cook. The complexity of theorem-proving procedures. STOC 1971, 151–158.

[HK2025] M. M. Hann, G. M. Keserű. The continuing importance of chemical intuition for the medicinal chemist in the era of Artificial Intelligence. Expert Opinion on Drug Discovery 20:137–140, 2025.

[LMEW2023] S. Luukkonen, H. W. van den Maagdenberg, M. T. M. Emmerich, G. J. P. van Westen. Artificial intelligence in multi-objective drug design. Current Opinion in Structural Biology 79:102537, 2023. doi:10.1016/j.sbi.2023.102537

[LL2004] J. G. Lombardino, J. A. Lowe III. The role of the medicinal chemist in drug discovery – then and now. Nature Reviews Drug Discovery 3:853–862, 2004.

[OEIS] OEIS A000602: Number of n-node unrooted quartic trees; number of n-carbon alkanes ignoring stereoisomers. oeis.org/A000602

[Polya1937] G. Pólya. Kombinatorische Anzahlbestimmungen für Gruppen, Graphen und chemische Verbindungen. Acta Mathematica 68:145–254, 1937.

[RDBR2012] L. Ruddigkeit, R. van Deursen, L. C. Blum, J.-L. Reymond. Enumeration of 166 billion organic small molecules in the chemical universe database GDB-17. Journal of Chemical Information and Modeling 52:2864–2875, 2012.

[SBBW2012] J. W. Scannell, A. Blanckley, H. Boldon, B. Warrington. Diagnosing the decline in pharmaceutical R&D efficiency. Nature Reviews Drug Discovery 11:191–200, 2012.

[Sylvester1878] J. J. Sylvester. Chemistry and algebra. Nature 17:284, 1878.

Leave a comment