OpenAI’s recent announcement for GPT-Rosalind — named after Rosalind Franklin — is either an homage or an irony, depending on how you feel about what follows. Franklin’s legacy was built on generating the hard, physical, indisputable X-ray diffraction data that made biological reasoning possible. GPT-Rosalind, conversely, is a “frontier reasoning model built to accelerate drug discovery, genomics analysis, and protein reasoning” — one that promises to reason its way to the finish line, without a complete picture of the underlying physical data that would make that reasoning valid. Amgen, Moderna, Thermo Fisher, and the Allen Institute are already on board.
I’ve been in this field for four decades. I’ve been to this rodeo. Combinatorial chemistry was going to change everything. Then solving the proteome. Enough time has passed that both now register as things people have “heard about.” Entire revolutions, complete with certainty and compressed timelines, have already come and gone. For many in the field today, they’re vague historical artifacts rather than lived experience. Biology, for its part, did what it always does. It refused to cooperate.
For a decade now, I’ve been hearing echoes of this exact claim — some from companies that didn’t survive the promise — that if we just embrace AI at scale, the long-standing complexity of the field will resolve into alchemy and cures.
Here’s what GPT-Rosalind — and every model before it — cannot fix: we don’t understand how the 20,000+ proteins in each human cell interact to create the most complex, most redundant, and most indecipherable machine the world has ever seen. The data aren’t there for the taking. They haven’t, in large part, been measured. AI cannot know what it has never seen, cannot predict what has never been suggested to it. And the literature it reasons over is what scientists thought they saw, not what was actually happening at the atomic level. No amount of reasoning over existing literature changes the fundamental shape of what we don’t know. This isn’t a reasoning problem. It’s a systems biology problem. We’ve observed perhaps 20% of protein–protein interactions in a human cell, much of it noisy or inferred. The rest hasn’t been measured. You can’t surface connections in a graph that hasn’t been built.
Drug discovery succeeds through rigor, intuition, and luck — not the luck of randomness, but the luck of the prepared mind encountering the right biology at the right moment, which is a different thing entirely and cannot be scheduled.
AI is a genuine complement to physics-based chemistry — I’ve argued this at length elsewhere, and I believe it. But complement is not replacement. It can’t close the last mile of affinity optimization, where the physics has to be right or the drug doesn’t bind. And affinity is the easy part. Wait until GPT-Rosalind meets ADMET — the absorption, distribution, metabolism, excretion, and toxicity profile that determines whether a molecule that works in a dish survives contact with an actual human body. That’s where most drugs die. That’s where the unmeasured biology lives. No interpolation engine fixes that.
Consider BenevolentAI. It was the industry’s first grand bet on mapping the “unseen” via a knowledge graph. It had its hero moment in 2020, identifying baricitinib for COVID-19 in ninety minutes. But repurposing an existing drug is just reasoning over a known map. When the platform moved into novel targets, it hit the systems biology wall — not because the AI failed to do what AI does, but because the map it was reasoning over was incomplete. The target was real. The redundancy of human immune pathways was not in the data.
It turns out that when the data isn’t there, AI just helps us optimize our ignorance.
We can do better, and probably will. But in ten years, people will still be getting sick and dying. AI or not. I hope somewhat fewer. I hope we can point to these models and nod at the value they contributed. I genuinely do.
The bet I’m making isn’t that AI fails. It’s that the disease stays harder than the announcement.
Prove me wrong. This is a bet I’d love to lose.



Very well articulated. Please see this book chapter which we tried to argue in 2020.
A Few Guiding Principles for Practical Applications of Machine Learning to Chemistry and Materials | Machine Learning in ChemistryThe Impact of Artificial Intelligence | Books Gateway | Royal Society of Chemistry
https://books.rsc.org/books/edited-volume/1902/chapter-abstract/2496236/A-Few-Guiding-Principles-for-Practical?redirectedFrom=fulltext
The map is not the territory because the territory is continuously recalculating itself.