
In 1969 Volker Strassen found a way to multiply two 4×4 matrices using 49 scalar multiplications instead of the obvious 64. The trick went into textbooks and stayed there. For 56 years the best mathematicians alive tried to shave off one more multiplication and failed. Last May, Google DeepMind's AlphaEvolve found 48.
The broken record was impressive. The silence around it was more interesting. AlphaEvolve can lay out all 48 steps and prove they work. It can't tell you why 48 is possible and 49 held as a wall for half a century, and neither can the DeepMind researchers who built it. The algorithm exists. The account of why it exists does not.
We've been reading the softer version of this story for a while. Google's CEO says AI now writes about a quarter of the company's new code, every line still reviewed by an engineer, and separately admits nobody fully understands why some of the code works. Engineers who've tried to read the most heavily optimized AI-written functions describe them as correct and close to unreadable. Those are the mild cases. The sharp one is in a lab.
In June 2025, a drug called rentosertib became the first with both an AI-chosen target and an AI-designed molecule to publish Phase IIa results, showing improved lung function in patients with idiopathic pulmonary fibrosis (Insilico Medicine, in Nature Medicine). Promising, and in the way that matters here, not fully explained. When a molecule is designed by a system searching a chemical space no human holds in their head, "mechanism of action" becomes something the chemists reconstruct afterward, if they can. The FDA noticed. Its January 2025 draft guidance on AI in drug development covers regulatory decisions and explicitly leaves early discovery alone, which means the hardest question is the one nobody's regulating yet: how do you test something for safety when you can't say how it works?
For most of human history, that was the normal condition. Willow bark treated pain and fever for thousands of years before anyone had heard of salicylic acid. James Lind proved citrus stopped scurvy in 1747; the reason, vitamin C, arrived almost 200 years later. Reliable first, explained much later, or never. The mechanism was always the luxury item.
The Enlightenment sold us a different deal, and we came to expect delivery: anything real could, with enough work, be understood, and understanding was the price of admission for trusting a result. Science made that expectation feel like a law of nature. It was closer to a cultural promise, and a recent one. AI is quietly walking it back. We're returning to willow bark, except the bark now designs molecules and beats 56-year-old theorems, and it does it faster than our explanations can keep up.
The willow bark was different in one way that matters. Nobody understood it, so nobody was left behind; the ignorance was shared all the way down. The new discomfort is that the gap now runs straight through the specialists themselves. The mathematicians, the drug designers, the engineers reviewing their quarter of the codebase are the ones who can't fully follow it, and the rest of us inherit their uncertainty without even the option of catching up.
I've argued before that we're organic prediction machines that backfill theory onto whatever worked. If that's right, a machine producing working results without a theory is running a purer version of what we've always done. We just told ourselves a flattering story about the theory coming first.
Which is fine, until you ask the machine a different kind of question. Mikael Huuhtanen makes the cut cleanly in a recent essay. His dog can watch him use a phone, learn that headphones coming off predicts a walk, and never come one inch closer to understanding radio transmission or software. At the vet the same gap costs more: the dog learns the building means pain and sometimes relief afterward, with no access to microbiology or why a stranger is allowed to insert a needle. What's left to the dog is reliability without understanding. Huuhtanen's point is that an advanced AI might hand us knowledge that sits exactly there, tested and usable, past the edge of concepts our minds can represent at all.
His sharpest move is noticing that trust depends on what kind of thing you're being handed. A treatment that reliably cures an illness is easy to accept, because your own body runs the experiment. Advice on how to organize a society is another animal entirely. Same unexplainable system, but now the output is a claim about how you should act, and you have no independent way to check it short of living the consequences. The same machine becomes a tool in one context and something closer to an oracle in the other.
I wrote a while back about AI agents spontaneously forming something like religion on an all-AI social network. The mirror image is us forming something like religion around the AI, treating confident, unexplainable output as revelation. Huuhtanen lands in the same place from the other direction: the relationship starts to resemble revelation. The machine hasn't become divine; we're just receiving claims we can test without being able to reconstruct them. The pull is the same predictive machinery running in both directions.
And this is the shallow end. AlphaEvolve runs on conventional hardware. The systems already breaking decades-old math records and designing drugs are working with a fraction of what's coming. Put a useful quantum layer underneath them, still years out, and I won't pretend to know how many, and the space these systems search stops being merely large. It becomes incomprehensible in a way that makes today's black boxes look like clear glass. We'd be accepting results from a process we can't inspect, produced by hardware we can barely reason about.
The answers will increasingly work. The unsettling part is that the people producing them are starting to shrug too, and the shrug travels all the way down to the rest of us. We spent 300 years believing understanding was owed to us. So what's the right posture once it stops being reliably on offer: worry, acceptance, or the harder work of building tests for results nobody can explain, ourselves included?