AI4Science
Agentic AI and the Automation of Scientific Discovery: Moving beyond task-specific ML: how multi-agent LLM systems are attempting to automate the full closed-loop cycle of hypothesis generation, code execution, data analysis, and laboratory validation.
In Robot-Proof I asked why "nobody is innovating on innovation”. The #AI4Science crowd is finally picking up speed and publishing breakthroughs…that all look weirdly the same. Now it seems like nobody's innovating on innovation innovation.
Three papers have landed in Nature recently. All three claim to automate scientific discovery. All three, if you squint, are the same abstract with the system name and the target problem swapped out.
I'm genuinely excited. Not only do I work in applied AI and this is the part of the field I actually care about, but my new company (Possibility Sciences, see below!) is innovating innovation in the field of innovations [1]. There aren’t just more well-posed benchmarks but machines that produced facts about the world that nobody knew the day before [2]. One of the projects found a drug for a disease that blinds people, and the drug had never been proposed for it before. That's decades of promises about AI finally being met.
But my excitement keeps getting tangled up with a nagging feeling that we've gotten very good at optimizing the last step of a process nobody has redesigned in 50 years. Reading these three papers back to back, I think we've now reached the recursive version of the problem. Nobody's innovating on innovation innovation.
Here's what I mean, paper by paper.
[1] Buffalo buffalo Buffalo buffalo buffalo buffalo Buffalo buffalo
[2] Well, it turns out people did know. But the AI didn’t know that they knew…fuck, this is getting complicated.
Research Roundup
Robin: Human-in-the-Loopdeeloop
The journal Nature just published a breakthrough #AI4Science system called “Robin”. [1]
Robin is a multi-agent system that generated a hypothesis for dry age-related macular degeneration (ADM, a common cause of blindness), proposed the experiment, interpreted the results, and generated the next hypothesis: ripasudil, a drug already in clinical use but never before proposed for dry AMD. It even confirmed efficacy in vitro.
Then it proposed a follow-up RNA-seq run that surfaced another possible novel target for treatment of ADM. Every hypothesis, figure, and analysis in the paper's main text came from the Robin…no humans gumming up the works and slowing things down.
That's a real result. The interesting design question isn't the architecture, however; it's where the designers cut the loop. Robin does everything except touch a pipette. Humans run the wet lab; the agent owns the reasoning on either side of it.
I don’t think that cut is a philosophical position about human oversight as much as optimization of a cost function. Wet-lab experiments are slow and expensive, so the loop closes where the money runs out. This system is shaped less by what its designers believe about intelligence than by what their experiments cost per iteration.
When you read an agentic science paper, don't ask how smart the agent is. Ask where the loop was cut and why.
[1] This footnote is just an inside joke for me and readers of my newsletter.
Co-Scientist: the value of brute force
The journal Nature just published a breakthrough #AI4Science system called “Co-Scientist”. [1]
Co-Scientist is Google's multi-agent system built on Gemini. Agents generate hypotheses, critique them, and refine them, with a tournament evolution process ranking candidates against each other and test-time compute scaling driving quality upward over time. [2]
At the end of the process, Co-scientist produced drug-repurposing candidates and synergistic combinations for acute myeloid leukemia and validated the hypothesis in vitro.
Co-Scientist's success argues that hypothesis quality is a search problem you can buy: spend more compute, get better ideas. I’d call that brute force, and I suspect the authors will at least partially agree. Much as with AlphaFold, brute force is the point.
I have no objection to brute force generating candidate ideas. The more millions of ideas the better, presuming you have agents to track them all.
My concern lives one layer down, in the tournament. To rank candidates you need a critic, and the critic agents share the same priors as the generator agents because they’re trained on the same literature. So, the massive search is then filtered by a consensus mechanism.
We built a machine to find ideas we already know but just haven’t yet realized we know them.
[1] Wait…wasn’t that the same opening line of yesterday’s post.
[2] Not shocking for an institution more famous for reinforcement learning than reasoning agents.
AI Scientist: the footnote
The journal Nature just published a breakthrough #AI4Science system called “the AI Scientist”. [1]
The AI Scientist [2] generates ideas, writes the code, runs the experiments, plots the data, writes the manuscript, and peer-reviews itself. One of its papers passed the first round of review at a workshop of a top-tier ML conference.
The authors also disclose, in the same breath, that the workshop had a 70% acceptance rate. The authors knew exactly what a reader would do with "passed peer review" and they took the air out of it themselves. More labs (and companies) should practice self-deflation [3].
Unlike the other biomedical AIs I’ve written about, this system can close the loop end-to-end because in computer science research the experiment is a script. Execution is nearly free and nearly instant. End-to-end automation shows up first in the one domain where the world model and the world are the same object.
(Which is why Agents for hacking are so terrifying. Defending agents will need to be similarly evolutionary, and the internet might quickly become more of an ecosystem than a technology.)
The rate limiter on automated science is not cognition; it's the cost of contact with reality. That cost is not falling at anything like the rate our models are improving.
[1] We’re trapped in an #AI4Science Groundhog Day! Aaaahhhhhhh!
[2] What's up with all the dull-ass AI4Science names: Co-Scientist, the AI Scientist…when Robin is the most original you have a creativity problem…which is problematic for AI all about creativity. Somebody get a geneticist in the marketing department!
[3] Or you might just find yourself in serious threat of losing an election to a dude with a trash bin on his head.
Media Mentions

𝐏𝐨𝐬𝐬𝐢𝐛𝐢𝐥𝐢𝐭𝐲 𝐒𝐜𝐢𝐞𝐧𝐜𝐞𝐬 is live! (possibilitysciences.org)
Nearly every tool for anticipating innovation works by counting: papers published, patents granted, dollars deployed. Counting tells you what's loud. It's terrible at telling you what's real.
So we stopped counting and started doing physics. Possibility Sciences maps the global research literature into a high-dimensional space where hidden emerging fields have measurable position, direction, and momentum. In this space you can watch ideas move rather than just tally how often it's mentioned.
What makes an idea heavy. 500 papers from a single lab is not 500 bets on a latent innovation trajectory. But 50 papers from 50 institutions across 9 countries and 3 unrelated disciplines is truly 50 independent bets. We measure the multiscale, multidimensional momentum of ideas by the independence of the people making them rather than by raw volume. Fields that looked like tipping points turn out to be closed communities talking to themselves at increasing frequency. Fields nobody was watching turn out to be quietly diffusing everywhere at once…and change the world.
The other thing Possibility Science does is let you run the counterfactual. Not "what do you think would happen if we funded this," but a computed projection of where the frontier actually moves when you push on it. Was that intervention causal or would this have happened anyway? Almost no organization can answer that about its own history.
And now our book of the same name is up for preorder: Possibility Sciences. (I contributed a bit to it myself!)
P.S. — I'm in this week's Men's Health, on what AI is doing to our brains, under the greatest headline I've been adjacent to in some time: “How Does AI Change Your Brain? 'It's Like TikTok and Fentanyl Had a Baby'“. I've spent much of the past few years exploring the cognitive costs of outsourcing your thinking, and this one gets at it about as bluntly as you can in a consumer magazine.
SciFi, Fantasy, & Me
Intelligence With Nobody Home. Every time I watch an agentic AI churn through a task, I think about the Scramblers, 9-limbed anaerobic things in Peter Watts' Blindsight. They out-think humanity at every turn despite nobody doing the thinking. The book suggests that self-awareness is a metabolic tax that a universe full of evolution rarely pays. Blindsight isn’t the only novel exploring the disconnect between intelligence and self-awareness:
The Tines (Vernor Vinge, A Fire Upon the Deep): These dog-like creatures, mindless alone, become a full person in packs of four to eight. Packs merge, split, get amputated. (This is the direct ancestor of the corvid pairs in Tchaikovsky's Children of Memory, who spend the whole book insisting they aren't anybody.)
The Machines (Stanisław Lem, The Invincible): A dead civilization's machines evolve for millions of years and what wins is a cloud of individually stupid micro-flies that swarms into something lethal.
It’s a phase (Karl Schroeder, Permanence) Consciousness isn't required for toolmaking; it's just a phase species pass through and out of.
Latency (Charles Stross, Accelerando) Selfhood is a latency problem, and the things that shed it eat the things that don't.
The Protagonist (Adrian Tchaikovsky, Service Model): The valet robot is absolutely certain it has no selfhood…or is that the selfhood talking.
p.s. Ditch the idea that your favorite bot has selfhood and awareness (it doesn't), and its behavior with make so much more sense.
Stage & Screen
- September 15, Amsterdam: How might AI change the world of investing?
- September 15, SF: Innovation Day with INSEAD!
- September 16, DC: AI and education–beyond dreams and dread.
- September 19, Phoenix: I'm giving the keynote for the Association of Science & Technology Centers annual conference.
- September 21, Stanford: We're still working on the details, but hopefully I'll be talking about my research on machine learning and neurodiversity for Stanford's Neurodiversity Project.
- September 24, UC Berkeley: It's my annual Berkeley Change-makers Lecture!
- September 24, NYC: Culture Shifting Deal Making Summit
- September 29, Cincinnati: Still baking...
- September 30, Irvine: Hybrid Intelligence for innovation!
- October 6, SF: UCSD Alumni Association
- October 6, SF: Giving a talk at the Draper Richards Kaplan Foundation
- October 6, Park City: It's Robot-Proof in the Rockies with setups.
- October 21-23, Warsaw: So much good stuff is in the works for my first visit to Poland
- October 27, Cologne: Maybe, maybe a visit to Germany!
- October, Toronto: The Future of Work...in the Future
- November 19, NYC: Secrets in the dark!