Thanks Olivia! First of all thanks so much for writing this! I learned and am so grateful you took the time to consider this so deeply and make a compelling piece! :)
I realize I should have made it much clearer that state actors were outside the scope of my piece! I was particularly focused on lone wolf scenarios, and the difficulty of engineered pandemics predominantly for lone wolf and non-state actors. Those are the ones that I am no longer worried about.
State actors and even experts in institutions could do great harm!
Diving into some of the points:
I definitely agree that automation and humanoid robots will drastically change what is easy. These were the conclusions of 2 of the 3 pieces. Piece 1: theoretical lab automations could make what is easy different, but that's not what's going on now, (so we should track lab automation) Piece 2: lab automations will make some steps easier, though those are good targets for oversight AND the steps between and after these are still hard.
Anthropic's willingness to make Fable unusuable isn't evidence in and of itself. They personify the approach that I'm questioning (hence why I focused on lone wolf / x-risk level engineered pandemic, not the things that the broader biosecurity community focus on). Anthropic been viewed as overstating their claim and being an echo chamber e.g. cyber practitioners couldn't replicate their China cyber agent attack report, for example. The same framework is also the underlying assumption when you highlight "when enabling technology is improving at such a high speed" which implies LLM speed is the thing to watch. I'm saying that we should really prioritize trying to get a better sense of the scale of the tacit knowledge barrier or lab automation.
On the "some groups meet the malicious bar": the existence of one group doesn't negate that the rule acts as a huge filter. Because biosec requires prioritization, natsec circles see these filters and weigh them.
I agree resources will get cheaper, but that's probably the thing to watch. I'm not sure how much AI will make equipment like pulverizers, fume hoods, etc. cheaper nor the labor associated with lack of tacit knowledge, etc. It's much more expensive than a bomb or chemical attack, and again, this point was mostly only geared for non-state actors. State actors have nearly unlimited funds sometimes!
On institutional access, you're pointing to what's easy to skirt (which I hope will be closed ofc) but you're not pointing at what's harder to skirt. Those bottlenecks become more important. Also, the fact that only tens of thousands can access these is why, realistically, why natsec circles worry less. Some worry should be focused on those tens of thousands, but there's also tradeoffs on whether making it hard for institutions to buy makes bioresilience efforts harder too since pretty much all the work is for vaccine/learning/etc.
You mention seeing Aum Shinrikyo as proof that a group that has all the things to commit a bio attack, though they failed 10 times and never succeeded. I think how one construes this evidence depends on one's pre-existing framework. (I think when I was at CSR, I would have only applied the biosec lens, but here I tried to learn the biosec/military-strategy lens) On whether they'd succeed today: the framework I'm explaining highlights that the tacit knowledge barrier is the thing to measure and appears not greatly changed. Given ActiveSite results, it seems unlikely that LLMs help much more than YouTube or paying a lab student (participants said they found YouTube more helpful than LLMs). A lot of the stuff Aum Shinrikyo would have needed to have the skill to do, to properly aerolize the spray, get the right concentration, is pretty advanced hands-on lab work, the kind that you can't really automate with lab automation out in the field nor with LLMs.
Pipetting seems to be hard partly because of what you're working with too. So I don't think someone successfully pippetting in a week means pipetting is mastered in all tasks nor all materials. All lab work takes time and it's pretty specific to the materials and processes you're doing. I agree (and point out in part 3) that some parts of the process can easily be done by lab students. Honestly, you can just buy a lab tech or probably a skilled lab student, though you don't need AI for that. Though cloud labs and robotic cells are specialized workflows and seem easier to have oversight on. I don't agree with the claim that because lab automation exists, that it will necessarily make things easier. It leaves out the fact that there's likely to be KYC type of process AND that it doesn't make life easier for those not using lab automation. It seems nuanced.
You say LLMs could get people to 80% success on tasks, but that is probably an example of the framework I was trying to counter. 80% success is a *huge* claim that would really need evidence - it rests on the assumption that all you need is someone telling you how to do something for you to do it right. It assumes there's no tacit knowledge or lab hands skill or real-life obstacles. That framework is the same reason why the SecureBio virology tests don't convince people in DC. The test cannot make claims about how much tacit knowledge or hands on skill there is to know. The size and importance of that aspect is the framework I'm explaining. Ideally the ActiveSite study had more time, I did advocate in my Part 3 for more studies like that. Though it's probably helpful that there was a realistic time frame, it's not like non-state actors have infinite time. Though notably, it was also much easier than real-life scenarios because students were handed the materials.
The 7% change is probably a figure that matters in the context of what you're looking at. Bacteriophages are much simpler, almost template-like from what I remember. I'm less confident here as that was a side debate that I didn't delve into deeply. Thanks for flagging about Arc Institute's goals, I didn't realize that their goal was narrower. BUT the difficulty lies in editing a virus that will then have to survive on everything (whether it be lettuce, lung or stomach lining, mucus, saliva, blood, air, etc.). Hence why there wasn't the same level of alarm outside of the AIxbioxXrisk community about bacteriophages. Even if it's a step, the hard part is making something that survives in all the different environments it has to survive in. So the size of the step to those who weight "survival in all conditions is incredible hard" is much smaller.
On the COVID-19, I only meant unedited COVID-19. This was my non-Gain of Function segment.
Again, sorry for not clarifying the narrowness of my claim. I do think part of it is that I try to "not write for my critics" which would probably defined here be as people who don't buy claims about LLMs are all that matter and are unhappy with the discourse. In terms of bio, it seems a lot of people are aware that we've made leaps in bioresilience that dont factor in risk headlines.
I didn't point out that some biologists see a kill-most-humans engineered pandemic as possible because 1. this seemed to be more of a minority opinion when I did interviews 2. I'm responding to claims that a kill-most-humans engineered pandemic is possible. It's a bit of a tradeoff to write succintly, which I notice you have to do to too :) E.g. when you say "Engineering can bypass the tradeoff" "by separating the transmission phase from the lethal phase, a technique which could also be engineered into a weapon". HIV nor smallpox are not existential risk level AND "engineering into a weapon" is theoretical only still (even x-risk people will clarify that there's no single gene-like or Lego-like on/off switch for "transmissibility" or "lethality" in viruses for humans.)
I don't name the participants so that I would get more off-the-record type of responses. I specifically prioritized getting people who actively worked in laboratories for 5-10+ years but also work on biosecurity, and who are not just from EA / X-risk / AI policy organizations. You mentioned that you're in this space because the evidence led you to believe this is urgent; this likely means your circles will be similar though. I run in X-risk / EA circles a lot hence I specifically got the the steelman of the other sides repeatedly, until their framework clicked. Happy to intro you to people. :) You seem very thoughtful and thorough. :)
Thank you for this! Biosecurity is one of my favorite topics, it's so important and I've added an update to clarify that biosecurity is still ridiculously important!
Thank you for such a thorough, kind response, and for adding the update to the original essay.
In responding to this comment, I would like to first focus on the question of scope. The piece is titled "Why AI-assisted bioweapons won't kill you." It says "let me first explain why I'm no longer worried." I think this is the bailey in a fairly textbook motte and bailey. These statements are not limited to the scope you are defending. If you want to make the narrower claim, I think it is important to only make the narrower claim. A reader who skims this piece, or even reads it thoroughly, is likely to come away with “AI-assisted bioweapons are not a threat” or “we should no longer be worried about AI-assisted bioweapons/bioweapons broadly,” not "lone-wolf bioterrorism is over-weighted relative to other biosecurity priorities.”
Below I have responded to some specific points:
Re: trendlines broadly. It seems like you agree that technology is better and things are getting easier. I am left slightly confused as to why you are no longer worried if this is the case. Even longer timelines for this to become easy should inspire some urgency as it takes time to build the kind of infrastructure we need.
Re: COVID: You wrote here, “On the COVID-19, I only meant unedited COVID-19. This was my non-Gain of Function segment.” This is hard for me to parse when the sentence I read is: “An *engineered* version of COVID-19 would have about the same effect as an infected person coughing in a crowded room.” What did you mean by engineered version here?
Re: frontier labs being worried does not mean I should infer the risk is real: I think this is a fine counterargument. I obviously still disagree, but this is not a provable claim, even if the risk is proven real.
Re: ActiveSite and the 60>80% comment: I do think an eight-week window is short for biotech, and I also think that they chose 153 participants with near-zero bench experience, mostly STEM undergrads, to perform an experimental design that ActiveSite described as deliberately demanding. This measures something, but I don’t think it accurately measures what the models can add over time, especially for more experienced personnel. Regarding my BOTE math, my only point was to illustrate that chains of events are sensitive in maybe surprising ways to changes in individual event success rates. However, I think AI increasing task likelihood of success by 25% is pretty reasonable, and it does not rest on the assumption that there's no tacit knowledge. Many things requiring tacit knowledge are improved by AI. (+tacit knowledge importance does shrink in the face of things like cloud labs, which I think will only have good KYC if we make sure people care about this.) A ~25% improvement on individual tasks is pretty expected from AI, even if biology turns out to be relatively harder than other technical domains. Here are documented productivity gains elsewhere:
-Software engineering: Up to 55% gains in task completion speed and overall productivity (GitHub)
-Customer support: 15–25% improvements in First Contact Resolution rates; AI-assisted agents move FCR from the 55–75% range up to 80–85%
-Surgery: 21–26% increase in total hospital production and 29% improvement in labor productivity. Some studies show ~20–25% gains in surgeon workflow efficiency and even greater reductions in length of stay + a decrease in postoperative visits
Re: existential risk: Specifically, you wrote: “I'm responding to claims that a kill-most-humans engineered pandemic is possible.” I think this is not out of scope. By many estimates, close to half of Europe died during the black plague. Here is a list of some species which have had much of their population wiped out by natural pathogens:
- Phocine Distemper Virus (PDV) – North Sea harbor seals: ~60% (1988); ~51% of remaining population (2002)
- Canine Distemper Virus (CDV) – Serengeti lions: ~30–33% of the population overall
- Avian malaria (protozoa) + Avian pox virus – Hawaiian honeycreepers: 68–98% population declines (average ~68%; some species higher)
-Pasteurella multocida (bacterium) – Saiga antelopes: >60% of the global population (2015)
Of course, these species all share special vulnerabilities and we can develop medical countermeasures against such a threat, but these are precisely the things you can engineer for/around through immune evasive techniques. Pathogen-agnostic measures like reusable respirators that can be implemented quickly remain very important. Of course there is “no single gene-like or Lego-like on/off switch for "transmissibility" or "lethality" in viruses for humans.” But we are getting better at this type of engineering and there are “Lego-like” things you can do to increase immune evasion. (I will not list them here for obvious reasons.)
Yes, some of it is theoretical, but it is important to consider things that have not been developed yet in a threat model.
Re: leaps in bioresilience –– would love to discuss further.
Re: succinctness. Unfortunately I make no such tradeoff, as you can probably tell from my 3,000 word response to your originally ~1,200 word piece and this comment.
Responded over DM on intros! (Reposted this so I could share on notes.)
On the title: The subtitle says that lone wolf attacks are the core of the piece. The intro before “I am not worried” only mentions lone wolf and X risk engineered pandemic. maybe other relevant context: I chose that title because people like Noah Smith have popularized the idea that “AI means Joe Schmoe can order a virus online and unleash death”. The non bio discourse about AIxBio is oddly filled with this specific scenario, hence why state actors didn’t come to mind much for me. I didn’t jump to state actors myself so I didn’t think to clarify that. I try to follow Scott Alexander advice to not write for critics or else you’ll caveat every line and not talk to your audience as you would in real life. I think this is usually underappreciated and underdone but I could have added bigger caveat sooner. I did add one to the end.
Why I don’t see technology getting better as a compelling risk here: automated labs are much easier to have oversight on, they’re specialized workflows and most likely to continue to be highly specialized given how many different bio materials need different setups/approaches. As this happens, we also get increases in bio resilience, understanding of science, more things like mRNA. The bioresilience is a pretty big part of the puzzle - more tech contributes almost entirely to defender side. Automated labs are well suited for KYC and their existence doesn’t make lone wolf work like manually looking through the millions of bacteria in a tablespoon of soil easier. Lone wolf still filtered by the 9 factors, maybe resources matters less if have institutional access.
Thanks for pointing out my COVID-19 misuse of word engineered!! I meant assembled or created in lab.
my guess is ActiveSite focused on near zero bench experience because it’s aiming to address lone wolf claims, which are getting disproportionate attention. It probably would be a different result with more experienced personnel, though these are fewer and even those are subject to the 9 factors that make work hard (do they have consistent institutional access? Can they execute all parts of the chain? The right type of motivated and competent yet malicious? Can they stay hidden from their institution the whole time?). My guess this is why ActiveSite didn’t structure it like that but definitely people should! I would love love love that.
I think increasing task likelihood by 25% still seems too high given ActiveSite results. Software and customer support seem entirely not applicable because those have no physical world aspect. The bottleneck is directly resolved by LLM. I’m very curious about the surgery one!! From what I understand , the improvement in surgical outcomes would be because AI compiles tacit knowledge that was otherwise hard to collect. for lone wolfs, I think they’d have much more gains from just paying a lab student for simpler tasks. The harder virology / bacteria tasks needed to get more extreme bio outcomes seem more specialized (very specific materials and very specific processing of these) than surgeons who all only work on the same specimen: human bodies. I think studies on LLMs getting someone to average-surgeon level but in biology would be helpful here.
Super interesting about the decimated populations but those don’t seem comparable. humans have more than just medical countermeasures. Humans have even basic ones too like coordination around quarantine and fear/awareness. I agree that reusable respirators and also stuff like vaccines for all the virus types would be amazing.
If we did get lego like understanding, we would also probably solve a LOT of open bio questions and get huge gains in bioresilience. The world where we have that is a different world than now - we have to weigh those positives too or else it’s an incomplete inaccurate model.
:) thanks so much for engaging so kindly and in good faith!!
Worth pointing out the most capable model in ActiveSite was Opus 4. Models have obviously improved a LOT since then. Also, the participants barely used the models. I do not think anything should be taken away from that study apart from lessons on how to do it better next time.
It seems like when you say tacit knowledge, you mean only somatic tacit knowledge, and actually almost exclusively mean pipetting ability. I am really very skeptical that a situation where pipetting technique is the main thing stopping bioterrorism is fine/sustainable. FWIW, the wet lab people I know are pretty skeptical about this as a barrier. I wonder if your experts were referring to the broader class of on-the-job experience often included within tacit knowledge, such as what it means when your cultures look like this or that - crucially, stuff that LLMs can uplift.
I also am very surprised you managed to come away with the impression that acceleration of bioresilience technologies is likely to outweigh offensive acceleration. There is a very strong consensus among the biosec people I talk to - by and large, the people pushing this acceleration - that offensive capabilities will open up long before we can harden the whole world against them, and what we're doing is racing to make that window smaller. I also think this is just intuitively obvious. The work needed to truly protect the world from a supervirus is absolutely enormous; the work needed to engineer that supervirus may end up being a few $m or less.
Separately, regarding the title/scoping, you seem to be using "don't write for your critics" as something close to "your writing is exempt from criticism". It's an odd position. You wrote something which plausibly increases risk - in some part because it relies on faulty reasoning, but likely in larger part because it makes much bolder/broader claims than it should. It's quite obvious the reason for that is engagement - bold, controversial title, engaging hook. And I see they have not been edited since. I don't think you get to shrug that off with a "well, I'm not writing for my critics!"
*Models weren't advanced enough* doesn't really engage with the results of the ActiveSite study: people didn't find the LLMs as useful as YouTube videos because the most helpful parts were demonstrations. Unless you're claim is that LLMs will make better demonstrations specifically so that it's more realistic than YouTube tutorials? If you really see nothing to take away from the ActiveSite Study, this seems like a very specific mindset. For a lot of the biosecurity community, this was proof that LLMs do not cause the uplift in unskilled lone wolfs because bio is physically difficult.
For an example of physical difficulty: one tablespoon of soil has anthrax but you have to manually sift through millions of bacteria yourself even if LLM gives you perfect information. Even if you find anthrax, you have to do extensive testing to see if it's even the varieties that are virulent. How does Fable make this much easier than Opus? This is an inherently manual process.
My Piece #2 was pretty clear that I only chose one skill for illustration purposes. The type of wet lab tacit knowledge was around the multifacted nature of bio work, including how pipetting can have unclear reasons of failure but that there's way more to troubleshoot than just what your culture look like.
If what the biologists said surprised you, I would take this a signal to find people outside of your circle to steelman their side better. Ideally, nothing I said surprises you and you can steelman the bioresilience-focus better than I can. If there is a "very strong consensus among the biosec people [you] talk to", I would steelman why typical wet lab people not in safety circles sees access control and a focus on superviruses very differently, including with safety circles. To be clear, I'm not saying "typical people" are right because they're more numerous, I'm saying nothing here should have been surprising.
For example, my Piece #3 goes into why many virologists and natsec scientists aren't concerned with superviruses because they don't seem possible because of inherent tradeoffs between catastrophic spread and lethality as well as how much harder modifying a virus is than a gene, and how much harder spread is via a virus is than a gene.
I did change the title...? I'd be happy to edit my any of the 4 pieces if you can give me specific faulty reasoning. But saying things like "but LLMs will get better" is not engaging with the physical-world-oriented cruxes.
LLMs probably will get great at making training videos. I would happily bet you X$ at strong odds that an AI can make a lab training video showing someone how to do something like PCR accurately, in a start-to-finish custom video, by 2028/2029. FWIW, I don't know a ton of folks in the biosecurity community who felt the ActiveSite study was definitive –– mostly just seemed to make people feel like we have a little more time. This is of course biased by who I know.
For demonstration of a physical process, a video is of course going to be more helpful than text, no matter how intelligent the thing that wrote the text. They address different parts of the problem and don't feel directly comparable. I'd guess these people were also just insufficiently motivated to succeed at the task, such that they'd really rather not go to the effort of parsing dense LLM text. And I suspect the scope and time limit on the task also makes things unrealistic - if someone was really engaging in this effort, they'd probably want to ask the LLM for ways to practice the muscle memory they lack.
I do not know *anyone* in the biosec community that looked closely at this study and took it as proof that LLMs do not uplift unskilled. I'd be very interested to know who you spoke to that had this view, and whether they'd actually read the study or just heard about it.
I do know an ERA AIxBio fellow whose current research project is trying to figure out if there's anything useful that can be pulled from the otherwise fairly unhelpful data in ActiveSite...
If pipetting was just one of many physical skills, what are the others required for viral recovery workflows? Everything else - aseptic technique and so on - seems pretty communicable via text & images, though some of it may take a bit of practice.
Re whether offensive or defensive acceleration comes first - I mean, I can go talk to a bunch of biologists who haven't thought much about this stuff before, but I don't know why that would be informative? Just about everyone who's thought seriously about it seems to think the former does. I would appreciate if you could point me to an example of a well-informed person making the argument that def acc will naturally win.
Unfortunately I'm unable to read piece #3 - the page errors, at least for me.
The title, unless it's just an old cache, is still "Why AI-assisted bioweapons won't kill us all" and the hook still says "Let me first explain why I’m no longer worried" - these are what I take issue with. If there was an even more egregious title before, well I guess thanks for changing it
This entire exchange is a platonic ideal of "republic of letters" style exchange, I'm delighted you're both on substack!
On the lone wolf to state actor axis, if I'm reading various views correctly (here and elsewhere) makes me come to the conclusion that advanced states could more or less already do it (and have chosen not to for various reasons) and lone wolves are still a ways off, but semi-state or mid sized organisation actors might have most to gain here? Actors like Al-Qaeda / ISIS, cults like Aum Shinrikiyo, or small unstable countries like Central African Republic. As in, access to reasonable amounts of capital (tens of millions of dollars), smuggling networks, and some industrial capacity, but lacking expertise.
Yes this seems true for even very bad weapons, that state actors could already do something like recreate 1918 virus.
I think the only thing I’d want to specific here is where engineered pathogens fall. If anyone is able to engineer a pandemic-capable virus, it would be state actors first. Though the more you want to edit it, the harder it is to do that edit, not just in the sense of whether the virus can exist, but whether it works on people and whether it doesn’t revert back to its default form.
Ironically, the discussion in here about how dangerous state bioweapons programs are is why I’m *less* worried about the impact of AI. State actors are already incredibly capable of producing bioweapons and have been for decades. AI doesn’t really move the needle on their ability to launch a devastating attack because they can already do it with existing techniques. This implies that 1) AI shouldn’t be that big of a threat update on this front, and 2) the strategic dynamics that currently constrain states’ use of bioweapons will not be significantly altered by AI, because merely adding marginally more capabilities doesn’t change the existing calculus.
As far as technological improvements enabling more lone wolves or unaffiliated groups, I think you are underestimating the scope of a successful bioweapons program and overestimating how much the technology will dissolve all the existing bottlenecks, but I would completely agree that future automated labs and other providers should have robust auditing and KYC. This would be an example of a threat that I think is legitimate and takes vigilance but is solidly manageable with normal governance responses.
On the work done by Arc that you cite, those viruses showed *less* variation than natural evolution and little evidence of functionally directed novelty (Black et al., 2026). Given how poor the scaling and generalization have been in gene language models and other bio foundation models (Jiang et al., 2026; Tzanakakis et al., 2026), I’d say it’s entirely non-obvious that those techniques will scale to producing truly novel pathogens (which also requires predicting pathogenicity, which in itself is a major challenge). This could turn out to be incorrect, and progress in the field is very much worth watching, but I’d say the early evidence suggests that the bio models are on nothing like the progress curve of LLMs.
Black, J.R.M., Maiwald, A., Pannu, J. & Crook, O.M. (2026). *Quantifying evolutionary novelty and design efficiency in generative genome design.* bioRxiv. https://doi.org/10.64898/2026.06.12.731871
Jiang, S., Liu, X. & Wang, Z.J. (2026). *Evaluating DNA Function Understanding in Genomic Language Models Using Evolutionarily Implausible Sequences.* ACS Synthetic Biology, 15(6), 2256–2263. https://doi.org/10.1021/acssynbio.6c00024
The problem here is motivation/goals. Though there are good reasons that state actors develop bioweapons and might use them, but they have significantly less motivation to do so than a group like Aum. That is, states are capability rich and motivation poor, groups are the opposite, and AI is an uplift on capability. Re: model development, bio models yes will be slower to improve than LLMs, but that doesn't mean they won't be fast to improve. I think saying that the variation was less than natural evolution is not the spirit of what was claimed in the paper, which only claims that it was within the bounds of natural variation. The time scales here boggle my mind. Variation like this in nature could take years to tens of thousands of years. Re: functionally directed novelty, the paper you cite explicitly notes that evolutionary novelty is distinct from functional novelty and does not make claims about functional novelty. Several of the AI phages from Arc outcompeted the wild ΦX174, w/ Evo-Φ69 growing 16-65x versus WT's 1.3-4.0x
It sounds like we mostly agree that for state actors the binding current constraint is strategic, not technical whereas NSAs are significantly more capability bound. A problem in your argument is that you slip between using state capabilities to justify those of potential terrorist groups. For example, here in your reply to point 4:
> “Your team must be both highly competent and completely secret. A contamination incident could destroy the pathogen or infect the team.”
The Soviets ran an effective 30,000-person bioweapons program for two decades.
We also don't disagree that AI will increase certain technological capabilities. Where I think you are too quick is the jump from “AI will increase the technological capacity for certain inputs” to therefore “end-to-end bioweapon capabilities will become accessible to a wide array of actors.” This is exactly where Abi and I disagree. We think that myriad practical, logistical, and institutional constraints will bottleneck the dispersion of real-world capabilities. I think we also both support strengthening existing monitoring and regulatory controls to ensure they continue to bind. An example of a place where you do this is when you translate decreasing costs to increasing capabilities. In many domains capital hasn't been the constraint for years.
On the comment about the novelty of the virus, there were two claims in my original sentence that I should probs have unpacked more. The first was that there was less variation than natural evolution. In particular, here I meant novel variation as measured by the difference from the existing phylogenetic manifold. The relevant results in the Black et al. paper are described in this section from the abstract:
> However, this efficiency derives largely from staying close to previously observed sequences rather than exploring novel sequence space, reflecting the combined performance of the model and additional filters that were applied to its outputs. Compared to baselines of random mutagenesis and serial passage, the model achieves substantial design efficiency while its outputs remain phylogenetically close to natural genomes. We conclude that the generative capabilities of Evo 2 warrant low to moderate biosecurity concern for de novo hazard creation, although the degree to which these findings generalise to larger or less constrained viral architectures is an open question.
Whether you want to say that a generative model sampling from a highly constrained phylogenetic repertoire produces 1–10,000 years of variation is ultimately a matter of taste, but to me it seems like a category error since it's variation of a distinctly different type.
The other claim was about functionally directed novelty. In the case of the phage synthesis experiments, their only goal was to preserve the tropism of a known virus. That is not demonstrating the ability to select for a novel phenotype, which would be a significantly more impressive and worrying sign. Showing that some phages outcompeted wild type doesn't really bear on that limitation, as increased replication was not something being selected for.
Finally, on whether or not the models will be fast to improve, I think there are strong reasons to believe the progress will be slow because biology is a much messier and more opaque domain with much longer feedback cycles. Most of the evidence so far points to that being the case, but like I said, it's something to monitor in case there is a big change.
Also, I should clarify that the reason I'm arguing about this is not because I think we should be totally complacent and ignore biosecurity risks. It's that correctly assessing the danger is extremely important for decisions like whether or not to pause AI development. I think both that we should harden our defenses and that the magnitude of the increased threat due to AI doesn't warrant dramatically slowing progress.
Well I don't think slowing down because of biorisk is really feasible/desirable regardless of the risk magnitude. I'd have to see the contours of what's being proposed though because slowing down could mean a lot of things.
And what is the specific concern around novelty, anyway? You can get significant gain of function without much change to the sequence, and screening evasion is a different beast entirely than simulated evolution (see our latest paper https://www.rand.org/pubs/research_reports/RRA4741-2.html)
Also, could you explain more why lowering of capability floors for bio isn't concerning for you?
Yeah, I used that as an example. I'm also not a fan of most pause proposals, but imo the risks should be part of that calculus.
The reason I was bringing up novelty here was because it relates to how useful the Evo models mentioned in the piece are for developing phenotypically novel viruses, and thereby how much the AI tools improve our ability to do de novo GOF. That paper you cited looks interesting; improving our screening against adversarially designed sequences is an important aspect of hardening our bio defenses.
It's not that I am totally unconcerned about lowering the floor; it's that I think AI only does so much to change the threat environment.
Yeah I'm theoretically concerned about AI enabled GOF, but I feel like we are pretty far away from it now and evo 2 doesn't really bring us closer in a meaningful way.
An analogy I sometimes use is being scared about how fast humans can go in a world with horses only. You design an AI that makes different horse breeds, some of which are faster and some of which are slower. That's slightly concerning, but probably not more concerning than selective breeding. The real issue comes when it can do something like design a car.
I adored this extended exchange between you and Abi. I just read both and my main conclusion is that the two of y’all should hang out. I don’t know enough about the space to opine on the several “how hard is executing lab work well actually” studies but I am basically persuaded that the popular sub community worry about lone wolves specifically seems overrated (but nonzero) and the risk of state actor engagement remains high. I am generally long transferrably on “real world tasks have counterintuitive combined rate limitations” and also “long-horizon prognostication is insanely difficult because of emergent phenomena” so that informs my priors.
I think we are going to! Abi is great. One of the problems with this exchange is I (and other scientists) can’t exactly explain how easy it is or talk about the several switches we have for improving immune evasion because it is a massive info-hazard. On some level, this piece could be a positive contribution to deterring lone wolves/small groups from acting. (I think this is another motte and bailey though, to go from all small groups to lone actors.)
Unfortunately, I think the sensationalism in the headline instead results in making people think this is not a concern, even though she agrees it is getting easier. It is better that we put infrastructure in place now, and not wait until it becomes a bigger issue. (Especially given that state actors at a minimum are already an issue.) I worry this piece mostly contributes to people thinking this is not urgent.
While the above article seems very intelligent, articulate and informed, I feel it's missing an important bottom line, as seems true of most of science and the public at large.
The bottom line problem is a mismatch between 1) the accelerating pace of knowledge and power development, and 2) the incremental (at best) pace of human maturity development. So long as that relationship remains in place, the gap between our power and maturity will continue to widen, probably at an accelerating pace.
Attempting to address threats one at a time as they roll off the end of the knowledge assembly line is a loser's game. It's only when we back up and look at the situation as a whole that there will be some hope.
Hi, I'll probably have more comments, but first an editorial comment. In #5, you have "...AS shows why biosecurity is hard." I think you meant *bioterrism* is hard, right?
Thanks Olivia! First of all thanks so much for writing this! I learned and am so grateful you took the time to consider this so deeply and make a compelling piece! :)
I realize I should have made it much clearer that state actors were outside the scope of my piece! I was particularly focused on lone wolf scenarios, and the difficulty of engineered pandemics predominantly for lone wolf and non-state actors. Those are the ones that I am no longer worried about.
State actors and even experts in institutions could do great harm!
Diving into some of the points:
I definitely agree that automation and humanoid robots will drastically change what is easy. These were the conclusions of 2 of the 3 pieces. Piece 1: theoretical lab automations could make what is easy different, but that's not what's going on now, (so we should track lab automation) Piece 2: lab automations will make some steps easier, though those are good targets for oversight AND the steps between and after these are still hard.
Anthropic's willingness to make Fable unusuable isn't evidence in and of itself. They personify the approach that I'm questioning (hence why I focused on lone wolf / x-risk level engineered pandemic, not the things that the broader biosecurity community focus on). Anthropic been viewed as overstating their claim and being an echo chamber e.g. cyber practitioners couldn't replicate their China cyber agent attack report, for example. The same framework is also the underlying assumption when you highlight "when enabling technology is improving at such a high speed" which implies LLM speed is the thing to watch. I'm saying that we should really prioritize trying to get a better sense of the scale of the tacit knowledge barrier or lab automation.
On the "some groups meet the malicious bar": the existence of one group doesn't negate that the rule acts as a huge filter. Because biosec requires prioritization, natsec circles see these filters and weigh them.
I agree resources will get cheaper, but that's probably the thing to watch. I'm not sure how much AI will make equipment like pulverizers, fume hoods, etc. cheaper nor the labor associated with lack of tacit knowledge, etc. It's much more expensive than a bomb or chemical attack, and again, this point was mostly only geared for non-state actors. State actors have nearly unlimited funds sometimes!
On institutional access, you're pointing to what's easy to skirt (which I hope will be closed ofc) but you're not pointing at what's harder to skirt. Those bottlenecks become more important. Also, the fact that only tens of thousands can access these is why, realistically, why natsec circles worry less. Some worry should be focused on those tens of thousands, but there's also tradeoffs on whether making it hard for institutions to buy makes bioresilience efforts harder too since pretty much all the work is for vaccine/learning/etc.
You mention seeing Aum Shinrikyo as proof that a group that has all the things to commit a bio attack, though they failed 10 times and never succeeded. I think how one construes this evidence depends on one's pre-existing framework. (I think when I was at CSR, I would have only applied the biosec lens, but here I tried to learn the biosec/military-strategy lens) On whether they'd succeed today: the framework I'm explaining highlights that the tacit knowledge barrier is the thing to measure and appears not greatly changed. Given ActiveSite results, it seems unlikely that LLMs help much more than YouTube or paying a lab student (participants said they found YouTube more helpful than LLMs). A lot of the stuff Aum Shinrikyo would have needed to have the skill to do, to properly aerolize the spray, get the right concentration, is pretty advanced hands-on lab work, the kind that you can't really automate with lab automation out in the field nor with LLMs.
Pipetting seems to be hard partly because of what you're working with too. So I don't think someone successfully pippetting in a week means pipetting is mastered in all tasks nor all materials. All lab work takes time and it's pretty specific to the materials and processes you're doing. I agree (and point out in part 3) that some parts of the process can easily be done by lab students. Honestly, you can just buy a lab tech or probably a skilled lab student, though you don't need AI for that. Though cloud labs and robotic cells are specialized workflows and seem easier to have oversight on. I don't agree with the claim that because lab automation exists, that it will necessarily make things easier. It leaves out the fact that there's likely to be KYC type of process AND that it doesn't make life easier for those not using lab automation. It seems nuanced.
You say LLMs could get people to 80% success on tasks, but that is probably an example of the framework I was trying to counter. 80% success is a *huge* claim that would really need evidence - it rests on the assumption that all you need is someone telling you how to do something for you to do it right. It assumes there's no tacit knowledge or lab hands skill or real-life obstacles. That framework is the same reason why the SecureBio virology tests don't convince people in DC. The test cannot make claims about how much tacit knowledge or hands on skill there is to know. The size and importance of that aspect is the framework I'm explaining. Ideally the ActiveSite study had more time, I did advocate in my Part 3 for more studies like that. Though it's probably helpful that there was a realistic time frame, it's not like non-state actors have infinite time. Though notably, it was also much easier than real-life scenarios because students were handed the materials.
The 7% change is probably a figure that matters in the context of what you're looking at. Bacteriophages are much simpler, almost template-like from what I remember. I'm less confident here as that was a side debate that I didn't delve into deeply. Thanks for flagging about Arc Institute's goals, I didn't realize that their goal was narrower. BUT the difficulty lies in editing a virus that will then have to survive on everything (whether it be lettuce, lung or stomach lining, mucus, saliva, blood, air, etc.). Hence why there wasn't the same level of alarm outside of the AIxbioxXrisk community about bacteriophages. Even if it's a step, the hard part is making something that survives in all the different environments it has to survive in. So the size of the step to those who weight "survival in all conditions is incredible hard" is much smaller.
On the COVID-19, I only meant unedited COVID-19. This was my non-Gain of Function segment.
Again, sorry for not clarifying the narrowness of my claim. I do think part of it is that I try to "not write for my critics" which would probably defined here be as people who don't buy claims about LLMs are all that matter and are unhappy with the discourse. In terms of bio, it seems a lot of people are aware that we've made leaps in bioresilience that dont factor in risk headlines.
I didn't point out that some biologists see a kill-most-humans engineered pandemic as possible because 1. this seemed to be more of a minority opinion when I did interviews 2. I'm responding to claims that a kill-most-humans engineered pandemic is possible. It's a bit of a tradeoff to write succintly, which I notice you have to do to too :) E.g. when you say "Engineering can bypass the tradeoff" "by separating the transmission phase from the lethal phase, a technique which could also be engineered into a weapon". HIV nor smallpox are not existential risk level AND "engineering into a weapon" is theoretical only still (even x-risk people will clarify that there's no single gene-like or Lego-like on/off switch for "transmissibility" or "lethality" in viruses for humans.)
I don't name the participants so that I would get more off-the-record type of responses. I specifically prioritized getting people who actively worked in laboratories for 5-10+ years but also work on biosecurity, and who are not just from EA / X-risk / AI policy organizations. You mentioned that you're in this space because the evidence led you to believe this is urgent; this likely means your circles will be similar though. I run in X-risk / EA circles a lot hence I specifically got the the steelman of the other sides repeatedly, until their framework clicked. Happy to intro you to people. :) You seem very thoughtful and thorough. :)
Thank you for this! Biosecurity is one of my favorite topics, it's so important and I've added an update to clarify that biosecurity is still ridiculously important!
Thank you for such a thorough, kind response, and for adding the update to the original essay.
In responding to this comment, I would like to first focus on the question of scope. The piece is titled "Why AI-assisted bioweapons won't kill you." It says "let me first explain why I'm no longer worried." I think this is the bailey in a fairly textbook motte and bailey. These statements are not limited to the scope you are defending. If you want to make the narrower claim, I think it is important to only make the narrower claim. A reader who skims this piece, or even reads it thoroughly, is likely to come away with “AI-assisted bioweapons are not a threat” or “we should no longer be worried about AI-assisted bioweapons/bioweapons broadly,” not "lone-wolf bioterrorism is over-weighted relative to other biosecurity priorities.”
Below I have responded to some specific points:
Re: trendlines broadly. It seems like you agree that technology is better and things are getting easier. I am left slightly confused as to why you are no longer worried if this is the case. Even longer timelines for this to become easy should inspire some urgency as it takes time to build the kind of infrastructure we need.
Re: COVID: You wrote here, “On the COVID-19, I only meant unedited COVID-19. This was my non-Gain of Function segment.” This is hard for me to parse when the sentence I read is: “An *engineered* version of COVID-19 would have about the same effect as an infected person coughing in a crowded room.” What did you mean by engineered version here?
Re: frontier labs being worried does not mean I should infer the risk is real: I think this is a fine counterargument. I obviously still disagree, but this is not a provable claim, even if the risk is proven real.
Re: ActiveSite and the 60>80% comment: I do think an eight-week window is short for biotech, and I also think that they chose 153 participants with near-zero bench experience, mostly STEM undergrads, to perform an experimental design that ActiveSite described as deliberately demanding. This measures something, but I don’t think it accurately measures what the models can add over time, especially for more experienced personnel. Regarding my BOTE math, my only point was to illustrate that chains of events are sensitive in maybe surprising ways to changes in individual event success rates. However, I think AI increasing task likelihood of success by 25% is pretty reasonable, and it does not rest on the assumption that there's no tacit knowledge. Many things requiring tacit knowledge are improved by AI. (+tacit knowledge importance does shrink in the face of things like cloud labs, which I think will only have good KYC if we make sure people care about this.) A ~25% improvement on individual tasks is pretty expected from AI, even if biology turns out to be relatively harder than other technical domains. Here are documented productivity gains elsewhere:
-Software engineering: Up to 55% gains in task completion speed and overall productivity (GitHub)
-Customer support: 15–25% improvements in First Contact Resolution rates; AI-assisted agents move FCR from the 55–75% range up to 80–85%
-Surgery: 21–26% increase in total hospital production and 29% improvement in labor productivity. Some studies show ~20–25% gains in surgeon workflow efficiency and even greater reductions in length of stay + a decrease in postoperative visits
Re: existential risk: Specifically, you wrote: “I'm responding to claims that a kill-most-humans engineered pandemic is possible.” I think this is not out of scope. By many estimates, close to half of Europe died during the black plague. Here is a list of some species which have had much of their population wiped out by natural pathogens:
- Phocine Distemper Virus (PDV) – North Sea harbor seals: ~60% (1988); ~51% of remaining population (2002)
- Canine Distemper Virus (CDV) – Serengeti lions: ~30–33% of the population overall
- Canine Distemper Virus (CDV) – Caspian seals: >50% of localized populations (2000)
- Avian malaria (protozoa) + Avian pox virus – Hawaiian honeycreepers: 68–98% population declines (average ~68%; some species higher)
-Pasteurella multocida (bacterium) – Saiga antelopes: >60% of the global population (2015)
Of course, these species all share special vulnerabilities and we can develop medical countermeasures against such a threat, but these are precisely the things you can engineer for/around through immune evasive techniques. Pathogen-agnostic measures like reusable respirators that can be implemented quickly remain very important. Of course there is “no single gene-like or Lego-like on/off switch for "transmissibility" or "lethality" in viruses for humans.” But we are getting better at this type of engineering and there are “Lego-like” things you can do to increase immune evasion. (I will not list them here for obvious reasons.)
Yes, some of it is theoretical, but it is important to consider things that have not been developed yet in a threat model.
Re: leaps in bioresilience –– would love to discuss further.
Re: succinctness. Unfortunately I make no such tradeoff, as you can probably tell from my 3,000 word response to your originally ~1,200 word piece and this comment.
Responded over DM on intros! (Reposted this so I could share on notes.)
(pasted from Note)
Thanks Olivia :) always great to hear from you!
On the title: The subtitle says that lone wolf attacks are the core of the piece. The intro before “I am not worried” only mentions lone wolf and X risk engineered pandemic. maybe other relevant context: I chose that title because people like Noah Smith have popularized the idea that “AI means Joe Schmoe can order a virus online and unleash death”. The non bio discourse about AIxBio is oddly filled with this specific scenario, hence why state actors didn’t come to mind much for me. I didn’t jump to state actors myself so I didn’t think to clarify that. I try to follow Scott Alexander advice to not write for critics or else you’ll caveat every line and not talk to your audience as you would in real life. I think this is usually underappreciated and underdone but I could have added bigger caveat sooner. I did add one to the end.
Why I don’t see technology getting better as a compelling risk here: automated labs are much easier to have oversight on, they’re specialized workflows and most likely to continue to be highly specialized given how many different bio materials need different setups/approaches. As this happens, we also get increases in bio resilience, understanding of science, more things like mRNA. The bioresilience is a pretty big part of the puzzle - more tech contributes almost entirely to defender side. Automated labs are well suited for KYC and their existence doesn’t make lone wolf work like manually looking through the millions of bacteria in a tablespoon of soil easier. Lone wolf still filtered by the 9 factors, maybe resources matters less if have institutional access.
Thanks for pointing out my COVID-19 misuse of word engineered!! I meant assembled or created in lab.
my guess is ActiveSite focused on near zero bench experience because it’s aiming to address lone wolf claims, which are getting disproportionate attention. It probably would be a different result with more experienced personnel, though these are fewer and even those are subject to the 9 factors that make work hard (do they have consistent institutional access? Can they execute all parts of the chain? The right type of motivated and competent yet malicious? Can they stay hidden from their institution the whole time?). My guess this is why ActiveSite didn’t structure it like that but definitely people should! I would love love love that.
I think increasing task likelihood by 25% still seems too high given ActiveSite results. Software and customer support seem entirely not applicable because those have no physical world aspect. The bottleneck is directly resolved by LLM. I’m very curious about the surgery one!! From what I understand , the improvement in surgical outcomes would be because AI compiles tacit knowledge that was otherwise hard to collect. for lone wolfs, I think they’d have much more gains from just paying a lab student for simpler tasks. The harder virology / bacteria tasks needed to get more extreme bio outcomes seem more specialized (very specific materials and very specific processing of these) than surgeons who all only work on the same specimen: human bodies. I think studies on LLMs getting someone to average-surgeon level but in biology would be helpful here.
Super interesting about the decimated populations but those don’t seem comparable. humans have more than just medical countermeasures. Humans have even basic ones too like coordination around quarantine and fear/awareness. I agree that reusable respirators and also stuff like vaccines for all the virus types would be amazing.
If we did get lego like understanding, we would also probably solve a LOT of open bio questions and get huge gains in bioresilience. The world where we have that is a different world than now - we have to weigh those positives too or else it’s an incomplete inaccurate model.
:) thanks so much for engaging so kindly and in good faith!!
Worth pointing out the most capable model in ActiveSite was Opus 4. Models have obviously improved a LOT since then. Also, the participants barely used the models. I do not think anything should be taken away from that study apart from lessons on how to do it better next time.
It seems like when you say tacit knowledge, you mean only somatic tacit knowledge, and actually almost exclusively mean pipetting ability. I am really very skeptical that a situation where pipetting technique is the main thing stopping bioterrorism is fine/sustainable. FWIW, the wet lab people I know are pretty skeptical about this as a barrier. I wonder if your experts were referring to the broader class of on-the-job experience often included within tacit knowledge, such as what it means when your cultures look like this or that - crucially, stuff that LLMs can uplift.
I also am very surprised you managed to come away with the impression that acceleration of bioresilience technologies is likely to outweigh offensive acceleration. There is a very strong consensus among the biosec people I talk to - by and large, the people pushing this acceleration - that offensive capabilities will open up long before we can harden the whole world against them, and what we're doing is racing to make that window smaller. I also think this is just intuitively obvious. The work needed to truly protect the world from a supervirus is absolutely enormous; the work needed to engineer that supervirus may end up being a few $m or less.
Separately, regarding the title/scoping, you seem to be using "don't write for your critics" as something close to "your writing is exempt from criticism". It's an odd position. You wrote something which plausibly increases risk - in some part because it relies on faulty reasoning, but likely in larger part because it makes much bolder/broader claims than it should. It's quite obvious the reason for that is engagement - bold, controversial title, engaging hook. And I see they have not been edited since. I don't think you get to shrug that off with a "well, I'm not writing for my critics!"
*Models weren't advanced enough* doesn't really engage with the results of the ActiveSite study: people didn't find the LLMs as useful as YouTube videos because the most helpful parts were demonstrations. Unless you're claim is that LLMs will make better demonstrations specifically so that it's more realistic than YouTube tutorials? If you really see nothing to take away from the ActiveSite Study, this seems like a very specific mindset. For a lot of the biosecurity community, this was proof that LLMs do not cause the uplift in unskilled lone wolfs because bio is physically difficult.
For an example of physical difficulty: one tablespoon of soil has anthrax but you have to manually sift through millions of bacteria yourself even if LLM gives you perfect information. Even if you find anthrax, you have to do extensive testing to see if it's even the varieties that are virulent. How does Fable make this much easier than Opus? This is an inherently manual process.
My Piece #2 was pretty clear that I only chose one skill for illustration purposes. The type of wet lab tacit knowledge was around the multifacted nature of bio work, including how pipetting can have unclear reasons of failure but that there's way more to troubleshoot than just what your culture look like.
If what the biologists said surprised you, I would take this a signal to find people outside of your circle to steelman their side better. Ideally, nothing I said surprises you and you can steelman the bioresilience-focus better than I can. If there is a "very strong consensus among the biosec people [you] talk to", I would steelman why typical wet lab people not in safety circles sees access control and a focus on superviruses very differently, including with safety circles. To be clear, I'm not saying "typical people" are right because they're more numerous, I'm saying nothing here should have been surprising.
For example, my Piece #3 goes into why many virologists and natsec scientists aren't concerned with superviruses because they don't seem possible because of inherent tradeoffs between catastrophic spread and lethality as well as how much harder modifying a virus is than a gene, and how much harder spread is via a virus is than a gene.
I did change the title...? I'd be happy to edit my any of the 4 pieces if you can give me specific faulty reasoning. But saying things like "but LLMs will get better" is not engaging with the physical-world-oriented cruxes.
LLMs probably will get great at making training videos. I would happily bet you X$ at strong odds that an AI can make a lab training video showing someone how to do something like PCR accurately, in a start-to-finish custom video, by 2028/2029. FWIW, I don't know a ton of folks in the biosecurity community who felt the ActiveSite study was definitive –– mostly just seemed to make people feel like we have a little more time. This is of course biased by who I know.
For demonstration of a physical process, a video is of course going to be more helpful than text, no matter how intelligent the thing that wrote the text. They address different parts of the problem and don't feel directly comparable. I'd guess these people were also just insufficiently motivated to succeed at the task, such that they'd really rather not go to the effort of parsing dense LLM text. And I suspect the scope and time limit on the task also makes things unrealistic - if someone was really engaging in this effort, they'd probably want to ask the LLM for ways to practice the muscle memory they lack.
I do not know *anyone* in the biosec community that looked closely at this study and took it as proof that LLMs do not uplift unskilled. I'd be very interested to know who you spoke to that had this view, and whether they'd actually read the study or just heard about it.
I do know an ERA AIxBio fellow whose current research project is trying to figure out if there's anything useful that can be pulled from the otherwise fairly unhelpful data in ActiveSite...
If pipetting was just one of many physical skills, what are the others required for viral recovery workflows? Everything else - aseptic technique and so on - seems pretty communicable via text & images, though some of it may take a bit of practice.
Re whether offensive or defensive acceleration comes first - I mean, I can go talk to a bunch of biologists who haven't thought much about this stuff before, but I don't know why that would be informative? Just about everyone who's thought seriously about it seems to think the former does. I would appreciate if you could point me to an example of a well-informed person making the argument that def acc will naturally win.
Unfortunately I'm unable to read piece #3 - the page errors, at least for me.
The title, unless it's just an old cache, is still "Why AI-assisted bioweapons won't kill us all" and the hook still says "Let me first explain why I’m no longer worried" - these are what I take issue with. If there was an even more egregious title before, well I guess thanks for changing it
This entire exchange is a platonic ideal of "republic of letters" style exchange, I'm delighted you're both on substack!
On the lone wolf to state actor axis, if I'm reading various views correctly (here and elsewhere) makes me come to the conclusion that advanced states could more or less already do it (and have chosen not to for various reasons) and lone wolves are still a ways off, but semi-state or mid sized organisation actors might have most to gain here? Actors like Al-Qaeda / ISIS, cults like Aum Shinrikiyo, or small unstable countries like Central African Republic. As in, access to reasonable amounts of capital (tens of millions of dollars), smuggling networks, and some industrial capacity, but lacking expertise.
Yes this seems true for even very bad weapons, that state actors could already do something like recreate 1918 virus.
I think the only thing I’d want to specific here is where engineered pathogens fall. If anyone is able to engineer a pandemic-capable virus, it would be state actors first. Though the more you want to edit it, the harder it is to do that edit, not just in the sense of whether the virus can exist, but whether it works on people and whether it doesn’t revert back to its default form.
Nice post, it’s great there is a healthy debate.
Ironically, the discussion in here about how dangerous state bioweapons programs are is why I’m *less* worried about the impact of AI. State actors are already incredibly capable of producing bioweapons and have been for decades. AI doesn’t really move the needle on their ability to launch a devastating attack because they can already do it with existing techniques. This implies that 1) AI shouldn’t be that big of a threat update on this front, and 2) the strategic dynamics that currently constrain states’ use of bioweapons will not be significantly altered by AI, because merely adding marginally more capabilities doesn’t change the existing calculus.
As far as technological improvements enabling more lone wolves or unaffiliated groups, I think you are underestimating the scope of a successful bioweapons program and overestimating how much the technology will dissolve all the existing bottlenecks, but I would completely agree that future automated labs and other providers should have robust auditing and KYC. This would be an example of a threat that I think is legitimate and takes vigilance but is solidly manageable with normal governance responses.
On the work done by Arc that you cite, those viruses showed *less* variation than natural evolution and little evidence of functionally directed novelty (Black et al., 2026). Given how poor the scaling and generalization have been in gene language models and other bio foundation models (Jiang et al., 2026; Tzanakakis et al., 2026), I’d say it’s entirely non-obvious that those techniques will scale to producing truly novel pathogens (which also requires predicting pathogenicity, which in itself is a major challenge). This could turn out to be incorrect, and progress in the field is very much worth watching, but I’d say the early evidence suggests that the bio models are on nothing like the progress curve of LLMs.
Black, J.R.M., Maiwald, A., Pannu, J. & Crook, O.M. (2026). *Quantifying evolutionary novelty and design efficiency in generative genome design.* bioRxiv. https://doi.org/10.64898/2026.06.12.731871
Jiang, S., Liu, X. & Wang, Z.J. (2026). *Evaluating DNA Function Understanding in Genomic Language Models Using Evolutionarily Implausible Sequences.* ACS Synthetic Biology, 15(6), 2256–2263. https://doi.org/10.1021/acssynbio.6c00024
Tzanakakis, G. et al. (2026). *[Independent evaluation of Evo 2 genomic sequence generation].* bioRxiv. https://doi.org/10.64898/2026.01.17.700093
The problem here is motivation/goals. Though there are good reasons that state actors develop bioweapons and might use them, but they have significantly less motivation to do so than a group like Aum. That is, states are capability rich and motivation poor, groups are the opposite, and AI is an uplift on capability. Re: model development, bio models yes will be slower to improve than LLMs, but that doesn't mean they won't be fast to improve. I think saying that the variation was less than natural evolution is not the spirit of what was claimed in the paper, which only claims that it was within the bounds of natural variation. The time scales here boggle my mind. Variation like this in nature could take years to tens of thousands of years. Re: functionally directed novelty, the paper you cite explicitly notes that evolutionary novelty is distinct from functional novelty and does not make claims about functional novelty. Several of the AI phages from Arc outcompeted the wild ΦX174, w/ Evo-Φ69 growing 16-65x versus WT's 1.3-4.0x
It sounds like we mostly agree that for state actors the binding current constraint is strategic, not technical whereas NSAs are significantly more capability bound. A problem in your argument is that you slip between using state capabilities to justify those of potential terrorist groups. For example, here in your reply to point 4:
> “Your team must be both highly competent and completely secret. A contamination incident could destroy the pathogen or infect the team.”
The Soviets ran an effective 30,000-person bioweapons program for two decades.
We also don't disagree that AI will increase certain technological capabilities. Where I think you are too quick is the jump from “AI will increase the technological capacity for certain inputs” to therefore “end-to-end bioweapon capabilities will become accessible to a wide array of actors.” This is exactly where Abi and I disagree. We think that myriad practical, logistical, and institutional constraints will bottleneck the dispersion of real-world capabilities. I think we also both support strengthening existing monitoring and regulatory controls to ensure they continue to bind. An example of a place where you do this is when you translate decreasing costs to increasing capabilities. In many domains capital hasn't been the constraint for years.
On the comment about the novelty of the virus, there were two claims in my original sentence that I should probs have unpacked more. The first was that there was less variation than natural evolution. In particular, here I meant novel variation as measured by the difference from the existing phylogenetic manifold. The relevant results in the Black et al. paper are described in this section from the abstract:
> However, this efficiency derives largely from staying close to previously observed sequences rather than exploring novel sequence space, reflecting the combined performance of the model and additional filters that were applied to its outputs. Compared to baselines of random mutagenesis and serial passage, the model achieves substantial design efficiency while its outputs remain phylogenetically close to natural genomes. We conclude that the generative capabilities of Evo 2 warrant low to moderate biosecurity concern for de novo hazard creation, although the degree to which these findings generalise to larger or less constrained viral architectures is an open question.
Whether you want to say that a generative model sampling from a highly constrained phylogenetic repertoire produces 1–10,000 years of variation is ultimately a matter of taste, but to me it seems like a category error since it's variation of a distinctly different type.
The other claim was about functionally directed novelty. In the case of the phage synthesis experiments, their only goal was to preserve the tropism of a known virus. That is not demonstrating the ability to select for a novel phenotype, which would be a significantly more impressive and worrying sign. Showing that some phages outcompeted wild type doesn't really bear on that limitation, as increased replication was not something being selected for.
Finally, on whether or not the models will be fast to improve, I think there are strong reasons to believe the progress will be slow because biology is a much messier and more opaque domain with much longer feedback cycles. Most of the evidence so far points to that being the case, but like I said, it's something to monitor in case there is a big change.
Also, I should clarify that the reason I'm arguing about this is not because I think we should be totally complacent and ignore biosecurity risks. It's that correctly assessing the danger is extremely important for decisions like whether or not to pause AI development. I think both that we should harden our defenses and that the magnitude of the increased threat due to AI doesn't warrant dramatically slowing progress.
Well I don't think slowing down because of biorisk is really feasible/desirable regardless of the risk magnitude. I'd have to see the contours of what's being proposed though because slowing down could mean a lot of things.
And what is the specific concern around novelty, anyway? You can get significant gain of function without much change to the sequence, and screening evasion is a different beast entirely than simulated evolution (see our latest paper https://www.rand.org/pubs/research_reports/RRA4741-2.html)
Also, could you explain more why lowering of capability floors for bio isn't concerning for you?
Yeah, I used that as an example. I'm also not a fan of most pause proposals, but imo the risks should be part of that calculus.
The reason I was bringing up novelty here was because it relates to how useful the Evo models mentioned in the piece are for developing phenotypically novel viruses, and thereby how much the AI tools improve our ability to do de novo GOF. That paper you cited looks interesting; improving our screening against adversarially designed sequences is an important aspect of hardening our bio defenses.
It's not that I am totally unconcerned about lowering the floor; it's that I think AI only does so much to change the threat environment.
Yeah I'm theoretically concerned about AI enabled GOF, but I feel like we are pretty far away from it now and evo 2 doesn't really bring us closer in a meaningful way.
An analogy I sometimes use is being scared about how fast humans can go in a world with horses only. You design an AI that makes different horse breeds, some of which are faster and some of which are slower. That's slightly concerning, but probably not more concerning than selective breeding. The real issue comes when it can do something like design a car.
I adored this extended exchange between you and Abi. I just read both and my main conclusion is that the two of y’all should hang out. I don’t know enough about the space to opine on the several “how hard is executing lab work well actually” studies but I am basically persuaded that the popular sub community worry about lone wolves specifically seems overrated (but nonzero) and the risk of state actor engagement remains high. I am generally long transferrably on “real world tasks have counterintuitive combined rate limitations” and also “long-horizon prognostication is insanely difficult because of emergent phenomena” so that informs my priors.
I think we are going to! Abi is great. One of the problems with this exchange is I (and other scientists) can’t exactly explain how easy it is or talk about the several switches we have for improving immune evasion because it is a massive info-hazard. On some level, this piece could be a positive contribution to deterring lone wolves/small groups from acting. (I think this is another motte and bailey though, to go from all small groups to lone actors.)
Unfortunately, I think the sensationalism in the headline instead results in making people think this is not a concern, even though she agrees it is getting easier. It is better that we put infrastructure in place now, and not wait until it becomes a bigger issue. (Especially given that state actors at a minimum are already an issue.) I worry this piece mostly contributes to people thinking this is not urgent.
While the above article seems very intelligent, articulate and informed, I feel it's missing an important bottom line, as seems true of most of science and the public at large.
The bottom line problem is a mismatch between 1) the accelerating pace of knowledge and power development, and 2) the incremental (at best) pace of human maturity development. So long as that relationship remains in place, the gap between our power and maturity will continue to widen, probably at an accelerating pace.
Attempting to address threats one at a time as they roll off the end of the knowledge assembly line is a loser's game. It's only when we back up and look at the situation as a whole that there will be some hope.
Hi, I'll probably have more comments, but first an editorial comment. In #5, you have "...AS shows why biosecurity is hard." I think you meant *bioterrism* is hard, right?
Yes, thank you! I did mean that. Fixed.