MoreRSS

site iconJeff KaufmanModify

A programmer living in the Boston area, working at the Nucleic Acid Observatory.
Please copy the RSS to your reader, or quickly subscribe to:

Inoreader Feedly Follow Feedbin Local Reader

Rss preview of Blog of Jeff Kaufman

State of Pandemic Early Warning

2026-09-24 21:00:00

Cross-posted from my SecureBio Notebook.

This is a lightly-edited version of a memo that I presented at the Summer 2026 Biosecurity Summit outside of DC. While others at SecureBio often see things similarly, I'm attempting to present my view and not a SecureBio "house view".

The Goal

We need to be robust to adversaries who want to cause very large-scale harm with biology. This includes actors (human or AI) who want to kill all humans, cause short-term incapacitation or long-term civilizational collapse, or who have strategies for sparing some while they harm others. There are multiple reasons an actor might have these targets aside from being directly omnicidal, such as reducing response capacity during an AI takeover.

That there is an attacker itself is a key constraint to any defensive system: it must be designed for adversarial attacks. The attacker can assess the state of the world's detection systems and plan accordingly. Taken to the extreme, this presents a "minimax" landscape: a system is only as good as its weakest link (the place where it is least sensitive). This is an important framing, and it correctly prioritizes getting some sensitivity towards a wide range of attacks over very high sensitivity towards just a few. On the other hand, (a) to the extent that gaps depend on non-public choices, you can maintain strategic ambiguity to prevent attackers from aiming for the gaps, and (b) reducing the number of gaps reduces attacker options.

The strongest form of success is to deter an attacker by denying them the ability to achieve their goal: an adversary who knows an attack wouldn't accomplish their goal will generally not try that attack. This deterrent effect is tightly coupled to the extent to which detection would indeed thwart the achievement of the attacker's goals. This means that achieving deterrence-by-denial means both building an effective system, from detection through to action, and making it known that you've built it.

The other main kind of deterrence is deterrence-by-punishment. If you develop strong attribution capabilities, an attacker risks identification and retaliation. The extent to which this would deter an adversary, however, depends a lot on what they have to lose. A state seeking strategic advantage might be deterred by the prospect of retaliation, while someone keen on killing everyone (including themselves) has little left to threaten.

Biosurveillance Applications

Where specifically does biosurveillance fit in? What are the threats, and where can it make the difference between an attacker succeeding and failing?

Initial Detection of Stealth Pandemics

A stealth pandemic is one where a pathogen spreads through most of the population unnoticed, with no or unremarkable symptoms, before causing very serious effects. [1] If it were subtle enough, people wouldn't realize how serious the situation was in time to respond effectively. While we call these "stealth pathogens", whether a given pathogen would cause a stealth pandemic depends on the interaction between the pathogen, the body, and humanity's many ways of noticing that something unusual is happening.

Whether it is possible to create a pathogen that would be sufficiently difficult to notice is an open question: I've heard different things from different experts. When considering (a) the significant advances in biological design tools and general biological understanding that we've been seeing with AI progress, and (b) the speed and unpredictability of the process by which unusual symptoms today lead to attention and action, I do think there's a significant chance that within the next five years many actors would be in a position to cause stealth pandemics.

Detection of suspicious sequencing reads may not be sufficient to estimate whether they represent an ongoing stealth pandemic. A pathogen may be constructed in a way that makes its potential for rapid spread and delayed harm obvious, but that is far from guaranteed. Assessing this likely requires additional scientific work: genome completion, estimating likely effects in the human body, and considering whether the genome suggests an intentional attack.

Triggering Initial Response

A pathogen doesn't have to be stealthy to be disastrous: it could simply be very hard to contain (a "wildfire pandemic"). Beyond its direct effects, such a pathogen could be intentionally released to reduce capacity at a critical time, such as during a coup or an AI takeover attempt.

Whether an outbreak is wildfire, stealth, or has aspects of both, the time from initial discovery to serious response is critical. Initial indications are generally ambiguous, and it is often difficult to understand the extent or trajectory of the threat. With the 1976 swine flu, we overreacted and vaccinated 45M people because we didn't have the monitoring to know it wasn't spreading widely. With 2014 ebola in West Africa, we underreacted and let it spread freely for three months because the extent wasn't recognized. Similarly, with 2026 ebola, we saw another three month delay, this time in part because field PCR tests couldn't see it. The case of 2009 H1N1, however, showed how a well functioning (though flu-specific) biosurveillance system could enable timely response. In today's COVID-weary climate where public health is deeply worried about losing credibility through false alarms, biosurveillance can help avoid a default of delaying response while waiting for more information.

Enabling Ongoing Suppression

Once an initial response is in motion, you need monitoring to effectively deploy mitigations and know whether they're working. How much value there is depends on how symptoms relate to infectiousness. If symptoms are absent or highly delayed, effective response is essentially impossible without solid monitoring. At the other extreme, if symptoms are highly visible and begin immediately, monitoring is moderately valuable: you know you have a problem, but with "fog of war" you don't fully know its extent or distribution. Large-scale monitoring allows you to compare locations and track trajectories to optimize resource deployment.

I give relatively little attention to this biosurveillance application in this memo, mainly because I think it's a place where what exists today is closest to what needs to exist, so this is a lower priority for additional work.

What does success look like?

There's no single threshold, where you win by building a system with a specific level of capability. Biosurveillance systems reduce the cost-benefit tradeoff of initiating a pandemic; increasingly capable systems decrease the likelihood that an attacker deems this worth their effort and reduce the harm if they decide to attack. Still, for each application, there are some 'sweet spots' where the cost-benefit ratio is maximized.

Initial Detection of Stealth Pandemics

If you imagine the most capable system that can be deployed for a given level of investment, it will be capable of averting some fraction of expected possible stealth harm. Here's how I see it:

  • A system that is too slow to beat status-quo detection might still help with triggering initial response, but doesn't provide initial detection benefit.

  • A system that is a bit faster but doesn't give enough time to act before most people are already infected provides some benefit, especially in cases where the harm can be mitigated if it's known soon enough, but most expected harm will still occur.

  • A system capable enough to flag an attack in time to allow us to protect enough workers, such that (a) civilization does not collapse and (b) medical countermeasures can be developed and deployed, provides significant benefit. Attacks could still be massively disruptive.

  • An extremely capable system, including a very broad sampling regime, could flag outbreaks when they were still small enough to contain. At this point an attack is still costly, but the disruption is limited to the areas where the attack was seeded.

None of these are hard boundaries. There are factors that 'smear' these thresholds across many capability levels: some of this is luck (ex: who contributes to what samples), while some is uncertainty about the world that both we and an attacker would share (ex: how much shedding a given pathogen would actually produce in a large population). Here's an illustrative chart:

I've intentionally left the x-axis vague. It's not "cumulative incidence at detection", because (a) that's horizontally smeared as described above and (b) it would imply that increased capacity is downstream of sensitivity only and not other important factors like what kind of pathogens you can detect at all or how quickly you can trigger response. Instead, the x-axis represents the level of resources invested.

A system that is sufficiently capable to flag pathogens early enough to protect enough workers to avert civilizational collapse is well worth the investment; additional sensitivity beyond this is valuable, but likely substantially less cost-effective.

This means that success looks like a pathogen-agnostic system that flags attacks in time to protect enough workers to prevent civilizational collapse and buy time for pathogen-specific mitigations. How early this needs to be depends on how quickly response can happen. If effective response takes a month from detection, you need to flag before ~0.05% of people have been infected. On the other hand, if it takes two weeks, you only need to flag before ~2% of people have been infected, and if you can get it down to seven days, then even flagging at 10% cumulative infections would be enough. See appendix for more detailed reasoning.

Note that the "fast enough to beat status-quo detection" regime would be a far higher bar for wildfire than stealth. This means that, on time scales rapid enough to factor into planning, I don't expect it to be economically feasible to build a system where biosurveillance would be your first indication of a wildfire pathogen.

Triggering Initial Response

We don't know very much about what would actually get decision-makers to take sufficiently prompt action. In a stealth scenario, this is extremely challenging, since response must begin before the main symptoms manifest, but even in a wildfire scenario, there's a huge difference between a response that follows immediately from when someone identifies the first cluster vs one where it kicks off in earnest only after deaths start to become highly visible.

Success looks like a very short time, ideally under a week, from when a catastrophic pathogen is flagged until the danger has been recognized, PPE has been distributed to essential workers, biohardening has been deployed or activated, lockdowns have been instituted, and development of rapid diagnostics and other medical countermeasures has begun. These are very costly actions, in economic terms but also via anteing political capital and institutional trust. Biosurveillance can contribute by giving decision-makers the information to determine whether those costs are worth paying. This looks like clarifying the extent of spread to date, estimating trajectory, and performing initial wet lab work such as genome completion.

Beyond the technical work, response requires trust. Before someone will act on an alert from a system they need to believe that it indicates something real, and that trust needs to be built over time. Non-catastrophic detections are key, showing you can track trends that match what other evidence shows, surface matters of public health concern, and turn up real engineered 'benign positives'.

Most of the work in reducing time to response, however, is outside biosurveillance. This could include helping government agencies develop better plans for how to handle various indications, running exercises that get the actual principals to experience the feeling of making these specific calls with realistically incomplete information, or streamlining inter-agency communication so available clinical and epidemiological data gets to the right people quickly.

Enabling Ongoing Suppression

Wastewater PCR was used by Australia, New Zealand, and Singapore (down to the building level), among others, as part of their suppression strategies to guide public health response during COVID-19. At the point when transmission has been diminished to some threshold, or if an attack is identified before the pathogen has spread widely, a sensitive biosurveillance system can tell you when and where you need to focus more expensive and intrusive detection methods and interventions.

This means that a system for successfully maintaining suppression looks like a larger-scale and more fine-grained implementation of the biosurveillance component for triggering initial response. Knowing that something is spreading in ten major cities around the US might be enough to spur rapid action, but you need a much more detailed picture if you're trying to maintain ongoing suppression.

What gets us there?

Initial detection of stealth pandemics is SecureBio Detection's focus, and where I have the most developed view. I'll walk through what I think is needed for this scenario, and then discuss how this changes for other scenarios. Please don't interpret this structure as a claim about the relative likelihood of stealth scenarios!

Sampling Strategies

Municipal wastewater is a very helpful sampling modality: in most cities, sewage is processed at a small number of locations, letting you track pathogens across hundreds of thousands of people from a single easily-collected sample. Since many pathogens don't shed heavily into wastewater, however, and it's a very noisy sample type, comprehensive initial detection at a reasonable cost likely requires tracking additional sample types. SecureBio runs a nasal swab program; beyond swabs I see air, blood, aircraft wastewater, and leftover material from clinical tests (clinical lab discards) as the strongest candidates for supplemental sampling. SecureBio has looked into these strategies in some depth (initial overview, blood, aircraft wastewater, clinical discards), but there's still a lot that could be learned here.

Lab Technology

You need technology that is pathogen agnostic: if you monitor only specific pathogens, the adversary can choose ones you don't monitor. With current and near-future tech, "pathogen agnostic" means sequencing. For viruses, it is practical to concentrate particles by size, and then perform untargeted and enriched metagenomic sequencing metagenomic sequencing. This lets you do the rest of the detection in the computer, for maximum flexibility and generality, and this is what SecureBio does today.

For bacteria, I'm pessimistic about adapting this approach directly, at least with wastewater, because the genomes are much larger and there is a rich background of sewer bacteria that can't be physically separated from potentially threatening bacteria before sequencing in the same way that viruses can. How to handle this is still an open question; see below.

For mirror life, the initial stages of spread might be very hard to recognize, and an important step would be learning that mirror life was spreading at all. It's likely that a highly sensitive and relatively cheap assay could be developed, but no one has started on this yet.

Computational Technology

Metagenomic sequencing moves much of the problem of initial detection from the lab into the computer. This means massively more sequencing reads than humans could evaluate (8+ orders of magnitude), so we need a detection system that can identify which reads indicate something concerning is happening. Some reads can easily be recognized as concerning (ex: nucleic acid subsequences unique to the smallpox genome should not be in wastewater) while others require very sophisticated processing (ex: understanding the complex background well enough to flag de novo genomes with no sequence similarity to anything currently existing). The general approaches are looking for sequences with one or more of the following features:

  • Dangerous. Sequences matching known pathogens that you would not expect to see in the sample, perhaps signaling that they've been introduced.
  • Modified. Partial sequence matches, perhaps signaling something engineered.
  • New. Sequences that you haven't seen before, perhaps because they were created de-novo.
  • Growing. Sequences that are becoming more common, perhaps signaling that they're spreading through the human population.

The computational approach is very difficult given the scale of the data, uncertainty over what an attack might look like, and the need for rapid analysis. On the other hand, these are the kinds of highly computational problems where I expect AI can be very productively applied, whereas many other aspects of this system require relatively slow real-world effort.

What exists today?

Pathogen-agnostic biosurveillance is still in its early stages. I know of four systems doing untargeted metagenomic sequencing for biosurveillance today:

  • CASPER (SecureBio + Marc Johnson's lab at the University of Missouri and other academic partners). This is wastewater, primarily municipal, from 49 facilities representing 24 US cities. It's virus-focused, and generated with very deep short-read sequencing. Lab work happens independently at MU and SecureBio, with bioinformatics at SecureBio.

  • Zephyr (SecureBio + Helena Solo-Gabriele's lab at the University of Miami). This is pooled nasal swabs from Boston and Miami, collected primarily in public places, bringing in swabs from about 1,000 people weekly. It's virus-focused, and unlike CASPER uses long-read sequencing.

  • ANTI-DOTE (DoW + PHC + SecureBio). This is wastewater from five US military facilities, which SecureBio processes under contract from PHC. These samples go through the same lab and bioinformatic processes as CASPER samples.

  • mSCAPE (UKGOV). This is bronchoalveolar lavage from UK hospitals, currently including relatively few samples, which limits sensitivity. Unlike the previous three, mSCAPE uses combined viral and bacterial sequencing. These samples are sequenced to a relatively low depth with long-read sequencing, and the data supports clinical practice, public health, and biodefense.

There are also several hybrid-capture sequencing projects that might detect an attack if the agent was similar enough to existing pathogens.

In addition to detection via pathogen-agnostic sequencing, there are also paths where an outbreak becomes visible via showing symptoms in a sufficiently large fraction of infected people. This could lead to suspicious clusters, and then to sequencing and noticing that a genome looked edited. This is not a well-developed path today, but (as discussed below) I'd like to see investment here.

What's missing?

Here's an overview of what I think most needs doing. It represents the current state of my thinking, but it's not as thoroughly considered as I wish it were: please don't overweight it in your own decision-making! While SecureBio Detection is exploring some of these, I think the ideal structure is a healthy ecosystem of organizations taking on different parts of the problem in parallel. SecureBio is often able to share samples, sequence prepared nucleic acids, or share data to help others make progress.

In roughly descending order of how valuable I estimate non-SecureBio work would be, representing a combination of both overall value and the value of the work happening independently:

  • [Other] Modeling. We need good estimates of the necessary system scale and ideal network design to support optimal initial detection, response triggering, and ongoing suppression. There are initial estimates, but because this is a huge question (essentially pulling the whole field together), current work is rough. Rigorous treatments would be valuable for planning, both for independent funding allocation and for governments. This is especially valuable to happen outside of SecureBio, as a way of checking our work, and doesn't need to all be executed by one group (groups can pick off subquestions).

  • [Comp] Red teaming. It's important to analyze existing systems to assess how well they would handle a range of adversarially designed attacks. This can be done with reference to the system implementation, or by treating the system as a black box.

  • [Other] Parallel orgs. SecureBio Detection has historically focused on the US. Setting up parallel orgs in other geographies, especially Europe and Asia, would be really valuable. We are happy to advise anyone interested in doing this work!

  • [Lab] Detection of stealth bacterial pathogens. Viral particles are small, which means that you can separate them from human and bacterial cells without specifying in advance what sequences you're interested in. Beyond this, pathogenic viruses represent enough of the total viral portion that just pulling it all into the computer to sort it out there is economical. This method can't be directly applied to bacteria, at least not in wastewater, because there is so much irrelevant bacterial genetic material. SecureBio hasn't done any work in this area yet, but some plausible approaches include aggressive depletion for things you know are not worrying, partially-targeted methods, and working with samples with a more favorable background. People have a wide range of estimates on how difficult this will be, but personally I expect it to be very hard.

  • [Comp] Parallel methods development. It would be great for others to be exploring other avenues in parallel. Since our main expertise is in traditional bioinformatics, I'd be especially excited to see people try other approaches, such as applying deep learning. We share much of our sequencing data publicly (PRJNA1247874, PRJNA1379685), in part to facilitate exactly this type of work.

  • [Lab] Sampling streams beyond wastewater and nasal swabs. Not everything sheds much into wastewater or the nose. A comprehensive system very likely needs a wider range of sample types. Blood, air, and clinical lab discards are all very promising here.

  • [Other] Response. Figure out what would get decision-makers to reliably act rapidly in a real emergency, and lay the groundwork so that happens. As noted above, this is probably mostly not biosurveillance. There are components (ex: developing escalation relationships) that need to be tightly coupled with biosurveillance, perhaps within SecureBio and parallel orgs, because of the limitations of information sharing. Other components (ex: policy recommendations, public outreach) make more sense as separate organizations. This is delicate work, however, and someone coming in noisily and insensitively could easily set the field back.

  • [Lab] Genome completion. A pathogen identified by a short-read biosurveillance system like CASPER would start as just a suspicious section of a genome. You can sometimes learn more about the genome through techniques such as outward assembly, but only if you happened to sequence the relevant reads. Lab work to assemble the whole genome lets you better understand what the pathogen would do in a human, which is likely on the critical path to both assess whether a response is warranted (and, if so, what that response should be), as well as to get people to act appropriately quickly.

  • [Comp] Analysis tooling. How do you go from "this genome looks worrying" to a good understanding of whether it's engineered and what effect it would have in humans?

  • [Other] Swab collection. Currently, Zephyr requires field samplers standing on street corners and interacting with the public, and a team brings in ~40 samples per hour. If instead workplaces or schools could be convinced to integrate sample provision into daily routines, this could be far more scalable.

  • [Other] Clinical sequencing. Sequencing of clinical samples will likely eventually be deployed broadly based on its clinical benefits alone, displacing a wide range of pathogen-specific tests, but by default, this displacement process is far too slow. Sequencing needs to be a standard option doctors can easily reach for when someone presents with unusual symptoms. I think the tech is ready, or could be ready with a small amount of R&D, and so this is a commercial opportunity.

  • [Lab] Sensitivity increases. Untargeted metagenomic sequencing for pathogen-agnostic initial detection is an early-stage field, with relatively few people exploring it, and so on priors I expect there are large sensitivity improvements waiting to be discovered.

  • [Comp] Deterrence tracking. You can ask models to estimate the probability of an attack achieving its goals. If models report low probability to defenders, they probably report low probability to attackers. The best models I have access to can't do a good job at this in response to a simple prompt, and if you need a series of prompts, you can't expect your answer to reflect what an attacker would get. I expect this to change quickly, however, and it would be good to have an automatically-updated tracker showing the range of answers LLMs give to the question of whether existing detection systems are sufficient to make an attack not worth an attacker's while. This would require some thought on which threat models to poll for and how to phrase the question in a way that is a good proxy for what an attacker would ask. It also closely ties into red teaming work.

  • [Lab] Partially targeted methods. Hybrid capture, tiled degenerate amplicon arrays, and other methods of mismatch-tolerant sequence-based enrichment could offer far higher sensitivity. While they're unlikely to be sufficiently general to address all threats, there is a good chance that it makes sense for a mature detection system to include a partially-targeted component as a way to effectively exclude large areas of the threat landscape.

  • [Lab] Cheaper protocols. Sequencing is still an expensive proposition, but as sequencing has become cheaper the operation of the sequencing machine is no longer the largest cost on a per-sample basis. Bringing those other costs down would allow more scale for a given budget.

  • [Lab] Shorter lab time. Current metagenomics protocols require about a day on the bench and about a day on the sequencer. There are already strong pressures for reducing sequencer run time, and new machines are coming out in the ~8 hour range, but there's a lot of work that could be done to speed up the rest of the processing.

  • [Lab] Pathogens outside of bacteria and viruses: fungi (and oomycetes), parasites, and prions. These are generally much more difficult for an attacker, especially for strategies that involve rapid spread.

  • [Lab] Mirror life. Current sequencing wouldn't detect mirror life at all. Municipal wastewater samples would be a good place to look, which allows you to re-use much of the collection infrastructure, but would require an assay developed specifically for mirror life. I have this low because I currently expect mirror life to take long enough to develop that there will be time later for assay development.

How does this change for accelerating response to a wildfire pandemic?

During the COVID-19 pandemic, many groups built out PCR-based targeted wastewater monitoring. In the US, wastewater monitoring networks include NWSS (CDC), WastewaterSCAN (philanthropic), and Biobot (private). EU member states, Canada, Australia, and other countries track wastewater as a standard part of their public health systems. There are maybe two dozen countries that have some form of wastewater PCR that could be retargeted to track a new pathogen in an emergency once its genome was known.

On the other hand, none of this would move quickly enough today to address a wildfire pandemic. There are delays throughout the process: some of this is technical (ex: stocking consumables), but most of it is organizational (ex: policies that permit setting production work aside, overtime budgeting, on-calls, deals with synthesis providers for rush orders). Even in a serious emergency, I think ten days is a good best-case estimate today. With good preparation, however, three days is possible. I think getting these existing networks to prepare for rapid turnaround emergency response is really valuable, and SecureBio has started to have some of these conversations.

This is also a place where the same metagenomic sequencing system you would build for stealth pandemic detection could help you cut off additional days in your response. The technical and organizational delays that slow down PCR-based detection are downstream from how changing targets requires making changes in the physical world. Untargeted sequencing lets you skip those steps, at the cost of much lower sensitivity. On the other hand, once you know what you're looking for, you can use approaches (like PCR) that are significantly more sensitive in the case of SCV2, by a factor of about 100. If you've built a sequencing system that can flag a stealth pandemic before 1% of people have been infected, then that factor of ~100 means it would be able to confirm a wildfire pandemic at ~0.01%. Still, it's not clear to me that even 0.01% is early enough. It's possible that you need 0.001% or even lower, at which point this argues either for a substantially larger investment in untargeted sequencing (to get enough data quickly) or giving up on sequencing for this application (because it can't economically reach the target sensitivity).

If you wanted to firmly decide whether to go with MGS or PCR to accelerate initial response to a wildfire pandemic the key thing you'd need would be estimates of (a) what fraction of the population would likely be infected when a wildfire pandemic was noticed, and then (b) how many doubling periods there would be before decision-makers took action in the absence of this system. On the other hand, I'm not sure a firm decision is needed, and instead lean towards different groups exploring these approaches in parallel.

How does this change for ongoing suppression?

The same PCR-based targeted wastewater monitoring that was built for COVID-19 and could potentially accelerate response to a wildfire pandemic could also be applied to ongoing monitoring to support suppression. This is already widely understood to be valuable, but there is still less investment here than there should be. In a legitimate emergency, I expect governments to be able to organize existing capacity and deploy it reasonably well, but not as quickly as would be ideal. My bigger worry is whether that existing capacity would be large enough. For example, you would ideally have monitoring at the neighborhood or building level, which means you'd need a very large number of in-manhole composite samplers. Since these are low-volume products built by a small number of manufacturers, it would be hard to build more quickly during a crisis.

The main work is building up capacity in advance that can be quickly deployed in an emergency. The best bet for such capacity is systems installed for ongoing public health monitoring: this ensures that they work, including as part of a larger system, and that lots of people know how they work. It also gives some ongoing benefit, which may make it an easier sell than stockpiling.

Appendix: Initial Detection Scale

There is a long chain of reasoning in estimating in what fraction of attacks a system would achieve the goal of protecting enough workers, and that chain involves several steps where our knowledge is limited. Still, we can make the best estimates we can. The three key parameters, are:

  • How quickly would the pathogen double as it spread through the population? This is a combination of the basic reproduction number and generation time. We've generally worked from a doubling period of ~3 days, which represents the high end for naturally occurring pathogens. You could argue that it should be lower, because a pathogen could be designed to spread much more quickly than existing pathogens, or higher, because of physical limits on how quickly a pathogen can spread without attracting attention.

  • What fraction of workers would need to be protected to allow the development of medical countermeasures and avert collapse? This is also not something we've focused on, but the difference between 35% and 70% would again represent only a factor of two in the required detection sensitivity. Here I've put it at 50%, following Patel et al. in Physical Approaches to Civilian Biodefense: "Protecting 100 percent of [Vital Workers] is likely not strictly necessary because [National Critical Function] operators likely have enough flex capacity to handle a small fraction of [Vital Workers] being absent, but protecting more than single-digit percentages of [Vital Workers] is likely necessary to keep [National Critical Functions] operational. We therefore chose 50 percent of [Vital Workers] as a convenient midrange protection target." [2]

  • How long would it take from initial detection until effective mitigations were in place to protect workers? For every doubling that happens during response, we need the detection system to be twice as sensitive: since required sensitivity grows exponentially with response lag, reducing that lag becomes the best use of marginal dollars after a relatively small initial investment. We've worked from a response time of ~15 days: this is not where the world is today but we think it's achievable.

Taking this all together, you get the target that SecureBio has been working towards for a while: a system sensitive enough to avert civilizational collapse would need to flag a pathogen before, very roughly, 1% of people had been infected: 1% * 215/3 = 32% < 50%.


[1] While mirror bacteria could spread through the environment instead of between humans, to the extent that they might still propagate widely before detection, I group them in with stealth.

[2] This is not dependent on which specific workers are considered "vital" or "essential" or even how many there are, though of course that has large impacts on the question of how to get them protected.

Comment via: facebook, lesswrong, the EA Forum, mastodon, bluesky, substack

Initial DIY Cleanroom Experimentation

2026-09-22 21:00:00

In It May Be Possible to Improvise A High Grade Bioshelter, Adin Richards discusses the possibility of improvising defenses against an environmental threat such as mirror bacteria. He gives an exploratory overview of why it might be possible to apply materials and equipment people often already have in their houses to pressurize all or part of a house with filtered air. It would be great if this were possible, but with all the ways for an improvised system to fail I'm pretty skeptical.

I decided to try a simpler version, testing how much I could positively pressurize a single room with a relatively powerful HEPA air purifier (AirFanta 3Pro. This is close to a best case, since most people won't have something as good as an AirFanta.

My first question was whether the AirFanta can pressurize much of anything, since it uses axial fans and these can stall when facing excessive pressure:

(By Prj1991, via Wikipedia)

To measure the pressure delta I was going to need some kind of manometer. Most cheap test tools are designed for very large pressure differentials, but I found the Testo 510i which looked to be just barely good enough with its 5 Pa rated accuracy.

The kids enjoyed testing it out:

I taped a trash bag around the top of the AirFanta and measured what pressure it could get to:

There were some small gaps in the taping, but it got to 83 Pa. Pretty good!

Then I tried pressurizing a room. Someone in a tight modern house could probably get close to that 83 Pa, but ours is old and leaky so I wasn't expecting anything in that range. I opened the bag the rest of the way at the top, taped it across the door, and used plastic sheeting to close off the rest of the doorway:

It billowed out, and I felt air escaping through the many tiny gaps in my taping. I forgot to take a bias measurement first to calibrate, though, so I don't have a pressure number for this setup.

The issue with putting something over the door is that every time you go in and out of the room you need to carefully seal and unseal it. Which uses a ton of tape! Unless your room has two doors, you'll quickly run out of tape. Clearly we should go through a wall.

I made a 13x13" hole in my wall [1], and mounted the AirFanta horizontally. It immediately fell apart, because it uses friction fits that aren't designed for sideways loading. I reassembled it, and used baling wire to keep it together:

I taped around the edges to prevent air from escaping:

It looks even sillier from outside the room:

I measured 12 Pa. I'm not sure how seriously to take this: the meter is only rated to 5 Pa. Moving the hose in and out of the room repeatedly it was consistently a 12 Pa delta, so maybe it's real? A professional meter that's very sensitive is an expensive proposition, but the raw sensor is cheap and a DIY project sounds fun. I've ordered a SDP810-500PA pressure sensor, and will wire it up with an Arduino. [2] With fine-grained pressure sensing I should be able to plug holes and see what makes the number go up, which also suggests that if at some point people are going to do this in a widespread DIY manner we'd want cheap phone-attached pressure sensors.

There's also a big question of what pressure delta you actually need. Unfortunately it's not a single number: Adin has estimates based on wind speed, but if the leaks are through cracks that run into the wall cavity or other parts of the building (ex: baseboard leaks) then that's much more wind-resistant than if they run directly to the outside (ex: window leaks). I haven't been able to think of a good way to measure this: maybe waiting for winter, running the fan in reverse (negative pressure) and using my fingers, a thermal camera, or a thermometer to get a sense of how cold the incoming air is?

I'd also like to measure how well this works in a realistic test, beyond extrapolating from pressure plus knowing that the AirFanta uses HEPA filters. I did some very rough testing where I made smoke in the kitchen and looked at levels in this room, but even with the fan off particle levels stayed very low as long as I kept the door closed. I could try lighting eastern Canada on fire again, but I think that wouldn't be worth it: the levels of smoke I need for a real test are impractically high. During the smokiest bit of the summer, an AirFanta fully within the house could already get PM 2.5 down to the limit of detection of my meter.

When I'm done I'll put something over the hole, leaving it available for potential emergency use. But for now, more testing!


[1] This was more work than eight words suggests.

[2] While I'm at it I'll also set up a PMS5003 particle monitor (which I've been eyeing since this post). I don't really know much about electronics, but Claude can walk me through.

Comment via: facebook, lesswrong, the EA Forum, mastodon, bluesky

Why I Stay Off Twitter

2026-09-19 21:00:00

I avoid Twitter (𝕏) for similar reasons to drugs: I think it would change me for the worse, and I would be unable to give it up.

After staying off Twitter reasonably successfully for years, I cross-posted my AI Tweets there a few weeks ago. I had something very Twitter-shaped to say, and I thought it was important to get out, so I do think this was worth it. And it all went well: none of this is complaining about the comments I got there.

Coming back a few times to check notifications, however, it's been very good at baiting me: Tweets that are confidently wrong in cases where I have relevant and uncommon knowledge. The pull to dive in and share what I know is very strong! Then this bleeds over to the far broader case where people are wrong, and you have a large potential time sink.

If it were just the time sink, I'd stop resisting. I spend some time on HN and Reddit, and to the extent that Twitter could substitute for that by showing me things I was more interested in, that wouldn't be an issue. The real problem is the culture.

Twitter has developed a culture of raising the temperature, and generally being mean. People dunk on each other, and are aggressively uncharitable. I see how the incentives push this way: fights and taking sides draw people in, and an engagement-based algorithm [1] will move the crowd in this direction. You internalize the drive to make the internet numbers representing your influence go up, and I know I'm very susceptible to this. But the practices this develops are bad for discourse, and becoming "good at Twitter" would strengthen harmful aspects of my personality and push me towards being a worse person. [2]

A lot of people swear by heavy use of blocking, lists, and the "Following" feed. I expect this to help, but not enough: the bait is in the replies and quote-tweets, and from the outside it looks like the culture seeps everywhere.

This makes me a big fan of Zvi's roundups (example), where he quotes relevant Tweets and gives context. Just as academic AI publications have moved from journals to preprints, a lot of the news has moved to Tweets. And AI news is both really important and poorly covered by traditional media: if Zvi didn't put out these roundups I might have to break down and read Twitter. I'm super thankful that this means I can stay informed without opening myself up too much to Twitter culture.

So I'll continue to accept that I'm missing out on a place where people have interesting public discussions on topics I care a lot about: engaging would expose me to an incentive gradient I don't think is worth the risk.


[1] Now that we have better algorithms, in the form of LLMs, we could have much better ranking signals. For example, you could write an equivalent of the HN guidelines or Duncan's Basics of Rationalist Discourse and then give more visibility to posts that did a good job of exhibiting the kind of discourse you'd like to see more of. Of course this requires the platform to want this: it likely lowers engagement, and hence revenue. But maybe there are paths that run via taking a long view on building a site that's more influential or where people feel better about the time they spend there.

[2] There are of course many people who remain wonderful people despite being popular on Twitter, through natural immunity or impressive resolve.

Comment via: facebook, lesswrong, mastodon, bluesky, substack

Teleoperated Humans

2026-09-13 21:00:00

When I look at why I expect the world to change a lot in the next few years, and why other people expect slower changes, I think a big component is disagreement on the extent to which AI will affect non-computer work. Sure, programming has sped up massively with Claude Code etc, and models like Astra seem posed to make similar changes to work with spreadsheets and other common business tools, but what about work that doesn't include computers at all?

The classic picture of AIs doing things in the world is robots, but I think a more realistic picture of the near future is computers telling people what to do. Leaning into the way the world has become very scifi, we could call this "teleoperating" people. Many things that are hard for robots are very easy for people, there are strong economic reasons that push towards teleoperation, and this bypasses many legal and social limitations on what AI can do. We should expect this to lead to large and rapid changes in the physical world.

One of the most widespread examples today is driving. I put my destination into the GPS, and it tells me what to do. I handle the low-level physical motions and responding to the local circumstances; the GPS has a broader view of the world and handles the strategy.

When I think about why this happened much earlier than the huge amount of "teleoperation" I expect to see soon, a few factors. Driving is a major human activity, so it was worth making navigation software at a time when AI wasn't very good yet, even though this meant a ton of human hours going into building the system. It was also a place where the strategic component was a very strong fit for automation. You can memorize the map with enough work, but even then you won't have real-time street-by-street traffic information. AI solved this problem so well we don't even call it "AI" anymore. On the other hand, driving is a realtime control problem in an unconstrained environment where people die if you screw up and you can't even always safely stop. This makes it hard to automate, but also would make it impractical for an AI to guide non-drivers through the process. The only reason Uber etc have been able to commodify driving as they have is that so many people already know how to drive.

Thinking about where else we might see this, most AI use today looks a lot like management. You figure out what you want it to do, and describe in detail. It asks you some questions up front and others while it works. After some churning you get some a work product to assess. Maybe there's more back-and-forth, or maybe it's good as is. You set strategy and give context; the AI handles the implementation. Today's AI is normally only applied to the implementation to the extent that the task can happen fully within the computer. In cases when the AI can't physically, legally, or intellectually do something, the most efficient path to completing the task will often be the for the AI to handle strategy while delegating to a human to fill these gaps.

To illustrate what this delegation pattern can look like, let's look at how I recently got my Whistle Synth app into the Mac App Store.

At a high level, I set the strategy: "Can you walk me through the process of getting this into the Mac App Store?" But everything after that was either handled by the AI or delegated back to me. It handled included figuring out what tasks needed to be done, modifying the implementation to be compatible with the App Store restrictions, building the app, and giving me instructions. And then it delegated to me to record a demo video involving whistling (physical), register as a Mac Developer (legal), and clean up its App Store description (intellectual).

This was mostly pure instruction-following on my part: I was being teleoperated. Here's one example:

Claude Code:
  1. Open the Profiles list https://developer.apple.com/account/resources/profiles/list — sign in with the account for [team].
  2. Click the blue + next to "Profiles". You land on "Register a New Provisioning Profile", a page of radio buttons grouped into Development and Distribution sections.
  3. Under Distribution, select "Mac App Store Connect". Not "Developer ID" — that's for distributing outside the store. Click Continue.
...

This was relatively mindless work for me. Just like being navigated through a city I don't expect to return to, I didn't bother trying to learn how this worked. I was loosely paying attention to make sure I wasn't doing anything dumb, but for future more capable systems I expect people to stop attending even that little.

Once it finished walking me through submission I had to wait a few days for review. It was accepted in the first round with no reviewer comments. This is a pretty big deal: App Store rules are notoriously complex, the reviewers very picky, and as a first-time amateur Mac developer there's no way I would have gotten this all right on the first attempt pre-AI.

Even though this was an almost entirely within-computers case, the important thing here is the pattern: by following AI instructions I did something that would have taken me a ton of work to learn how to do alone.

Note that in this case I was both doing the high level strategy ("put this in the app store") and filling in gaps for the AI (clicking a blue plus in App Store Connect). As "teleoperation" becomes more common I expect some of this, as people automate away parts of their jobs. Other times I expect it will look like, for example, a highly AI-pilled startup founder directing AIs that direct employees. A lot like gig workers "below the API" today. I expect early iterations of these jobs to be frustrating, with the AI not delegating well. Then, as AIs get sufficiently good at directing and anticipating, they'll be pretty mindless, for better or worse, as you stop needing to think for yourself at all.

What sort of jobs might switch to being teleoperation? The top candidates are any where the physical motions are relatively straightforward, timing is not critical, and people today are paid a lot for their knowledge and judgement. If you had an expert looking over your shoulder and telling you what to do, I expect most of you could do most of the work of an electrician. In fact, that's the bulk of how electricians learn their trade: through apprenticeship. Same goes for mechanics, healthcare technicians, inspectors, etc: they combine physical and intellectual components, where it's the knowledge that keeps a random person off the street from being able to do the job. People wearing glasses with built-in cameras, connected to today's strongest AIs could already do a lot with a bit of scaffolding.

To have a large impact, teleoperated workers wouldn't need to be able to do 100% of an existing job category. As long as the parts that can and can't be done this way can be easily separated, 90% could be done by teleoperated novices, while some of the former professionals spend their time on the remaining 10%. When I think about how these other jobs are likely to go, I expect we start with ones without regulatory barriers: HVAC techs (typically unlicensed) before electricians (licensed) before surgeons (licensed + heavily regulated + realtime + high stakes). [1]

So, teleoperation is probably very economically productive. Is it a good thing? I think mostly no, for several reasons. The big one is that I expect it to speed up the rate at which AI advances turn into additional AI advances. This shortens the time our society has to figure out what to do about these massive changes, and increases the risk that immature technology is rolled out widely. Rushed deployment is more likely to lead to disaster, and there are many ways this could go extremely wrong. And by "extremely wrong" I mean "AI kills everyone wrong". Creating minds smarter than ourselves is the most consequential thing humanity has ever done or will ever do, and we have to get it right.

Which is why I'm heartened to see a lot of support, including from the CEOs of Anthropic and OpenAI, for managing the pace at which these systems become increasingly capable. But even if we held constant at the capabilities of models publicly available today (let alone trained but not yet released) I think widespread teleoperation is still very likely. I expect this to be a massive disruption, one very difficult to integrate into our existing societal system.

The first issue is just that I expect these to be unpleasant jobs with low negotiating power. Since there are many tasks that almost anyone could do if expertly advised, and the employer can easily filter out the people who can't or won't, there's very little to keep wages or working conditions up. Then add in competition from laid-off knowledge workers, and I expect unprecedented unemployment.

So even if we can avoid the large risks of losing control of the future, falling into AI-enabled authoritarianism, facilitating bioattacks, etc, how we handle a world in which most people can't find work that pays them enough to live on will be an serious challenge. I expect this will require very large scale redistribution. [2] I'm not sure this happens by default, but I think it's achievable with significant effort. And as a very small fraction of spending in a vastly larger economy it would be a much easier sell.


[1] For a future post:

$ echo "[redacted]" | sha512sum
7820a2ecae1fcab8d7a29fe4f98f56b96c403cdfb9a6833fad0198070e118233490a078c9230de0c6ed5cf1c6cf557fef493aeb33118e75b401deabbf3a1aae4 -

[2] Looking at what there is already, the US does less than most rich countries, but even here we have medicaid, EITC, CTC, WIC, SNAP, SSI, TANF, Section 8, LIHEAP. We spend maybe 3-5% of GDP on means-tested programs. Then ~7-10% of GDP goes to things like universal public education and medicare which aren't directed specifically at the poor but are still effectively redistributive. Internationally there's been some of this, but much less; until recently the US was spending maybe 0.04% of GDP on the kind of foreign aid (ex: PEPFAR) that is really about helping the world's poorest, and then the private sector (Gates etc) adding maybe 0.1% of GDP.

Comment via: facebook, lesswrong, hacker news, mastodon, bluesky, X, substack

Replace Net Metering With Batteries

2026-09-11 21:00:00

We should end net metering, including for existing solar installations, and compensate people with a one-time subsidy for battery purchase. I'm going to give an argument from grid efficiency, which I expect to be the main consideration for most people, though the benefit that makes me enthusiastic about this change is actually increasing societal resilience.

Net metering is a common form of solar subsidy where you only pay for your "net" usage: the difference between how much you consume and how much you produce. At first glance this doesn't even seem like a subsidy: if you take 800 kWh and put back 800 kWh, then did you really use any? But it's not like a bank account: the kWh you put back are usually much less useful than the kWh you used.

Say I started a solar farm, putting out a lot of panels somewhere out in the less populated part of the state (MA), and sold the power to the grid. Averaging over the year, the electricity market might pay me $0.05/kWh. On the other hand, when the panels on my house send power back to the grid I get $0.32/kWh. [1] There are several factors that pull these apart, but I think the most illuminating one is how the value of electricity varies over time.

The $0.05/kWh that the solar farm might receive in direct market compensation is an average. It's a market-based system: when supply is high relative to demand you don't make much, and vice versa. In the summer, you might see lows of ~$0.02/kWh in the middle of the night (low power usage) or middle of the day (lots of solar), and highs of $0.08/kWh in the late afternoons and early evenings (solar diminishing; lots of AC).

In the winter the mismatch between what solar can supply and when power is demanded is even more stark, perhaps a high of ~$0.20/kWh in the mornings and evenings when solar isn't producing. As people install more solar and heat pumps, this supply-demand delta will continue shifting towards these times when solar isn't producing: on a cold winter morning the sun isn't up yet, but the heat pumps are working very hard.

Which is a long way of saying that if I send kWh to the grid when it's convenient for me (lots of sun) and draw kWh from the grid when it's convenient for me (no sun) the kWh I send are significantly less valuable to others than the kWh I draw. Then add in the large cost of maintaining the grid, and it's really very strange that my electric bill treats them the same. More than strange: when I described this system to a UK friend who has thought a lot about power, they assessed net metering as "completely insane". In MA, ratepayers are spending somewhere in the $150M to $300M range annnually [3] subsidizing households with solar.

So how did we get here? Net metering started out as a very simple technical solution. In 1978, after the oil shocks, congress passed PURPA. It required utilities to compensate based on (what today would be) the market value of their production:

the cost to the electric utility of the electric energy which, but for the purchase from such cogenerator or small power producer, such utility would generate or purchase from another source.

Residential solar installations back then were rare and small, and this number was hard to calculate. Collecting the information you'd need to get to the actual number would have been very hard, while letting the meter run in reverse was very easy, so net metering came about through technological expedience. When solar was a tiny part of the overall generation mix the overall effect was tiny. [2]

Over time it became practical to use other metering systems, but solar advocates fought to keep net metering: it's unusually politically acceptable for the scale of the subsidy, and really gets solar installed. But at the cost of making power more expensive for everyone else. Many states have stopped allowing new net metering customers, but discontinuing net metering for existing installs is much more controversial: people bought expensive systems or signed long-term leases under the assumption that net metering would continue.

Technology has changed a lot in other ways since the 1970s, and a big one is that batteries are also far cheaper. You can charge at times of low demand, and discharge a few hours later when demand is higher. When you can't do net metering, residential rooftop solar is often still worth it as long as you also install batteries. Instead of using the grid as a giant battery, drawing and exporting kWh as needed, you do it with an actual battery. This doesn't fully solve the incentives problem, because the right to draw as many kWh as you want whenever you want it is underpriced, but it does help.

The other advantage of batteries, which is the big reason I'm interested in this, is a battery is the expensive part of making a system that produces power when the grid is down. Regular residential grid-tied solar is useless in power outages: it shuts down and produces nothing. If people install batteries, however, making the house operate as an "island" during a blackout is standard.

I think people in places where the grid has been reliable are massively underrating the benefit of having power during blackouts. Living in Somerville it's been decades since we had an outage long enough to even spoil food in fridges. [4] If the power grid maintained this level of reliability, backup power would resolve an inconvenience at best. Looking at other countries, however, grids have become unreliable through natural disasters, war, and state mismanagement. And looking forward, I'm especially concerned about how the recklessly rapid pace of AI development increases the risk of all kinds of instability. Electricity is so useful for so many things that I see a lot of value in a distributed and resilient power system that can continue to make even small amounts of power available in many places if the grid goes down.

My proposal is that we end net metering, and instead of counting exported energy 1:1 against later consumption it's compensated based on the utility's "avoided cost". This is a lot like what CA did with NEM 3.0, and they saw large increases in battery installations. Unlike CA, where they allowed in existing installs to continue to use net metering for up to 20y, I propose we end it for everyone but partially buy out the subsidy with a credit you can use towards island-capable battery systems.

How big that credit should be is not something I have strong feelings about, but I expect it would be very controversial because there are big winners and losers here depending on the shape of the policy. At one end of the spectrum you could size the payment to attempt to fully compensate owners for the net present value of their foregone subsidy [5]; at the other you give them a token amount that's just sufficient to get many of them to install a battery. My big question here is whether there's enough of a constituency for any point along the spectrum that this could actually become law. This is unfortunately not a free lunch: while the batteries do save money, they don't save enough to pay for themselves and someone, whether solar owners or general ratepayers, would need to pay the bill.

(While any subsidy would apply to us, since we have solar without batteries, I think the benefits of batteries here are large enough that we're planning to go ahead and install them regardless, so we wouldn't qualify for a subsidy.)


[1] Both of these numbers exclude state incentives for solar production, beyond net metering. At maybe $0.04/kWh these help much more for solar farms than for net metering installs, but they're not enough to appreciably change the net metering picture. All numbers for MA since I live here.

[2] Solar growing in the mix is a big part of why the marginal exported solar kWh today isn't that valuable. A while ago, in sunny places, people would use a lot of AC when the sun was shining, which meant solar production was reasonably well timed. But today there's so much solar going to the grid already that power when there's no sun is disproportionately valuable.

[3] Very roughly, MA has ~1.5 GW of residential solar, producing ~1.7 TWh/y. A little under half of residential production is typically self-consumed, so figure 0.9 TWh/y in exports. These are credited at retail (my bill is $0.32/kWh) but the value to the grid is more like $0.06/kWh on average (wholesale energy, plus a little for avoiding line losses and capacity increases). This comes to $234M, but with wide error bars.

[4] I don't remember this happening and tried to look it up, but didn't find much. Even the Northeast Blackout of 2003 probably wouldn't have qualified, since power in most places was restored in 2-6hr, plus it didn't affect this part of MA.

[5] A fully "make-whole" payment would need to be sized to the net present value of the delta between the value of the current system over the remainder of a 20y operation window, and the value they'd get from the battery (less its purchase price). Penciling this out, if someone is averaging 15 kWh/day at a marginal cost of $0.32/kWh and has 5kW of solar on their roof, and received permission to operate 5y ago, the net present value of their future net metering credits, less avoided cost compensation, over the 15y remainder of the 20y window, would be ~10k. A battery (let's say 13 kWh) would regain ~$6k of that from increased self-consumption and another ~$5k from battery-operation incentives (ConnectedSolutions in MA; assuming drops to ~0 after 5y). This means you about break even (~+$1k) until you get into the cost of the battery and installation. Which is unfortunately a lot more than $1k; I see quotes for $16k, though this doesn't fully reflect how much improvements in battery tech should be bringing the price down.

Comment via: facebook, lesswrong, mastodon, bluesky

Allow Babywearing Carriers on Planes

2026-09-08 21:00:00

The FAA has one of my favorite examples of thoughtful rulemaking. They haven't banned flying with a baby on your lap, because the extra cost would mean many parents would drive instead. Since driving is far less safe than flying, a ban would lead to more deaths. I'd love to see more of this "all things considered" thinking around bans.

In fact, one specific place where I'd like to see this thinking applied is adjacent to this rule: babywearing carriers on planes. When our babies were little, carriers were massively helpful in flying. The baby likes it, and your arms are free. But for takeoff and landing, the FAA requires you to take your baby out of the carrier and hold them in your arms.

The rule is that if an under-two is going to ride on a lap, they must not "occupy or use any restraining device," (14 CFR 121.311.b.1) and FAA guidance to parents is clear: "Baby carriers ...are not allowed to be used during ground movement, take-off, or landing."

This goes back to 1995, and if you look at the notice of proposed rulemaking it says:

This notice proposes to withdraw FAA approval for the use of booster seats and vest- and harness-type child restraint systems in aircraft during takeoff, landing, and movement on the surface. ... The FAA believes that, during an aircraft crash, the banned devices may put children in a potentially worse situation than the allowable alternatives.

This is based on their 1994 study, where the FAA compared child restraint options. They looked at "booster seats, forward facing carriers, aft facing carriers, a harness device, a belly belt, and passenger seat lap belts." The "carriers" here are car seats; they didn't test babywearing carriers. I don't think that's a major flaw in the study, however, since I do expect babywearing carriers would have performed poorly. [1] The real problem is that they didn't test the most common alternative, holding the baby in your lap. So "may put children in a potentially worse situation than the allowable alternatives" seems clearly wrong to me.

In the 1996 final rule [2] they discuss why they're making a different decision than the UK (CAA) and Europe (JAA, predates EASA), and they avoid the obvious comparison. They say belly belts can be dangerous, acknowledge that lap-holding has risks, predict parents will buy a second seat, and then also say they don't want to require a second seat because people will drive. This isn't completely nuts, since it's possible that allowing belly belts would have caused some parents to choose them over a second seat. It's pretty unlikely for the effects to balance out in just this way, however, and I don't see any indication that they tried to do this balancing.

I expect that, if fairly evaluated, modern babywearing carriers would prove much safer than arms during turbulence, crashes, and evacuations. And I think this is likely enough that if we're not going to do these tests we should default to not banning carriers. In 2024 there was a bill proposing to do something like this (Rep Bill Posey's HR 8972), but it went to the aviation subcommittee and died without a vote.

I don't think the harm of the current rule is very large in the scheme of things: flying is very safe, even for unrestrained lap infants. Most of the harm is probably the inconvenience of waking happily sleeping babies. Still, it bugs me as a clear example of incoherent rulemaking. The FAA sensibly considered substitution behavior in deciding not to ban lap infants, but failed to balance it here.


[1] The closest thing to a carrier they tested was a 'belly belt': "This belt is designed to be buckled around the child's abdomen and is secured to an adult's abdomen with the adult's safety belt by routing the safety belt through a small loop of webbing sewn on the belly belt." They did not perform well: "In the test, these systems allowed the anthropomorphic test dummy to make severe contact with the back of the seat in the row in front of the test dummy. The child also may be crushed by the forward bending motion of the adult to whom the child is attached."

[2] Here's the section, if the PDF is hard to load or read:

CAA and JAA state that they permit the belly belt on the grounds that it provides a measure of protection to children and/or other passengers versus lap holding a child.

FAA Response: The FAA would like to emphasize that belly belts are not permitted under current regulations. Even if belly belts do provide some measure or protection, the CAMI study found that belly belts allowed the test dummy to make severe contact with the back of the seat in the row in front of the test dummy and that a child may be crushed by the forward bending motion of the adult to whom the child is attached. Consideration of revising this current prohibition is beyond the scope of the notice.

The JAA also stated that in a crash or severe air turbulence, parents are often unable to keep a lap-held child in their arms.

FAA Response: As discussed earlier, the FAA has determined that mandating child restraint devices could cause more deaths and injuries than it would prevent. However, the FAA does not encourage lap-holding children. The FAA expects, with its education campaign providing clear guidance on child restraint devices, parents will choose an approved device, rather than lap holding their children, in order to provide the safest traveling environment for their children. The two members of the APCS Working Group submitted identical letters that discussed the need to mandate restraints for children. In addition, they stated that the FAA's argument that the extra cost to families caused by mandating child restraint devices would force them to less safe road travel is invalid since the same cost situation arises when the child is 3 or 4 or 10 years old.

FAA Response: The APCS Working Group's argument is that the extra cost to families of mandating child restraint devices is no more of a deterrent to air travel than the price of a ticket for a child of any age. However, the FAA notes that this argument does not take into account that ordinarily there is no charge for a lap-held child, whereas certificate holders very often do charge if a seat is requested for this infant. Thus, many people would switch to less safe automobile travel as a result of mandating child restraint usage because unlike most rulemakings where the compliance costs are passed along to all travelers, mandatory use of child restraint would impose compliance costs only on families with infants.

Comment via: facebook, lesswrong, hacker news, mastodon, bluesky