A group of US government agencies, a research philanthropy and three big tech AI groups have committed $1.8 billion to one unglamorous problem: there is not yet enough of the right biological data to train AI models that can predict how living cells behave.
The announcement, made on 7 October 2026 by Biohub with the US Department of Energy (DOE), the National Institutes of Health (NIH) and new partners, is confirmed in Biohub’s press release [1]. Google DeepMind, Isomorphic Labs and Meta are together putting in $300 million [1].
What has been announced?
Biohub calls it a major expansion of its Virtual Biology Initiative, first announced in April 2026. The aim is to generate open, “AI-ready” biological data that researchers anywhere can use to build predictive models of biology [1].
The pieces, as Biohub sets them out [1]:
- Biohub: a founding commitment of $500 million. Of that, $400 million supports new measurement technology, including cryo-electron tomography (near-atomic detail inside cells), microscopy that can image millions to billions of cells in living tissue, and tools to build and perturb biology. A further $100 million funds research outside Biohub.
- US Department of Energy: more than $500 million over five years for lab measurement, modelling and computation, through its cross-agency Genesis Mission. DOE says this draws on exascale supercomputers, X-ray and neutron facilities, cryo-electron microscopy and autonomous laboratories in the National Laboratory system.
- National Institutes of Health: coordinating datasets, repositories and knowledge bases built through more than $500 million in prior federal investment. Biohub will work with NIH to standardise them for AI training.
- Google DeepMind, Isomorphic Labs and Meta: $300 million collectively, for technologies and multi-modal datasets.
Biohub describes the combined total as $1.8 billion “in funding, data, computation, and new measurement technology”, and calls it the largest coordinated commitment to generating AI-ready biological data to date [1].
Is that all new money?
Not entirely, and Biohub’s own wording makes that clear. The $1.8 billion explicitly includes data and computation as well as cash [1]. The NIH portion, in particular, is described as contributing existing datasets developed through more than $500 million of prior federal investment, not as a new $500 million grant [1].
That does not make the commitment small. It does mean the headline is best read as the total value being pooled, not a single new fund of $1.8 billion.
What is a “virtual cell”, and why does it need this?
The long-term goal is a computer model accurate enough that scientists could run some experiments digitally: change something about a cell, and predict what happens, before going to the lab.
“An accurate predictive model of biology could dramatically accelerate scientific discovery by enabling scientists to perform experiments digitally,” said Biohub Head of Science Alex Rives. He called the creation of a virtual cell “one of the most important challenges for the next era of science” [1].
AI models learn from examples. Language models had the internet. Biology has a great deal of data, but much of it was collected in different ways, in different formats, for different questions. What is missing, the partners argue, is large, consistent data on how many types of cells respond to many kinds of interventions. Biohub says the initiative will expand “cell response data to interventions across far more cell types and conditions than have yet been studied” [1].
NIH’s Nicole Kleinstreuer put the aim as models that can “predict how any cell responds to an intervention” [1]. Google DeepMind’s Pushmeet Kohli said the challenge cannot be solved “without open, experimental biological data at an unprecedented scale” [1].
Who else is involved?
Biohub lists a set of research institutions committed to working together on the initiative: the Allen Institute, the Broad Institute, the Gladstone Institutes, the Human Cell Atlas, the Human Protein Atlas and the UK’s Wellcome Sanger Institute [1]. NVIDIA will support it with computing infrastructure, software and expertise, and Renaissance Philanthropy is helping expand funding for data generation [1].
Biohub says it will build the layer that makes these datasets work together, “shared standards, common identifiers, and a single point of access”, drawing on projects it already runs, such as CELLxGENE and the CryoET Data Portal [1].
Why does it matter beyond biology labs?
Two reasons.
First, it is a clear statement that in AI for science, data is now the bottleneck, not just model size or computing power. DOE’s Darío Gil framed it as combining DOE’s computing and measurement assets with Biohub’s AI models and data capabilities to set “a new standard for open science” [1].
Second, the result is meant to be open. Biohub says the output will be “an open resource for the research community” [1]. With Google DeepMind, Isomorphic Labs and Meta all contributing, the same open data could feed both academic work and commercial drug-discovery models.
What this does not prove
- That a working virtual cell is close. The announcement funds data generation and tools. It does not claim a predictive model of a whole cell exists.
- That $1.8 billion is new cash. It includes data and computation, and the NIH share is existing datasets [1].
- How the $300 million from tech companies is split. Biohub gives only the collective figure for Google DeepMind, Isomorphic Labs and Meta [1].
- Timelines for data releases. Apart from DOE’s five-year horizon, the release does not set dates for when new datasets will become available [1].
- Terms of access. “Open” is the stated goal; detailed licensing and access rules were not part of the announcement.
The Bottom Line
The Virtual Biology Initiative pools $1.8 billion in funding, data, computing and measurement technology from Biohub, the US Department of Energy, the NIH and, collectively, $300 million from Google DeepMind, Isomorphic Labs and Meta [1]. Its goal is the open, consistent biological data that AI models need before anyone can credibly build a “virtual cell”. Part of the headline is existing data rather than new money, and the hard science is still ahead. But as a bet on data as the real constraint in AI for biology, it is large and specific.
Sources
- Biohub, “International, cross-sector collaboration commits nearly $2 billion to build foundational data for AI models to predict and treat disease” (press release; page title “AI-ready biological data: $1.8 billion global commitment”), 7 October 2026. https://biohub.org/news/virtual-biology-initiative-expansion/

