ReutersMeta, Google DeepMind and drug discovery startup Isomorphic Labs are jointly investing $300 million. The Department of Energy will invest more than $500 million over five years in laboratory measurement, modeling and computation. The National Institutes of Health will coordinate datasets and repositories built with more than $500 million in earlier federal funding, which Biohub will standardize for AI training.
The commitments follow the $500 million that Biohub - a philanthropic venture of Meta CEO Mark Zuckerberg and his wife, Dr. Priscilla Chan - put into the project in April.
The investments fund the Virtual Biology Initiative, which aims to measure how cells respond to changes across far more conditions than scientists have so far studied, and use that data to build predictive models that could compress drug development timelines that currently take years.
"Biology has been just sort of a clever discovery-based science until this point," Chan said in an interview. "We have always held this as a community asset, not just for one group, so that it can build upon itself over time."
The datasets will eventually be released publicly, but companies that fund them get a head start, according to Biohub's head of science Alex Rives.
Discover the stories of your interest
"With commercial funders we have embargo periods where there's a period of time where the groups can work on the data, and then it becomes available as a public scientific resource," Rives said.
The arrangement is how Biohub is drawing private money into a project it describes as open science. Rives said the government-funded work running in parallel will carry no such restrictions, and that Biohub plans to approach pharmaceutical companies and philanthropies next.
Current cell datasets run to hundreds of millions of cells, Rives said, while an accurate predictive model will require billions and eventually trillions. Biohub's goal is to close that gap.
"We need to capture the language of biology, we need to capture the language of the cell. And that doesn't exist today," he said.
The data will come from techniques including spatial transcriptomics, which maps molecular activity inside intact tissue, and screens that record how cells respond to changes in environment. Much of it has never been generated in a coordinated way.
Rives said the work would normally take decades and that the partners aim to compress it into five years, with a first dataset ready in about a year. He expects accurate predictive models within five years.
Other AI labs are pursuing similar initiatives. Anthropic doubled down on its biology efforts with a wet lab, while the OpenAI Foundation started a more than $125 million grant program to fund biological and medical datasets for AI research.



