“AI is extraordinarily capable at finding patterns in biological data, but it can only work with the observations scientists have actually collected,” Jacob Trefethen, head of Life Sciences and Curing Diseases at the OpenAI Foundation, said.
Trefethen said the Foundation expects the value of scientific data to increase as AI becomes more capable and more widely used in life sciences. Public Data for Health is therefore targeting not only existing information that needs to be organized or preserved, but also entirely new biological measurements.
Accessibility will be a major requirement, although that will not mean every dataset can simply be published online. Information involving patients will require privacy protections and consent safeguards, while other types of biological data could potentially be released much more openly.
“The shared thread, I think, is accessibility,” Trefethen said. “Making sure that the data that we fund is as accessible as possible to as many researchers as possible is going to be a real big focus.”
The Foundation said institutions handling patient information will follow established practices for protecting sensitive data. As the funder, it will also have a governance role in ensuring those standards are followed.
One of the first projects will support the University of North Carolina in work related to generative immunotherapy and personalized cancer vaccines. These vaccines can draw on characteristics of a patient's tumor to identify abnormal proteins that the immune system might recognize and target.
Machine learning already contributes to predicting which tumor characteristics could provoke an immune response. The UNC work will go further by connecting tumor sequencing with measurements of proteins displayed by cancer cells and information about how a patient's immune system reacts.
“By creating that linked multimodal data, we think that we’ll be able to move forward the whole field,” Trefethen said.
That project reflects what the Foundation describes as connected data: combining different biological measurements to give researchers a fuller view of a problem. The program is also interested in scarce datasets that are difficult to produce or could disappear, as well as direct measurements closely tied to biological and clinical outcomes.
Another grant will go to OpenADMET, which includes researchers at the University of California, San Francisco. The project will generate datasets related to absorption, distribution, metabolism, excretion and toxicity, characteristics that play a role in whether prospective drugs succeed or fail.
OpenADMET also plans to use the data in public competitions that allow researchers to test predictive approaches against common benchmarks. Trefethen said a previous competition from the group attracted more than 300 participants.
“They’re going to collect really high-quality data related to ADMET properties, and they’re going to run these blind challenges, and then anyone can enter,” he said.
CTD Commons, another initial recipient, will work on preserving and organizing regulatory and drug-development knowledge. That gives Public Data for Health a broader scope than generating new laboratory measurements alone: the program will also back efforts to keep useful scientific information from remaining fragmented or disappearing.
The Foundation has intentionally left its definition of valuable data relatively broad. Scarcity is one factor. A rigorous dataset covering something that is poorly measured today may offer more value than adding observations to an area where researchers already have extensive information.
The organization is also interested in tools that could create new kinds of scientific observations in the first place. That could involve instruments, imaging systems or experimental methods capable of measuring aspects of biology that researchers cannot adequately capture today.
“Fundamentally, progress requires observation,” Trefethen said. “We do get stuck at the equipment that we’re using.”
Public Data for Health is part of a much larger expansion of the OpenAI Foundation's grantmaking. Trefethen said the Foundation plans to award more than $1 billion across its programs during its first year, with that level expected to grow. It has also committed $25 billion toward curing diseases and strengthening resilience around AI.
The health initiative follows the Foundation's AI for Alzheimer's program, which applies advanced AI to research around the disease.
Public Data for Health grants will generally run for about two years, with some extending to three. Projects will undergo formal reviews quarterly or every six months depending on their size, while the Foundation also plans ongoing contact with recipients.
The money comes from an organization with an unusual relationship to OpenAI's commercial business. Following OpenAI's restructuring, the nonprofit OpenAI Foundation controls OpenAI Group PBC, appoints its directors and owns a substantial stake in the company.
Trefethen said the Foundation's grants are managed separately from OpenAI's commercial operations. Recipients are also free to use competing AI systems rather than OpenAI products.
“This is an independent philanthropic endeavor, and our grantees can use any AI models they want—not just OpenAI models,” he said.
That means Public Data for Health is not being presented as a program for creating private training material exclusively for OpenAI. Its stated goal is to finance scientific resources that researchers at other institutions can access and use regardless of which AI models they choose.
With more than $125 million committed to the initial program, the Foundation is putting significant resources behind the premise that better AI alone will not be enough to accelerate medical research. Its bet is that progress will also depend on producing better observations, connecting information that currently sits apart and making high-quality scientific data available to the researchers and models capable of learning from it.
This analysis is based on reporting from Inside Precision Medicine.
Image courtesy of OpenAI Foundation.
This article was generated with AI assistance and reviewed for accuracy and quality.