
Anthony Costa
· 4 min read
How Open Science Can Help Researchers Prepare for the Next Pandemic
When COVID-19 emerged, scientists had a crucial advantage: Decades of prior research on coronaviruses meant they understood the virus’ key proteins well enough to design vaccines in record time. The next pandemic may not offer the same head start.
To help improve the odds, NVIDIA has joined a coalition of global research organizations, including Google DeepMind and the European Molecular Biology Laboratory’s European Bioinformatics Institute (EMBL-EBI), to release predicted 3D structures for the protein complexes of more than 2,800 viruses — openly available to any scientist, anywhere, through the AlphaFold Database.
The structures in the newly released dataset were inferred using AlphaFold2 — Google DeepMind’s AI model for predicting how proteins fold into 3D shapes — with optimization from NVIDIA BioNeMo Inference Runtime. This allowed the team to scale inference to thousands of viral proteomes, predicting the complexes, or groups of interacting proteins, encoded within each virus.
“Our ambition with the AlphaFold Database has always been to democratize access to foundational biology at scale,” said Risha Patel, life sciences partnerships manager at Google DeepMind. “This collaboration to bring thousands of viral complexes into the database will equip scientists around the world with insights they need to help prepare for future outbreaks.”
NVIDIA is also openly releasing the BioNeMo Structure Prediction Pipeline, the GPU-accelerated workflow used to generate the dataset, so researchers can go from protein sequence to predicted 3D structure for their own targets.
Preparation for the next pandemic must begin now. An analysis by the Center for Global Development estimates a roughly 50% chance of the world facing a pandemic as severe as COVID-19 by 2050.
“When the next pandemic happens, there may be something that comes out of the blue, and we’ll be lacking the knowledge we had for COVID,” said Joe Grove, professor of molecular virology at the Medical Research Council-University of Glasgow Centre for Virus Research and a collaborator on the project. “What we’re trying to do is stockpile some of that knowledge ahead of time.”
About 30% of the protein interactions being added to the database are completely new to science, showing interaction shapes that have never been documented in the Protein Data Bank, the main repository of experimentally determined protein structures. This translates to new insights for the biological community to explore and harness to generate new knowledge.
“This database is an engine for hypothesis generation,” said Chris Dallago, applied research science team lead in digital biology at NVIDIA. “We’re enabling biologists and the AI community to investigate protein interactions, not just as single molecules but as complexes, so the whole field can move forward.”
Predicting Complex Protein Structures
Most proteins don’t work alone — they come together in complexes of multiple molecules to perform sophisticated functions. Those structures are often what a vaccine or drug must target to disrupt viral function.
Understanding the 3D structure of the COVID-19 virus’ spike protein, for example, proved foundational to vaccine design. For thousands of other viruses, no such structural knowledge exists today. This dataset begins to fill that gap.
Traditional methods for determining protein structures — crystallizing proteins and shooting X-rays at them — can take years and cost thousands of dollars per structure. AlphaFold2, which was optimized with NVIDIA BioNeMo to efficiently run on NVIDIA GPUs, predicts a structure in minutes and can be run in bulk. Scientists can then verify high-confidence predictions through experimental methods.
For this project, the team systematically worked through the protein structures of viral families known to infect humans, from common-cold viruses to emerging threats like Mpox.
A Global Collaboration With Global Access
The collaboration spans the Coalition for Epidemic Preparedness Innovations, EMBL-EBI, Google DeepMind, NVIDIA, Seoul National University, Sungkyunkwan University, the Swiss Institute of Bioinformatics and the University of Glasgow.
The dataset release — coinciding with a United Nations General Assembly meeting convened by the World Economic Forum on pandemic prevention, preparedness and response taking place this week in New York City — contributes to the AlphaFold Database, which now holds more than 260 million protein and protein complex predictions covering nearly every cataloged protein known to science.
“Making this data open is critical for understanding viral diagnostics and developing treatments and vaccines,” said Jo McEntyre, interim director of EMBL-EBI. “The dataset also covers lesser-studied viruses and lowers the barriers for scientists in low-resource settings who are confronting outbreaks firsthand.”
Predictions in the open dataset are labeled by confidence. The structures show what viral complexes may look like and how individual proteins might interact within a viral proteome.
Overall, the new data represents a major contribution to the information available for scientists across digital biology and disease research.
“When I did my Ph.D., there were no structures for any of the proteins we were investigating. It was like working in the dark — we had to guess what was going on,” said Grove. “This dataset is a powerful tool for all the researchers doing their Ph.D.s now, giving them high-quality structural data that’s going to accelerate fundamental science.”
Explore the viral protein complex dataset on the AlphaFold Database Pandemic Preparedness Portal, predict structures for protein targets with the BioNeMo Structure Prediction Pipeline, and learn more about NVIDIA BioNeMo.
Original source
This story was published by NVIDIA Blog and written by Anthony Costa. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on blogs.nvidia.com


