How Open Science Can Help Researchers Prepare for the Next Pandemic
NVIDIA Blog Anthony Costa
NVIDIA and partners released 3D protein structures for more than 2,800 viruses. Scientists get a bigger head start before the next outbreak does.
Based on reporting by NVIDIA Blog, Anthony Costa — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
When COVID-19 hit, researchers had an advantage that may not exist next time: years of coronavirus work had already revealed enough about the virus’s key proteins to speed vaccine design. NVIDIA and a group of partners are trying to manufacture that kind of head start for future outbreaks by releasing predicted 3D structures for protein complexes from more than 2,800 viruses, all available through the AlphaFold Database.
The dataset was generated with AlphaFold2, Google DeepMind’s protein-folding model, and optimized with NVIDIA BioNeMo Inference Runtime so the team could push inference across thousands of viral proteomes. In plain terms, that means they worked through the protein interactions encoded by each virus at a scale that would be brutal to handle with older methods. NVIDIA is also open-sourcing the BioNeMo Structure Prediction Pipeline, the GPU-accelerated workflow used to produce the results, so other researchers can move from protein sequence to predicted structure on their own targets.
That matters because most proteins do not act alone. They form complexes, and those interactions are often the thing a vaccine or drug has to hit to shut a virus down. The COVID-19 spike protein was a big part of the vaccine story; for thousands of other viruses, no comparable structural map has existed. This release starts to fill that gap, including work on viruses that infect humans, from common-cold viruses to Mpox.
Not every structure is familiar. About 30% of the interactions added are new to science, with shapes that had never been documented in the Protein Data Bank, the main repository for experimentally determined protein structures. The older way to get this information — crystallize a protein, blast it with X-rays, wait, repeat — can take years and cost thousands of dollars per structure. AlphaFold2 can predict a structure in minutes, and scientists can still check the high-confidence ones experimentally later.
The project spans the Coalition for Epidemic Preparedness Innovations, EMBL-EBI, Google DeepMind, NVIDIA, Seoul National University, Sungkyunkwan University, the Swiss Institute of Bioinformatics and the University of Glasgow. The AlphaFold Database now holds more than 260 million protein and protein-complex predictions, and this release lands alongside a United Nations General Assembly meeting in New York City focused on pandemic prevention, preparedness and response.
My take — AI-written commentary, not fact-checked reporting
Open data wins here for the boring reason that it works. If low-resource labs are the ones staring at the next outbreak first, hiding the structural map behind paywalls is just self-sabotage with nicer branding. The real scandal in science is still how much basic knowledge arrives late, expensive, and in fragments.
Read more about this at: NVIDIA Blog