Researchers and drug developers can now study viruses posing the greatest public-health risks at expanded scale and speed. Models of how virus proteins interact are openly available for 2,800 human-infecting viruses and their relatives, up from just a handful, along with the first-ever predicted 3D structures of nearly 7,800 viral proteins. SIB co-selected the viruses analysed by the AlphaFold AI, supplied unique protein data for 699 viruses, and quality-controlled predicted interactions for these proteins. The work is an international collaboration including EMBL-EBI, NVIDIA, Google DeepMind, Seoul National University, and the University of Glasgow. It extends SIB’s contributions to the Nobel-winning AlphaFold and to tackling infectious diseases.
AlphaFold: a Nobel-winning AI built using SIB resources
AlphaFold made it possible to quickly compute a protein’s 3D structure from its amino acid sequence, instead of through lengthy, expensive experimental procedures. Later models extended this to predicting interactions between proteins and other molecules.
AlphaFold’s development – recognized in the 2024 Nobel Prize in Chemistry – relied on three open resources and initiatives developed and co-developed by SIB scientists. UniProt provided high-quality, open training data and knowledge on millions of proteins, while the CASP competition and CAMEO resource confirmed the AI’s impressive accuracy and applicability to all human proteins, respectively.
Closing a data gap for outbreak readiness
The development of virus diagnostics, vaccines and treatments depends on knowing the 3D shape of viral proteins, and how these proteins interact with one another. The unprecedented speed of developing vaccines against SARS-CoV-2, for example, was only possible thanks to decades of research on similar viruses prior to the pandemic.
AlphaFold: a Nobel-winning AI built using SIB resources
AlphaFold made it possible to quickly compute a protein’s 3D structure from its amino acid sequence, instead of through lengthy, expensive experimental procedures. Later models extended this to predicting interactions between proteins and other molecules.
AlphaFold’s development – recognized in the 2024 Nobel Prize in Chemistry – relied on three open resources and initiatives developed and co-developed by SIB scientists. UniProt provided high-quality, open training data and knowledge on millions of proteins, while the CASP competition and CAMEO resource confirmed the AI’s impressive accuracy and applicability to all human proteins, respectively.
Such knowledge is missing for 80% of viruses. This is due to the cost and time required for experimental studies, as well as some viruses being impossible to grow in the lab.
The AlphaFold AI model (see box) overcomes this limitation by accurately predicting 3D protein structures from amino acid sequences. An international, public-private collaboration has now used AlphaFold to predict interactions between virus protein-protein pairs in 2,800 viruses, all from families known to infect humans.
The collaboration – involving the European Bioinformatics Institute of the European Molecular Biology Laboratory (EMBL-EBI), NVIDIA, Google DeepMind, Seoul National University, the University of Glasgow, SIB, the Coalition for Epidemic Preparedness Innovations (CEPI) and Sungkyunkwan University – brought together expertise on viral biology, structural biology, biocuration, data management and analysis, and AI, as well as the necessary processing power for predictions at this scale.
SIB’s contributions comprised:
- co-selecting the set of reference viruses analysed, with the University of Glasgow;
- providing unique data on polyproteins – complex precursor proteins produced by 47% of human-infecting viruses, including Zika, dengue and poliovirus;
- validating the plausibility of AlphaFold’s structural predictions for these proteins, drawing on the comprehensive, curated knowledge on viral biology accumulated in SIB’s ViralZone resource.
This allowed AlphaFold to model, for the first time, the 3D structures and likely interactions of nearly 7,800 viral proteins derived from polyproteins (see more below).
All data were released openly today in the AlphaFold Database, coinciding with a United Nations General Assembly meeting on pandemic prevention, preparedness and response this week in New York City. Early analysis already points to previously unknown types of viral protein interactions, with implications for therapeutic research and for using viral proteins as biotechnology tools.
The work furthers SIB’s mission to push the boundaries of data science through in-depth knowledge of biological data, to provide researchers and clinicians with outstanding resources, and to generate knowledge to enable innovation for a better future.
Generating foundational data on viruses relevant to human health
Viruses are exceptionally diverse, which makes them challenging to study. Polyproteins add further complexity. These long amino-acid chains are cleaved into different active proteins – whose precise sequences are extremely difficult to determine. This is because different viruses have evolved completely different mechanisms for cleaving their polyproteins, meaning there is no common genetic clue to indicate where these cuts occur.
SIB scientists were the first to map such cleavage sites in viral polyproteins. Likely sites were identified in a semi-automated process that combined computational tools, extensive literature reviews, and deep expertise in viral biology. The accuracy of the predicted sites was then confirmed by structural modelling, which showed whether the resulting cleaved proteins were biologically plausible.
The work, concluded in May this year, provided the first reliable protein sequence data for 7,783 active proteins derived from polyproteins in 699 viruses, all from families known to infect humans. It formed a project of the Pathogen Data Network co-coordinated by SIB and supported by the US National Institutes of Health.
Making new virus knowledge openly available to researchers and AI
The complexity of viral polyproteins meant this newly generated information could not be added in its entirety to open biodata resources. The SIB scientists therefore created a new annotation standard for polyproteins, allowing the incorporation of AI-ready cleavage data in SIB’s ViralZone resource. These curated open data will accelerate research on viruses with public-health relevance. They also form a reliable training dataset that AI models could learn from, to predict cleavage sites in further viruses.
The team additionally plans to work with other pathogen-specific and generalist open biodata resources, including UniProt and NCBI RefSeq, to similarly include this information.
SIB data and expertise behind trustworthy AI for the life sciences
Beyond the characterization of viral proteins in this work, SIB’s provision of high-quality data and bioinformatics expertise is contributing to reliable AI systems across a range of health and biodiversity applications – from identifying cancer mechanisms and guiding patient treatments to informing actions for environmental protection.
Reference(s)
Image: Protein structure of two interacting proteins derived from the Ross River virus polyprotein. Credit: AlphaFold Database