Leila T Alexander, Janani Durairaj, Andriy Kryshtafovych, Luciano A Abriata, Yusupha Bayo, Gira Bhabha, Cécile Breyton, Simon G Caulton, James Chen, Séraphine Degroux, Damian C Ekiert, Benedikte S Erlandsen, Lydia Freddolino, Dominic Gilzer, Chris Greening, Jonathan M Grimes, Rhys Grinter, Manickam Gurusaran, Marcus D Hartmann, Charlie J Hitchman, Jeremy R Keown, Ashleigh Kropp, Petri Kursula, Andrew L Lovering, Bruno Lemaitre, Andrea Lia, Shiheng Liu, Maria Logotheti, Shuze Lu, Sigurbjörn Markússon, Mitchell D Miller, George Minasov, Hartmut H Niemann, Felipe Opazo, George N Phillips Jr, Owen R Davies, Samuel Rommelaere, Monica Rosas‐Lemus, Pietro Roversi, Karla Satchell, Nathan Smith, Mark A Wilson, Kuan‐Lin Wu, Xian Xia, Han Xiao, Wenhua Zhang, Z Hong Zhou, Krzysztof Fidelis, Maya Topf, John Moult, Torsten Schwede
Proteins: Structure, Function, and Bioinformatics · 91 · 1571-1599
Abstract
Abstract We present an in‐depth analysis of selected CASP15 targets, focusing on their biological and functional significance. The authors of the structures identify and discuss key protein features and evaluate how effectively these aspects were captured in the submitted predictions. While the overall ability to predict three‐dimensional protein structures continues to impress, reproducing uncommon features not previously observed in experimental structures is still a challenge. Furthermore, instances with conformational flexibility and large multimeric complexes highlight the need for novel scoring strategies to better emphasize biologically relevant structural regions. Looking ahead, closer integration of computational and experimental techniques will play a key role in determining the next challenges to be unraveled in the field of structural molecular biology.
Andriy Kryshtafovych, Maciej Antczak, Marta Szachniuk, Tomasz Zok, Rachael C Kretsch, Ramya Rangan, Phillip Pham, Rhiju Das, Xavier Robin, Gabriel Studer, Janani Durairaj, Jerome Eberhardt, Aaron Sweeney, Maya Topf, Torsten Schwede, Krzysztof Fidelis, John Moult
Proteins: Structure, Function, and Bioinformatics · 91 · 1550-1557
Abstract
Abstract Prediction categories in the Critical Assessment of Structure Prediction (CASP) experiments change with the need to address specific problems in structure modeling. In CASP15, four new prediction categories were introduced: RNA structure, ligand‐protein complexes, accuracy of oligomeric structures and their interfaces, and ensembles of alternative conformations. This paper lists technical specifications for these categories and describes their integration in the CASP data management system.
Xavier Robin, Gabriel Studer, Janani Durairaj, Jerome Eberhardt, Torsten Schwede, W Patrick Walters
Proteins: Structure, Function, and Bioinformatics · 91 · 1811-1821
Abstract
Abstract CASP15 introduced a new category, ligand prediction, where participants were provided with a protein or nucleic acid sequence, SMILES line notation, and stoichiometry for ligands and tasked with generating computational models for the three‐dimensional structure of the corresponding protein–ligand complex. These models were subsequently compared with experimental structures determined by x‐ray crystallography or cryoEM. To assess these predictions, two novel scores were developed. The Binding‐Site Superposed, Symmetry‐Corrected Pose Root Mean Square Deviation (BiSyRMSD) evaluated the absolute deviations of the models from the experimental structures. At the same time, the Local Distance Difference Test for Protein–Ligand Interactions (lDDT‐PLI) assessed the ability of models to reproduce the protein–ligand interactions in the experimental structures. The ligands evaluated in this challenge range from single‐atom ions to large flexible organic molecules. More than 1800 submissions were evaluated for their ability to predict 23 different protein–ligand complexes. Overall, the best models could faithfully reproduce the geometries of more than half of the prediction targets. The ligands' size and flexibility were the primary factors influencing the predictions' quality. Small ions and organic molecules with limited flexibility were predicted with high fidelity, while reproducing the binding poses of larger, flexible ligands proved more challenging.
Michael W Mullowney, Katherine R Duncan, Somayah S Elsayed, Neha Garg, Justin JJ van der Hooft, Nathaniel I Martin, David Meijer, Barbara R Terlouw, Friederike Biermann, Kai Blin, Janani Durairaj, Marina Gorostiola González, Eric JN Helfrich, Florian Huber, Stefan Leopold-Messer, Kohulan Rajan, Tristan de Rond, Jeffrey A van Santen, Maria Sorokina, Marcy J Balunas, Mehdi A Beniddir, Doris A van Bergeijk, Laura M Carroll, Chase M Clark, Djork-Arné Clevert, Chris A Dejong, Chao Du, Scarlet Ferrinho, Francesca Grisoni, Albert Hofstetter, Willem Jespers, Olga V Kalinina, Satria A Kautsar, Hyunwoo Kim, Tiago F Leao, Joleen Masschelein, Evan R Rees, Raphael Reher, Daniel Reker, Philippe Schwaller, Marwin Segler, Michael A Skinnider, Allison S Walker, Egon L Willighagen, Barbara Zdrazil, Nadine Ziemert, Rebecca JM Goss, Pierre Guyomard, Andrea Volkamer, William H Gerwick, Hyun Uk Kim, Rolf Müller, Gilles P van Wezel, Gerard JP van Westen, Anna KH Hirsch, Roger G Linington, Serina L Robinson, Marnix H Medema
Nature Reviews Drug Discovery · 22 · 895-916
Abstract
Developments in computational omics technologies have provided new means to access the hidden diversity of natural products, unearthing new potential for drug discovery. In parallel, artificial intelligence approaches such as machine learning have led to exciting developments in the computational drug design field, facilitating biological activity prediction and de novo drug design for molecular targets of interest. Here, we describe current and future synergies between these developments to effectively identify drug candidates from the plethora of molecules produced by nature. We also discuss how to address key challenges in realizing the potential of these synergies, such as the need for high-quality datasets to train deep learning algorithms and appropriate strategies for algorithm validation.
Michèle Leemann, Ander Sagasta, Jerome Eberhardt, Torsten Schwede, Xavier Robin, Janani Durairaj
Proteins: Structure, Function, and Bioinformatics
Abstract
Abstract The prediction of protein‐ligand complexes (PLC), using both experimental and predicted structures, is an active and important area of research, underscored by the inclusion of the Protein‐Ligand Interaction category in the latest round of the Critical Assessment of Protein Structure Prediction experiment CASP15. The prediction task in CASP15 consisted of predicting both the three‐dimensional structure of the receptor protein as well as the position and conformation of the ligand. This paper addresses the challenges and proposed solutions for devising automated benchmarking techniques for PLC prediction. The reliability of experimentally solved PLC as ground truth reference structures is assessed using various validation criteria. Similarity of PLC to previously released complexes are employed to judge PLC diversity and the difficulty of a PLC as a prediction target. We show that the commonly used PDBBind time‐split test‐set is inappropriate for comprehensive PLC evaluation, with state‐of‐the‐art tools showing conflicting results on a more representative and high quality dataset constructed for benchmarking purposes. We also show that redocking on crystal structures is a much simpler task than docking into predicted protein models, demonstrated by the two PLC‐prediction‐specific scoring metrics created. Finally, we introduce a fully automated pipeline that predicts PLC and evaluates the accuracy of the protein structure, ligand pose, and protein–ligand interactions.
Janani Durairaj, Andrew M Waterhouse, Toomas Mets, Tetiana Brodiazhenko, Minhal Abdullah, Gabriel Studer, Gerardo Tauriello, Mehmet Akdel, Antonina Andreeva, Alex Bateman, Tanel Tenson, Vasili Hauryliuk, Torsten Schwede, Joana Pereira
Nature · 622 · 646-653
Abstract
We are now entering a new era in protein sequence and structure annotation, with hundreds of millions of predicted protein structures made available through the AlphaFold database 1 . These models cover nearly all proteins that are known, including those challenging to annotate for function or putative biological role using standard homology-based approaches. In this study, we examine the extent to which the AlphaFold database has structurally illuminated this ‘dark matter’ of the natural protein universe at high predicted accuracy. We further describe the protein diversity that these models cover as an annotated interactive sequence similarity network, accessible at https://uniprot3d.org/atlas/AFDB90v4 . By searching for novelties from sequence, structure and semantic perspectives, we uncovered the β-flower fold, added several protein families to Pfam database 2 and experimentally demonstrated that one of these belongs to a new superfamily of translation-targeting toxin–antitoxin systems, TumE–TumA. This work underscores the value of large-scale efforts in identifying, annotating and prioritizing new protein families. By leveraging the recent deep learning revolution in protein bioinformatics, we can now shed light into uncharted areas of the protein universe at an unprecedented scale, paving the way to innovations in life sciences and biotechnology.
Gianluca Palermo, Gianmarco Accordi, Davide Gadioli, Emanuele Vitali, Cristina Silvano, Bruno Guindani, Danilo Ardagna, Andrea Rosario Beccari, Domenico Bonanni, Carmine Talarico, F Lughini, Jan Martinovic, Paulo Silva, Ada Bohm, Jakub Beranek, Jan Krenek, Branislav Jansik, Biagio Cosenza, Luigi Crisci, Peter Thoman, Philip Salzmann, Thomas Fahringer, Leila T Alexander, Gerardo Tauriello, Torsten Schwede, Janani Durairaj, Andrew Emerson, Federico Ficarelli, Sebastian Wingbermühle, E Lindhal, Daniele Gregori, Emanuele Sana, Silvano Coletti, Philipp Gschwandtner
Proceedings of the 20th ACM International Conference on Computing Frontiers · 272-278
Abstract
Today digital revolution is having a dramatic impact on the pharmaceutical industry and the entire healthcare system. The implementation of machine learning, extreme-scale computer simulations, and big data analytics in the drug design and development process offers an excellent opportunity to lower the risk of investment and reduce the time to the patient.
Within the LIGATE project, we aim to integrate, extend, and co-design best-in-class European components to design Computer-Aided Drug Design (CADD) solutions exploiting today's high-end supercomputers and tomorrow's Exascale resources, fostering European competitiveness in the field.
The proposed LIGATE solution is a fully integrated workflow that enables to deliver the result of a virtual screening campaign for drug discovery with the highest speed along with the highest accuracy. The full automation of the solution and the possibility to run it on multiple supercomputing centers at once permit to run an extreme scale in silico drug discovery campaign in few days to respond promptly for example to a worldwide pandemic crisis.
Janani Durairaj, Leila Tamara Alexander, Gabriel Studer, Gerardo Tauriello, Ingrid Guarnetti Prandi, Rosalba Lepore, Giovanni Chillemi, Torsten Schwede
Exscalate4CoV: High-Performance Computing for COVID Drug Discovery · 41-50
Abstract
The large amount of genetic, phenotypic, and structural data from diverse conditions and environments offers opportunities for new groundbreaking research. Today, the major scientific task is to interpret the vast number of genetic variants within these data. As described in this chapter, identifying relevant variants and connecting them with the associated protein structural and environmental information is a powerful approach to biological discoveries. The unified view of the data brings us a step closer to understanding genetic variation, which is also fundamental for achieving the goals of personalized medicine and the planet’s environment.