Publications

Google Scholar ↗Updated 9 Sep 2026

Selected papers

Publications

26 entries

2026

Stoic : fast and accurate protein stoichiometry prediction

Daniil Litvinov, Lorenzo Pantolini, Peter Škrinjar, Gerardo Tauriello, Caitlyn L McCafferty, Benjamin D Engel, Torsten Schwede, Janani Durairaj

Bioinformatics · 42 · btag448

Abstract

Abstract Motivation Protein complexes are central to cellular function, but experimental determination of their structures remains challenging. Structure prediction methods require prior knowledge of stoichiometry—the number of copies of each protein entity within a complex. Current approaches rely on computationally expensive brute-force methods that run structure prediction on multiple stoichiometry combinations, often with limited accuracy. Results We introduce Stoic, a method that uses protein language model embeddings to predict protein complex stoichiometry. Our approach learns to identify interface residues that participate in protein-protein interactions, rather than relying on global sequence features. By integrating these interface-aware embeddings into a graph neural network, Stoic achieves fast and accurate stoichiometry prediction for both homomeric and heteromeric targets. Availability Source code for inference and training along with web versions are available in the repository at https://github.com/PickyBinders/stoic.

Designing CD8 T-cell mRNA vaccines targeting the viral large T antigen to prevent BK polyomavirus disease after kidney transplantation

Maud Wilhelm, Océane M Follonier, Anne Geng, Kristiana Mebelli, Janani Durairaj, Marion Wernli, Torsten Schwede, Hans H Hirsch

American Journal of Transplantation

Abstract

BK polyomavirus (BKPyV) continues to threaten kidney transplantation outcomes by directly or indirectly causing premature allograft failure. Insufficient BKPyV-specific immunity underlies the onset and duration of BKPyV-DNAemia and nephropathy. BKPyV-specific T cells have been correlated with protection and include cytotoxic CD8 T cells targeting immunodominant 9-mer epitopes encoded in the viral large tumor antigen (LTag). To develop vaccines increasing LTag-specific T cells, we reduced the 695 amino acid-long LTag to 143 residues or less, comprising 4 clusters of immunodominant 9-mers presented by at least 54 human leukocyte antigen alleles. Prediction algorithms assisted in improved global coverage, alanine-alanine-tyrosine-linker placement, proteasomal processing, and solubility. Using a human cell culture vaccination model, transfection of adherent monocytes with large T antigen-coding nucleotide sequence 1-messenger RNA (LTAG1-mRNA) and LTAG2-mRNA demonstrated the expected 13.9 and 15.2 kD immunogens, respectively. LTAG2-mRNA induced higher CD8 T cell responses than LTAG1-mRNA, with an interferon-γ-positive CD62L-CD45RA-effector memory phenotype and cytotoxicity for BKPyV-replicating primary human renal tubular epithelial cells. We tested 3 additional candidates (LTAG3-mRNA, LTAG4-mRNA, and LTAG5-mRNA) by changing alanine-alanine-tyrosine linkers and extending immunogenic regions. The encoded immunogens were expressed and increased upon proteasome inhibitor treatment. LTag5-expanded T cells showed BKPyV-specific interferon-γ expression and cytotoxicity. Our results demonstrated novel approaches to designing modular mRNA-based vaccine candidates for clinical development of nonsecreted antigens using BKPyV as a paradigm of nonenveloped DNA viruses.

Capabilities, specificity gaps and training-data dependence of AlphaFold3 across diverse application areas

Océane Follonier, Yan Liu, Pablo Campomanes, Luc Lafrenaye, Julien Racle, Daniel Alvarez, Julian van Gerwen, Ricardo Heinzmann, Jürgen Jänes, Eric Kummelstedt, Janani Durairaj, David Gfeller, Stefano Vanni, Pedro Beltrao

bioRxiv · Preprint

Abstract

Abstract Structure prediction models have moved from single proteins to assemblies that include diverse biomolecules and their modifications. AlphaFold3 (AF3) and related models extended structural modelling via an all-atom framework, opening many new potential applications in structural biology. We evaluate how well the new capabilities of AF3 translate into application tasks in diverse areas: prediction of ubiquitinated protein structures, T-cell receptor (TCR)-epitope recognition, antibody-antigen complexes, protein-RNA and protein-lipid interactions. We find that, while AF3 can perform well in favourable settings, this performance is uneven across applications. In RNA-target predictions, the model confidence fails to separate genuine from decoy interaction partners and in several tasks accuracy depends on the presence of related complexes in the training set. Taken together, our assessment is more cautious than for AF2, whose gains in modelling monomers and complexes were clear and broadly generalisable. AF3’s extension to new biomolecule types shows less consistent performance and generalisation. AF3 can be a powerful tool for hypothesis generation and prioritisation, but its predictions and use of confidence metrics will depend strongly on the specific application area and must be interpreted with respect to training-set overlap. We expect that the benchmarks provided here will serve for testing of future developments in the structure prediction field.

Evaluating generalization in protein–ligand cofolding methods

Peter Škrinjar, Jérôme Eberhardt, Gabriel Studer, Gerardo Tauriello, Torsten Schwede, Janani Durairaj

Nature Structural & Molecular Biology · 33 · 1-13

Abstract

Deep learning has driven major breakthroughs in protein structure prediction; however, one of the next critical steps forward is accurately predicting how proteins interact with small-molecule ligands, to enable real-world applications such as drug discovery. Recent cofolding methods aim to address this challenge, but evaluating their performance has been inconclusive because of the lack of relevant benchmarking datasets. Here we present a comprehensive evaluation of four leading all-atom cofolding methods using our newly introduced benchmark dataset, Runs N' Poses. Runs N' Poses comprises 2,600 high-resolution protein-ligand systems released after the training cutoff used by these methods. We demonstrate that current cofolding approaches largely memorize ligand poses from their training data, hindering their use for de novo drug design. With this assessment and benchmark dataset, we aim to accelerate progress in the field by allowing for a more realistic assessment of the current state-of-the-art deep learning methods for predicting protein-ligand interactions.

Reading TEA leaves for de novo protein design

Lorenzo Pantolini, Janani Durairaj

bioRxiv · Preprint

Abstract

De novo protein design expands the functional protein universe beyond natural evolution, offering vast therapeutic and industrial potential. Monte Carlo sampling in protein design is under-explored due to the typically long simulation times required or prohibitive time requirements of current structure prediction oracles. Here we make use of a 20-letter structure-inspired alphabet derived from protein language model embeddings to score random mutagenesis-based Metropolis sampling of amino acid sequences. This facilitates fast template-guided and unconditional design, generating sequences that satisfy in silico designability criteria without known homologues. Ultimately, this unlocks a new path to fast and de novo protein design.

A fully automated benchmarking suite to compare macromolecular complexes

Gabriel Studer, Xavier Robin, Stefan Bienert, Janani Durairaj, Peter Škrinjar, Gerardo Tauriello, Andrew Mark Waterhouse, Torsten Schwede

Nature methods · 23 · 387-394

Abstract

Abstract Protein structure prediction has a long history of benchmarking efforts such as critical assessment of structure prediction, continuous automated model evaluation and critical assessment of prediction of interactions. With the rise of artificial intelligence-based methods for prediction of macromolecular complexes, benchmarking with large datasets and robust, unsupervised scores to compare predictions against a reference has become essential. Also, the increasing size and complexity of experimentally determined reference structures by crystallography or cryogenic electron microscopy poses challenges for structure comparison methods. Here we review the current state of the art in scoring methodologies, identify existing limitations and present more suitable approaches for scoring of tertiary and quaternary structures, protein–protein interfaces and protein–ligand complexes. Our methods are designed to scale efficiently, enabling the assessment of large, complex systems. All developments are available in the structure benchmarking framework of OpenStructure. OpenStructure is open source software and available for free at https://openstructure.org/ .

Beyond Single Chains: Benchmarking Macromolecular Complex Prediction Methods With the Continuous Automated Model EvaluatiOn (CAMEO)

Xavier Robin, Peter Škrinjar, Andrew M Waterhouse, Gabriel Studer, Gerardo Tauriello, Janani Durairaj, Torsten Schwede

Proteins: Structure, Function, and Bioinformatics · 94 · 403-413

Abstract

ABSTRACT Independent, blind assessment of structure prediction methods is essential for establishing state‐of‐the‐art performance, identifying limitations, and guiding future developments. The Continuous Automated Model EvaluatiOn (CAMEO) platform provides weekly, automated benchmarking of structure prediction servers, complementing the biennial Critical Assessment of Structure Prediction (CASP) experiments.

Assessment of pharmaceutical protein–ligand pose and affinity predictions in CASP16

Michael K Gilson, Jerome Eberhardt, Peter Škrinjar, Janani Durairaj, Xavier Robin, Andriy Kryshtafovych

Proteins: Structure, Function, and Bioinformatics · 94 · 249-266

Abstract

ABSTRACT The protein–ligand component of the 16th Critical Assessment of Structure Prediction (CASP16) challenged participants to predict both binding poses and affinities of small molecules to protein targets, with a focus on drug‐like compounds from pharmaceutical discovery projects. Thirty research groups submitted predictions for 229 protein–ligand pose targets and 140 affinity targets across five protein systems. Among the submitted predictions, template‐based pose‐prediction methods did particularly well, with the best groups achieving mean LDDT‐PLI values of 0.69 (scale of 0–1 with 1 best). For comparison, we also ran a set of automated baseline pose‐prediction methods, including ones using deep neural networks. Of these, AlphaFold 3 did particularly well, with a mean LDDT‐PLI of 0.8, thus outscoring the best CASP16 predictor. The CASP affinity predictions showed modest correlation with experimental data (maximum Kendall's τ = 0.42), well below the theoretical maximum possible given experimental uncertainty (~0.73). As seen in prior challenges, providing experimental structures did not improve affinity predictions in the second stage of the challenge, suggesting that the scoring functions used here are a key limiting factor. Overall, the accuracy achieved by CASP participants is similar to that observed in the prior Drug Design Data Resource (D3R) blinded prediction challenges. The present results highlight the progress and persistent challenges in computational protein–ligand modeling and provide valuable benchmarks for the field of computer‐aided drug design.

MotifCraft: scalable functional protein binder design with AlphaFold2 hallucination

Océane Follonier, Torsten Schwede, Janani Durairaj

Preprint

Abstract

Motif scaffolding is one promising direction in computational protein design where de novo protein scaffolds are designed to maintain a set of functional protein coordinates – a motif. Here we explore an AlphaFold-Multimer hallucination workflow, MotifCraft, for motif scaffolding and present approaches that improve efficiency while maintaining or improving protein generation accuracy. Our approach produces higher in silico success rates than existing diffusion-based scaffolding methods. Furthermore, we show that scaffolding a binder interface motif results in designs scoring as well, or better, according to established in silico interface metrics when we include the target protein in the workflow. Finally, we conclude that cropping the target protein to the minimal binding interface of a target protein is an effective method for scalable binder design against large proteins. MotifCraft thus enables fast, accurate, and scalable motif scaffolding as well as binder design against large protein targets.

2025

Rewriting protein alphabets with language models

Lorenzo Pantolini, Gabriel Studer, Laura Engist, Ieva Pudžiuvelytė, Florian Pommerening, Andrew Mark Waterhouse, Stefan Bienert, Gerardo Tauriello, Martin Steinegger, Torsten Schwede, Janani Durairaj

bioRxiv · Preprint

Abstract

Detecting remote homology with speed and sensitivity is crucial for tasks like function annotation and structure prediction. We introduce a novel approach using contrastive learning to convert protein language model embeddings into a new 20-letter alphabet, TEA, enabling highly efficient large-scale protein homology searches. Searching with our alphabet performs on par with and complements structure-based methods without requiring any structural information, and with the speed of sequence search. Ultimately, we bring the exciting advances in protein language model representation learning to the plethora of sequence bioinformatics algorithms developed over the past half-century, offering a powerful new tool for biological discovery.

The Viral AlphaFold Database of monomers and homodimers reveals conserved protein folds in viruses of bacteria, archaea, and eukaryotes

Roni Odai, Michèle Leemann, Tamim Al-Murad, Minhal Abdullah, Lena Shyrokova, Tanel Tenson, Vasili Hauryliuk, Janani Durairaj, Joana Pereira, Gemma C Atkinson

Science Advances · 11 · eadz8560

Abstract

Viruses are the most abundant and genetically diverse entities on Earth, yet the functions and evolution of most viral proteins remain poorly understood. Their rapid evolution often obscures evolutionary relationships, limiting the ability to assign functions using sequence-based methods. Although the conservation of protein fold can reveal deep homologies, viral proteins remain underrepresented in structural databases. We address this by clustering viral sequences from RefSeq and predicting the structures of ~27,000 representative proteins using AlphaFold2 to create the Viral AlphaFold Database (VAD). We uncover conserved folds in diverse viruses infecting bacteria, archaea, and eukaryotes. We predict homodimers and make comparisons to the Protein Data Bank, providing data on oligomerization potential. We reveal considerable functional darkness in the viral protein universe and report the discovery and validation of an uncharacterized toxin-antitoxin system. The VAD provides a foundation for exploring viral structure-function relationships, including ancient folds shaping viral interactions across all life.

Large-scale protein clustering in the age of deep learning

Joana Pereira, Lorenzo Pantolini, Janani Durairaj, Torsten Schwede

Current Opinion in Structural Biology · 94 · 103078

Abstract

Proteins within a family sharing sequence and structure similarity due to a common evolutionary origin often also share functional similarities. Clustering of proteins therefore offers valuable insights, enabling the transfer of features and annotations from well-studied proteins to less-investigated ones. On a local scale, clustering helps identify patterns within specific protein families. On a larger scale, it provides insights into the entire protein universe, showcasing relationships that may not be immediately apparent. Traditionally, this was done at the sequence level or with the use of experimentally resolved protein structures, but the advent of deep learning in protein bioinformatics has brought new options to the table, increasing the breadth, depth, and diversity of similarity metrics and clustering approaches.

Structural basis for cooperative ssDNA binding by bacteriophage protein filament P12

Lena K Träger, Morris Degen, Joana Pereira, Janani Durairaj, Raphael Dias Teixeira, Sebastian Hiller, Nicolas Huguenin-Dezot

Nucleic Acids Research · 53 · gkaf132

Abstract

Abstract Protein-primed DNA replication is a unique mechanism, bioorthogonal to other known DNA replication modes. It relies on specialised single-stranded DNA (ssDNA)-binding proteins (SSBs) to stabilise ssDNA intermediates by unknown mechanisms. Here, we present the structural and biochemical characterisation of P12, an SSB from bacteriophage PRD1. High-resolution cryo-electron microscopy reveals that P12 forms a unique, cooperative filament along ssDNA. Each protomer binds the phosphate backbone of 6 nucleotides in a sequence-independent manner, protecting ssDNA from nuclease degradation. Filament formation is driven by an intrinsically disordered C-terminal tail, facilitating cooperative binding. We identify residues essential for ssDNA interaction and link the ssDNA-binding ability of P12 to toxicity in host cells. Bioinformatic analyses place the P12 fold as a distinct branch within the OB-like fold family. This work offers new insights into protein-primed DNA replication and lays a foundation for biotechnological applications.

2024

Identification of a unique, de novo MYCBP2 variant in an individual with highly superior autobiographical memory

Andreas Papassotiropoulos, Jana Petrovska, Andreas Arnold, Aurora KR LePort, Pavlina Mastrandreas, Melanie Neutzner, Virginie Freytag, Dmytro Nesterenko, Vaibhav Gharat, Nathalie Schicktanz, Vanja Vukojevic, David Coynel, Attila Stetak, Noëlle Burri, Navid Ghaffari, Claudia Riva, Janani Durairaj, Torsten Schwede, Oliver Bieri, Johannes Gräff, Efthimios MC Skoulakis, Katharina Henke, Sven Cichon, Verdon Taylor, Craig EL Stark, James L McGaugh, Camin Dean, Dominique J-F de Quervain

medRxiv

Abstract

Abstract Highly Superior Autobiographical Memory (HSAM) is an extremely rare condition characterized by an individual’s unparallelled ability to recall personal past events with exceptional detail and accuracy, including exact dates and days of the week, spanning many decades 1–3 . The molecular underpinnings of HSAM are unknown. Here, we investigated an individual with HSAM through neuropsychological testing, structural brain imaging, and genetic analyses. HSAM was confirmed as an isolated exceptional cognitive ability, with brain imaging revealing exceptionally large volumes of regions within the hippocampal formation, which have been previously linked to autobiographical memory. Using whole exome sequencing of the HSAM individual and their unaffected parents, we identified a unique de novo missense variant in MYCBP2 , which encodes an E3 ubiquitin-protein ligase 4,5 . To explore the potential behavioral consequences of this variant, we introduced the homologous variant into C. elegans , which resulted in reduced forgetting and increased membrane-bound glutamate receptor in relevant neuronal cells. These findings show that the studied HSAM individual carries a unique, de novo missense variant in MYCBP2 , which reduces forgetting in a model organism. The identification of functionally relevant genetic variants in individuals with superior memory traits has the potential to inform future research into memory-modulating therapies.

PLINDER: The protein-ligand interactions dataset and evaluation resource

Janani Durairaj, Yusuf Adeshina, Zhonglin Cao, Xuejin Zhang, Vladas Oleinikovas, Thomas Duignan, Zachary McClure, Xavier Robin, Gabriel Studer, Daniel Kovtun, Emanuele Rossi, Guoqing Zhou, Srimukh Veccham, Clemens Isert, Yuxing Peng, Prabindh Sundareson, Mehmet Akdel, Gabriele Corso, Hannes Stärk, Gerardo Tauriello, Zachary Carpenter, Michael Bronstein, Emine Kucukbenli, Torsten Schwede, Luca Naef

bioRxiv · Preprint

Abstract

Protein-ligand interactions (PLI) are foundational to small molecule drug design. With computational methods striving towards experimental accuracy, there is a critical demand for a well-curated and diverse PLI dataset. Existing datasets are often limited in size and diversity, and commonly used evaluation sets suffer from training information leakage, hindering the realistic assessment of method generalization capabilities. To address these shortcomings, we present PLIN-DER, the largest and most annotated dataset to date, comprising 449,383 PLI systems, each with over 500 annotations, similarity metrics at protein, pocket, interaction and ligand levels, and paired unbound ( apo ) and predicted structures. We propose an approach to generate training and evaluation splits that minimizes task-specific leakage and maximizes test set quality, and compare the resulting performance of DiffDock when retrained with different kinds of splits.

LIGATE-LIgand Generator and portable drug discovery platform AT Exascale

G Palermo, G Accordi, D Gadioli, Y Zhang, E Vitali, B Guindani, D Ardagna, C Silvano, AR Beccari, D Bonanni, C Talarico, F Lughini, AD Biswas, J Martinovic, M Golasowski, P Silva, A Bohm, J Beranek, J Krenek, B Jansik, P Thoman, P Salzmann, T Fahringer, T Schwede, LT Alexander, G Tauriello, J Durairaj, B Cosenza, L Crisci, A Emerson, F Ficarelli, E Lindahl, S Wingbermühle, D Gregori, S Coletti, E Sana, P Gschwandtner

Proceedings of the 21st ACM International Conference on Computing Frontiers Workshops and Special Sessions · 107-109

Abstract

The COVID-19 pandemic demonstrates that a top priority for society, now and in the future, is to be able to respond quickly to diseases with effective treatments. Among the new tools that pharmaceutical industries and researchers have in their hands nowadays, there are the extensive computer simulations capable of evaluating in-silico the interaction between possible drugs and the target proteins. The central goal of the LIGATE project is to create and validate a leading application solution for drug discovery in High-Performance Computing (HPC) systems up to the exascale level. The overall project purpose is the automation of the drug design process, which is currently performed with substantial human effort throughout the different phases of the process: preparation of input parameters, management of data sets with billions of molecules, interaction with HPC queue management systems to handle jobs, and optimization of scoring function parameters and thresholds.

Structural implications of BK polyomavirus sequence variations in the major viral capsid protein Vp1 and large T-antigen: a computational study

Janani Durairaj, Océane M Follonier, Karoline Leuzinger, Leila T Alexander, Maud Wilhelm, Joana Pereira, Caroline A Hillenbrand, Fabian H Weissbach, Torsten Schwede, Hans H Hirsch

Msphere · 9 · e00799-23

Abstract

ABSTRACT BK polyomavirus (BKPyV) is a double-stranded DNA virus causing nephropathy, hemorrhagic cystitis, and urothelial cancer in transplant patients. The BKPyV-encoded capsid protein Vp1 and large T-antigen (LTag) are key targets of neutralizing antibodies and cytotoxic T-cells, respectively. Our single-center data suggested that variability in Vp1 and LTag may contribute to failing BKPyV-specific immune control and impact vaccine design. We, therefore, analyzed all available entries in GenBank (1516 VP1 ; 742 LTAG ) and explored potential structural effects using computational approaches. BKPyV-genotype (gt)1 was found in 71.18% of entries, followed by BKPyV-gt4 (19.26%), BKPyV-gt2 (8.11%), and BKPyV-gt3 (1.45%), but rates differed according to country and specimen type. Vp1-mutations matched a serotype different than the assigned one or were serotype-independent in 43%, 18% affected more than one amino acid. Notable Vp1-mutations altered antibody-binding domains, interactions with sialic acid receptors, or were predicted to change conformation. LTag-sequences were more conserved, with only 16 mutations detectable in more than one entry and without significant effects on LTag-structure or interaction domains. However, LTag changes were predicted to affect HLA-class I presentation of immunodominant 9mers to cytotoxic T-cells. These global data strengthen single center observations and specifically our earlier findings revealing mutant 9mer epitopes conferring immune escape from HLA-I cytotoxic T cells. We conclude that variability of BKPyV-Vp1 and LTag may have important implications for diagnostic assays assessing BKPyV-specific immune control and for vaccine design. IMPORTANCE Type and rate of amino acid variations in BKPyV may provide important insights into BKPyV diversity in human populations and an important step toward defining determinants of BKPyV-specific immunity needed to protect vulnerable patients from BKPyV diseases. Our analysis of BKPyV sequences obtained from human specimens reveals an unexpectedly high genetic variability for this double-stranded DNA virus that strongly relies on host cell DNA replication machinery with its proof reading and error correction mechanisms. BKPyV variability and immune escape should be taken into account when designing further approaches to antivirals, monoclonal antibodies, and vaccines for patients at risk of BKPyV diseases.

Embedding-based alignment: combining protein language models with dynamic programming alignment to detect structural similarities in the twilight-zone

Lorenzo Pantolini, Gabriel Studer, Joana Pereira, Janani Durairaj, Gerardo Tauriello, Torsten Schwede

Bioinformatics · btad786

Abstract

Abstract Motivation Language models are routinely used for text classification and generative tasks. Recently, the same architectures were applied to protein sequences, unlocking powerful new approaches in the bioinformatics field. Protein language models (pLMs) generate high-dimensional embeddings on a per-residue level and encode a “semantic meaning” of each individual amino acid in the context of the full protein sequence. These representations have been used as a starting point for downstream learning tasks and, more recently, for identifying distant homologous relationships between proteins. Results In this work, we introduce a new method that generates embedding-based protein sequence alignments (EBA) and show how these capture structural similarities even in the twilight zone, outperforming both classical methods as well as other approaches based on pLMs. The method shows excellent accuracy despite the absence of training and parameter optimization. We demonstrate that the combination of pLMs with alignment methods is a valuable approach for the detection of relationships between proteins in the twilight-zone. Availability and implementation The code to run EBA and reproduce the analysis described in this article is available at: https://git.scicore.unibas.ch/schwede/EBA and https://git.scicore.unibas.ch/schwede/eba_benchmark.

2023

Protein target highlights in CASP15: Analysis of models by structure providers

Leila T Alexander, Janani Durairaj, Andriy Kryshtafovych, Luciano A Abriata, Yusupha Bayo, Gira Bhabha, Cécile Breyton, Simon G Caulton, James Chen, Séraphine Degroux, Damian C Ekiert, Benedikte S Erlandsen, Lydia Freddolino, Dominic Gilzer, Chris Greening, Jonathan M Grimes, Rhys Grinter, Manickam Gurusaran, Marcus D Hartmann, Charlie J Hitchman, Jeremy R Keown, Ashleigh Kropp, Petri Kursula, Andrew L Lovering, Bruno Lemaitre, Andrea Lia, Shiheng Liu, Maria Logotheti, Shuze Lu, Sigurbjörn Markússon, Mitchell D Miller, George Minasov, Hartmut H Niemann, Felipe Opazo, George N Phillips Jr, Owen R Davies, Samuel Rommelaere, Monica Rosas‐Lemus, Pietro Roversi, Karla Satchell, Nathan Smith, Mark A Wilson, Kuan‐Lin Wu, Xian Xia, Han Xiao, Wenhua Zhang, Z Hong Zhou, Krzysztof Fidelis, Maya Topf, John Moult, Torsten Schwede

Proteins: Structure, Function, and Bioinformatics · 91 · 1571-1599

Abstract

Abstract We present an in‐depth analysis of selected CASP15 targets, focusing on their biological and functional significance. The authors of the structures identify and discuss key protein features and evaluate how effectively these aspects were captured in the submitted predictions. While the overall ability to predict three‐dimensional protein structures continues to impress, reproducing uncommon features not previously observed in experimental structures is still a challenge. Furthermore, instances with conformational flexibility and large multimeric complexes highlight the need for novel scoring strategies to better emphasize biologically relevant structural regions. Looking ahead, closer integration of computational and experimental techniques will play a key role in determining the next challenges to be unraveled in the field of structural molecular biology.

New prediction categories in CASP15

Andriy Kryshtafovych, Maciej Antczak, Marta Szachniuk, Tomasz Zok, Rachael C Kretsch, Ramya Rangan, Phillip Pham, Rhiju Das, Xavier Robin, Gabriel Studer, Janani Durairaj, Jerome Eberhardt, Aaron Sweeney, Maya Topf, Torsten Schwede, Krzysztof Fidelis, John Moult

Proteins: Structure, Function, and Bioinformatics · 91 · 1550-1557

Abstract

Abstract Prediction categories in the Critical Assessment of Structure Prediction (CASP) experiments change with the need to address specific problems in structure modeling. In CASP15, four new prediction categories were introduced: RNA structure, ligand‐protein complexes, accuracy of oligomeric structures and their interfaces, and ensembles of alternative conformations. This paper lists technical specifications for these categories and describes their integration in the CASP data management system.

Assessment of protein–ligand complexes in CASP15

Xavier Robin, Gabriel Studer, Janani Durairaj, Jerome Eberhardt, Torsten Schwede, W Patrick Walters

Proteins: Structure, Function, and Bioinformatics · 91 · 1811-1821

Abstract

Abstract CASP15 introduced a new category, ligand prediction, where participants were provided with a protein or nucleic acid sequence, SMILES line notation, and stoichiometry for ligands and tasked with generating computational models for the three‐dimensional structure of the corresponding protein–ligand complex. These models were subsequently compared with experimental structures determined by x‐ray crystallography or cryoEM. To assess these predictions, two novel scores were developed. The Binding‐Site Superposed, Symmetry‐Corrected Pose Root Mean Square Deviation (BiSyRMSD) evaluated the absolute deviations of the models from the experimental structures. At the same time, the Local Distance Difference Test for Protein–Ligand Interactions (lDDT‐PLI) assessed the ability of models to reproduce the protein–ligand interactions in the experimental structures. The ligands evaluated in this challenge range from single‐atom ions to large flexible organic molecules. More than 1800 submissions were evaluated for their ability to predict 23 different protein–ligand complexes. Overall, the best models could faithfully reproduce the geometries of more than half of the prediction targets. The ligands' size and flexibility were the primary factors influencing the predictions' quality. Small ions and organic molecules with limited flexibility were predicted with high fidelity, while reproducing the binding poses of larger, flexible ligands proved more challenging.

Artificial intelligence for natural product drug discovery

Michael W Mullowney, Katherine R Duncan, Somayah S Elsayed, Neha Garg, Justin JJ van der Hooft, Nathaniel I Martin, David Meijer, Barbara R Terlouw, Friederike Biermann, Kai Blin, Janani Durairaj, Marina Gorostiola González, Eric JN Helfrich, Florian Huber, Stefan Leopold-Messer, Kohulan Rajan, Tristan de Rond, Jeffrey A van Santen, Maria Sorokina, Marcy J Balunas, Mehdi A Beniddir, Doris A van Bergeijk, Laura M Carroll, Chase M Clark, Djork-Arné Clevert, Chris A Dejong, Chao Du, Scarlet Ferrinho, Francesca Grisoni, Albert Hofstetter, Willem Jespers, Olga V Kalinina, Satria A Kautsar, Hyunwoo Kim, Tiago F Leao, Joleen Masschelein, Evan R Rees, Raphael Reher, Daniel Reker, Philippe Schwaller, Marwin Segler, Michael A Skinnider, Allison S Walker, Egon L Willighagen, Barbara Zdrazil, Nadine Ziemert, Rebecca JM Goss, Pierre Guyomard, Andrea Volkamer, William H Gerwick, Hyun Uk Kim, Rolf Müller, Gilles P van Wezel, Gerard JP van Westen, Anna KH Hirsch, Roger G Linington, Serina L Robinson, Marnix H Medema

Nature Reviews Drug Discovery · 22 · 895-916

Abstract

Developments in computational omics technologies have provided new means to access the hidden diversity of natural products, unearthing new potential for drug discovery. In parallel, artificial intelligence approaches such as machine learning have led to exciting developments in the computational drug design field, facilitating biological activity prediction and de novo drug design for molecular targets of interest. Here, we describe current and future synergies between these developments to effectively identify drug candidates from the plethora of molecules produced by nature. We also discuss how to address key challenges in realizing the potential of these synergies, such as the need for high-quality datasets to train deep learning algorithms and appropriate strategies for algorithm validation.

Automated benchmarking of combined protein structure and ligand conformation prediction

Michèle Leemann, Ander Sagasta, Jerome Eberhardt, Torsten Schwede, Xavier Robin, Janani Durairaj

Proteins: Structure, Function, and Bioinformatics

Abstract

Abstract The prediction of protein‐ligand complexes (PLC), using both experimental and predicted structures, is an active and important area of research, underscored by the inclusion of the Protein‐Ligand Interaction category in the latest round of the Critical Assessment of Protein Structure Prediction experiment CASP15. The prediction task in CASP15 consisted of predicting both the three‐dimensional structure of the receptor protein as well as the position and conformation of the ligand. This paper addresses the challenges and proposed solutions for devising automated benchmarking techniques for PLC prediction. The reliability of experimentally solved PLC as ground truth reference structures is assessed using various validation criteria. Similarity of PLC to previously released complexes are employed to judge PLC diversity and the difficulty of a PLC as a prediction target. We show that the commonly used PDBBind time‐split test‐set is inappropriate for comprehensive PLC evaluation, with state‐of‐the‐art tools showing conflicting results on a more representative and high quality dataset constructed for benchmarking purposes. We also show that redocking on crystal structures is a much simpler task than docking into predicted protein models, demonstrated by the two PLC‐prediction‐specific scoring metrics created. Finally, we introduce a fully automated pipeline that predicts PLC and evaluates the accuracy of the protein structure, ligand pose, and protein–ligand interactions.

Uncovering new families and folds in the natural protein universe

Janani Durairaj, Andrew M Waterhouse, Toomas Mets, Tetiana Brodiazhenko, Minhal Abdullah, Gabriel Studer, Gerardo Tauriello, Mehmet Akdel, Antonina Andreeva, Alex Bateman, Tanel Tenson, Vasili Hauryliuk, Torsten Schwede, Joana Pereira

Nature · 622 · 646-653

Abstract

We are now entering a new era in protein sequence and structure annotation, with hundreds of millions of predicted protein structures made available through the AlphaFold database 1 . These models cover nearly all proteins that are known, including those challenging to annotate for function or putative biological role using standard homology-based approaches. In this study, we examine the extent to which the AlphaFold database has structurally illuminated this ‘dark matter’ of the natural protein universe at high predicted accuracy. We further describe the protein diversity that these models cover as an annotated interactive sequence similarity network, accessible at https://uniprot3d.org/atlas/AFDB90v4 . By searching for novelties from sequence, structure and semantic perspectives, we uncovered the β-flower fold, added several protein families to Pfam database 2 and experimentally demonstrated that one of these belongs to a new superfamily of translation-targeting toxin–antitoxin systems, TumE–TumA. This work underscores the value of large-scale efforts in identifying, annotating and prioritizing new protein families. By leveraging the recent deep learning revolution in protein bioinformatics, we can now shed light into uncharted areas of the protein universe at an unprecedented scale, paving the way to innovations in life sciences and biotechnology.

Tunable and portable extreme-scale drug discovery platform at exascale: the lIGATE approach

Gianluca Palermo, Gianmarco Accordi, Davide Gadioli, Emanuele Vitali, Cristina Silvano, Bruno Guindani, Danilo Ardagna, Andrea Rosario Beccari, Domenico Bonanni, Carmine Talarico, F Lughini, Jan Martinovic, Paulo Silva, Ada Bohm, Jakub Beranek, Jan Krenek, Branislav Jansik, Biagio Cosenza, Luigi Crisci, Peter Thoman, Philip Salzmann, Thomas Fahringer, Leila T Alexander, Gerardo Tauriello, Torsten Schwede, Janani Durairaj, Andrew Emerson, Federico Ficarelli, Sebastian Wingbermühle, E Lindhal, Daniele Gregori, Emanuele Sana, Silvano Coletti, Philipp Gschwandtner

Proceedings of the 20th ACM International Conference on Computing Frontiers · 272-278

Abstract

Today digital revolution is having a dramatic impact on the pharmaceutical industry and the entire healthcare system. The implementation of machine learning, extreme-scale computer simulations, and big data analytics in the drug design and development process offers an excellent opportunity to lower the risk of investment and reduce the time to the patient. Within the LIGATE project, we aim to integrate, extend, and co-design best-in-class European components to design Computer-Aided Drug Design (CADD) solutions exploiting today's high-end supercomputers and tomorrow's Exascale resources, fostering European competitiveness in the field. The proposed LIGATE solution is a fully integrated workflow that enables to deliver the result of a virtual screening campaign for drug discovery with the highest speed along with the highest accuracy. The full automation of the solution and the possibility to run it on multiple supercomputing centers at once permit to run an extreme scale in silico drug discovery campaign in few days to respond promptly for example to a worldwide pandemic crisis.

From Genomes to Variant Interpretations Through Protein Structures

Janani Durairaj, Leila Tamara Alexander, Gabriel Studer, Gerardo Tauriello, Ingrid Guarnetti Prandi, Rosalba Lepore, Giovanni Chillemi, Torsten Schwede

Exscalate4CoV: High-Performance Computing for COVID Drug Discovery · 41-50

Abstract

The large amount of genetic, phenotypic, and structural data from diverse conditions and environments offers opportunities for new groundbreaking research. Today, the major scientific task is to interpret the vast number of genetic variants within these data. As described in this chapter, identifying relevant variants and connecting them with the associated protein structural and environmental information is a powerful approach to biological discoveries. The unified view of the data brings us a step closer to understanding genetic variation, which is also fundamental for achieving the goals of personalized medicine and the planet’s environment.