Refine
Year of publication
Document Type
- Article (32)
Has Fulltext
- yes (32)
Is part of the Bibliography
- no (32)
Keywords
- Virtual Screening (2)
- AF4–MLL (1)
- Compound Database (1)
- Gaussian Process (1)
- Identical Topology (1)
- Lead Structure (1)
- Multiple Kernel (1)
- Oncoprotein activation (1)
- Pairwise Sequence Alignment (1)
- Support Vector Regression (1)
Shape complementarity is a compulsory condition for molecular recognition. In our 3D ligand-based virtual screening approach called SQUIRREL, we combine shape-based rigid body alignment with fuzzy pharmacophore scoring. Retrospective validation studies demonstrate the superiority of methods which combine both shape and pharmacophore information on the family of peroxisome proliferator-activated receptors (PPARs). We demonstrate the real-life applicability of SQUIRREL by a prospective virtual screening study, where a potent PPARalpha agonist with an EC50 of 44 nM and 100-fold selectivity against PPARgamma has been identified...
Background: The human pathogen Helicobacter pylori (H. pylori) is a main cause for gastric inflammation and cancer. Increasing bacterial resistance against antibiotics demands for innovative strategies for therapeutic intervention. Methodology/Principal Findings: We present a method for structure-based virtual screening that is based on the comprehensive prediction of ligand binding sites on a protein model and automated construction of a ligand-receptor interaction map. Pharmacophoric features of the map are clustered and transformed in a correlation vector (‘virtual ligand’) for rapid virtual screening of compound databases. This computer-based technique was validated for 18 different targets of pharmaceutical interest in a retrospective screening experiment. Prospective screening for inhibitory agents was performed for the protease HtrA from the human pathogen H. pylori using a homology model of the target protein. Among 22 tested compounds six block E-cadherin cleavage by HtrA in vitro and result in reduced scattering and wound healing of gastric epithelial cells, thereby preventing bacterial infiltration of the epithelium. Conclusions/Significance: This study demonstrates that receptor-based virtual screening with a permissive (‘fuzzy’) pharmacophore model can help identify small bioactive agents for combating bacterial infection.
Spherical harmonics coeffcients for ligand-based virtual screening of cyclooxygenase inhibitors
(2011)
Background: Molecular descriptors are essential for many applications in computational chemistry, such as ligand-based similarity searching. Spherical harmonics have previously been suggested as comprehensive descriptors of molecular structure and properties. We investigate a spherical harmonics descriptor for shape-based virtual screening. Methodology/Principal Findings: We introduce and validate a partially rotation-invariant three-dimensional molecular shape descriptor based on the norm of spherical harmonics expansion coefficients. Using this molecular representation, we parameterize molecular surfaces, i.e., isosurfaces of spatial molecular property distributions. We validate the shape descriptor in a comprehensive retrospective virtual screening experiment. In a prospective study, we virtually screen a large compound library for cyclooxygenase inhibitors, using a self-organizing map as a pre-filter and the shape descriptor for candidate prioritization. Conclusions/Significance: 12 compounds were tested in vitro for direct enzyme inhibition and in a whole blood assay. Active compounds containing a triazole scaffold were identified as direct cyclooxygenase-1 inhibitors. This outcome corroborates the usefulness of spherical harmonics for representation of molecular shape in virtual screening of large compound collections. The combination of pharmacophore and shape-based filtering of screening candidates proved to be a straightforward approach to finding novel bioactive chemotypes with minimal experimental effort.
Background: Pathogenic bacteria infecting both animals as well as plants use various mechanisms to transport virulence factors across their cell membranes and channel these proteins into the infected host cell. The type III secretion system represents such a mechanism. Proteins transported via this pathway (‘‘effector proteins’’) have to be distinguished from all other proteins that are not exported from the bacterial cell. Although a special targeting signal at the N-terminal end of effector proteins has been proposed in literature its exact characteristics remain unknown. Methodology/Principal Findings: In this study, we demonstrate that the signals encoded in the sequences of type III secretion system effectors can be consistently recognized and predicted by machine learning techniques. Known protein effectors were compiled from the literature and sequence databases, and served as training data for artificial neural networks and support vector machine classifiers. Common sequence features were most pronounced in the first 30 amino acids of the effector sequences. Classification accuracy yielded a cross-validated Matthews correlation of 0.63 and allowed for genome-wide prediction of potential type III secretion system effectors in 705 proteobacterial genomes (12% predicted candidates protein), their chromosomes (11%) and plasmids (13%), as well as 213 Firmicute genomes (7%). Conclusions/Significance: We present a signal prediction method together with comprehensive survey of potential type III secretion system effectors extracted from 918 published bacterial genomes. Our study demonstrates that the analyzed signal features are common across a wide range of species, and provides a substantial basis for the identification of exported pathogenic proteins as targets for future therapeutic intervention. The prediction software is publicly accessible from our web server ( www.modlab.org ).
We developed the Pharmacophore Alignment Search Tool (PhAST), a text-based technique for rapid hit and lead structure searching in large compound databases. For each molecule, a two-dimensional graph of potential pharmacophoric points (PPPs) is created, which has an identical topology as the original molecule with implicit hydrogen atoms. Each vertex is coloured by a symbol representing the corresponding PPP. The vertices of the graph are canonically labelled. The symbols associated with the vertices are combined to a so-called PhAST-Sequence beginning with the vertex with the lowest canonical label. Due to the canonical labelling the created PhAST-Sequence is characteristic for each molecule. For similarity assessment, PhAST-Sequences are compared using the sequence identity in their global pairwise alignment. The alignment score lies between 0 (no similarity) and 1 (identical PhAST-Sequences). In order to use global pairwise sequence alignment, a score matrix for pharmacophoric symbols was developed and gap penalties were optimized. PhAST performed comparably and sometimes superior to other similarity search tools (CATS2D, MOE pharmacophore quadruples) in retrospective virtual screenings using the COBRA collection of drugs and lead structures. Most importantly, the PhAST alignment technique allows for the computation of significance estimates that help prioritize a virtual hit list.
The representation of small molecules as molecular graphs is a common technique in various fields of cheminformatics. This approach employs abstract descriptions of topology and properties for rapid analyses and comparison. Receptor-based methods in contrast mostly depend on more complex representations impeding simplified analysis and limiting the possibilities of property assignment. In this study we demonstrate that ligand-based methods can be applied to receptor-derived binding site analysis. We introduce the new method PocketGraph that translates representations of binding site volumes into linear graphs and enables the application of graph-based methods to the world of protein pockets. The method uses the PocketPicker algorithm for characterization of binding site volumes and employs a Growing Neural Gas procedure to derive graph representations of pocket topologies. Self-organizing map (SOM) projections revealed a limited number of pocket topologies. We argue that there is only a small set of pocket shapes realized in the known ligand-receptor complexes.
Wie findet man einen neuen Wirkstoff? Die pharmazeutisch-chemische Forschung steht mit diesem Vorhaben vor einer scheinbar unlösbaren Aufgabe, denn der "chemische Raum" aller wirkstoffartigen Moleküle ist unvorstellbar groß. So wurde geschätzt, dass man prinzipiell aus 1060 bis 10100 verschiedenen Verbindungen die geeigneten Kandidaten auswählen kann. Zum Vergleich: Seit dem Urknall sollen "nur" etwa 10 hoch 18 Sekunden, etwa 14 Milliarden Jahre, vergangen sein. Dies bedeutet, dass der chemische Raum praktisch unendlich ist. Aus dieser Überlegung lassen sich zumindest zwei Schlussfolgerungen ziehen: Zum einen gibt es die begründete Hoffnung, dass ein Molekül mit der gewünschten Aktivität existiert, zum anderen stellt sich die Frage, wie diese unvorstellbar große Zahl chemischer Verbindungen systematisch durchmustert werden kann? Doch die Situation ist nicht so hoffnungslos, wie sie auf den ersten Blick erscheint. Dies zeigt die erfolgreiche Entwicklung immer neuer Medikamente. Das Forschungsgebiet der Chemieinformatik befasst sich mit der Entwicklung von intelligenten Lösungsansätzen, die Chemikern bei dieser Suche nach den "Nadeln im riesigen Heuhaufen" helfen können.
For a virtual screening study, we introduce a combination of machine learning techniques, employing a graph kernel, Gaussian process regression and clustered cross-validation. The aim was to find ligands of peroxisome-proliferator activated receptor gamma (PPAR-y). The receptors in the PPAR family belong to the steroid-thyroid-retinoid superfamily of nuclear receptors and act as transcription factors. They play a role in the regulation of lipid and glucose metabolism in vertebrates and are linked to various human processes and diseases. For this study, we used a dataset of 176 PPAR-y agonists published by Ruecker et al. ...
A new method to bridge the gap between ligand and receptor-based methods in virtual screening (VS) is presented. We introduce a structure-derived virtual ligand (VL) model as an extension to a previously published pseudo-ligand technique [1]: LIQUID [2] fuzzy pharmacophore virtual screening is combined with grid-based protein binding site predictions of PocketPicker [3]. This approach might help reduce bias introduced by manual selection of binding site residues and introduces pocket shape information to the VL. It allows for a combination of several protein structure models into a single "fuzzy" VL representation, which can be used to scan screening compound collections for ligand structures with a similar potential pharmacophore. PocketPicker employs an elaborate grid-based scanning procedure to determine buried cavities and depressions on the protein's surface. Potential binding sites are represented by clusters of grid probes characterizing the shape and accessibility of a cavity. A rule-based system is then applied to project reverse pharmacophore types onto the grid probes of a selected pocket. The pocket pharmacophore types are assigned depending on the properties and geometry of the protein residues surrounding the pocket with regard to their relative position towards the grid probes. LIQUID is used to cluster representative pocket probes by their pharmacophore types describing a fuzzy VL model. The VL is encoded in a correlation vector, which can then be compared to a database of pre-calculated ligand models. A retrospective screening using the fuzzy VL and several protein structures was evaluated by ten fold cross-validation with ROC-AUC and BEDROC metrics, obtaining a significant enrichment of actives. Future work will be devoted to prospective screening using a novel protein target of Helicobacter pylori and compounds from commercial providers.
Two methods for the fast, fragment-based combinatorial molecule assembly were developed. The software COLIBREE® (Combinatorial Library Breeding) generates candidate structures from scratch, based on stochastic optimization [1]. Result structures of a COLIBREE design run are based on a fixed scaffold and variable linkers and side-chains. Linkers representing virtual chemical reactions and side-chain building blocks obtained from pseudo-retrosynthetic dissection of large compound databases are exchanged during optimization. The process of molecule design employs a discrete version of Particle Swarm Optimization (PSO) [2]. Assembled compounds are scored according to their similarity to known reference ligands. Distance to reference molecules is computed in the space of the topological pharmacophore descriptor CATS [3]. In a case study, the approach was applied to the de novo design of potential peroxisome proliferator-activated receptor (PPAR gamma) selective agonists. In a second approach, we developed the formal grammar Reaction-MQL [4] for the in silico representation and application of chemical reactions. Chemical transformation schemes are defined by functional groups participating in known organic reactions. The substructures are specified by the linear Molecular Query Language (MQL) [5]. The developed software package contains a parser for Reaction-MQL-expressions and enables users to design, test and virtually apply chemical reactions. The program has already been used to create combinatorial libraries for virtual screening studies. It was also applied in fragmentation studies with different sets of retrosynthetic reactions and various compound libraries.
There is a renewed interest in pseudoreceptor models which enable computational chemists to bridge the gap of ligand- and receptor-based drug design. We developed a pseudoreceptor model for the histamine H4 receptor (H4R) based on five potent antagonists representing different chemotypes. Here we present the selection of potential ligand binding pockets that occur during molecular dynamics (MD) simulations of a homology-based receptor model. We present a method for prioritizing receptor models according to their match with the consensus ligand-binding mode represented by the pseudoreceptor. In this way, ligand information can be transferred to receptor-based modelling. We use Geometric Hashing to match three-dimensional points in Cartesion space. This allows for the rapid translation- and rotation-free comparison of atom coordinates, which also permits partial matching. The only prerequisite is a hash table, which uses distance triplets as hash keys. Each time a distance triplet occurring in the candidate point set which corresponds to an existing key, the match is represented by a vote of the respective key. Finally, the global match of both point sets can be easily extracted by selection of voted distance triplets. The results revealed a preferred ligand-binding pocket in H4R, which would not have been identified using an unrefined homology model of the protein. The key idea was to rely on ligand information by pseudoreceptor modelling.
Chemical language models enable de novo drug design without the requirement for explicit molecular construction rules. While such models have been applied to generate novel compounds with desired bioactivity, the actual prioritization and selection of the most promising computational designs remains challenging. Herein, we leveraged the probabilities learnt by chemical language models with the beam search algorithm as a model-intrinsic technique for automated molecule design and scoring. Prospective application of this method yielded novel inverse agonists of retinoic acid receptor-related orphan receptors (RORs). Each design was synthesizable in three reaction steps and presented low-micromolar to nanomolar potency towards RORγ. This model-intrinsic sampling technique eliminates the strict need for external compound scoring functions, thereby further extending the applicability of generative artificial intelligence to data-driven drug discovery.
The repertoire of natural products offers tremendous opportunities for chemical biology and drug discovery. Natural product-inspired synthetic molecules represent an ecologically and economically sustainable alternative to the direct utilization of natural products. De novo design with machine intelligence bridges the gap between the worlds of bioactive natural products and synthetic molecules. On employing the compound Marinopyrrole A from marine Streptomyces as a design template, the algorithm constructs innovative small molecules that can be synthesized in three steps, following the computationally suggested synthesis route. Computational activity prediction reveals cyclooxygenase (COX) as a putative target of both Marinopyrrole A and the de novo designs. The molecular designs are experimentally confirmed as selective COX-1 inhibitors with nanomolar potency. X-ray structure analysis reveals the binding of the most selective compound to COX-1. This molecular design approach provides a blueprint for natural product-inspired hit and lead identification for drug discovery with machine intelligence.
Protein kinases are targets for drug development. Dysregulation of kinase activity leads to various diseases, e.g. cancer, inflammation, diabetes. Human polo-like kinase 1 (Plk1), a serine/threonine kinase, is a cancer-relevant gene and a potential drug target which attracts increasing attention in the field of cancer therapy. Plk1 is a key player in mitosis and modulates entry into mitosis and the spindle checkpoint at the meta-/anaphase transition. Plk1 overexpression is observed in various human tumors, and it is a negative prognostic factor for cancer patients. The same catalytical mechanism and the same co-substrate (ATP) lead to the problem of inhibitor selectivity. A strategy to solve this problem is represented by targeting the inactive conformation of kinases. Kinases undergo conformational changes between active and inactive conformation and thus an additional hydrophobic pocket is created in the inactive conformation where the surrounding amino acids are less conserved. A "homology model" of the inactive conformation of Plk1 was constructed, as the crystal structure in its inactive conformation is unknown. A crystal structure of Aurora A kinase served as template structure. With this homology model a receptor-based pharmacophore search was performed using SYBYL7.3 software. The raw hits were filtered using physico-chemical properties. The resulting hits were docked using Gold3.2 software, and 13 candidates for biological testing were manually selected. Three compounds of the 13 tested exhibit anti-proliferative effects in HeLa cancer cells. The most potent inhibitor, SBE13, was further tested in various other cancer cell lines of different origins and displayed EC50 values between 12 microM and 39 microM. Cancer cells incubated with SBE13 showed induction of apoptosis, detected by PARP (Poly-Adenosyl-Ribose-Polymerase) cleavage, caspase 9 activation and DAPI staining of apoptotic nuclei.
Poster presentation at 5th German Conference on Cheminformatics: 23. CIC-Workshop Goslar, Germany. 8-10 November 2009 Protein kinases are important targets for drug development. The almost identical protein folding of kinases and the common co-substrate ATP leads to the problem of inhibitor selectivity. Type II inhibitors, targeting the inactive conformation of kinases, occupy a hydrophobic pocket with less conserved surrounding amino acids. Human polo-like kinase 1 (Plk1) represents a promising target for approaches to identify new therapeutic agents. Plk1 belongs to a family of highly conserved serine/threonine kinases, and is a key player in mitosis, where it modulates the spindle checkpoint at metaphase/anaphase transition. Plk1 is over-expressed in all today analyzed human tumors of different origin and serves as a negative prognostic marker in cancer patients. The newly identified inhibitor, SBE13, a vanillin derivative, targets Plk1 in its inactive conformation. This leads to selectivity within the Plk family and towards Aurora A. This selectivity can be explained by docking studies of SBE13 into the binding pocket of homology models of Plk1, Plk2 and Plk3 in their inactive conformation. SBE13 showed anti-proliferative effects in cancer cell lines of different origins with EC50 values between 5 microM and 39 microM and induced apoptosis. Increasing concentrations of SBE13 result in increasing amounts of cells in G2/M phase 13 hours after double thymidin block of HeLa cells. The kinase activity of Plk1 was inhibited with an IC50 of 200 pM. Taken together, we could show that carefully designed structure-based virtual screening is well-suited to identify selective type II kinase inhibitors targeting Plk1 as potential anti-cancer therapeutics.
Targeting signals direct proteins to their extra- or intracellular destination such as the plasma membrane or cellular organelles. Here we investigated the structure and function of exceptionally long signal peptides encompassing at least 40 amino acid residues. We discovered a two-domain organization ("NtraC model") in many long signals from vertebrate precursor proteins. Accordingly, long signal peptides may contain an N-terminal domain (N-domain) and a C-terminal domain (C-domain) with different signal or targeting capabilities, separable by a presumably turn-rich transition area (tra). Individual domain functions were probed by cellular targeting experiments with fusion proteins containing parts of the long signal peptide of human membrane protein shrew-1 and secreted alkaline phosphatase as a reporter protein. As predicted, the N-domain of the fusion protein alone was shown to act as a mitochondrial targeting signal, whereas the C-domain alone functions as an export signal. Selective disruption of the transition area in the signal peptide impairs the export efficiency of the reporter protein. Altogether, the results of cellular targeting studies provide a proof-of-principle for our NtraC model and highlight the particular functional importance of the predicted transition area, which critically affects the rate of protein export. In conclusion, the NtraC approach enables the systematic detection and prediction of cryptic targeting signals present in one coherent sequence, and provides a structurally motivated basis for decoding the functional complexity of long protein targeting signals.
We performed a bioinformatical analysis of protein export elements (PEXEL) in the putative proteome of the malaria parasite Plasmodium falciparum. A protein family-specific conservation of physicochemical residue profiles was found for PEXEL-flanking sequence regions. We demonstrate that the family members can be clustered based on the flanking regions only and display characteristic hydrophobicity patterns. This raises the possibility that the flanking regions may contain additional information for a family-specific role of PEXEL. We further show that signal peptide cleavage results in a positional alignment of PEXEL from both proteins with, and without, a signal peptide.
Bacterial autotransporters represent a diverse family of proteins that autonomously translocate across the inner membrane of Gram-negative bacteria via the Sec complex and across the outer bacterial membrane. They often possess exceptionally long N-terminal signal sequences. We analyzed 90 long signal sequences of bacterial autotransporters and members of the two-partner secretion pathway in silico and describe common domain organization found in 79 of these sequences. The domains are in agreement with previously published experimental data. Our algorithmic approach allows for the systematic identification of functionally different domains in long signal sequences. Keywords: bacterial autotransporter, sequence analysis, pattern, protein targeting, signal peptide, protein trafficking