Cheminformatics

Following a successful high-throughput screening (HTS) campaign, the large volume of data is analyzed using advanced cheminformatics and medicinal chemistry toolkits to identify the most promising compound scaffolds. To ensure specificity and relevance, non-specific or promiscuous compounds (“frequent hitters”) are filtered out using established substructure filters and comprehensive cross-analysis with existing bioactivity data from the COMAS database. An analog search can also be performed within the COMAS compound collection to identify related molecules, recover potential false negatives, and further refine hit selection. Validated hit structures are then explored for structure-activity relationships (SAR) through the identification and testing of commercially available analogs. Where available, protein structural data is integrated to guide compound selection and prioritization, enabling a more targeted and rational approach to hit optimization.

When appropriate, virtual screening methods can be implemented as accompanying or main screening approaches.

The ChemInformatics and MedChem toolchain used by COMAS includes: Pipeline Pilot, KNIME, RDKit, Jupyter, DataWarrior, Molecular Docking using Smina or Gnina, ligand-based virtual screening using scikit-learn, PyTorch, DeepChem

Go to Editor View