{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":67356,"databundleVersionId":8006601,"sourceType":"competition"},{"sourceId":175791383,"sourceType":"kernelVersion"}],"dockerImageVersionId":30698,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-05-07T22:20:59.031410Z","iopub.execute_input":"2024-05-07T22:20:59.031800Z","iopub.status.idle":"2024-05-07T22:21:00.327958Z","shell.execute_reply.started":"2024-05-07T22:20:59.031771Z","shell.execute_reply":"2024-05-07T22:21:00.327029Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# The RDKit infrastructure to be used in manipulating and displaying molecules\n# Note: Currently Pandas 2.2 does not play with the RDKit package that is installed from PyPi\n!pip install rdkit\n","metadata":{"execution":{"iopub.status.busy":"2024-05-07T22:21:35.756759Z","iopub.execute_input":"2024-05-07T22:21:35.758312Z","iopub.status.idle":"2024-05-07T22:21:52.457508Z","shell.execute_reply.started":"2024-05-07T22:21:35.758271Z","shell.execute_reply":"2024-05-07T22:21:52.455975Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Analyzing the DEL Kaggle Dataset\n\nAs I  ([Bernhard Rohde's](https://www.kaggle.com/bernhardrohde)) found while competing in [1st EUOS/SLAS Joint Challenge: Compound Solubility](https://www.kaggle.com/competitions/euos-slas/overview), the first step in any such exercise is to analyze the peculiarities of the datasets to evaluate what can an cannot be learned from them. In the solubility challenge, it turned out that the acquisition structure (plate layout) of the dataset accounted for a large part of the attainable signal, and guessing this from the compound IDs and including it in the model got me close to the top of the list. Eventually, I was disqualified, because I was not just using structure derived descriptor in the competition.\n\nHowever, I do think it should pay off to first find out what one is up to by descriptively analyzing the dataset, which I'm trying to do here. I will not include the [description](https://www.kaggle.com/competitions/leash-BELKA/data) by the challenge providers here but assume it to be known.","metadata":{}},{"cell_type":"markdown","source":"# Support Methods For Analysis\n\nThe folowing cell declares a number of methods that will be useful in further analsis and visualization.","metadata":{}},{"cell_type":"code","source":"import csv\nimport random\n\nfrom rdkit import Chem\nfrom rdkit.Chem.Draw import IPythonConsole\nfrom rdkit.Chem import Draw\nfrom rdkit.Chem.MolStandardize import rdMolStandardize\nIPythonConsole.ipython_useSVG=True  # set this to False if you want PNGs instead of SVGs\n# currently broken from rdkit.Chem import PandasTools\nfrom rdkit import RDLogger\nimport time\n\nRDLogger.DisableLog('rdApp.info') # Disable most of the non-fatal RDKit warnings\n\ndef standardize(smiles, noTautomers=True):\n    # Adapted from https://bitsilla.com/blog/2021/06/standardizing-a-molecule-using-rdkit/\n    # which follows the steps in\n    # https://github.com/greglandrum/RSC_OpenScience_Standardization_202104/blob/main/MolStandardize%20pieces.ipynb\n    # as described **excellently** (by Greg) in\n    # https://www.youtube.com/watch?v=eWTApNX8dJQ\n    mol = Chem.MolFromSmiles(smiles)\n     \n    # removeHs, disconnect metal atoms, normalize the molecule, reionize the molecule\n    clean_mol = rdMolStandardize.Cleanup(mol) \n     \n    # if many fragments, get the \"parent\" (the actual mol we are interested in) \n    parent_clean_mol = rdMolStandardize.FragmentParent(clean_mol)\n         \n    # try to neutralize molecule\n    uncharger = rdMolStandardize.Uncharger() # annoying, but necessary as no convenience method exists\n    uncharged_parent_clean_mol = uncharger.uncharge(parent_clean_mol)\n    if noTautomers : return Chem.MolToSmiles(uncharged_parent_clean_mol),uncharged_parent_clean_mol\n     \n    # note that no attempt is made at reionization at this step\n    # nor at ionization at some pH (rdkit has no pKa caculator)\n    # the main aim to to represent all molecules from different sources\n    # in a (single) standard way, for use in ML, catalogue, etc.\n     \n    te = rdMolStandardize.TautomerEnumerator() # idem\n    taut_uncharged_parent_clean_mol = te.Canonicalize(uncharged_parent_clean_mol)\n     \n    return Chem.MolToSmiles(taut_uncharged_parent_clean_mol), taut_uncharged_parent_clean_mol\n\ndef addMainFragmentColumns(ds, smilesColumn='smiles', mainFragmentColumn=None, molCol='MainMOL') :\n    # returns an updated dataframe containing an additional column with the corresponding standardized main fragments\n    # and a column with the corresponding RDKit Molecule objects\n    # the RDKit molecule can be more easily visualized.\n    major_SMILES = []\n    major_MOL = []\n    for smi in ds[smilesColumn].values :\n        nsmi, nmol = standardize(smi)\n        major_SMILES.append(nsmi)\n        major_MOL.append(nmol)\n    if mainFragmentColumn is not None : ds[mainFragmentColumn] = major_SMILES\n    if molCol is not None : ds[molCol] = major_MOL\n    return ds\n\ndef sample_csv_by_reservoir_sampling(filename, sample_size=100, max_lines=None, lov_indices=None, protein_indexes=None, collect_activity_counts=True) :\n    # Uses reservoir sampling (https://en.wikipedia.org/wiki/Reservoir_sampling) to extract a random sample of sample_size rows\n    # from a potentially large CSV and returns it as a pandas data frame.\n    # It currently assumes pivoted data\n    # To save time, this method can also return the dataframe, an array with lists of values for columns with indices in lov_indices, and dictionary of BB activity counts\n    sample_array = []\n    if lov_indices:\n        lists_of_values = [[] for _ in range(len(lov_indices))]\n        bb_stats = [{} for _ in range(len(lov_indices))]\n    with open(filename, \"r\") as f:\n        reader = csv.reader(f, delimiter=\",\")\n        for i, line in enumerate(reader):\n            if i == 0 :\n                header = line\n                protein_names = [line[_] for _ in protein_indexes]\n                continue\n            if max_lines is not None and i >= max_lines : break\n            if i % 100000 == 0 : print ('.',end='')\n            if i % 10000000 == 0 : print ('\\t',i)\n            if lov_indices :\n                for j,idx in enumerate(lov_indices) : # loop over BB positions\n                    bb = line[idx]\n                    lists_of_values[j].append(bb)\n                    if collect_activity_counts :\n                        stats = bb_stats[j]\n                        if bb not in stats :\n                            stats[bb] = {\"totalActive\":0, \"totalAll\":0}\n                            for protein in protein_names :\n                                stats[bb][protein] = 0\n                                stats[bb][protein+\"All\"] = 0\n                        for ip,protein in zip(protein_indexes,protein_names) :\n                            activity = float(line[ip])\n                            if activity and activity > 0 :\n                                stats[bb][protein] += 1\n                                stats[bb][\"totalActive\"] += 1\n                            stats[bb][\"totalAll\"] += 1\n                            stats[bb][protein+\"All\"] += 1\n                    if i % 100 == 0 : lists_of_values[j] = list(set(lists_of_values[j])) # remove duplicates\n            if i <= sample_size : # initialize with first rows\n                sample_array.append(line)\n            else :\n                i_replace = random.randint(0, i-1)\n                if i_replace < sample_size :\n                    sample_array[i_replace] = line\n        print ('\\t',i)\n    # Now, we have a random sample of rows and the header of the data frame to be returned\n    if lov_indices :\n        for j in range(len(lov_indices)) : # final removal of duplicates\n            lists_of_values[j] = list(set(lists_of_values[j]))\n        if collect_activity_counts : # also return BB statistics Dataframe\n            return pd.DataFrame(sample_array, columns=header), lists_of_values, bb_stats\n        else :\n            return pd.DataFrame(sample_array, columns=header), lists_of_values\n    else :\n        return pd.DataFrame(sample_array, columns=header)\n\ndef drawIntructionsForMoleculeSamples(ds):\n    # Draws a handcrafted table of the building blocks and molecules in the given dataset. Uses a pseudo molecule for table heading\n    # Step throuch rows and collect cells for tabular representation\n    # The following is a hack to get a title row. Unfortunately, the PandasTools don't work reliably with pandas 2.2 and thus, I need handcraft a table\n    mols = [Chem.MolFromSmiles('[Rn]'), Chem.MolFromSmiles('[Rn]'), Chem.MolFromSmiles('[Rn]'), Chem.MolFromSmiles('[Rn]')] # Least intrusive molecule that still can carry a legend \n    legends = ['BB1','BB2','BB3','MOL']\n    highlightAtomLists = [None, None, None, None]\n    triazine = Chem.MolFromSmiles('c1ncncn1') # Scaffold from training set\n    suzuki = Chem.MolFromSmarts('OBO')        # BB indicating library generated by Suzuki coupling\n    for irow, (bb1, bb2, bb3, mol) in enumerate(zip(ds['BB1'].values,ds['BB2'].values,ds['BB3'].values,ds['MOL'].values)) :\n        mols.append(Chem.MolFromSmiles(bb1))\n        legends.append('')\n        highlightAtomLists.append(None)\n        #\n        mMol = Chem.MolFromSmiles(bb2)\n        mols.append(mMol)\n        match = mMol.GetSubstructMatch(suzuki)\n        highlightAtomLists.append(match)\n        if match : legends.append('Suzuki') # Label this building block as a Suzuki reactant\n        else : legends.append('')\n        #\n        mols.append(Chem.MolFromSmiles(bb3))\n        legends.append('')\n        highlightAtomLists.append(None)\n        mMol = Chem.MolFromSmiles(mol)\n        mols.append(mMol)\n        match = mMol.GetSubstructMatch(triazine)\n        if match : legends.append('Triazine') # Label this molecule as a triazine, like the ones from the training set\n        else : legends.append('Suzuki')\n        highlightAtomLists.append(match)\n\n    # return display components\n    return mols, legends, highlightAtomLists","metadata":{"execution":{"iopub.status.busy":"2024-05-07T22:22:09.587286Z","iopub.execute_input":"2024-05-07T22:22:09.587738Z","iopub.status.idle":"2024-05-07T22:22:09.840693Z","shell.execute_reply.started":"2024-05-07T22:22:09.587705Z","shell.execute_reply":"2024-05-07T22:22:09.839448Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Analyzing the Training Set\n\nWe use [Reservoir Sampling](https://en.wikipedia.org/wiki/Reservoir_sampling) to convert the input file <code>train.csv</code> into a Dataframe containing a random sample of 20 rows. The utility method also collects complete lists of unique values of the building block reagents to describe the actual dimensions.","metadata":{}},{"cell_type":"code","source":"import pickle\nif not os.path.exists(\"train.pickle\") :\n    startTime = time.time()\n    ds,lists_of_values,bb_stats = sample_csv_by_reservoir_sampling(\"/kaggle/input/pivot-input-files/pivoted_train.csv\", sample_size=2000, max_lines=None, lov_indices=[0,1,2], protein_indexes=[4,5,6], collect_activity_counts=True)\n    print (\"Sampling training set structures took\",time.time()-startTime,\"seconds\")\n\n    ds = addMainFragmentColumns(ds, smilesColumn='buildingblock1_smiles', mainFragmentColumn='BB1', molCol=None)\n    ds = addMainFragmentColumns(ds, smilesColumn='buildingblock2_smiles', mainFragmentColumn='BB2', molCol=None)\n    ds = addMainFragmentColumns(ds, smilesColumn='buildingblock3_smiles', mainFragmentColumn='BB3', molCol=None)\n    ds = addMainFragmentColumns(ds, smilesColumn='molecule_smiles', mainFragmentColumn='MOL')\n    # Save sample and statistics\n    ds.to_pickle(\"train.pickle\")\n    with open(\"BB1List.pickle\", \"wb\") as fp:\n        pickle.dump(lists_of_values[0], fp)\n    with open(\"BB2List.pickle\", \"wb\") as fp:\n        pickle.dump(lists_of_values[1], fp)\n    with open(\"BB3List.pickle\", \"wb\") as fp:\n        pickle.dump(lists_of_values[2], fp)\n    with open(\"bb_stats.pickle\", \"wb\") as fp:\n        pickle.dump(bb_stats, fp)\nelse :\n    ds = pd.read_pickle(\"train.pickle\")\n    lists_of_values = [None,None,None]\n    with open(\"BB1List.pickle\", \"rb\") as fp:\n        lists_of_values[0] = pickle.load(fp)\n    with open(\"BB2List.pickle\", \"rb\") as fp:\n        lists_of_values[1] = pickle.load(fp)\n    with open(\"BB3List.pickle\", \"rb\") as fp:\n        lists_of_values[2] = pickle.load(fp)\n    with open(\"bb_stats.pickle\", \"rb\") as fp:\n        bb_stats = pickle.load(fp)\n\nprint (\"BB1 has\",len(set(lists_of_values[0])),\"different building blocks\")\nprint (\"BB2 has\",len(lists_of_values[1]),\"different building blocks\")\nprint (\"BB3 has\",len(lists_of_values[2]),\"different building blocks\")\nds[['BB1','BB2', 'BB3', 'MOL', 'BRD4', 'HSA', 'sEH']]\n","metadata":{"execution":{"iopub.status.busy":"2024-05-07T22:22:19.694291Z","iopub.execute_input":"2024-05-07T22:22:19.694868Z","iopub.status.idle":"2024-05-07T22:47:00.914026Z","shell.execute_reply.started":"2024-05-07T22:22:19.694823Z","shell.execute_reply":"2024-05-07T22:47:00.912143Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The training set is spanned by <b>271</b> protected aminoacids for BB1, <b>693</b> amines for BB2, and <b>872</b> amines for BB3. The DEL setup forms a triazine connected to the DNA code via the acid of BB1 from these reagents.\n\nThe following cell displays the building block reagents as well as the final molecules from 25 of the sampled rows. The common triazine ring is highlighted in the MOL column. (Note: The current combination of RDKit and Pandas version makes it impossible to used Dataframe rendering here.)","metadata":{}},{"cell_type":"code","source":"mols, legends, highlightAtomLists = drawIntructionsForMoleculeSamples(ds)\nDraw.MolsToGridImage(mols, legends=legends, highlightAtomLists=highlightAtomLists, maxMols=104, molsPerRow=4, subImgSize=(350,250))","metadata":{"execution":{"iopub.status.busy":"2024-05-07T22:53:33.561491Z","iopub.execute_input":"2024-05-07T22:53:33.563121Z","iopub.status.idle":"2024-05-07T22:53:36.117689Z","shell.execute_reply.started":"2024-05-07T22:53:33.563056Z","shell.execute_reply":"2024-05-07T22:53:36.115768Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 'Failed' and 'Promiscuous' Building Block\n\nThe <code>sample_csv_by_reservoir_sampling()</code> method also returned statistics of the number of times a BB was member of a hit for a certain target. The next cell displays for each building block position the 6 least and 6 most active building blocks. The least frequently active one could have had problems during library synthesis and could, thus, be considered possible 'failures'. The most frequently active one could be called promiscuous as they might be the single reason for binding independent of the other fragments.\n\nGiven the number of training rows, BB1 has 320 K molecules per building block alternative, BB2 has 140 K, and BB3 has 110 K.\n\nThe most promiscuous BB at position 1 hits at least one of the targets for 80 % of the molecules containg it, while the poorest alternative only hits 0.00375 % and only for HSA. BB position 1 seems to have problems for aromatic building blocks, which might either indicate a synthesis issue or a negative SAR for aromatics close to the DNA-tag. Interestingly, the second most promiscuous BB reagent on this position is actually a stereospecific version of the most promiscuous one. The reason might be that it was used as a diastereomeric mixture of the cis- and trans-isomer and thus might hit more frequently since only one of the two iseomers needs to bind sEH to count as a hit. In addition, those two building block reagents are 4.5 times more frequent hitters than the other alternatives.\n\nReviewing the other examples, it also looks as if stereochemistry is either not reported or not known for many of the reagents. \n\nThe most promiscuous BB at position 2 hits at least one of the targets for 40 % of the molecules containg it, while the poorest alternative only hits 0.044 % for BRD4. BB position 2 has a rather diverse and also not very pronounced set of failures and promiscuous reagents.\n\nThe most promiscuous BB at position 3 hits at least one of the targets for 30 % of the molecules containg it, while the poorest alternative hits no target with any BB1/BB2 partners in the training set. So BB position 3 has a few problem reagents and no really promiscuous ones.\n\nThese observations might indicate that for some building blocks the specifics of the triazine library formation might interfere with any structural model translating to another library with a different chemistry.","metadata":{}},{"cell_type":"code","source":"PER_ROW = 4\nmols = []\nlegends = []\nfor i in range(3) :\n    for TARGET in [\"sEH\",\"BRD4\",\"HSA\"] :\n        pos = \"BB\"+str(i+1)+\":\"\n        stat = bb_stats[i]\n        counted_keys = [(_,stat[_][TARGET]) for _ in stat.keys() if TARGET in stat[_]]\n        counted_keys = sorted(counted_keys, key = lambda x: x[1])\n        for bb,cnt in counted_keys[:PER_ROW] :\n            mols.append(Chem.MolFromSmiles(bb))\n            legends.append(pos + \": \" + str({\"total\":stat[bb][\"totalActive\"], TARGET:stat[bb][TARGET], \"PercentActive\":'{:.4%}'.format(((stat[bb][TARGET])/stat[bb][TARGET+\"All\"]))}))\n            # legends.append(pos + \": \" + str(stat[bb]))\n        for bb,cnt in counted_keys[-PER_ROW:] :\n            mols.append(Chem.MolFromSmiles(bb))\n            legends.append(pos + \": \" + str({\"total\":stat[bb][\"totalActive\"], TARGET:stat[bb][TARGET], \"PercentActive\":'{:.2%}'.format(((stat[bb][TARGET])/stat[bb][TARGET+\"All\"]))}))\nDraw.MolsToGridImage(mols, legends=legends, maxMols=100, molsPerRow=PER_ROW, subImgSize=(450,250))","metadata":{"execution":{"iopub.status.busy":"2024-05-07T22:54:04.873032Z","iopub.execute_input":"2024-05-07T22:54:04.874464Z","iopub.status.idle":"2024-05-07T22:54:06.078674Z","shell.execute_reply.started":"2024-05-07T22:54:04.874406Z","shell.execute_reply":"2024-05-07T22:54:06.077430Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Analyzing the Test Set\n\nNow, we also process and sample the test set. It turns out to contain not just triazines with possibly new building blocks, but also a sample from a new library created via [Suzuki Coupling](https://en.wikipedia.org/wiki/Suzuki_reaction). This makes scoring the test set much more challenging since any model cannot be based on the building block fragments (the pieces of the building block reagent that end up in the final product) but needs to generalize to arbitrary molecules. As noted above, 'failed' and 'promiscuous' building blocks might also interfere with the generalization of a model to the unseen library.","metadata":{}},{"cell_type":"code","source":"startTime = time.time()\ndsTest,lists_of_values_test = sample_csv_by_reservoir_sampling(\"/kaggle/input/pivot-input-files/pivoted_test.csv\", sample_size=2000, max_lines=None, lov_indices=[0,1,2], protein_indexes=[4,5,6], collect_activity_counts=False)\nprint (\"Sampling test set structures took\",time.time()-startTime,\"seconds\")\n\nstartTime = time.time()\nsuzuki = Chem.MolFromSmarts('OBO')\ndsTest[\"isSuzuki\"] = [\"B\" in _ and Chem.MolFromSmiles(_).HasSubstructMatch(suzuki) for _ in dsTest[\"buildingblock2_smiles\"]] # Check if BB2 is a Suzuki coupling reagent, string needs to contain a 'B' to match\nprint (\"Adding Suzuki label took\",time.time()-startTime,\"seconds\")\n\nstartTime = time.time()\ndsTest = addMainFragmentColumns(dsTest, smilesColumn='buildingblock1_smiles', mainFragmentColumn='BB1', molCol=None) # standardized column w/o counter ions\ndsTest[\"New1\"] = [_ not in bb_stats[0] for _ in dsTest['buildingblock1_smiles'].values]\ndsTest = addMainFragmentColumns(dsTest, smilesColumn='buildingblock2_smiles', mainFragmentColumn='BB2', molCol=None) # standardized column w/o counter ions\ndsTest[\"New2\"] = [_ not in bb_stats[1] for _ in dsTest['buildingblock1_smiles'].values]\ndsTest = addMainFragmentColumns(dsTest, smilesColumn='buildingblock3_smiles', mainFragmentColumn='BB3', molCol=None) # standardized column w/o counter ions\ndsTest[\"New3\"] = [_ not in bb_stats[2] for _ in dsTest['buildingblock1_smiles'].values]\ndsTest = addMainFragmentColumns(dsTest, smilesColumn='molecule_smiles', mainFragmentColumn='MOL') # probably redundant\ndsTest[\"#New\"] = (0+dsTest['New1'].values)+dsTest['New2'].values+dsTest['New3'].values\nprint (\"Checking for new building block reagents label took\",time.time()-startTime,\"seconds\")\n\nprint (\"BB1 has\",len(lists_of_values_test[0]),\"different building blocks of which\", len(set(lists_of_values_test[0]).intersection(lists_of_values[0])), \"are also in the training set\")\nprint (\"BB2 has\",len(lists_of_values_test[1]),\"different building blocks of which\", len(set(lists_of_values_test[1]).intersection(lists_of_values[1])), \"are also in the training set\")\nprint (\"BB3 has\",len(lists_of_values_test[2]),\"different building blocks of which\", len(set(lists_of_values_test[2]).intersection(lists_of_values[2])), \"are also in the training set\")\nprint (\"Test rows with 1 new building block:\",(dsTest[\"#New\"]==1).sum())\nprint (\"Test rows with 2 new building blocks:\",(dsTest[\"#New\"]==2).sum())\nprint (\"Test rows with 3 new building blocks:\",(dsTest[\"#New\"]==3).sum())\nprint (\"Test rows from Suzuki library:\",(dsTest[\"isSuzuki\"]==True).sum())\nprint (\"Test rows from Triazine library:\",(dsTest[\"isSuzuki\"]==False).sum())\n\ndsTest[['buildingblock1_smiles','buildingblock2_smiles', 'buildingblock3_smiles', 'molecule_smiles', \"New1\", \"New2\", \"New3\", \"#New\", \"isSuzuki\", 'BRD4', 'HSA', 'sEH']]","metadata":{"execution":{"iopub.status.busy":"2024-05-07T22:54:25.647763Z","iopub.execute_input":"2024-05-07T22:54:25.648174Z","iopub.status.idle":"2024-05-07T22:54:55.236066Z","shell.execute_reply.started":"2024-05-07T22:54:25.648141Z","shell.execute_reply":"2024-05-07T22:54:55.234485Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The test set contains <b>341</b> building block reagents BB1, <b>1140</b> at BB2, and <b>1390</b>, which most likely form disjoint subsets for the triazine and Suzuki designs.","metadata":{}},{"cell_type":"code","source":"mols, legends, highlightAtomLists = drawIntructionsForMoleculeSamples(dsTest)\nDraw.MolsToGridImage(mols, legends=legends, highlightAtomLists=highlightAtomLists, maxMols=100, molsPerRow=4, subImgSize=(350,250))","metadata":{"execution":{"iopub.status.busy":"2024-05-07T22:55:32.999121Z","iopub.execute_input":"2024-05-07T22:55:32.999602Z","iopub.status.idle":"2024-05-07T22:55:35.261374Z","shell.execute_reply.started":"2024-05-07T22:55:32.999568Z","shell.execute_reply":"2024-05-07T22:55:35.256382Z"},"trusted":true},"execution_count":null,"outputs":[]}]}