{"cells":[{"metadata":{"trusted":true},"cell_type":"code","source":"!conda install -y -c rdkit rdkit","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Generating New Data\n\nThis notebook shows how to create additional image/inchi pairs using RDKit"},{"metadata":{"trusted":true},"cell_type":"code","source":"import numpy as np\nfrom skimage import color\nimport matplotlib.pyplot as plt\nfrom rdkit import Chem\nfrom rdkit.Chem import Draw\nfrom rdkit.Chem.Draw import IPythonConsole\nfrom PIL import Image\nimport io","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Here we load the molecule as a smiles string and convert to an inchi"},{"metadata":{"trusted":true},"cell_type":"code","source":"smile = 'Fc(c1)ccc-2c1C(=O)N(C)Cc3n2cnc3C(=O)OCC'\nmol = Chem.MolFromSmiles(smile)\ninchi = Chem.inchi.MolToInchi(mol)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"inchi","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We can then use RDKit to generate a black and white image similar to the dataset"},{"metadata":{"trusted":true},"cell_type":"code","source":"IPythonConsole.drawOptions.useBWAtomPalette()\nim = Draw.MolsToGridImage([mol], molsPerRow=1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"im = Image.open(io.BytesIO(im.data))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"im","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"im.save('new_im.png')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"With a little post-processing, the compound image could be made to look more like the scanned images in the dataset.\n\nThe ability to generate new inchi/image pairs adds an interesting element to this competition. A sufficiently motivated person could download 1 billion compounds from the [Zinc Database](https://zinc.docking.org/tranches/home/) to create a massive dataset of image/inchi pairs.\n\nEven if the generated data doesn't quite match the challenge data, the sheer volume of data that can be created allows for pre-training a model on the generated data before fine-tuning on the actual challenge data"},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}