{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Introduction\n\n**Main Topic**\n\nThis notebook is for **Generate Molecular Image using [PubChem](https://pubchem.ncbi.nlm.nih.gov/)** \n\n**References**\n\n[**PubChem Official Docs**](https://pubchemdocs.ncbi.nlm.nih.gov/about)\n\n[**Generate SMILES Molecular Image(Korean)**](https://dacon.io/competitions/official/235640/codeshare/1630?dtype=recent)\n"},{"metadata":{},"cell_type":"markdown","source":"# Install PubChem from scratch"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"!conda install -y -c rdkit rdkit;","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Download PubChem Compound ID(CDI) for InChI\n\nWe can download tons of Molecular Images from https://ftp.ncbi.nlm.nih.gov/pubchem\n\nThere are index of ftp at /pubchem/Compound/Extras, and I'm going to download **CID-InChI-Key.gz**\n\n![](https://drive.google.com/uc?export=view&id=1kgOTcGQnZFchzyQvZV4HbVdxEtf5uME2)"},{"metadata":{"trusted":true},"cell_type":"code","source":"! wget https://ftp.ncbi.nlm.nih.gov/pubchem/Compound/Extras/CID-InChI-Key.gz","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Set up environment¶"},{"metadata":{"trusted":true},"cell_type":"code","source":"import cv2\nimport os\nimport gzip\nimport numpy as np\nimport pandas as pd\nfrom tqdm import tqdm\nimport matplotlib.pyplot as plt\n\n\nimport rdkit\nfrom rdkit import Chem\nfrom rdkit.Chem import Draw","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Open gz file\n\nThere are tons of InChI Component in **CID-InChI-Key.gz** so I'll just extract 500 Components from on it.\n\n**- note -**\n\nComponents are stored with similar components in order, So It would be better to select randomly if you want to use this datasets  "},{"metadata":{"trusted":true},"cell_type":"code","source":"length = 500\nwith gzip.open('CID-InChI-Key.gz', 'r') as InChIs:\n    data = [InChIs.readline().decode() for _ in tqdm(range(length))]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"Chem.MolFromInchi(data[0].split('\\t')[1])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"InChI_dict = {'InChI':[]}\nfor i, d in tqdm(enumerate(data)):\n    InChI = d.split('\\t')[1]\n    m = Chem.MolFromInchi(InChI)\n    if m != None:\n        InChI_dict['InChI'].append(InChI)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Generate DataFrame"},{"metadata":{"trusted":true},"cell_type":"code","source":"save_path = './images/'\ntrain = pd.DataFrame(InChI_dict)\ntrain['file_name'] = 'train_' + train.index.astype('str') + '.png'\ntrain['file_path'] = save_path + train['file_name']\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Save Images"},{"metadata":{"trusted":true},"cell_type":"code","source":"if not (os.path.isdir(save_path)):\n    os.makedirs(os.path.join(save_path))\n    \nfor idx, row in tqdm(train.iterrows()):\n    file = row['file_path']\n    InChI = row['InChI']\n    m = Chem.MolFromInchi(InChI)\n    if m != None:\n        img = Draw.MolToImage(m, size=(300,300))\n        img.save(file)    ","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"image_paths = train.file_path","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Show Images"},{"metadata":{"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(20, 18))\nfor i in range(20):\n    img = cv2.imread(image_paths[i])\n    plt.subplot(5, 4, i+1)\n    plt.imshow(img)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Next Step\n\nAs we know, Competition dataset is low resolution images.\nSo it would be better to matching resolution using Image Processing.\n\n-low resolution --> high resolution\n\n-high resolution --> low resolution\n\nI don't know which one is better now. But we can figure it out :)\n\nHope to be helpful this NB."}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}