{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Intro\nWelcome to the [Bristol-Myers Squibb – Molecular Translation](https://www.kaggle.com/c/bms-molecular-translation/overview) Competition:\n\n![](https://storage.googleapis.com/kaggle-competitions/kaggle/22422/logos/header.png)\n\nFor informations about the International Chemical Identifier we recommend this [link](https://en.wikipedia.org/wiki/International_Chemical_Identifier).\n\n<span style=\"color: royalblue;\">Please vote the notebook up if it helps you. Feel free to leave a comment above the notebook. Thank you. </span>","metadata":{}},{"cell_type":"markdown","source":"# Libraries","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport os\nimport matplotlib.pyplot as plt\nimport cv2\nimport random\nimport re\n\nimport warnings\nwarnings.filterwarnings(\"ignore\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Path","metadata":{}},{"cell_type":"code","source":"path = '/kaggle/input/bms-molecular-translation/'\nos.listdir(path)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Functions\nWe define some helper functions for visualizations.","metadata":{}},{"cell_type":"code","source":"def plot_example(image_id):\n    fig = plt.figure(figsize=(12, 7))\n    ax = fig.add_subplot(111)\n    filename = train_data.loc[image_id, 'image_id']\n    path_img ='/'.join([path, 'train', filename[0], filename[1], filename[2]])\n    path_img = path_img+'/'\n    img = cv2.imread(path_img+filename+'.png')\n    ax.imshow(img)\n    ax.set_title(train_data.loc[image_id, 'InChI'])\n    ax.set_xticklabels([])\n    ax.set_yticklabels([])\n    plt.show()\n\ndef plot_examples(list_IDs):\n    fig, axs = plt.subplots(5, 5, figsize=(25, 14))\n    fig.subplots_adjust(hspace = .2, wspace=.2)\n    axs = axs.ravel()\n    for i in range(25):\n        filename = train_data.loc[list_IDs[i], 'image_id']\n        path_img ='/'.join([path, 'train', filename[0], filename[1], filename[2]])\n        path_img = path_img+'/'\n        img = cv2.imread(path_img+filename+'.png')\n        axs[i].imshow(img)\n        axs[i].set_title(train_data.loc[list_IDs[i], 'InChI'][0:20]+'...')\n        axs[i].set_xticklabels([])\n        axs[i].set_yticklabels([])\n    plt.show()","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load Data","metadata":{}},{"cell_type":"code","source":"train_data = pd.read_csv(path+'train_labels.csv')\nsamp_subm = pd.read_csv(path+'sample_submission.csv')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Overview","metadata":{}},{"cell_type":"code","source":"print('Number train samples:', len(train_data.index))\nprint('Number submission samples:', len(samp_subm.index))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Exploratory Data Analysis","metadata":{}},{"cell_type":"markdown","source":"## Images\nWe plot some images and a part of the InChi string as title:","metadata":{}},{"cell_type":"code","source":"list_IDs = random.sample(list(train_data.index), 25)\nplot_examples(list_IDs)","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Labels\nWe count the lenght of the labels. Every label starts with the string InChI=, so we subtract 6 of the whole label lenght:","metadata":{}},{"cell_type":"code","source":"def lenght_label(s):\n    return len(s)-6\n\ntrain_data['lenght_label'] = train_data['InChI'].apply(lenght_label)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = plt.figure(figsize=(8, 5))\ntrain_data['lenght_label'].hist(bins=100)\nplt.title('Distribution of label lenght', loc='left')\nplt.xlabel('Lenght of label')\nplt.ylabel('Frequency')\nplt.show()","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The labels consists of layers and sublaysers which are aseparated by the delimiter \"/\" and start with a characteristic prefix letter.\n\nThe six layers with important sublayers are:\n\n1) Main layer\n\n* Chemical formula (no prefix). This is the only sublayer that must occur in every InChI.\n* Atom connections (prefix: \"c\"). The atoms in the chemical formula (except for hydrogens) are numbered in sequence; this sublayer describes which atoms are connected by bonds to which other ones.\n* Hydrogen atoms (prefix: \"h\"). Describes how many hydrogen atoms are connected to each of the other atoms.\n\n2) Charge layer\n   \n* charge sublayer (prefix: \"q\")\n* proton sublayer (prefix: \"p\" for \"protons\")\n\n3) Stereochemical layer\n   \n* double bonds and cumulenes (prefix: \"b\")\n* tetrahedral stereochemistry of atoms and allenes (prefixes: \"t\", \"m\")\n* type of stereochemistry information (prefix: \"s\")\n4) Isotopic layer (prefixes: \"i\", \"h\", as well as \"b\", \"t\", \"m\", \"s\" for isotopic stereochemistry)\n\n5) Fixed-H layer (prefix: \"f\"); contains some or all of the above types of layers except atom connections; may end with \"o\" sublayer; never included in standard InChI\n\n6) Reconnected layer (prefix: \"r\"); contains the whole InChI of a structure with reconnected metal atoms; never included in standard InChI","metadata":{}},{"cell_type":"code","source":"def number_layer(s):\n    return len(s.split('/'))-1\n\ntrain_data['number_layer'] = train_data['InChI'].apply(number_layer)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = plt.figure(figsize=(8, 5))\ntrain_data['number_layer'].hist(bins=12)\nplt.title('Distribution number of layers')\nplt.xlabel('Number layer')\nplt.ylabel('Frequency')\nplt.show()","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Focus On Examples\nFor molecular informations we recommend the [page](https://pubchem.ncbi.nlm.nih.gov/). Next we consider 3 examples with different number of layers.","metadata":{}},{"cell_type":"markdown","source":"### 1-[5-methyl-2-(2-methylpropylsulfanyl)phenyl]ethanol\nWe focus on the first train data sample. For more informations look [here](https://pubchem.ncbi.nlm.nih.gov/compound/82265033).","metadata":{}},{"cell_type":"code","source":"plot_example(0)","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.loc[0, 'InChI']","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This is the topology of the example: <br>\n1) Main Layer<br>\n* Formular C<sub>13</sub>H<sub>20</sub>OS\n* Atom connections: c1-9(2)8-15-13-6-5-10(3)7-12(13)11(4)14\n* Hydrogen atoms: h5-7,9,11,14H,8H2,1-4H3","metadata":{}},{"cell_type":"markdown","source":"### (E)-3-(2-Bromophenyl)-2-(4-oxo-4aH-quinazolin-2-yl)prop-2-enenitrile\nWe focus on example on index 6. For more informations look [here](https://pubchem.ncbi.nlm.nih.gov/compound/133560351).","metadata":{}},{"cell_type":"code","source":"plot_example(6)","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\ntrain_data.loc[6, 'InChI']","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This is the topology of the example: <br>\n1) Main Layer<br>\n* Formular C<sub>17</sub>H<sub>10</sub>BrN<sub>3O</sub>\n* Atom connections: c18-14-7-3-1-5-11(14)9-12(10-19)16-20-15-8-4-2-6-13(15)17(22)21-16\n* Hydrogen atoms: h1-9,13H\n\n3) Stereochemical layer\n   \n* double bonds and cumulenes: b12-9+","metadata":{}},{"cell_type":"markdown","source":"### N-[(1R,2Z)-2-[(3Ar,5S,6aR)-5-[(4R)-2,2-dimethyl-1,3-dioxolan-4-yl]-2,2-dimethyl-3a,6a-dihydrofuro[2,3-d][1,3]dioxol-6-ylidene]-1-deuterioethyl]-2,2,2-trichloroacetamide\n\nWe focus on example on index 774,948 with 10 layers. For more informations look [here](https://pubchem.ncbi.nlm.nih.gov/compound/134870524).","metadata":{}},{"cell_type":"code","source":"plot_example(774948)","metadata":{"_kg_hide-output":false,"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.loc[774948, 'InChI']","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This is the topology of the example: <br>\n1) Main Layer<br>\n* Formular C<sub>16</sub>H<sub>22</sub>Cl<sub>3</sub>NO<sub>6</sub>\n* Atom connections: c1-14(2)22-7-9(24-14)10-8(5-6-20-13(21)16(17,18)19)11-12(23-10)26-15(3,4)25-11\n* Hydrogen atoms: h5,9-12H,6-7H2,1-4H3,(H,20,21)\n\n3) Stereochemical layer\n   \n* double bonds and cumulenes: b8-5-\n* tetrahedral stereochemistry of atoms and allenes 1: t9-,10+,11-,12-\n* tetrahedral stereochemistry of atoms and allenes 1: m1\n* type of stereochemistry information: s1\n\n4) Isotopic layer\n\n* #1: i6D\n* #2: t6-,9+,10-,11+,12+\n* #3: m0","metadata":{}},{"cell_type":"markdown","source":"# Main Layer - Formular\nThe formular layer must occure in every InChI. To focus on the formular we extract the first layer.","metadata":{}},{"cell_type":"code","source":"def split_formular_layer(s):\n    return s.split('/')[1]\n\ntrain_data['formular'] = train_data['InChI'].apply(split_formular_layer)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are a lot of duplicates:","metadata":{}},{"cell_type":"code","source":"train_data['formular'].value_counts()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In total there are 329,768 different formulars.\n\nThe next step is to split the formular into the atoms by there symbol and number of atomes.","metadata":{}},{"cell_type":"code","source":"def split_formular(formular):\n    dict_formular = {k: int(v) if v else 1 for k,v in\n                     re.findall(r\"([A-Z][a-z]?)(\\d+)?\", formular)}\n    return dict_formular\n\n# Test the function\ntest_formular = train_data.loc[0, 'formular']\nprint(test_formular)\nsplit_formular(test_formular)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_formular = pd.DataFrame(columns = ['C', 'H', 'Br','Cl','I', 'F', 'N', 'O', 'S', 'Si', 'P'])\nlist_formulars = list(train_data['formular'].value_counts().keys())\nfor formular in list_formulars[0:100]:\n    dict_formular = split_formular(formular)\n    temp = pd.DataFrame.from_dict(dict_formular, orient='index').T\n    df_formular = pd.concat([df_formular, temp])\ndf_formular.index=list_formulars[0:100]\ndf_formular.fillna(0, inplace=True)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_formular.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To be continued ...","metadata":{}},{"cell_type":"markdown","source":"# Data Generator\nWe define a data generator to load the image data on demand.\n\n*Comming Soon*","metadata":{}},{"cell_type":"markdown","source":"# Export","metadata":{}},{"cell_type":"code","source":"output = samp_subm\noutput.to_csv('submission.csv', index=False)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"output.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}