{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# SUMMARY\n\nThis notebook reproduces my submission that scores **1.31** on the private LB and reaches the **47th place**. The notebook implements ensemble of CNN-LSTM models using model predictions saved as Kaggle datasets. \n- solution summary is published [in this discussion topic](https://www.kaggle.com/c/bms-molecular-translation/discussion/243845)\n- complete training codes are available [in this GitHub repo](https://github.com/kozodoi/BMS_Molecular_Translation)\n\nThe table with the main model parameters and CV performance (before beam searchg and normalization) is provided below.\n![models](https://i.postimg.cc/cLrTp1Pc/Screen-2021-06-04-at-10-17-02.jpg)","metadata":{}},{"cell_type":"markdown","source":"# PREPARATIONS","metadata":{}},{"cell_type":"code","source":"##### LIBRARIES\n\nimport numpy as np \nimport pandas as pd \nfrom tqdm import tqdm ","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2021-06-08T09:57:14.100917Z","iopub.execute_input":"2021-06-08T09:57:14.101578Z","iopub.status.idle":"2021-06-08T09:57:14.108093Z","shell.execute_reply.started":"2021-06-08T09:57:14.101435Z","shell.execute_reply":"2021-06-08T09:57:14.106589Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Below we define a list with seven base models. For each of these models, test molecule predictions are stored as Kaggle datasets after beam-search with `k = 5` (thanks @tugstugi) and `RDKit`-based normalization (thanks @nofreewill). The models are sorted by their performance after beam search and normalization in the ascending order (the first model performs best).","metadata":{}},{"cell_type":"code","source":"##### BASE MODELS\n\nmodel_list = ['/kaggle/input/bms-norm-v22/submission_norm.csv',\n              '/kaggle/input/bms-norm-v17/submission_norm.csv',\n              '/kaggle/input/bms-normalization-v21/submission_norm.csv',\n              '/kaggle/input/bms-normalization-v2733/submission_norm.csv',\n              '/kaggle/input/bms-normalization-v20/submission_norm.csv',\n              '/kaggle/input/bms-normalization-v6/submission_norm.csv',\n              '/kaggle/input/bms-normalization-public/submission_norm.csv']","metadata":{"execution":{"iopub.status.busy":"2021-06-08T09:57:14.114739Z","iopub.execute_input":"2021-06-08T09:57:14.115399Z","iopub.status.idle":"2021-06-08T09:57:14.124783Z","shell.execute_reply.started":"2021-06-08T09:57:14.115344Z","shell.execute_reply":"2021-06-08T09:57:14.123571Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"##### PREPARE MODEL PREDICTIONS\n\nmodels = []\n\nfor model in model_list:\n    sub = pd.read_csv(model)\n    sub = sub.sort_values('image_id').reset_index(drop = True)\n    print('- {}: {}'.format(model, sub.shape))\n    models.append(sub)","metadata":{"execution":{"iopub.status.busy":"2021-06-08T09:57:14.127097Z","iopub.execute_input":"2021-06-08T09:57:14.127487Z","iopub.status.idle":"2021-06-08T09:58:29.016489Z","shell.execute_reply.started":"2021-06-08T09:57:14.127449Z","shell.execute_reply":"2021-06-08T09:58:29.015459Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I also import another model that produces partial predictions for 273k molecules that proved to be more challenging to translate in my previous experiments. The predictions are done using a beam search with a larger `k`.","metadata":{}},{"cell_type":"code","source":"##### IMPORT PARTIAL PREDICTIONS\n\npart_sub = pd.read_csv('/kaggle/input/bms-normalization-bad-27/submission_norm.csv')\npart_sub = part_sub.sort_values('image_id').reset_index(drop = True)\nprint(part_sub.shape)","metadata":{"execution":{"iopub.status.busy":"2021-06-08T09:58:29.017824Z","iopub.execute_input":"2021-06-08T09:58:29.018102Z","iopub.status.idle":"2021-06-08T09:58:30.878199Z","shell.execute_reply.started":"2021-06-08T09:58:29.018075Z","shell.execute_reply":"2021-06-08T09:58:30.877364Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Following the idea of @nofreewill, I store three possible events in the process of InChI normalization: `['valid', 'none', 'error']`. Value `valid` means that RDKit was able to convert prediction to a valid InChI string.  ","metadata":{}},{"cell_type":"code","source":"##### CHECK PREDICTION FORMAT\n\nsub = models[0].copy()\ndisplay(sub.head())\nprint('\\nEvents:')\ndisplay(sub['event'].value_counts())","metadata":{"execution":{"iopub.status.busy":"2021-06-08T09:58:30.880199Z","iopub.execute_input":"2021-06-08T09:58:30.880520Z","iopub.status.idle":"2021-06-08T09:58:31.410835Z","shell.execute_reply.started":"2021-06-08T09:58:30.880489Z","shell.execute_reply":"2021-06-08T09:58:31.409655Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# ENSEMBLING\n\nThe ensembling is done in the following way:\n1. If 4/7 models have the same output, set the final prediction to this mode value.\n2. Else:\n    - go through each model starting from the best-performing one and set prediction to a first model with valid output\n    - if there are now valid outputs, set prediction to the output of the lowest-CV model\n    - if output of the partial model is available and valid, overwrite prediction for this molecule","metadata":{}},{"cell_type":"code","source":"##### ENSEMBLING\n\n# placeholders\nnum_equals  = []\nmodel_preds = []\n\n# loop through test molecules\nfor i in tqdm(range(len(sub))):\n    \n    # extract base model predictions and mode\n    preds     = [model.iloc[i]['InChI'] for model in models]\n    mode      = max(set(preds), key = preds.count)\n    num_equal = preds.count(mode)\n    num_equals.append(num_equal)\n    \n    # set prediction to mode\n    if num_equal >= 4:\n        sub.loc[i, 'InChI'] = mode\n        model_preds.append('mode')\n        \n    else:\n        \n        # look for valid pred from all models\n        valid_pred = False\n        for m in range(len(models)):\n            if models[m].loc[i, 'event'] == 'valid':\n                sub.loc[i, 'InChI'] = models[m].loc[i, 'InChI']\n                model_preds.append(model_list[m])\n                valid_pred = True\n                break\n                \n        # set preds to lowest-CV model\n        if not valid_pred:\n            sub.loc[i, 'InChI'] = models[0].loc[i, 'InChI']\n            model_preds.append(model_list[0])\n                \n        # set preds to better model if possible\n        if not valid_pred:\n            image_id = sub.loc[i, 'image_id']\n            if image_id in list(part_sub['image_id'].values):\n                if part_sub.loc[part_sub['image_id'] == image_id]['event'].item() == 'valid':\n                    sub.loc[i, 'InChI'] = part_sub.loc[part_sub['image_id'] == image_id, 'InChI'].item()\n                    model_preds.append('part_model')","metadata":{"execution":{"iopub.status.busy":"2021-06-08T09:58:31.412654Z","iopub.execute_input":"2021-06-08T09:58:31.412983Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"##### CHECK NO. EQUAL PREDS DISTRIBUTION\n\npd.Series(num_equals).value_counts()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In many cases, our models agree quite well with each other.","metadata":{}},{"cell_type":"code","source":"##### CHECK MODEL PREDS DISTRIBUTION\n\npd.Series(model_preds).value_counts()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As expected, most predictions are coming from model `v22`. Very few molecules are translated by the last models in a row, suggesting that usually at least one of the models is able to provide valid edictions.","metadata":{}},{"cell_type":"markdown","source":"# SUBMISSION","metadata":{}},{"cell_type":"code","source":"##### EXPORT SUBMISSION\n\nsub = sub[['image_id', 'InChI']]\nsub.to_csv('submission.csv', index = False)\nsub.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}