{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Hello fellow Kagglers,\n\nGetting in the medal zone of the leaderboard was a challenge for me, as I was constantly positioned ~125.\n\nNormalizing the prediction gave a nice LB improvement of 0.12 (2.12->2.00), however this was not enough to get in the medal zone.\nNormalizing the predicted InChI's was adapted from [this](https://www.kaggle.com/wuliaokaola/bmsmt-0331-normalize-your-predictions) notebook.\n\nIn the last few days of the competition I was inspired by [this](https://www.kaggle.com/c/bms-molecular-translation/discussion/242082) discussion about assembling techniques.\n\nI started working on this notebook and got a massive LB improvement of 0.19, which was enough to get me in the medal zone!\n\nHope you enjoy this notebook and see you in the next competition :D","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\n\nimport re\n\nfrom tqdm.notebook import tqdm\n\n# initialize pandas apply progress bar\ntqdm.pandas()","metadata":{"execution":{"iopub.status.busy":"2021-06-04T08:15:10.692148Z","iopub.execute_input":"2021-06-04T08:15:10.692524Z","iopub.status.idle":"2021-06-04T08:15:10.698897Z","shell.execute_reply.started":"2021-06-04T08:15:10.692483Z","shell.execute_reply":"2021-06-04T08:15:10.697895Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"5 different submission are used in this assemble, which are all submission from 1 model on different training checkpoints. The concept behind assembling is simple, training a model will improve the overall score, however there will be mistakes introduced which weren't made earlier. I.e. there will exists prediction in the 2.12 LB submission which are predicted correctly in the 2.21 LB submission. Those are the prediction we are interested in and want to find.","metadata":{}},{"cell_type":"code","source":"# names of the submission files, which indicate the LB score\nNAMES = ['212_v2', '212', '218', '219', '221']","metadata":{"execution":{"iopub.status.busy":"2021-06-04T08:15:10.700237Z","iopub.execute_input":"2021-06-04T08:15:10.700703Z","iopub.status.idle":"2021-06-04T08:15:10.712709Z","shell.execute_reply.started":"2021-06-04T08:15:10.700659Z","shell.execute_reply":"2021-06-04T08:15:10.711692Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Read Submissions","metadata":{}},{"cell_type":"markdown","source":"Here all submission files are read and the InChI column is renamed accordingly to the LB score. The $N$ indicates the InChI's are normalized. Throughout this notebook the predicted InChI will refer to the InChI predicted by the model and the normalized InChI will refer to the normalized predicted InChI.","metadata":{}},{"cell_type":"code","source":"# debugging variable to select a subset of each submission\nn = int(10e10)\n\nprint('Reading 212 v2...')\nsubmission_212_v2 = pd.read_csv('../input/bms-submissions/submission_2.12_v2.csv').head(n)\nsubmission_212_v2.rename({ 'InChI': 'InChI212_v2' }, axis=1, inplace=True)\nsubmission_212_v2N = pd.read_csv('../input/bms-submissions/submission_norm_2.12_v2.csv').head(n)\nsubmission_212_v2N.rename({ 'InChI': 'InChI212_v2N' }, axis=1, inplace=True)\n\nprint('Reading 212...')\nsubmission_212 = pd.read_csv('../input/bms-submissions/submission_2.12.csv').head(n)\nsubmission_212.rename({ 'InChI': 'InChI212' }, axis=1, inplace=True)\nsubmission_212N = pd.read_csv('../input/bms-submissions/submission_norm_2.12.csv').head(n)\nsubmission_212N.rename({ 'InChI': 'InChI212N' }, axis=1, inplace=True)\n\nprint('Reading 218...')\nsubmission_218 = pd.read_csv('../input/bms-submissions/submission_2.18.csv').head(n)\nsubmission_218.rename({ 'InChI': 'InChI218' }, axis=1, inplace=True)\nsubmission_218N = pd.read_csv('../input/bms-submissions/submission_norm_2.18.csv').head(n)\nsubmission_218N.rename({ 'InChI': 'InChI218N' }, axis=1, inplace=True)\n\nprint('Reading 219...')\nsubmission_219 = pd.read_csv('../input/bms-submissions/submission_2.19.csv').head(n)\nsubmission_219.rename({ 'InChI': 'InChI219' }, axis=1, inplace=True)\nsubmission_219N = pd.read_csv('../input/bms-submissions/submission_norm_2.19.csv').head(n)\nsubmission_219N.rename({ 'InChI': 'InChI219N' }, axis=1, inplace=True)\n\nprint('Reading 221...')\nsubmission_221 = pd.read_csv('../input/bms-submissions/submission_2.21.csv').head(n)\nsubmission_221.rename({ 'InChI': 'InChI221' }, axis=1, inplace=True)\nsubmission_221N = pd.read_csv('../input/bms-submissions/submission_norm_2.21.csv').head(n)\nsubmission_221N.rename({ 'InChI': 'InChI221N' }, axis=1, inplace=True)\n\nprint('Done Reading')","metadata":{"execution":{"iopub.status.busy":"2021-06-04T08:15:10.714278Z","iopub.execute_input":"2021-06-04T08:15:10.714693Z","iopub.status.idle":"2021-06-04T08:16:20.993624Z","shell.execute_reply.started":"2021-06-04T08:15:10.714664Z","shell.execute_reply":"2021-06-04T08:16:20.992615Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Merge Submission","metadata":{}},{"cell_type":"markdown","source":"Merge all submission files to a single submission file. Merging is done on image\\_id","metadata":{}},{"cell_type":"code","source":"submission = pd.DataFrame({ 'image_id': submission_218['image_id'] })\n\n# Adding 212_v2\nsubmission = submission.merge(submission_212_v2, on='image_id')\nsubmission = submission.merge(submission_212_v2N, on='image_id')\n\n# Adding 212\nsubmission = submission.merge(submission_212, on='image_id')\nsubmission = submission.merge(submission_212N, on='image_id')\n\n# Adding 218\nsubmission = submission.merge(submission_218, on='image_id')\nsubmission = submission.merge(submission_218N, on='image_id')\n\n# Adding 219\nsubmission = submission.merge(submission_219, on='image_id')\nsubmission = submission.merge(submission_219N, on='image_id')\n\n# Adding 221\nsubmission = submission.merge(submission_221, on='image_id')\nsubmission = submission.merge(submission_221N, on='image_id')","metadata":{"execution":{"iopub.status.busy":"2021-06-04T08:16:20.994966Z","iopub.execute_input":"2021-06-04T08:16:20.995400Z","iopub.status.idle":"2021-06-04T08:16:47.623851Z","shell.execute_reply.started":"2021-06-04T08:16:20.995356Z","shell.execute_reply":"2021-06-04T08:16:47.622786Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# For each submission file the original and normalized InChI are shown\ndisplay(submission.head())","metadata":{"execution":{"iopub.status.busy":"2021-06-04T08:16:47.625091Z","iopub.execute_input":"2021-06-04T08:16:47.625380Z","iopub.status.idle":"2021-06-04T08:16:47.655743Z","shell.execute_reply.started":"2021-06-04T08:16:47.625342Z","shell.execute_reply":"2021-06-04T08:16:47.654738Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check if we indeed have 1616107 rows\ndisplay(submission.info())","metadata":{"execution":{"iopub.status.busy":"2021-06-04T08:16:47.657095Z","iopub.execute_input":"2021-06-04T08:16:47.657538Z","iopub.status.idle":"2021-06-04T08:16:49.413774Z","shell.execute_reply.started":"2021-06-04T08:16:47.657496Z","shell.execute_reply":"2021-06-04T08:16:49.412628Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This next column decleration is important and will be used later on. The boolean indicates whether all normalized InChI's are equal in each submission, thus if all predictions are normalized to the same InChI.","metadata":{}},{"cell_type":"code","source":"submission['equal'] = (\n    (submission['InChI212_v2N'] == submission['InChI212N']) &\n    (submission['InChI212_v2N'] == submission['InChI218N']) &\n    (submission['InChI212_v2N'] == submission['InChI219N']) &\n    (submission['InChI212_v2N'] == submission['InChI221N'])\n)","metadata":{"execution":{"iopub.status.busy":"2021-06-04T08:23:35.106505Z","iopub.execute_input":"2021-06-04T08:23:35.106844Z","iopub.status.idle":"2021-06-04T08:23:36.462097Z","shell.execute_reply.started":"2021-06-04T08:23:35.106816Z","shell.execute_reply":"2021-06-04T08:23:36.461174Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Submission Statistics","metadata":{}},{"cell_type":"code","source":"# Percentage where all submission are equal\npercentage_equal = submission['equal'].sum() / len(submission) * 100\nprint(f'percentage_equal: {percentage_equal:.3f}%')","metadata":{"execution":{"iopub.status.busy":"2021-06-04T08:16:50.771985Z","iopub.execute_input":"2021-06-04T08:16:50.772277Z","iopub.status.idle":"2021-06-04T08:16:50.779239Z","shell.execute_reply.started":"2021-06-04T08:16:50.772247Z","shell.execute_reply":"2021-06-04T08:16:50.778234Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"These next percentages indicate how many InChI's were not valid and could therefore not be normalized. When normalizing the InChI's a non-valid InChI's would normally default to the predicted InChI. However for this method to work, non-valid InChI's need to be normalized to \"error\".","metadata":{}},{"cell_type":"code","source":"# Error rate in each submission\nfor n in NAMES:\n    error_rate = len(submission.loc[submission[f'InChI{n}N'] == 'error']) / len(submission) * 100\n    print(f'{n.ljust(6)} error rate: {error_rate:.3f}%')","metadata":{"execution":{"iopub.status.busy":"2021-06-04T08:16:50.781119Z","iopub.execute_input":"2021-06-04T08:16:50.781584Z","iopub.status.idle":"2021-06-04T08:16:53.457107Z","shell.execute_reply.started":"2021-06-04T08:16:50.781543Z","shell.execute_reply":"2021-06-04T08:16:53.456041Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This next percentage shows where not all predicted InChI's could not be normalized, but at least one could. This is an indication of the expected improvement. Intuively, a normalizable InChI has a higher chance of being the correct prediction, as a non-normalizable InChI is by definition a non-existing molecule and therefore incorrect. A normalizable InChI could be the correct InChI.","metadata":{}},{"cell_type":"code","source":"# Submission where some, but not all, have an error\nerror_sums = (\n    (submission['InChI212_v2N'] == 'error').astype(int) +\n    (submission['InChI212N'] == 'error').astype(int) +\n    (submission['InChI218N'] == 'error').astype(int) +\n    (submission['InChI219N'] == 'error').astype(int) +\n    (submission['InChI221N'] == 'error').astype(int)\n)\nnon_overlapping_error_ratio = sum((error_sums < len(NAMES)) & (error_sums > 0)) / len(submission) * 100\nprint(f'Non overlapping error ratio: {non_overlapping_error_ratio:.3f}%')","metadata":{"execution":{"iopub.status.busy":"2021-06-04T08:16:53.458662Z","iopub.execute_input":"2021-06-04T08:16:53.459119Z","iopub.status.idle":"2021-06-04T08:16:54.778352Z","shell.execute_reply.started":"2021-06-04T08:16:53.459076Z","shell.execute_reply":"2021-06-04T08:16:54.777344Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Submission Selection Process","metadata":{}},{"cell_type":"markdown","source":"This next loop is the beating heart of the submission assembling where the selection process takes place. The base case is when all predictions are equal and can be normalized. In this case it doesn't matter which prediction is selected, as they are all equal. If they are not all equal, the normalizable prediction from the best scoring submission is selected. If no prediction could be normalized the prediction from the best scoring submission is used.","metadata":{}},{"cell_type":"code","source":"submission_final_dict = dict()\n# Selection Statistics\nselection_stats = dict({\n    '12_v2': 0,\n    '12': 0,\n    '18': 0,\n    '19': 0,\n    '21': 0,\n    'e': 0,\n})\n\nfor idx, row in tqdm(submission.iterrows(), total=len(submission)):\n    # All Equal and not error, use submission\n    if row['equal'] and row['InChI212_v2N'] != 'error':\n        submission_final_dict[row['image_id']] = row['InChI212_v2N']\n    # else, choose best submission without error\n    else:\n        # Choose best normalized submission without error\n        if row['InChI212_v2N'] != 'error':\n            selection_stats['12_v2'] += 1\n            submission_final_dict[row['image_id']] = row['InChI212_v2N']\n        elif row['InChI212N'] != 'error':\n            selection_stats['12'] += 1\n            submission_final_dict[row['image_id']] = row['InChI212N']\n        elif row['InChI218N'] != 'error':\n            selection_stats['18'] += 1\n            submission_final_dict[row['image_id']] = row['InChI218N']\n        elif row['InChI219N'] != 'error':\n            selection_stats['19'] += 1\n            submission_final_dict[row['image_id']] = row['InChI219N']\n        elif row['InChI221N'] != 'error':\n            selection_stats['21'] += 1\n            submission_final_dict[row['image_id']] = row['InChI221N']\n        # if none could be normalized use best submission\n        else:\n            selection_stats['e'] += 1\n            submission_final_dict[row['image_id']] = row['InChI212_v2']","metadata":{"execution":{"iopub.status.busy":"2021-06-04T08:16:54.779896Z","iopub.execute_input":"2021-06-04T08:16:54.780324Z","iopub.status.idle":"2021-06-04T08:19:49.242551Z","shell.execute_reply.started":"2021-06-04T08:16:54.780282Z","shell.execute_reply":"2021-06-04T08:19:49.241522Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This pie chart shown the selection statistics. As expected, the selection process prefers high scoring submission. This can be explained by the order of the selection. Around half of the selections default to the best scoring prediction, as no prediction could be normalized. This indicates there is room for improvement, as these predictions are incorrect.","metadata":{}},{"cell_type":"code","source":"# Show Selection Statistics\nplt.figure(figsize=(10,10))\npd.Series(selection_stats, name='fill counts').plot.pie(y='fill count', title='Selection Statistics', legend=False, autopct='%1.1f%%')\npass","metadata":{"execution":{"iopub.status.busy":"2021-06-04T08:20:29.725315Z","iopub.execute_input":"2021-06-04T08:20:29.725694Z","iopub.status.idle":"2021-06-04T08:20:29.885233Z","shell.execute_reply.started":"2021-06-04T08:20:29.725657Z","shell.execute_reply":"2021-06-04T08:20:29.884479Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(pd.Series(selection_stats).to_frame(name='fill counts'))","metadata":{"execution":{"iopub.status.busy":"2021-06-04T08:20:32.842722Z","iopub.execute_input":"2021-06-04T08:20:32.843232Z","iopub.status.idle":"2021-06-04T08:20:32.852431Z","shell.execute_reply.started":"2021-06-04T08:20:32.843201Z","shell.execute_reply":"2021-06-04T08:20:32.851448Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Make DataFrame","metadata":{}},{"cell_type":"code","source":"# Transform dictionary to dataframe\nsubmission_final = pd.DataFrame.from_dict(submission_final_dict, orient='index', columns=['InChI'])\n# Set the index_name to 'image_id' as required by the submission format\nsubmission_final.index.name = 'image_id'","metadata":{"execution":{"iopub.status.busy":"2021-06-04T08:20:37.893678Z","iopub.execute_input":"2021-06-04T08:20:37.894044Z","iopub.status.idle":"2021-06-04T08:20:38.552293Z","shell.execute_reply.started":"2021-06-04T08:20:37.894014Z","shell.execute_reply":"2021-06-04T08:20:38.551179Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Remove \\p and \\q from submission","metadata":{}},{"cell_type":"markdown","source":"The normalized InChI's contained InChI parts which were not present in the training data, namely the /p and /q part. Assuming the test dataset also didn't contain these parts an improvement could be expected by removing these parts. No LB improvement was however observed, which is probably due to the low number of occurances of these parts. It would however be interesting to know if this removal did actually improve the score.","metadata":{}},{"cell_type":"code","source":"removed_pq_parts = 0\n\ndef remove_pq(InChI, debug=False):\n    global removed_pq_parts\n    \n    InChI_Original = InChI\n    # P\n    substr_p = re.findall('(?<=\\/p)[^\\/]*', InChI)\n    if len(substr_p) > 0:\n        for substr in substr_p:\n            removed_pq_parts += 1\n            InChI = InChI.replace(f'/p{substr}', '')\n    # Q\n    substr_q = re.findall('(?<=\\/q)[^\\/]*', InChI)\n    if len(substr_q) > 0:\n        for substr in substr_q:\n            removed_pq_parts += 1\n            InChI = InChI.replace(f'/q{substr}', '')\n            \n    if debug and (len(substr_p) > 0 or len(substr_q) > 0):\n        print('=' * 50)\n        print(InChI_Original)\n        print(InChI)\n        print('=' * 50)\n        \n    return InChI","metadata":{"execution":{"iopub.status.busy":"2021-06-04T08:20:39.204642Z","iopub.execute_input":"2021-06-04T08:20:39.204985Z","iopub.status.idle":"2021-06-04T08:20:39.212354Z","shell.execute_reply.started":"2021-06-04T08:20:39.204957Z","shell.execute_reply":"2021-06-04T08:20:39.211314Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Remove /p and /q parts as they are not present in the training set\nsubmission_final['InChI'] = submission_final['InChI'].progress_apply(remove_pq)","metadata":{"execution":{"iopub.status.busy":"2021-06-04T08:20:40.215173Z","iopub.execute_input":"2021-06-04T08:20:40.215543Z","iopub.status.idle":"2021-06-04T08:21:02.663201Z","shell.execute_reply.started":"2021-06-04T08:20:40.215510Z","shell.execute_reply":"2021-06-04T08:21:02.662351Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show statistics about /p /q removal\nprint(f'Total number of removed parts: {removed_pq_parts}')\nremoved_pq_parts_percentage = removed_pq_parts / len(submission_final) * 100\nprint(f'Mean percentage of InChI\\'s with /p or /q part removed: {removed_pq_parts_percentage:.3f}%')","metadata":{"execution":{"iopub.status.busy":"2021-06-04T08:21:37.253735Z","iopub.execute_input":"2021-06-04T08:21:37.254082Z","iopub.status.idle":"2021-06-04T08:21:37.259027Z","shell.execute_reply.started":"2021-06-04T08:21:37.254054Z","shell.execute_reply":"2021-06-04T08:21:37.258132Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Submission to CSV","metadata":{}},{"cell_type":"code","source":"# Show submission head\ndisplay(submission_final.head())","metadata":{"execution":{"iopub.status.busy":"2021-06-04T08:21:37.745504Z","iopub.execute_input":"2021-06-04T08:21:37.745838Z","iopub.status.idle":"2021-06-04T08:21:37.757485Z","shell.execute_reply.started":"2021-06-04T08:21:37.745809Z","shell.execute_reply":"2021-06-04T08:21:37.756433Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check if there are 1616107 rows!!!\ndisplay(submission_final.info())","metadata":{"execution":{"iopub.status.busy":"2021-06-04T08:21:38.042844Z","iopub.execute_input":"2021-06-04T08:21:38.043194Z","iopub.status.idle":"2021-06-04T08:21:38.218348Z","shell.execute_reply.started":"2021-06-04T08:21:38.043166Z","shell.execute_reply":"2021-06-04T08:21:38.217618Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# write submission file to CSV\nsubmission_final.to_csv('submission.csv', index=True)","metadata":{"execution":{"iopub.status.busy":"2021-06-04T08:21:40.331438Z","iopub.execute_input":"2021-06-04T08:21:40.331769Z","iopub.status.idle":"2021-06-04T08:21:49.128353Z","shell.execute_reply.started":"2021-06-04T08:21:40.331742Z","shell.execute_reply":"2021-06-04T08:21:49.127392Z"},"trusted":true},"execution_count":null,"outputs":[]}]}