{
  "id": 232685,
  "title": "Improve score of already trained models!",
  "url": "/competitions/bms-molecular-translation/discussion/232685",
  "author_name": "Achille Nazaret",
  "post_date": "2021-04-14T21:35:47.282000",
  "votes": 31,
  "comment_count": 20,
  "views": 0,
  "content": "<p>The InChI format is very rigid and structured, so we can check if our predictions are <strong>valid InChIs</strong>.</p>\n<p>Then, if we have two different models, we can check all InChIs and <strong>merge the different models' predictions</strong> by selecting the valid inchis of each one.<br>\nWith two medium models: </p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>LB Score</th>\n<th>Comment</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Model 1</td>\n<td>9.6</td>\n<td>Only a few epochs</td>\n</tr>\n<tr>\n<td>Model 2</td>\n<td>7.5</td>\n<td>Same model, more epochs</td>\n</tr>\n<tr>\n<td>Merge both</td>\n<td>7.2</td>\n<td>Even a bad model can improve a better model</td>\n</tr>\n</tbody>\n</table>\n<p>Now let’s push this way further and talk about how to improve a single model <strong>using InChI validation and kbeam</strong>.</p>\n<p>When we use kbeam, our model makes multiple suggestions &gt;= K and we usually take the one with the best score. However, sometimes the first predictions are not a valid InChI while the 10th prediction is.</p>\n<p>Here are the results of my experiment:</p>\n<ul>\n<li>Run kbeam with a large K</li>\n<li>Sort the predictions by score but keep all of them</li>\n<li>Select a limit T</li>\n<li>For each test image, successively test the validity of the T first predictions returned by KBeam and select the first valid InChI to submit.</li>\n<li>If none of the first T them were valid, submit the first one.</li>\n</ul>\n<p>Here are the results for one trained model. I tried various T on the same output:</p>\n<table>\n<thead>\n<tr>\n<th>T</th>\n<th>LB Score</th>\n<th># Valid InChI</th>\n<th>Comment</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>4.24</td>\n<td>1,269,671</td>\n<td>This is just taking the first InChI returned by KBeam</td>\n</tr>\n<tr>\n<td>10</td>\n<td>3.77</td>\n<td>1,436,191 (+166,520 )</td>\n<td>We have a large improvement</td>\n</tr>\n<tr>\n<td>30</td>\n<td>3.73</td>\n<td>1,450,371 (+14,180)</td>\n<td>Even with a large T, it is still beneficial</td>\n</tr>\n</tbody>\n</table>\n<p>(Total test set has 1,616,107 images).<br>\nWe can see that it is always better to take a valid inchi. </p>\n<p>It was quite slow and buggy to check for valid InChI with RdKit so I coded something else that can check ~10,000 inchis per second per cpu which is quite necessary when we have ~30*10^6 inchis strings! I recommend implementing your own method rather than RdKit (unless you know how to debug it :) ).</p>\n<p>Let me know if you are using similar methods to improve the predictions of a trained model, and how they perform!</p>",
  "messages": [
    {
      "id": 1273998,
      "postDate": "2021-04-14T21:35:47.283Z",
      "content": "<p>The InChI format is very rigid and structured, so we can check if our predictions are <strong>valid InChIs</strong>.</p>\n<p>Then, if we have two different models, we can check all InChIs and <strong>merge the different models' predictions</strong> by selecting the valid inchis of each one.<br>\nWith two medium models: </p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>LB Score</th>\n<th>Comment</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Model 1</td>\n<td>9.6</td>\n<td>Only a few epochs</td>\n</tr>\n<tr>\n<td>Model 2</td>\n<td>7.5</td>\n<td>Same model, more epochs</td>\n</tr>\n<tr>\n<td>Merge both</td>\n<td>7.2</td>\n<td>Even a bad model can improve a better model</td>\n</tr>\n</tbody>\n</table>\n<p>Now let’s push this way further and talk about how to improve a single model <strong>using InChI validation and kbeam</strong>.</p>\n<p>When we use kbeam, our model makes multiple suggestions &gt;= K and we usually take the one with the best score. However, sometimes the first predictions are not a valid InChI while the 10th prediction is.</p>\n<p>Here are the results of my experiment:</p>\n<ul>\n<li>Run kbeam with a large K</li>\n<li>Sort the predictions by score but keep all of them</li>\n<li>Select a limit T</li>\n<li>For each test image, successively test the validity of the T first predictions returned by KBeam and select the first valid InChI to submit.</li>\n<li>If none of the first T them were valid, submit the first one.</li>\n</ul>\n<p>Here are the results for one trained model. I tried various T on the same output:</p>\n<table>\n<thead>\n<tr>\n<th>T</th>\n<th>LB Score</th>\n<th># Valid InChI</th>\n<th>Comment</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>4.24</td>\n<td>1,269,671</td>\n<td>This is just taking the first InChI returned by KBeam</td>\n</tr>\n<tr>\n<td>10</td>\n<td>3.77</td>\n<td>1,436,191 (+166,520 )</td>\n<td>We have a large improvement</td>\n</tr>\n<tr>\n<td>30</td>\n<td>3.73</td>\n<td>1,450,371 (+14,180)</td>\n<td>Even with a large T, it is still beneficial</td>\n</tr>\n</tbody>\n</table>\n<p>(Total test set has 1,616,107 images).<br>\nWe can see that it is always better to take a valid inchi. </p>\n<p>It was quite slow and buggy to check for valid InChI with RdKit so I coded something else that can check ~10,000 inchis per second per cpu which is quite necessary when we have ~30*10^6 inchis strings! I recommend implementing your own method rather than RdKit (unless you know how to debug it :) ).</p>\n<p>Let me know if you are using similar methods to improve the predictions of a trained model, and how they perform!</p>",
      "rawMarkdown": "The InChI format is very rigid and structured, so we can check if our predictions are **valid InChIs**.\n\nThen, if we have two different models, we can check all InChIs and **merge the different models' predictions** by selecting the valid inchis of each one.\nWith two medium models: \n| Model | LB Score | Comment\n| --- | --- | -- |\n| Model 1 | 9.6 | Only a few epochs |\n| Model 2 | 7.5 | Same model, more epochs |\n| Merge both | 7.2 | Even a bad model can improve a better model |\n\nNow let’s push this way further and talk about how to improve a single model **using InChI validation and kbeam**.\n\nWhen we use kbeam, our model makes multiple suggestions >= K and we usually take the one with the best score. However, sometimes the first predictions are not a valid InChI while the 10th prediction is.\n\nHere are the results of my experiment:\n- Run kbeam with a large K\n- Sort the predictions by score but keep all of them\n- Select a limit T\n- For each test image, successively test the validity of the T first predictions returned by KBeam and select the first valid InChI to submit.\n- If none of the first T them were valid, submit the first one.\n\nHere are the results for one trained model. I tried various T on the same output:\n|  T      | LB Score | # Valid InChI | Comment |\n| ----- | --------- | ----------- | -- |\n| 0 | 4.24 | 1,269,671  |This is just taking the first InChI returned by KBeam |\n| 10 | 3.77 | 1,436,191 (+166,520 ) | We have a large improvement |\n| 30 | 3.73 | 1,450,371 (+14,180) | Even with a large T, it is still beneficial |\n\n(Total test set has 1,616,107 images).\nWe can see that it is always better to take a valid inchi. \n\nIt was quite slow and buggy to check for valid InChI with RdKit so I coded something else that can check ~10,000 inchis per second per cpu which is quite necessary when we have ~30*10^6 inchis strings! I recommend implementing your own method rather than RdKit (unless you know how to debug it :) ).\n\nLet me know if you are using similar methods to improve the predictions of a trained model, and how they perform!",
      "votes": 30
    },
    {
      "id": 1275709,
      "postDate": "2021-04-16T15:49:15.563Z",
      "content": "<p>This is wonderful. Thank you for this information.</p>\n<p>This might be a dumb ask… but… is this allowed? This level of manual check? It makes sense that it would be because why would you limit yourself to ML solutions only when you can also use clever solutions. Work smarter not harder and whatnot… </p>\n<p>With that being said, I just wanted to double-check? Sometimes Kaggle rules can be a bit odd.</p>\n<p>Thanks again!</p>",
      "rawMarkdown": "This is wonderful. Thank you for this information.\n\nThis might be a dumb ask... but... is this allowed? This level of manual check? It makes sense that it would be because why would you limit yourself to ML solutions only when you can also use clever solutions. Work smarter not harder and whatnot... \n\nWith that being said, I just wanted to double-check? Sometimes Kaggle rules can be a bit odd.\n\nThanks again!",
      "votes": 3,
      "replies": [
        {
          "id": 1275722,
          "postDate": "2021-04-16T15:57:08.497Z",
          "content": "<p>I think (correct me if I misunderstand) we established that compuationally checking that an InChI is valid is fine, but (to my disapointment) checking whether it corresponds to a known compound is not. In this sense, valid means that a graph-to-InChI transformation is correctly carried out according to the official InChI generation algorithm, not that it corresponds to the particular molecular structure in a given image. In the context of an InChI derived directly from an image by ML, it means that the same InChI could have been derived by the algorithm for some molecular graph (but not necessarily the one pictured).  </p>",
          "rawMarkdown": "I think (correct me if I misunderstand) we established that compuationally checking that an InChI is valid is fine, but (to my disapointment) checking whether it corresponds to a known compound is not. In this sense, valid means that a graph-to-InChI transformation is correctly carried out according to the official InChI generation algorithm, not that it corresponds to the particular molecular structure in a given image. In the context of an InChI derived directly from an image by ML, it means that the same InChI could have been derived by the algorithm for some molecular graph (but not necessarily the one pictured).  ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1274133,
      "postDate": "2021-04-15T03:02:30.903Z",
      "content": "<p>\"already trained models!\"<br>\n i have a few tricks for that:</p>\n<ol>\n<li>use SWA (Stochastic Weight Averaging) to make an average model, you should see improvement of around 0.05</li>\n<li>finetune your model on focal loss (to take care of long tail problem), then apply SWA</li>\n</ol>\n<p>i am interested in with improved SWA models, what is the effect of K-beams?<br>\n(this can be estimated from the top-K results  graph you have shown)</p>\n<p>waht is the best way to combine the models for better k-search? e.g. average? or average of p**0.5 ?<br>\nyou can see my post of results at : <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/231190\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/231190</a></p>",
      "rawMarkdown": "\"already trained models!\"\n i have a few tricks for that:\n\n1. use SWA (Stochastic Weight Averaging) to make an average model, you should see improvement of around 0.05\n2. finetune your model on focal loss (to take care of long tail problem), then apply SWA\n\ni am interested in with improved SWA models, what is the effect of K-beams?\n(this can be estimated from the top-K results  graph you have shown)\n\nwaht is the best way to combine the models for better k-search? e.g. average? or average of p**0.5 ?\nyou can see my post of results at : https://www.kaggle.com/c/bms-molecular-translation/discussion/231190",
      "votes": 3,
      "replies": [
        {
          "id": 1279852,
          "postDate": "2021-04-21T09:22:45.993Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1274111,
      "postDate": "2021-04-15T02:14:13.630Z",
      "content": "<p><img src=\"https://i.ibb.co/8r45Z1z/download-1.png\" alt=\"\"> </p>\n<p>Here is a distribution of the number of test images for which the Xth prediction (by KBeam) was the first to be valid. -1 means none of the predictions were valid.</p>",
      "rawMarkdown": "![](https://i.ibb.co/8r45Z1z/download-1.png) \n\nHere is a distribution of the number of test images for which the Xth prediction (by KBeam) was the first to be valid. -1 means none of the predictions were valid.",
      "votes": 3
    },
    {
      "id": 1274009,
      "postDate": "2021-04-14T22:14:30.747Z",
      "content": "<p>In principle, yes I am checking InChI validity. In practice, it is slow. Today I also encountered an InChI bomb that crashed the process - don't ask me how or why. I am essentially ensembling models and preferring (i) valid InChIs, (ii) better scoring models.</p>",
      "rawMarkdown": "In principle, yes I am checking InChI validity. In practice, it is slow. Today I also encountered an InChI bomb that crashed the process - don't ask me how or why. I am essentially ensembling models and preferring (i) valid InChIs, (ii) better scoring models.",
      "votes": 1,
      "replies": [
        {
          "id": 1274178,
          "postDate": "2021-04-15T04:39:51.200Z",
          "content": "<p>Ah yes, there are several bombs in the data set - I am compiling a short list so they can be easily excluded from validation with RDkit. </p>",
          "rawMarkdown": "Ah yes, there are several bombs in the data set - I am compiling a short list so they can be easily excluded from validation with RDkit. ",
          "votes": 2
        },
        {
          "id": 1277075,
          "postDate": "2021-04-18T11:25:31.497Z",
          "content": "<p>Hi! Could I please ask you which method from RDkit you are calling to validate inchis?</p>",
          "rawMarkdown": "Hi! Could I please ask you which method from RDkit you are calling to validate inchis?"
        },
        {
          "id": 1277145,
          "postDate": "2021-04-18T13:09:37.553Z",
          "content": "<p>Chem.MolFromInchi(inchi) -&gt; Chem.MolToInchi(mol)</p>\n<p>Check the following notebook: <a href=\"https://www.kaggle.com/wuliaokaola/bmsmt-0331-normalize-your-predictions\" target=\"_blank\">https://www.kaggle.com/wuliaokaola/bmsmt-0331-normalize-your-predictions</a></p>",
          "rawMarkdown": "Chem.MolFromInchi(inchi) -> Chem.MolToInchi(mol)\n\nCheck the following notebook: https://www.kaggle.com/wuliaokaola/bmsmt-0331-normalize-your-predictions",
          "votes": 1
        },
        {
          "id": 1277213,
          "postDate": "2021-04-18T14:26:55.487Z",
          "content": "<p>Thank you!</p>",
          "rawMarkdown": "Thank you!"
        }
      ]
    },
    {
      "id": 1274006,
      "postDate": "2021-04-14T22:06:30.043Z",
      "content": "<p>Combining beam search and inchi validation is a super smart idea! Thanks for sharing!</p>",
      "rawMarkdown": "Combining beam search and inchi validation is a super smart idea! Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 1274024,
      "postDate": "2021-04-14T22:59:38.053Z",
      "content": "<p>With LB 3.73, how many of them are valid InChI?</p>",
      "rawMarkdown": "With LB 3.73, how many of them are valid InChI?",
      "votes": 2,
      "replies": [
        {
          "id": 1274029,
          "postDate": "2021-04-14T23:12:59.187Z",
          "content": "<p>That is a good question. I will compute this metric for each T and add it to the post.</p>",
          "rawMarkdown": "That is a good question. I will compute this metric for each T and add it to the post.",
          "votes": 1
        },
        {
          "id": 1274123,
          "postDate": "2021-04-15T02:37:28.517Z",
          "content": "<p>Saw your update. Impressive there are ~90% valid InChI in your 3.73 submission. You can make a test submission where the rest 10% is substituted with 500 \"X\" (their LD are always 500) to estimate the accuracy from the valid InChI subset. I am curious how much gain InChI validation could bring. Hope this makes sense.</p>\n<p>Btw, you also mentioned that you coded <em>something else</em> rather than RDKit to speed up the validation. Can you share a bit more information? Thanks. </p>",
          "rawMarkdown": "Saw your update. Impressive there are ~90% valid InChI in your 3.73 submission. You can make a test submission where the rest 10% is substituted with 500 \"X\" (their LD are always 500) to estimate the accuracy from the valid InChI subset. I am curious how much gain InChI validation could bring. Hope this makes sense.\n\nBtw, you also mentioned that you coded *something else* rather than RDKit to speed up the validation. Can you share a bit more information? Thanks. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1285289,
      "postDate": "2021-04-26T18:38:50.167Z",
      "content": "<p><code>merge the different models predictions by selecting the valid inchis of each one.</code> What is the approaching for doing so ? </p>",
      "rawMarkdown": "`merge the different models predictions by selecting the valid inchis of each one.` What is the approaching for doing so ? ",
      "replies": [
        {
          "id": 1288903,
          "postDate": "2021-04-30T12:59:18.477Z",
          "content": "<p>I guess, one approach could be doing a majority vote for each token in the InchIs.<br>\nMeaning you have e.g. 3 models and then choose the token which was predicted by at least 2 of them. If none match you use the one from the best model. If your InChIs differ in length you would have to find matching token groups and do a voting on those.</p>\n<p>You also should have a threshold, where you just choose the InChI of the model which is best on average, IF there are &gt; x differences in the InChIs of your (e.g. 3) models. The latter makes sense if you have a very difficult image which leads to a high variance of predictions where the combined InChI would be even worse. But that is something to be checked in practice.</p>\n<p>You should also check if the combined InChI is still a valid one.</p>",
          "rawMarkdown": "I guess, one approach could be doing a majority vote for each token in the InchIs.\nMeaning you have e.g. 3 models and then choose the token which was predicted by at least 2 of them. If none match you use the one from the best model. If your InChIs differ in length you would have to find matching token groups and do a voting on those.\n\nYou also should have a threshold, where you just choose the InChI of the model which is best on average, IF there are > x differences in the InChIs of your (e.g. 3) models. The latter makes sense if you have a very difficult image which leads to a high variance of predictions where the combined InChI would be even worse. But that is something to be checked in practice.\n\nYou should also check if the combined InChI is still a valid one.\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 1277928,
      "postDate": "2021-04-19T11:34:28.720Z",
      "content": "<p>\"Combining beam search and inchi validation is a super smart idea! Thanks for sharing!\"</p>\n<p>there are yet other tricks like atomic number only appears once and only one in inchi sublayer.<br>\nalso all numbers must appear. sum of different atoms must follow predicted formula substraing, etc …</p>",
      "rawMarkdown": "\"Combining beam search and inchi validation is a super smart idea! Thanks for sharing!\"\n\nthere are yet other tricks like atomic number only appears once and only one in inchi sublayer.\nalso all numbers must appear. sum of different atoms must follow predicted formula substraing, etc ...",
      "replies": [
        {
          "id": 1278022,
          "postDate": "2021-04-19T13:45:14.147Z",
          "content": "<p><code>there are yet other tricks</code><br>\nIsn't that also InChI validation what you are talking about? Or do you mean fixing the InChI string based on a set of rules?</p>",
          "rawMarkdown": "`there are yet other tricks`\nIsn't that also InChI validation what you are talking about? Or do you mean fixing the InChI string based on a set of rules?"
        },
        {
          "id": 1278037,
          "postDate": "2021-04-19T14:04:26.623Z",
          "content": "<p>I think these tricks can be used when you are developing a custom inchi validator. A custom validator can speed up the process as suggested in the main post by <a href=\"https://www.kaggle.com/anazaret\" target=\"_blank\">@anazaret</a> .</p>",
          "rawMarkdown": "I think these tricks can be used when you are developing a custom inchi validator. A custom validator can speed up the process as suggested in the main post by @anazaret ."
        }
      ]
    },
    {
      "id": 1277396,
      "postDate": "2021-04-18T17:31:26.420Z",
      "content": "<p>Thanks for the information!</p>",
      "rawMarkdown": "Thanks for the information!"
    }
  ],
  "comments": [
    {
      "id": 1275709,
      "author_name": "Darien Schettler",
      "author_url": "",
      "post_date": "2021-04-16T15:49:15.563000",
      "content": "<p>This is wonderful. Thank you for this information.</p>\n<p>This might be a dumb ask… but… is this allowed? This level of manual check? It makes sense that it would be because why would you limit yourself to ML solutions only when you can also use clever solutions. Work smarter not harder and whatnot… </p>\n<p>With that being said, I just wanted to double-check? Sometimes Kaggle rules can be a bit odd.</p>\n<p>Thanks again!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1275722,
          "author_name": "John Mitchell",
          "author_url": "",
          "post_date": "2021-04-16T15:57:08.497000",
          "content": "<p>I think (correct me if I misunderstand) we established that compuationally checking that an InChI is valid is fine, but (to my disapointment) checking whether it corresponds to a known compound is not. In this sense, valid means that a graph-to-InChI transformation is correctly carried out according to the official InChI generation algorithm, not that it corresponds to the particular molecular structure in a given image. In the context of an InChI derived directly from an image by ML, it means that the same InChI could have been derived by the algorithm for some molecular graph (but not necessarily the one pictured).  </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1274133,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-04-15T03:02:30.903000",
      "content": "<p>\"already trained models!\"<br>\n i have a few tricks for that:</p>\n<ol>\n<li>use SWA (Stochastic Weight Averaging) to make an average model, you should see improvement of around 0.05</li>\n<li>finetune your model on focal loss (to take care of long tail problem), then apply SWA</li>\n</ol>\n<p>i am interested in with improved SWA models, what is the effect of K-beams?<br>\n(this can be estimated from the top-K results  graph you have shown)</p>\n<p>waht is the best way to combine the models for better k-search? e.g. average? or average of p**0.5 ?<br>\nyou can see my post of results at : <a href=\"https://www.kaggle.com/c/bms-molecular-translation/discussion/231190\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation/discussion/231190</a></p>",
      "votes": 3,
      "replies": [
        {
          "id": 1279852,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-04-21T09:22:45.993000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1274111,
      "author_name": "Achille Nazaret",
      "author_url": "",
      "post_date": "2021-04-15T02:14:13.630000",
      "content": "<p><img src=\"https://i.ibb.co/8r45Z1z/download-1.png\" alt=\"\"> </p>\n<p>Here is a distribution of the number of test images for which the Xth prediction (by KBeam) was the first to be valid. -1 means none of the predictions were valid.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1274009,
      "author_name": "John Mitchell",
      "author_url": "",
      "post_date": "2021-04-14T22:14:30.747000",
      "content": "<p>In principle, yes I am checking InChI validity. In practice, it is slow. Today I also encountered an InChI bomb that crashed the process - don't ask me how or why. I am essentially ensembling models and preferring (i) valid InChIs, (ii) better scoring models.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1274178,
          "author_name": "Charles",
          "author_url": "",
          "post_date": "2021-04-15T04:39:51.200000",
          "content": "<p>Ah yes, there are several bombs in the data set - I am compiling a short list so they can be easily excluded from validation with RDkit. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1277075,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2021-04-18T11:25:31.497000",
          "content": "<p>Hi! Could I please ask you which method from RDkit you are calling to validate inchis?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1277145,
          "author_name": "Murat Cihan Sorkun",
          "author_url": "",
          "post_date": "2021-04-18T13:09:37.553000",
          "content": "<p>Chem.MolFromInchi(inchi) -&gt; Chem.MolToInchi(mol)</p>\n<p>Check the following notebook: <a href=\"https://www.kaggle.com/wuliaokaola/bmsmt-0331-normalize-your-predictions\" target=\"_blank\">https://www.kaggle.com/wuliaokaola/bmsmt-0331-normalize-your-predictions</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1277213,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2021-04-18T14:26:55.487000",
          "content": "<p>Thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1274006,
      "author_name": "Murat Cihan Sorkun",
      "author_url": "",
      "post_date": "2021-04-14T22:06:30.043000",
      "content": "<p>Combining beam search and inchi validation is a super smart idea! Thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1274024,
      "author_name": "human intelligence",
      "author_url": "",
      "post_date": "2021-04-14T22:59:38.053000",
      "content": "<p>With LB 3.73, how many of them are valid InChI?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1274029,
          "author_name": "Achille Nazaret",
          "author_url": "",
          "post_date": "2021-04-14T23:12:59.187000",
          "content": "<p>That is a good question. I will compute this metric for each T and add it to the post.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1274123,
          "author_name": "human intelligence",
          "author_url": "",
          "post_date": "2021-04-15T02:37:28.517000",
          "content": "<p>Saw your update. Impressive there are ~90% valid InChI in your 3.73 submission. You can make a test submission where the rest 10% is substituted with 500 \"X\" (their LD are always 500) to estimate the accuracy from the valid InChI subset. I am curious how much gain InChI validation could bring. Hope this makes sense.</p>\n<p>Btw, you also mentioned that you coded <em>something else</em> rather than RDKit to speed up the validation. Can you share a bit more information? Thanks. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1285289,
      "author_name": "Rony",
      "author_url": "",
      "post_date": "2021-04-26T18:38:50.167000",
      "content": "<p><code>merge the different models predictions by selecting the valid inchis of each one.</code> What is the approaching for doing so ? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1288903,
          "author_name": "Gabriel Lindenmaier",
          "author_url": "",
          "post_date": "2021-04-30T12:59:18.477000",
          "content": "<p>I guess, one approach could be doing a majority vote for each token in the InchIs.<br>\nMeaning you have e.g. 3 models and then choose the token which was predicted by at least 2 of them. If none match you use the one from the best model. If your InChIs differ in length you would have to find matching token groups and do a voting on those.</p>\n<p>You also should have a threshold, where you just choose the InChI of the model which is best on average, IF there are &gt; x differences in the InChIs of your (e.g. 3) models. The latter makes sense if you have a very difficult image which leads to a high variance of predictions where the combined InChI would be even worse. But that is something to be checked in practice.</p>\n<p>You should also check if the combined InChI is still a valid one.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1277928,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-04-19T11:34:28.720000",
      "content": "<p>\"Combining beam search and inchi validation is a super smart idea! Thanks for sharing!\"</p>\n<p>there are yet other tricks like atomic number only appears once and only one in inchi sublayer.<br>\nalso all numbers must appear. sum of different atoms must follow predicted formula substraing, etc …</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1278022,
          "author_name": "Gabriel Lindenmaier",
          "author_url": "",
          "post_date": "2021-04-19T13:45:14.147000",
          "content": "<p><code>there are yet other tricks</code><br>\nIsn't that also InChI validation what you are talking about? Or do you mean fixing the InChI string based on a set of rules?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1278037,
          "author_name": "Murat Cihan Sorkun",
          "author_url": "",
          "post_date": "2021-04-19T14:04:26.623000",
          "content": "<p>I think these tricks can be used when you are developing a custom inchi validator. A custom validator can speed up the process as suggested in the main post by <a href=\"https://www.kaggle.com/anazaret\" target=\"_blank\">@anazaret</a> .</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1277396,
      "author_name": "Santosh kumar",
      "author_url": "",
      "post_date": "2021-04-18T17:31:26.420000",
      "content": "<p>Thanks for the information!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1273998": "The InChI format is very rigid and structured, so we can check if our predictions are **valid InChIs**.\n\nThen, if we have two different models, we can check all InChIs and **merge the different models' predictions** by selecting the valid inchis of each one.\nWith two medium models: \n| Model | LB Score | Comment\n| --- | --- | -- |\n| Model 1 | 9.6 | Only a few epochs |\n| Model 2 | 7.5 | Same model, more epochs |\n| Merge both | 7.2 | Even a bad model can improve a better model |\n\nNow let’s push this way further and talk about how to improve a single model **using InChI validation and kbeam**.\n\nWhen we use kbeam, our model makes multiple suggestions >= K and we usually take the one with the best score. However, sometimes the first predictions are not a valid InChI while the 10th prediction is.\n\nHere are the results of my experiment:\n- Run kbeam with a large K\n- Sort the predictions by score but keep all of them\n- Select a limit T\n- For each test image, successively test the validity of the T first predictions returned by KBeam and select the first valid InChI to submit.\n- If none of the first T them were valid, submit the first one.\n\nHere are the results for one trained model. I tried various T on the same output:\n|  T      | LB Score | # Valid InChI | Comment |\n| ----- | --------- | ----------- | -- |\n| 0 | 4.24 | 1,269,671  |This is just taking the first InChI returned by KBeam |\n| 10 | 3.77 | 1,436,191 (+166,520 ) | We have a large improvement |\n| 30 | 3.73 | 1,450,371 (+14,180) | Even with a large T, it is still beneficial |\n\n(Total test set has 1,616,107 images).\nWe can see that it is always better to take a valid inchi. \n\nIt was quite slow and buggy to check for valid InChI with RdKit so I coded something else that can check ~10,000 inchis per second per cpu which is quite necessary when we have ~30*10^6 inchis strings! I recommend implementing your own method rather than RdKit (unless you know how to debug it :) ).\n\nLet me know if you are using similar methods to improve the predictions of a trained model, and how they perform!",
    "1275709": "This is wonderful. Thank you for this information.\n\nThis might be a dumb ask... but... is this allowed? This level of manual check? It makes sense that it would be because why would you limit yourself to ML solutions only when you can also use clever solutions. Work smarter not harder and whatnot... \n\nWith that being said, I just wanted to double-check? Sometimes Kaggle rules can be a bit odd.\n\nThanks again!",
    "1274133": "\"already trained models!\"\n i have a few tricks for that:\n\n1. use SWA (Stochastic Weight Averaging) to make an average model, you should see improvement of around 0.05\n2. finetune your model on focal loss (to take care of long tail problem), then apply SWA\n\ni am interested in with improved SWA models, what is the effect of K-beams?\n(this can be estimated from the top-K results  graph you have shown)\n\nwaht is the best way to combine the models for better k-search? e.g. average? or average of p**0.5 ?\nyou can see my post of results at : https://www.kaggle.com/c/bms-molecular-translation/discussion/231190",
    "1274111": "![](https://i.ibb.co/8r45Z1z/download-1.png) \n\nHere is a distribution of the number of test images for which the Xth prediction (by KBeam) was the first to be valid. -1 means none of the predictions were valid.",
    "1274009": "In principle, yes I am checking InChI validity. In practice, it is slow. Today I also encountered an InChI bomb that crashed the process - don't ask me how or why. I am essentially ensembling models and preferring (i) valid InChIs, (ii) better scoring models.",
    "1274006": "Combining beam search and inchi validation is a super smart idea! Thanks for sharing!",
    "1274024": "With LB 3.73, how many of them are valid InChI?",
    "1285289": "`merge the different models predictions by selecting the valid inchis of each one.` What is the approaching for doing so ? ",
    "1277928": "\"Combining beam search and inchi validation is a super smart idea! Thanks for sharing!\"\n\nthere are yet other tricks like atomic number only appears once and only one in inchi sublayer.\nalso all numbers must appear. sum of different atoms must follow predicted formula substraing, etc ...",
    "1277396": "Thanks for the information!"
  }
}