{
  "id": 242082,
  "title": "Ensembling Techniques",
  "url": "/competitions/bms-molecular-translation/discussion/242082",
  "author_name": "",
  "post_date": "2021-05-27T12:47:35.886084700Z",
  "votes": 6,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Did anyone try any ensembling techniques other than voting?</p>\n<p>Here are few updates from comments:</p>\n<ol>\n<li>Ensembling at each time step during prediction. It needs the same vocabulary for all models. <br>\nthanks <a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a>   </li>\n<li>Using rdkit validator to combine the results of different models. <br>\nthanks <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </li>\n</ol>",
  "messages": [
    {
      "id": "1325020",
      "postDate": "05/27/2021 12:47:35",
      "content": "<p>Did anyone try any ensembling techniques other than voting?</p>\n<p>Here are few updates from comments:</p>\n<ol>\n<li>Ensembling at each time step during prediction. It needs the same vocabulary for all models. <br>\nthanks <a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a>   </li>\n<li>Using rdkit validator to combine the results of different models. <br>\nthanks <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </li>\n</ol>",
      "rawMarkdown": "Did anyone try any ensembling techniques other than voting?\n\nHere are few updates from comments:\n\n1. Ensembling at each time step during prediction. It needs the same vocabulary for all models. \n thanks @serigne   \n2. Using rdkit validator to combine the results of different models. \nthanks @hengck23",
      "votes": null
    },
    {
      "id": "1325273",
      "postDate": "05/27/2021 16:03:40",
      "content": "<p>Easy ensembling can be done at each time step during prediction.  But you will need to use the same vocabulary for all models. </p>",
      "rawMarkdown": "Easy ensembling can be done at each time step during prediction.  But you will need to use the same vocabulary for all models.",
      "votes": null
    },
    {
      "id": "1325964",
      "postDate": "05/28/2021 06:15:06",
      "content": "<p>sounds interesting. So did you apply this and got any improvement? </p>",
      "rawMarkdown": "sounds interesting. So did you apply this and got any improvement?",
      "votes": null
    },
    {
      "id": "1326015",
      "postDate": "05/28/2021 06:59:07",
      "content": "<p>I did and it helps. Also I train different models with different aspect ratios.<br>\nA few models, each above CV of 1.00 gets down to 0.80 with just ensembling and down to 0.76 with beam size of 16. Normalization pushes that down to 0.72.<br>\nI'm more and more sceptical about that gold, people are just going mad with their scores. :D<br>\nEspecially <a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> with a solo 0.65, amazing work!<br>\nThere must be some more clever tricks, but I've ran out of ideas.</p>\n<p>I could've worked more on big molecules, as my models don't do well on them, and I could have worked out the logic behind /m layer to design a model for it that's sole purpose is to predict that, as none of my models know what they're doing when predicting m0 / m1</p>",
      "rawMarkdown": "I did and it helps. Also I train different models with different aspect ratios.\nA few models, each above CV of 1.00 gets down to 0.80 with just ensembling and down to 0.76 with beam size of 16. Normalization pushes that down to 0.72.\nI'm more and more sceptical about that gold, people are just going mad with their scores. :D\nEspecially @fergusoci with a solo 0.65, amazing work!\nThere must be some more clever tricks, but I've ran out of ideas.\n\nI could've worked more on big molecules, as my models don't do well on them, and I could have worked out the logic behind /m layer to design a model for it that's sole purpose is to predict that, as none of my models know what they're doing when predicting m0 / m1",
      "votes": null
    },
    {
      "id": "1326253",
      "postDate": "05/28/2021 10:59:57",
      "content": "<p>Great tricks. Thank you so much <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> </p>",
      "rawMarkdown": "Great tricks. Thank you so much @nofreewill",
      "votes": null
    },
    {
      "id": "1326316",
      "postDate": "05/28/2021 11:51:03",
      "content": "<p>\"I could've worked more on big molecules, as my models don't do well on them,\"<br>\nyou should do something like this to find the hotspot:</p>\n<pre><code>seq len   count (percent)   lb_score    lb_score*percent\n---------------------------------------------------------------\n 0 :         0 (0.000)   : 0.00000   : 0.00000  \n20 :       446 (0.000)   : 0.81596   : 0.00036  \n40 :     59697 (0.060)   : 0.16666   : 0.00995  \n60 :    307855 (0.308)   : 0.13051   : 0.04018  \n80 :    306389 (0.306)   : 0.30007   : 0.09195  \n100 :    210539 (0.211)   : 0.71796   : 0.15117  \n120 :     80914 (0.081)   : 1.96946   : 0.15937  \n140 :     25564 (0.026)   : 4.39088   : 0.11226  \n160 :      6094 (0.006)   : 11.44041   : 0.06972  \n180 :      1230 (0.001)   : 34.99797   : 0.04305  \n200 :       619 (0.001)   : 70.19545   : 0.04346  \n220 :       397 (0.000)   : 78.71569   : 0.03125  \n240 :       146 (0.000)   : 86.23546   : 0.01259  \n260 :        17 (0.000)   : 96.78363   : 0.00165  \n                               sum      0.767004   \n</code></pre>\n<p>ps: seq length is correlated to num of black pixels (after filtering the pepper noise). you can use this to estimate test seq length</p>",
      "rawMarkdown": "\"I could've worked more on big molecules, as my models don't do well on them,\"\n\nyou should do something like this to find the hotspot:\n\n```\n\nseq len   count (percent)   lb_score    lb_score*percent\n---------------------------------------------------------------\n  0 :         0 (0.000)   : 0.00000   : 0.00000  \n 20 :       446 (0.000)   : 0.81596   : 0.00036  \n 40 :     59697 (0.060)   : 0.16666   : 0.00995  \n 60 :    307855 (0.308)   : 0.13051   : 0.04018  \n 80 :    306389 (0.306)   : 0.30007   : 0.09195  \n100 :    210539 (0.211)   : 0.71796   : 0.15117  \n120 :     80914 (0.081)   : 1.96946   : 0.15937  \n140 :     25564 (0.026)   : 4.39088   : 0.11226  \n160 :      6094 (0.006)   : 11.44041   : 0.06972  \n180 :      1230 (0.001)   : 34.99797   : 0.04305  \n200 :       619 (0.001)   : 70.19545   : 0.04346  \n220 :       397 (0.000)   : 78.71569   : 0.03125  \n240 :       146 (0.000)   : 86.23546   : 0.01259  \n260 :        17 (0.000)   : 96.78363   : 0.00165  \n                                sum      0.767004   \n```\n\nps: seq length is correlated to num of black pixels (after filtering the pepper noise). you can use this to estimate test seq length",
      "votes": null
    },
    {
      "id": "1326345",
      "postDate": "05/28/2021 12:13:36",
      "content": "<p>I don't understand this table.</p>\n<p>Is this really lb_score, or was that supposed to be ld_score (cv). I don't think so, as it's almost a million samples. ?</p>\n<p>Or did you submit an all empty string, and then took your model's prediction on molecules that are ie. 0-20 long and submit only those, and figure out the LD of these based on this and the empty submission?</p>",
      "rawMarkdown": "I don't understand this table.\n\nIs this really lb_score, or was that supposed to be ld_score (cv). I don't think so, as it's almost a million samples. ?\n\nOr did you submit an all empty string, and then took your model's prediction on molecules that are ie. 0-20 long and submit only those, and figure out the LD of these based on this and the empty submission?",
      "votes": null
    },
    {
      "id": "1326581",
      "postDate": "05/28/2021 14:51:38",
      "content": "<p>you can use train samples or their augmentation (this is upper limit since there is no generalization error), etc. (you also have extra csv if you know how to render them)</p>\n<p>large images score poorly, but there aren't many of them.</p>",
      "rawMarkdown": "you can use train samples or their augmentation (this is upper limit since there is no generalization error), etc. (you also have extra csv if you know how to render them)\n\nlarge images score poorly, but there aren't many of them.",
      "votes": null
    },
    {
      "id": "1326709",
      "postDate": "05/28/2021 16:10:31",
      "content": "<p>\"There must be some more clever tricks, but I've ran out of ideas.\"</p>\n<p>one can think of: how much improvement can i get</p>\n<ul>\n<li>if  i improve the network architecture?</li>\n<li>if i ensemble or post process?</li>\n<li>if i add data, augment, etc …</li>\n</ul>\n<hr>\n<p>one have to understand that \"ensemble of 4 vgg = one resnet50\"<br>\neven with better network architecture, you can improvement of only 1 to 2 %.</p>\n<p>one can spend more time to perturb the model or image, then think of a way to combine them.<br>\nthis remove generalisation, so limit is from validation to train score</p>\n<p>remember, you have a rdkit validator (or beam search or other post processing). use this to help in combining results etc.</p>\n<hr>\n<p>many people think of how to improve.<br>\ni would urge people to \"estimate of the limit (aka upper bound) of each method\" too.<br>\nif you got time, think of the reason for this limit too</p>\n<hr>\n<p>top teams have estimated the top score is about 0.60 a long time ago. this is not a random guess or coincidence.</p>",
      "rawMarkdown": "\"There must be some more clever tricks, but I've ran out of ideas.\"\n\none can think of: how much improvement can i get\n- if  i improve the network architecture?\n- if i ensemble or post process?\n- if i add data, augment, etc ...\n\n\n---\n\none have to understand that \"ensemble of 4 vgg = one resnet50\"\neven with better network architecture, you can improvement of only 1 to 2 %.\n\none can spend more time to perturb the model or image, then think of a way to combine them.\nthis remove generalisation, so limit is from validation to train score\n\nremember, you have a rdkit validator (or beam search or other post processing). use this to help in combining results etc.\n\n---\n\nmany people think of how to improve.\ni would urge people to \"estimate of the limit (aka upper bound) of each method\" too.\nif you got time, think of the reason for this limit too\n\n---\n\ntop teams have estimated the top score is about 0.60 a long time ago. this is not a random guess or coincidence.",
      "votes": null
    },
    {
      "id": "1327007",
      "postDate": "05/28/2021 19:31:28",
      "content": "<p>Yeah. Thanks for the inputs. These small tricks really help in improving the score. </p>",
      "rawMarkdown": "Yeah. Thanks for the inputs. These small tricks really help in improving the score.",
      "votes": null
    },
    {
      "id": "1327095",
      "postDate": "05/28/2021 22:26:05",
      "content": "<p><a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a>, given the quote below from the InChI trust's <a href=\"https://www.inchi-trust.org/technical-faq-2/\" target=\"_blank\">technical manual</a>, it might be that it's almost impossible to decide which is m0 and which is m1</p>\n<p>\"It is not evident how the mark m0 or m1 is assigned in the stereochemistry /t sub-layer… so are these marks of any interest?<br>\nThe only significant thing from a practical viewpoint is that the enantiomers of a chiral molecule do have the same ‘/t’ layer but different /m indicators, either /m0 or /m1.</p>\n<p>For example, the two enantiomers of bromochlorofluoromethane are represented as the following Standard InChIs:</p>\n<p>InChI=1S/CHBrClF/c2-1(3)4/h1H/t1-/m0/s1<br>\nInChI=1S/CHBrClF/c2-1(3)4/h1H/t1-/m1/s1<br>\nActually, the first string is generated for the (R)- and the second for the (S)-stereoisomer. However, there is no simple relation between InChI parities +/- and R/S configurations of stereocenters; InChI does not use CIP rules and deduces parities from its own canonical numbers of atoms.</p>\n<p>For diastereomers, InChI strings will differ by {/t../m..} sub-strings. Among those diastereomers, stereoisomers in each enantiomeric pair will have the same {/t..} sub-string but different {/m..}. For example, the four stereoisomers of (F)(Cl)C-C(Br)(I) are represented as:</p>\n<p>(R,R) InChI=1S/C2H2BrClFI/c3-1(6)2(4)5/h1-2H/t1-,2+/m1/s1<br>\n(S,S) InChI=1S/C2H2BrClFI/c3-1(6)2(4)5/h1-2H/t1-,2+/m0/s1<br>\n(R,S) InChI=1S/C2H2BrClFI/c3-1(6)2(4)5/h1-2H/t1-,2-/m0/s1<br>\n(S,R) InChI=1S/C2H2BrClFI/c3-1(6)2(4)5/h1-2H/t1-,2-/m1/s1<br>\n(again, the RS-convention is used just for convenience)\"</p>",
      "rawMarkdown": "nofreewill, given the quote below from the InChI trust's [technical manual](https://www.inchi-trust.org/technical-faq-2/), it might be that it's almost impossible to decide which is m0 and which is m1\n\n\"It is not evident how the mark m0 or m1 is assigned in the stereochemistry /t sub-layer… so are these marks of any interest?\nThe only significant thing from a practical viewpoint is that the enantiomers of a chiral molecule do have the same ‘/t’ layer but different /m indicators, either /m0 or /m1.\n\nFor example, the two enantiomers of bromochlorofluoromethane are represented as the following Standard InChIs:\n\nInChI=1S/CHBrClF/c2-1(3)4/h1H/t1-/m0/s1\nInChI=1S/CHBrClF/c2-1(3)4/h1H/t1-/m1/s1\nActually, the first string is generated for the (R)- and the second for the (S)-stereoisomer. However, there is no simple relation between InChI parities +/- and R/S configurations of stereocenters; InChI does not use CIP rules and deduces parities from its own canonical numbers of atoms.\n\nFor diastereomers, InChI strings will differ by {/t../m..} sub-strings. Among those diastereomers, stereoisomers in each enantiomeric pair will have the same {/t..} sub-string but different {/m..}. For example, the four stereoisomers of (F)(Cl)C-C(Br)(I) are represented as:\n\n(R,R) InChI=1S/C2H2BrClFI/c3-1(6)2(4)5/h1-2H/t1-,2+/m1/s1\n(S,S) InChI=1S/C2H2BrClFI/c3-1(6)2(4)5/h1-2H/t1-,2+/m0/s1\n(R,S) InChI=1S/C2H2BrClFI/c3-1(6)2(4)5/h1-2H/t1-,2-/m0/s1\n(S,R) InChI=1S/C2H2BrClFI/c3-1(6)2(4)5/h1-2H/t1-,2-/m1/s1\n(again, the RS-convention is used just for convenience)\"",
      "votes": null
    },
    {
      "id": "1327116",
      "postDate": "05/28/2021 23:29:23",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> How would you use RDKit Validator to combine results?</p>",
      "rawMarkdown": "hengck23 How would you use RDKit Validator to combine results?",
      "votes": null
    },
    {
      "id": "1330579",
      "postDate": "06/01/2021 01:09:47",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <code>remember, you have a rdkit validator</code><br>\nI only used it for normalization. Thanks for mentioning that, you made me think and kept me going!!</p>",
      "rawMarkdown": "hengck23 `remember, you have a rdkit validator`\nI only used it for normalization. Thanks for mentioning that, you made me think and kept me going!!",
      "votes": null
    },
    {
      "id": "1330621",
      "postDate": "06/01/2021 02:29:47",
      "content": "<p>everything has already been mentioned in the forum long time ago by different kagglers.</p>\n<p>one can dig and read all previous post in the kaggle forum for details.</p>\n<p>when many kagglers obtained much lower score than you are, it is a open secret.<br>\nit must have been mentioned somewhere in the forum/code at some time …</p>",
      "rawMarkdown": "everything has already been mentioned in the forum long time ago by different kagglers.\n \none can dig and read all previous post in the kaggle forum for details.\n\nwhen many kagglers obtained much lower score than you are, it is a open secret.\nit must have been mentioned somewhere in the forum/code at some time ...",
      "votes": null
    },
    {
      "id": "1330794",
      "postDate": "06/01/2021 05:33:57",
      "content": "<p>Yes, it was mentioned, I've read it, too. But I still missed it somehow. :D<br>\nSo no, really! Thank you:)</p>",
      "rawMarkdown": "Yes, it was mentioned, I've read it, too. But I still missed it somehow. :D\nSo no, really! Thank you:)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1325273,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "05/27/2021 16:03:40",
      "content": "<p>Easy ensembling can be done at each time step during prediction.  But you will need to use the same vocabulary for all models. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1325964,
          "author_name": "vikrant06",
          "author_url": "",
          "post_date": "05/28/2021 06:15:06",
          "content": "<p>sounds interesting. So did you apply this and got any improvement? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1326015,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "05/28/2021 06:59:07",
          "content": "<p>I did and it helps. Also I train different models with different aspect ratios.<br>\nA few models, each above CV of 1.00 gets down to 0.80 with just ensembling and down to 0.76 with beam size of 16. Normalization pushes that down to 0.72.<br>\nI'm more and more sceptical about that gold, people are just going mad with their scores. :D<br>\nEspecially <a href=\"https://www.kaggle.com/fergusoci\" target=\"_blank\">@fergusoci</a> with a solo 0.65, amazing work!<br>\nThere must be some more clever tricks, but I've ran out of ideas.</p>\n<p>I could've worked more on big molecules, as my models don't do well on them, and I could have worked out the logic behind /m layer to design a model for it that's sole purpose is to predict that, as none of my models know what they're doing when predicting m0 / m1</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1326253,
          "author_name": "vikrant06",
          "author_url": "",
          "post_date": "05/28/2021 10:59:57",
          "content": "<p>Great tricks. Thank you so much <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1326316,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "05/28/2021 11:51:03",
          "content": "<p>\"I could've worked more on big molecules, as my models don't do well on them,\"<br>\nyou should do something like this to find the hotspot:</p>\n<pre><code>seq len   count (percent)   lb_score    lb_score*percent\n---------------------------------------------------------------\n 0 :         0 (0.000)   : 0.00000   : 0.00000  \n20 :       446 (0.000)   : 0.81596   : 0.00036  \n40 :     59697 (0.060)   : 0.16666   : 0.00995  \n60 :    307855 (0.308)   : 0.13051   : 0.04018  \n80 :    306389 (0.306)   : 0.30007   : 0.09195  \n100 :    210539 (0.211)   : 0.71796   : 0.15117  \n120 :     80914 (0.081)   : 1.96946   : 0.15937  \n140 :     25564 (0.026)   : 4.39088   : 0.11226  \n160 :      6094 (0.006)   : 11.44041   : 0.06972  \n180 :      1230 (0.001)   : 34.99797   : 0.04305  \n200 :       619 (0.001)   : 70.19545   : 0.04346  \n220 :       397 (0.000)   : 78.71569   : 0.03125  \n240 :       146 (0.000)   : 86.23546   : 0.01259  \n260 :        17 (0.000)   : 96.78363   : 0.00165  \n                               sum      0.767004   \n</code></pre>\n<p>ps: seq length is correlated to num of black pixels (after filtering the pepper noise). you can use this to estimate test seq length</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1326345,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "05/28/2021 12:13:36",
          "content": "<p>I don't understand this table.</p>\n<p>Is this really lb_score, or was that supposed to be ld_score (cv). I don't think so, as it's almost a million samples. ?</p>\n<p>Or did you submit an all empty string, and then took your model's prediction on molecules that are ie. 0-20 long and submit only those, and figure out the LD of these based on this and the empty submission?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1326581,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "05/28/2021 14:51:38",
          "content": "<p>you can use train samples or their augmentation (this is upper limit since there is no generalization error), etc. (you also have extra csv if you know how to render them)</p>\n<p>large images score poorly, but there aren't many of them.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1326709,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "05/28/2021 16:10:31",
          "content": "<p>\"There must be some more clever tricks, but I've ran out of ideas.\"</p>\n<p>one can think of: how much improvement can i get</p>\n<ul>\n<li>if  i improve the network architecture?</li>\n<li>if i ensemble or post process?</li>\n<li>if i add data, augment, etc …</li>\n</ul>\n<hr>\n<p>one have to understand that \"ensemble of 4 vgg = one resnet50\"<br>\neven with better network architecture, you can improvement of only 1 to 2 %.</p>\n<p>one can spend more time to perturb the model or image, then think of a way to combine them.<br>\nthis remove generalisation, so limit is from validation to train score</p>\n<p>remember, you have a rdkit validator (or beam search or other post processing). use this to help in combining results etc.</p>\n<hr>\n<p>many people think of how to improve.<br>\ni would urge people to \"estimate of the limit (aka upper bound) of each method\" too.<br>\nif you got time, think of the reason for this limit too</p>\n<hr>\n<p>top teams have estimated the top score is about 0.60 a long time ago. this is not a random guess or coincidence.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1327007,
          "author_name": "vikrant06",
          "author_url": "",
          "post_date": "05/28/2021 19:31:28",
          "content": "<p>Yeah. Thanks for the inputs. These small tricks really help in improving the score. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1327095,
          "author_name": "jbomitchell",
          "author_url": "",
          "post_date": "05/28/2021 22:26:05",
          "content": "<p><a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a>, given the quote below from the InChI trust's <a href=\"https://www.inchi-trust.org/technical-faq-2/\" target=\"_blank\">technical manual</a>, it might be that it's almost impossible to decide which is m0 and which is m1</p>\n<p>\"It is not evident how the mark m0 or m1 is assigned in the stereochemistry /t sub-layer… so are these marks of any interest?<br>\nThe only significant thing from a practical viewpoint is that the enantiomers of a chiral molecule do have the same ‘/t’ layer but different /m indicators, either /m0 or /m1.</p>\n<p>For example, the two enantiomers of bromochlorofluoromethane are represented as the following Standard InChIs:</p>\n<p>InChI=1S/CHBrClF/c2-1(3)4/h1H/t1-/m0/s1<br>\nInChI=1S/CHBrClF/c2-1(3)4/h1H/t1-/m1/s1<br>\nActually, the first string is generated for the (R)- and the second for the (S)-stereoisomer. However, there is no simple relation between InChI parities +/- and R/S configurations of stereocenters; InChI does not use CIP rules and deduces parities from its own canonical numbers of atoms.</p>\n<p>For diastereomers, InChI strings will differ by {/t../m..} sub-strings. Among those diastereomers, stereoisomers in each enantiomeric pair will have the same {/t..} sub-string but different {/m..}. For example, the four stereoisomers of (F)(Cl)C-C(Br)(I) are represented as:</p>\n<p>(R,R) InChI=1S/C2H2BrClFI/c3-1(6)2(4)5/h1-2H/t1-,2+/m1/s1<br>\n(S,S) InChI=1S/C2H2BrClFI/c3-1(6)2(4)5/h1-2H/t1-,2+/m0/s1<br>\n(R,S) InChI=1S/C2H2BrClFI/c3-1(6)2(4)5/h1-2H/t1-,2-/m0/s1<br>\n(S,R) InChI=1S/C2H2BrClFI/c3-1(6)2(4)5/h1-2H/t1-,2-/m1/s1<br>\n(again, the RS-convention is used just for convenience)\"</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1330579,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "06/01/2021 01:09:47",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <code>remember, you have a rdkit validator</code><br>\nI only used it for normalization. Thanks for mentioning that, you made me think and kept me going!!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1330621,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "06/01/2021 02:29:47",
          "content": "<p>everything has already been mentioned in the forum long time ago by different kagglers.</p>\n<p>one can dig and read all previous post in the kaggle forum for details.</p>\n<p>when many kagglers obtained much lower score than you are, it is a open secret.<br>\nit must have been mentioned somewhere in the forum/code at some time …</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1330794,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "06/01/2021 05:33:57",
          "content": "<p>Yes, it was mentioned, I've read it, too. But I still missed it somehow. :D<br>\nSo no, really! Thank you:)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1327116,
      "author_name": "andrewshao05",
      "author_url": "",
      "post_date": "05/28/2021 23:29:23",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> How would you use RDKit Validator to combine results?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1325020": "Did anyone try any ensembling techniques other than voting?\n\nHere are few updates from comments:\n\n1. Ensembling at each time step during prediction. It needs the same vocabulary for all models. \n thanks @serigne   \n2. Using rdkit validator to combine the results of different models. \nthanks @hengck23",
    "1325273": "Easy ensembling can be done at each time step during prediction.  But you will need to use the same vocabulary for all models.",
    "1325964": "sounds interesting. So did you apply this and got any improvement?",
    "1326015": "I did and it helps. Also I train different models with different aspect ratios.\nA few models, each above CV of 1.00 gets down to 0.80 with just ensembling and down to 0.76 with beam size of 16. Normalization pushes that down to 0.72.\nI'm more and more sceptical about that gold, people are just going mad with their scores. :D\nEspecially @fergusoci with a solo 0.65, amazing work!\nThere must be some more clever tricks, but I've ran out of ideas.\n\nI could've worked more on big molecules, as my models don't do well on them, and I could have worked out the logic behind /m layer to design a model for it that's sole purpose is to predict that, as none of my models know what they're doing when predicting m0 / m1",
    "1326253": "Great tricks. Thank you so much @nofreewill",
    "1326316": "\"I could've worked more on big molecules, as my models don't do well on them,\"\n\nyou should do something like this to find the hotspot:\n\n```\n\nseq len   count (percent)   lb_score    lb_score*percent\n---------------------------------------------------------------\n  0 :         0 (0.000)   : 0.00000   : 0.00000  \n 20 :       446 (0.000)   : 0.81596   : 0.00036  \n 40 :     59697 (0.060)   : 0.16666   : 0.00995  \n 60 :    307855 (0.308)   : 0.13051   : 0.04018  \n 80 :    306389 (0.306)   : 0.30007   : 0.09195  \n100 :    210539 (0.211)   : 0.71796   : 0.15117  \n120 :     80914 (0.081)   : 1.96946   : 0.15937  \n140 :     25564 (0.026)   : 4.39088   : 0.11226  \n160 :      6094 (0.006)   : 11.44041   : 0.06972  \n180 :      1230 (0.001)   : 34.99797   : 0.04305  \n200 :       619 (0.001)   : 70.19545   : 0.04346  \n220 :       397 (0.000)   : 78.71569   : 0.03125  \n240 :       146 (0.000)   : 86.23546   : 0.01259  \n260 :        17 (0.000)   : 96.78363   : 0.00165  \n                                sum      0.767004   \n```\n\nps: seq length is correlated to num of black pixels (after filtering the pepper noise). you can use this to estimate test seq length",
    "1326345": "I don't understand this table.\n\nIs this really lb_score, or was that supposed to be ld_score (cv). I don't think so, as it's almost a million samples. ?\n\nOr did you submit an all empty string, and then took your model's prediction on molecules that are ie. 0-20 long and submit only those, and figure out the LD of these based on this and the empty submission?",
    "1326581": "you can use train samples or their augmentation (this is upper limit since there is no generalization error), etc. (you also have extra csv if you know how to render them)\n\nlarge images score poorly, but there aren't many of them.",
    "1326709": "\"There must be some more clever tricks, but I've ran out of ideas.\"\n\none can think of: how much improvement can i get\n- if  i improve the network architecture?\n- if i ensemble or post process?\n- if i add data, augment, etc ...\n\n\n---\n\none have to understand that \"ensemble of 4 vgg = one resnet50\"\neven with better network architecture, you can improvement of only 1 to 2 %.\n\none can spend more time to perturb the model or image, then think of a way to combine them.\nthis remove generalisation, so limit is from validation to train score\n\nremember, you have a rdkit validator (or beam search or other post processing). use this to help in combining results etc.\n\n---\n\nmany people think of how to improve.\ni would urge people to \"estimate of the limit (aka upper bound) of each method\" too.\nif you got time, think of the reason for this limit too\n\n---\n\ntop teams have estimated the top score is about 0.60 a long time ago. this is not a random guess or coincidence.",
    "1327007": "Yeah. Thanks for the inputs. These small tricks really help in improving the score.",
    "1327095": "nofreewill, given the quote below from the InChI trust's [technical manual](https://www.inchi-trust.org/technical-faq-2/), it might be that it's almost impossible to decide which is m0 and which is m1\n\n\"It is not evident how the mark m0 or m1 is assigned in the stereochemistry /t sub-layer… so are these marks of any interest?\nThe only significant thing from a practical viewpoint is that the enantiomers of a chiral molecule do have the same ‘/t’ layer but different /m indicators, either /m0 or /m1.\n\nFor example, the two enantiomers of bromochlorofluoromethane are represented as the following Standard InChIs:\n\nInChI=1S/CHBrClF/c2-1(3)4/h1H/t1-/m0/s1\nInChI=1S/CHBrClF/c2-1(3)4/h1H/t1-/m1/s1\nActually, the first string is generated for the (R)- and the second for the (S)-stereoisomer. However, there is no simple relation between InChI parities +/- and R/S configurations of stereocenters; InChI does not use CIP rules and deduces parities from its own canonical numbers of atoms.\n\nFor diastereomers, InChI strings will differ by {/t../m..} sub-strings. Among those diastereomers, stereoisomers in each enantiomeric pair will have the same {/t..} sub-string but different {/m..}. For example, the four stereoisomers of (F)(Cl)C-C(Br)(I) are represented as:\n\n(R,R) InChI=1S/C2H2BrClFI/c3-1(6)2(4)5/h1-2H/t1-,2+/m1/s1\n(S,S) InChI=1S/C2H2BrClFI/c3-1(6)2(4)5/h1-2H/t1-,2+/m0/s1\n(R,S) InChI=1S/C2H2BrClFI/c3-1(6)2(4)5/h1-2H/t1-,2-/m0/s1\n(S,R) InChI=1S/C2H2BrClFI/c3-1(6)2(4)5/h1-2H/t1-,2-/m1/s1\n(again, the RS-convention is used just for convenience)\"",
    "1327116": "hengck23 How would you use RDKit Validator to combine results?",
    "1330579": "hengck23 `remember, you have a rdkit validator`\nI only used it for normalization. Thanks for mentioning that, you made me think and kept me going!!",
    "1330621": "everything has already been mentioned in the forum long time ago by different kagglers.\n \none can dig and read all previous post in the kaggle forum for details.\n\nwhen many kagglers obtained much lower score than you are, it is a open secret.\nit must have been mentioned somewhere in the forum/code at some time ...",
    "1330794": "Yes, it was mentioned, I've read it, too. But I still missed it somehow. :D\nSo no, really! Thank you:)"
  },
  "source": "meta"
}