{
  "id": 243102,
  "title": "rdkit is more crucial",
  "url": "/competitions/bms-molecular-translation/discussion/243102",
  "author_name": "nofreewill42",
  "post_date": "2021-06-01T06:59:24.088000",
  "votes": 28,
  "comment_count": 35,
  "views": 0,
  "content": "<p>If you have multiple submissions, just modify the <a href=\"https://www.kaggle.com/nofreewill/normalize-your-predictions\" target=\"_blank\">normalization script</a> so that you note what happened with each InChI (valid, modified, mol_is_none, exception, segfault) and choose valid/modified ones based on some strategy.</p>\n<p>You can also mine the molecules for which you have no valid and/or modified predictions and predict only those with more resources, i.e. with more beams.</p>\n<p>I don't know how I missed this so far… it's all there throughout the discussions<br>\n<a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> thank you for opening my eyes!</p>",
  "messages": [
    {
      "id": 1330914,
      "postDate": "2021-06-01T06:59:24.090Z",
      "content": "<p>If you have multiple submissions, just modify the <a href=\"https://www.kaggle.com/nofreewill/normalize-your-predictions\" target=\"_blank\">normalization script</a> so that you note what happened with each InChI (valid, modified, mol_is_none, exception, segfault) and choose valid/modified ones based on some strategy.</p>\n<p>You can also mine the molecules for which you have no valid and/or modified predictions and predict only those with more resources, i.e. with more beams.</p>\n<p>I don't know how I missed this so far… it's all there throughout the discussions<br>\n<a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> thank you for opening my eyes!</p>",
      "rawMarkdown": "If you have multiple submissions, just modify the [normalization script](https://www.kaggle.com/nofreewill/normalize-your-predictions) so that you note what happened with each InChI (valid, modified, mol_is_none, exception, segfault) and choose valid/modified ones based on some strategy.\n\nYou can also mine the molecules for which you have no valid and/or modified predictions and predict only those with more resources, i.e. with more beams.\n\nI don't know how I missed this so far... it's all there throughout the discussions\n@hengck23 thank you for opening my eyes!",
      "votes": 28
    },
    {
      "id": 1331028,
      "postDate": "2021-06-01T08:25:57.320Z",
      "content": "<p>to be exact \"validation of inchi  is more crucial\"</p>\n<p>rdkit validation doesn't means correct (although it is mostly correct). If you assume valid = correct, then :</p>\n<p>your score is upper bounded by error of prediction that is valid. </p>\n<hr>\n<p>you may need to train a score predictor at the end of your decoder, or sum up the prob of each element of the seq.</p>\n<p>given a set of trained models, you are using heuristics to determine if the prediction is correct., e.g. valid inchi of rdkit. train a model to predict the LD score (given multiple predictions from different models and rdkit) may be better</p>",
      "rawMarkdown": "to be exact \"validation of inchi ~~rdkit~~ is more crucial\"\n\nrdkit validation doesn't means correct (although it is mostly correct). If you assume valid = correct, then :\n\nyour score is upper bounded by error of prediction that is valid. \n\n---\n\nyou may need to train a score predictor at the end of your decoder, or sum up the prob of each element of the seq.\n\ngiven a set of trained models, you are using heuristics to determine if the prediction is correct., e.g. valid inchi of rdkit. train a model to predict the LD score (given multiple predictions from different models and rdkit) may be better",
      "votes": 5,
      "replies": [
        {
          "id": 3140688,
          "postDate": "2025-03-04T20:26:45.493Z",
          "content": "<p>\"rdkit validation doesn't means correct\" - I haven't realised back then that this was also a low hanging fruit I haven't touched.<br>\nThere were so many possibilities for small improvements here-and-there.</p>\n<p>Also, to this day I have flashbacks that I basically robbed someone from their solo gold and with that to get the kaggle comp gm title.</p>",
          "rawMarkdown": "\"rdkit validation doesn't means correct\" - I haven't realised back then that this was also a low hanging fruit I haven't touched.\nThere were so many possibilities for small improvements here-and-there.\n\nAlso, to this day I have flashbacks that I basically robbed someone from their solo gold and with that to get the kaggle comp gm title."
        }
      ]
    },
    {
      "id": 1332519,
      "postDate": "2021-06-02T06:53:15.360Z",
      "content": "<p>Thanks. It did help me to decrease LB from 1.08 to 1.03. </p>",
      "rawMarkdown": "Thanks. It did help me to decrease LB from 1.08 to 1.03. ",
      "votes": 1,
      "replies": [
        {
          "id": 1332566,
          "postDate": "2021-06-02T07:30:11.483Z",
          "content": "<p>Glad to hear that, thank you for letting me know! :)</p>",
          "rawMarkdown": "Glad to hear that, thank you for letting me know! :)",
          "votes": -1
        },
        {
          "id": 1332574,
          "postDate": "2021-06-02T07:37:50.063Z",
          "content": "<p>We had an improvement (maybe a big) with this too before. not sure sure about it now. But you are idea is good and amazing. and also crucial</p>",
          "rawMarkdown": "We had an improvement (maybe a big) with this too before. not sure sure about it now. But you are idea is good and amazing. and also crucial",
          "votes": 1
        },
        {
          "id": 1332876,
          "postDate": "2021-06-02T11:07:49.210Z",
          "content": "<p>It's not my idea, there are a couple mentions of it in the discussions. I just made it more available.</p>",
          "rawMarkdown": "It's not my idea, there are a couple mentions of it in the discussions. I just made it more available.",
          "votes": -1
        }
      ]
    },
    {
      "id": 1332477,
      "postDate": "2021-06-02T06:29:00.053Z",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a></p>",
      "rawMarkdown": "Thanks for sharing @nofreewill",
      "votes": 1,
      "replies": [
        {
          "id": 1332490,
          "postDate": "2021-06-02T06:37:17.200Z",
          "content": "<p>Did it help you? You're welcome! :))</p>",
          "rawMarkdown": "Did it help you? You're welcome! :))",
          "votes": -1
        },
        {
          "id": 1332506,
          "postDate": "2021-06-02T06:44:34.610Z",
          "content": "<p>Not yet, still trying our best to figure out how to use it , trying to decipher your words haha</p>",
          "rawMarkdown": "Not yet, still trying our best to figure out how to use it , trying to decipher your words haha"
        },
        {
          "id": 1332564,
          "postDate": "2021-06-02T07:27:33.660Z",
          "content": "<p>In my <strong>validation</strong>:</p>\n<ul>\n<li>I predicted each molecule with all of my models one-by-one</li>\n<li>Normalized all the predictions</li>\n<li>Searched those that ended up with the same normalized form</li>\n<li>This is about 77% of all molecules, this part has LD of 0.145</li>\n</ul>\n<p>Then</p>\n<ul>\n<li>I predicted the remaining (23%) molecules with ensembled beam search -&gt; 2.7 LD (normalized, only the 23% part)</li>\n<li>Combined the two, got 0.727 LD</li>\n</ul>\n<p>Then</p>\n<ul>\n<li>I searched the molecules from the 23% that were not valid/modified</li>\n<li>for these I searched for valid/modified predictions from all of my models' predictions, and if there was one, I took that</li>\n<li>this got me to 0.706 LD</li>\n</ul>\n<p>But now I'm stuck here, I think that my models are just not good enough (as all of them have CV higher than 1 LD)<br>\nI may be able to work out some sophisticated strategies on how to choose between all the predictions if there are multiple ones being the same, but sadly I just have no time for that despite the extension.</p>",
          "rawMarkdown": "In my **validation**:\n\n - I predicted each molecule with all of my models one-by-one\n - Normalized all the predictions\n - Searched those that ended up with the same normalized form\n - This is about 77% of all molecules, this part has LD of 0.145\n\nThen\n - I predicted the remaining (23%) molecules with ensembled beam search -> 2.7 LD (normalized, only the 23% part)\n - Combined the two, got 0.727 LD\n\nThen\n - I searched the molecules from the 23% that were not valid/modified\n - for these I searched for valid/modified predictions from all of my models' predictions, and if there was one, I took that\n - this got me to 0.706 LD\n\nBut now I'm stuck here, I think that my models are just not good enough (as all of them have CV higher than 1 LD)\nI may be able to work out some sophisticated strategies on how to choose between all the predictions if there are multiple ones being the same, but sadly I just have no time for that despite the extension."
        },
        {
          "id": 1332602,
          "postDate": "2021-06-02T07:54:25.300Z",
          "content": "<p><a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> its humble of you to share but you should not have done it before the deadline , this might be unfair to people who had figured this out on their own . This could prove vital for a win and would have required good amount of digging </p>",
          "rawMarkdown": "@nofreewill its humble of you to share but you should not have done it before the deadline , this might be unfair to people who had figured this out on their own . This could prove vital for a win and would have required good amount of digging ",
          "votes": 6
        },
        {
          "id": 1332761,
          "postDate": "2021-06-02T09:27:03.553Z",
          "content": "<p>That's a good amount of downvotes.<br>\nWhat I said is all out in the discussions, so I don't agree with that hate.</p>",
          "rawMarkdown": "That's a good amount of downvotes.\nWhat I said is all out in the discussions, so I don't agree with that hate."
        },
        {
          "id": 1332877,
          "postDate": "2021-06-02T11:10:51.107Z",
          "content": "<p>There are also a lot of nuances that needs a lot of work.<br>\nI.e. how does beam size effect LD of shorter and longer inchis? How do you select your prediction if there are more inchis that are valid, and how do you take into account the beam size effect? How do you take into account if your models are trained on different aspect ratios? And many other things that needs a lot of work to optimize.</p>",
          "rawMarkdown": "There are also a lot of nuances that needs a lot of work.\nI.e. how does beam size effect LD of shorter and longer inchis? How do you select your prediction if there are more inchis that are valid, and how do you take into account the beam size effect? How do you take into account if your models are trained on different aspect ratios? And many other things that needs a lot of work to optimize.",
          "votes": 5
        },
        {
          "id": 1332904,
          "postDate": "2021-06-02T11:30:58.650Z",
          "content": "<p>So I guess that the downvotes come from teams that enjoyed that this idea is less available and didn't want to work harder on these issues.</p>",
          "rawMarkdown": "So I guess that the downvotes come from teams that enjoyed that this idea is less available and didn't want to work harder on these issues.",
          "votes": 2
        },
        {
          "id": 1333385,
          "postDate": "2021-06-02T17:14:21.663Z",
          "content": "<p>don't worry that can be normalised with upvotes ;)</p>",
          "rawMarkdown": "don't worry that can be normalised with upvotes ;)",
          "votes": 2
        },
        {
          "id": 1333397,
          "postDate": "2021-06-02T17:29:39.847Z",
          "content": "<p>haha, thank you! :D I'm keeping an eye on the number of votes, it's a roller coaster.. :D<br>\nI get why one might be angry, but really.. this is such a simple idea that it would be a shame if 3 month of one's hard work would go into garbage because he missed it but others did not.</p>",
          "rawMarkdown": "haha, thank you! :D I'm keeping an eye on the number of votes, it's a roller coaster.. :D\nI get why one might be angry, but really.. this is such a simple idea that it would be a shame if 3 month of one's hard work would go into garbage because he missed it but others did not.",
          "votes": -1
        },
        {
          "id": 1333419,
          "postDate": "2021-06-02T17:42:50.350Z",
          "content": "<p>I appreciate your humble sharing in the competition and It only worked in my favour that you shared exactly what you did and saved us time .<br>\nI just said on a general note that it might hurt people , there is a reason why all the gms and top teams don't share their secret sauce right untill the end of competition no matter how easy it is 😉</p>\n<p>Just to put things into perpective , the first competition I participated in people found a very easy post-processing like this that could be utilized but no one shared untill the end of the competition</p>",
          "rawMarkdown": "I appreciate your humble sharing in the competition and It only worked in my favour that you shared exactly what you did and saved us time .\nI just said on a general note that it might hurt people , there is a reason why all the gms and top teams don't share their secret sauce right untill the end of competition no matter how easy it is 😉\n\nJust to put things into perpective , the first competition I participated in people found a very easy post-processing like this that could be utilized but no one shared untill the end of the competition",
          "votes": 5
        },
        {
          "id": 1333481,
          "postDate": "2021-06-02T18:54:46.103Z",
          "content": "<p>You are right <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a>, the teams that are at the top which has already done it and found it useful to improve their score might think as a spoiler alert and downvotes it. </p>",
          "rawMarkdown": "You are right @tanulsingh077, the teams that are at the top which has already done it and found it useful to improve their score might think as a spoiler alert and downvotes it. "
        },
        {
          "id": 1333499,
          "postDate": "2021-06-02T19:02:23.203Z",
          "content": "<p>Don't worry. It is just a sensitive moment.  Eventually, people only remember your method and generosity.</p>",
          "rawMarkdown": "Don't worry. It is just a sensitive moment.  Eventually, people only remember your method and generosity.",
          "votes": 1
        },
        {
          "id": 1333501,
          "postDate": "2021-06-02T19:12:49.517Z",
          "content": "<p><a href=\"https://www.kaggle.com/ybwu01\" target=\"_blank\">@ybwu01</a> wow, thanks for your kind words! :)<br>\nIt's a little funny how by downvoting a thing you know that works you actually confirm that it does work and encourage others to go onto that path.</p>",
          "rawMarkdown": "@ybwu01 wow, thanks for your kind words! :)\nIt's a little funny how by downvoting a thing you know that works you actually confirm that it does work and encourage others to go onto that path."
        },
        {
          "id": 1333733,
          "postDate": "2021-06-03T02:44:28.490Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1334020,
          "postDate": "2021-06-03T08:10:59.360Z",
          "content": "<p>This happens all the time and I believe someone summed it up best in another competition as \"people already using the trick will be unhappy while people new to it will be happy\".</p>\n<p>I personally think it's Kaggle etiquette not to share anything that might have a material impact on the LB late in the competition - the caution on high-scoring notebooks is another facet of it.</p>",
          "rawMarkdown": "This happens all the time and I believe someone summed it up best in another competition as \"people already using the trick will be unhappy while people new to it will be happy\".\n\nI personally think it's Kaggle etiquette not to share anything that might have a material impact on the LB late in the competition - the caution on high-scoring notebooks is another facet of it.",
          "votes": 4
        },
        {
          "id": 1334141,
          "postDate": "2021-06-03T09:41:52.520Z",
          "content": "<p>With the recent changes on the leaderboard we at least know who already used the trick and who didn't.<br>\nQuite surprising for me to see so interrupting changes. But we will see if carries to the private board.</p>\n<p>I myself have decided to not upload a submission at all. I have done a lot of experiments to reproduce some papers and trying out stuff on myself. It was fun and I learned a lot. Meanwhile I never went for the highest-most score, just an efficient model (more on the encoder side). In the end I run out of time to train for long enough to be competitive (enough for my taste).</p>\n<p>I think I will make a post after the submission end, detailing my most interesting findings.</p>\n<p>Edit: Cheers to <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a> for keeping up with the teams in the gold-region so far.</p>",
          "rawMarkdown": "With the recent changes on the leaderboard we at least know who already used the trick and who didn't.\nQuite surprising for me to see so interrupting changes. But we will see if carries to the private board.\n\nI myself have decided to not upload a submission at all. I have done a lot of experiments to reproduce some papers and trying out stuff on myself. It was fun and I learned a lot. Meanwhile I never went for the highest-most score, just an efficient model (more on the encoder side). In the end I run out of time to train for long enough to be competitive (enough for my taste).\n\nI think I will make a post after the submission end, detailing my most interesting findings.\n\nEdit: Cheers to @charmq for keeping up with the teams in the gold-region so far.",
          "votes": 2
        },
        {
          "id": 1334958,
          "postDate": "2021-06-03T23:37:02.507Z",
          "content": "<p>I have to wonder how much of the shake-up in the gold zone was driven by this reveal (or maybe not at all).</p>\n<p>I would not be happy if I was in one of the teams being punished… </p>",
          "rawMarkdown": "I have to wonder how much of the shake-up in the gold zone was driven by this reveal (or maybe not at all).\n\nI would not be happy if I was in one of the teams being punished... ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1331111,
      "postDate": "2021-06-01T09:26:30.843Z",
      "content": "<p>Wow! This idea is very very important! Thanks a lot for sharing! </p>",
      "rawMarkdown": "Wow! This idea is very very important! Thanks a lot for sharing! ",
      "votes": 1,
      "replies": [
        {
          "id": 1331120,
          "postDate": "2021-06-01T09:37:33.890Z",
          "content": "<p>It was worth sharing, then! :) You're welcome!</p>",
          "rawMarkdown": "It was worth sharing, then! :) You're welcome!"
        }
      ]
    },
    {
      "id": 1330928,
      "postDate": "2021-06-01T07:09:02.070Z",
      "content": "<p>Well, then I hope that is enough for you to achieve the solo gold you wanted :)</p>",
      "rawMarkdown": "Well, then I hope that is enough for you to achieve the solo gold you wanted :)",
      "votes": 1,
      "replies": [
        {
          "id": 1330936,
          "postDate": "2021-06-01T07:14:44.117Z",
          "content": "<p>Thank you! :)<br>\nI still have doubts, but now I have a little hope, at least.</p>\n<p>But I guess that a lot of teams left their submissions to the end, so… :/ :D</p>",
          "rawMarkdown": "Thank you! :)\nI still have doubts, but now I have a little hope, at least.\n\nBut I guess that a lot of teams left their submissions to the end, so... :/ :D",
          "votes": 1
        },
        {
          "id": 1331031,
          "postDate": "2021-06-01T08:27:02.570Z",
          "content": "<p>i estimate gold level should be better 0.65~0.62</p>",
          "rawMarkdown": "i estimate gold level should be better 0.65~0.62",
          "votes": 1
        },
        {
          "id": 1331136,
          "postDate": "2021-06-01T09:47:54.717Z",
          "content": "<p><a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> i think you were hiding your score for a month. I think you have a better model and will win gold I think.</p>\n<p>I estimate gold will be &lt;0.62</p>",
          "rawMarkdown": "@nofreewill i think you were hiding your score for a month. I think you have a better model and will win gold I think.\n\nI estimate gold will be <0.62",
          "votes": 1
        },
        {
          "id": 1331142,
          "postDate": "2021-06-01T09:53:10.190Z",
          "content": "<p>I also think gold will be less than 0.62<br>\nBut sadly, no, my best score I can achieve is 0.70 right now..<br>\nI don't think I'll have gold, but I'm trying my best :)</p>",
          "rawMarkdown": "I also think gold will be less than 0.62\nBut sadly, no, my best score I can achieve is 0.70 right now..\nI don't think I'll have gold, but I'm trying my best :)",
          "votes": 1
        },
        {
          "id": 1331244,
          "postDate": "2021-06-01T11:10:23.743Z",
          "content": "<p>OT: Has here anybody enough experience to confidently say, if the public leader board compared to the private + public (complete) leaderboard may either be:</p>\n<ol>\n<li>A random sub-sample</li>\n<li>A stratified sub-sample</li>\n<li>A purposefully easier sub-sample.</li>\n</ol>\n<p>In other words: Is the public test-set usually representative of the private one or not? Or does this strictly vary with the competition?</p>",
          "rawMarkdown": "OT: Has here anybody enough experience to confidently say, if the public leader board compared to the private + public (complete) leaderboard may either be:\n\n1. A random sub-sample\n2. A stratified sub-sample\n3. A purposefully easier sub-sample.\n\nIn other words: Is the public test-set usually representative of the private one or not? Or does this strictly vary with the competition?"
        },
        {
          "id": 1331607,
          "postDate": "2021-06-01T15:17:26.760Z",
          "content": "<p>Good Question.. Ideally, it should be a stratified sub sample but based on my bast experience most of the times it's random. That's why shakeup happens</p>",
          "rawMarkdown": "Good Question.. Ideally, it should be a stratified sub sample but based on my bast experience most of the times it's random. That's why shakeup happens"
        },
        {
          "id": 1331677,
          "postDate": "2021-06-01T16:15:17.263Z",
          "content": "<p><a href=\"https://www.kaggle.com/cepheidq\" target=\"_blank\">@cepheidq</a> sampling of the private set from whole test set can any of your mention and it varies with each competition.<br>\nIn this competition it might be random sampled or length wise sampling as <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> did for his CV.<br>\npurposeful maybe but lower probability.</p>",
          "rawMarkdown": "@cepheidq sampling of the private set from whole test set can any of your mention and it varies with each competition.\nIn this competition it might be random sampled or length wise sampling as @nofreewill did for his CV.\npurposeful maybe but lower probability.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1333714,
      "postDate": "2021-06-03T02:12:39.673Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1331028,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-06-01T08:25:57.320000",
      "content": "<p>to be exact \"validation of inchi  is more crucial\"</p>\n<p>rdkit validation doesn't means correct (although it is mostly correct). If you assume valid = correct, then :</p>\n<p>your score is upper bounded by error of prediction that is valid. </p>\n<hr>\n<p>you may need to train a score predictor at the end of your decoder, or sum up the prob of each element of the seq.</p>\n<p>given a set of trained models, you are using heuristics to determine if the prediction is correct., e.g. valid inchi of rdkit. train a model to predict the LD score (given multiple predictions from different models and rdkit) may be better</p>",
      "votes": 5,
      "replies": [
        {
          "id": 3140688,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2025-03-04T20:26:45.493000",
          "content": "<p>\"rdkit validation doesn't means correct\" - I haven't realised back then that this was also a low hanging fruit I haven't touched.<br>\nThere were so many possibilities for small improvements here-and-there.</p>\n<p>Also, to this day I have flashbacks that I basically robbed someone from their solo gold and with that to get the kaggle comp gm title.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1332519,
      "author_name": "Yibing Wu",
      "author_url": "",
      "post_date": "2021-06-02T06:53:15.360000",
      "content": "<p>Thanks. It did help me to decrease LB from 1.08 to 1.03. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1332566,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-06-02T07:30:11.483000",
          "content": "<p>Glad to hear that, thank you for letting me know! :)</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1332574,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2021-06-02T07:37:50.063000",
          "content": "<p>We had an improvement (maybe a big) with this too before. not sure sure about it now. But you are idea is good and amazing. and also crucial</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1332876,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-06-02T11:07:49.210000",
          "content": "<p>It's not my idea, there are a couple mentions of it in the discussions. I just made it more available.</p>",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 1332477,
      "author_name": "Mr_KnowNothing",
      "author_url": "",
      "post_date": "2021-06-02T06:29:00.053000",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 1332490,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-06-02T06:37:17.200000",
          "content": "<p>Did it help you? You're welcome! :))</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1332506,
          "author_name": "Mr_KnowNothing",
          "author_url": "",
          "post_date": "2021-06-02T06:44:34.610000",
          "content": "<p>Not yet, still trying our best to figure out how to use it , trying to decipher your words haha</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1332564,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-06-02T07:27:33.660000",
          "content": "<p>In my <strong>validation</strong>:</p>\n<ul>\n<li>I predicted each molecule with all of my models one-by-one</li>\n<li>Normalized all the predictions</li>\n<li>Searched those that ended up with the same normalized form</li>\n<li>This is about 77% of all molecules, this part has LD of 0.145</li>\n</ul>\n<p>Then</p>\n<ul>\n<li>I predicted the remaining (23%) molecules with ensembled beam search -&gt; 2.7 LD (normalized, only the 23% part)</li>\n<li>Combined the two, got 0.727 LD</li>\n</ul>\n<p>Then</p>\n<ul>\n<li>I searched the molecules from the 23% that were not valid/modified</li>\n<li>for these I searched for valid/modified predictions from all of my models' predictions, and if there was one, I took that</li>\n<li>this got me to 0.706 LD</li>\n</ul>\n<p>But now I'm stuck here, I think that my models are just not good enough (as all of them have CV higher than 1 LD)<br>\nI may be able to work out some sophisticated strategies on how to choose between all the predictions if there are multiple ones being the same, but sadly I just have no time for that despite the extension.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1332602,
          "author_name": "Mr_KnowNothing",
          "author_url": "",
          "post_date": "2021-06-02T07:54:25.300000",
          "content": "<p><a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> its humble of you to share but you should not have done it before the deadline , this might be unfair to people who had figured this out on their own . This could prove vital for a win and would have required good amount of digging </p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1332761,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-06-02T09:27:03.553000",
          "content": "<p>That's a good amount of downvotes.<br>\nWhat I said is all out in the discussions, so I don't agree with that hate.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1332877,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-06-02T11:10:51.107000",
          "content": "<p>There are also a lot of nuances that needs a lot of work.<br>\nI.e. how does beam size effect LD of shorter and longer inchis? How do you select your prediction if there are more inchis that are valid, and how do you take into account the beam size effect? How do you take into account if your models are trained on different aspect ratios? And many other things that needs a lot of work to optimize.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1332904,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-06-02T11:30:58.650000",
          "content": "<p>So I guess that the downvotes come from teams that enjoyed that this idea is less available and didn't want to work harder on these issues.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1333385,
          "author_name": "Vikrant",
          "author_url": "",
          "post_date": "2021-06-02T17:14:21.663000",
          "content": "<p>don't worry that can be normalised with upvotes ;)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1333397,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-06-02T17:29:39.847000",
          "content": "<p>haha, thank you! :D I'm keeping an eye on the number of votes, it's a roller coaster.. :D<br>\nI get why one might be angry, but really.. this is such a simple idea that it would be a shame if 3 month of one's hard work would go into garbage because he missed it but others did not.</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1333419,
          "author_name": "Mr_KnowNothing",
          "author_url": "",
          "post_date": "2021-06-02T17:42:50.350000",
          "content": "<p>I appreciate your humble sharing in the competition and It only worked in my favour that you shared exactly what you did and saved us time .<br>\nI just said on a general note that it might hurt people , there is a reason why all the gms and top teams don't share their secret sauce right untill the end of competition no matter how easy it is 😉</p>\n<p>Just to put things into perpective , the first competition I participated in people found a very easy post-processing like this that could be utilized but no one shared untill the end of the competition</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1333481,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2021-06-02T18:54:46.103000",
          "content": "<p>You are right <a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a>, the teams that are at the top which has already done it and found it useful to improve their score might think as a spoiler alert and downvotes it. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1333499,
          "author_name": "Yibing Wu",
          "author_url": "",
          "post_date": "2021-06-02T19:02:23.203000",
          "content": "<p>Don't worry. It is just a sensitive moment.  Eventually, people only remember your method and generosity.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1333501,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-06-02T19:12:49.517000",
          "content": "<p><a href=\"https://www.kaggle.com/ybwu01\" target=\"_blank\">@ybwu01</a> wow, thanks for your kind words! :)<br>\nIt's a little funny how by downvoting a thing you know that works you actually confirm that it does work and encourage others to go onto that path.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1333733,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-06-03T02:44:28.490000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1334020,
          "author_name": "Chun Ming Lee",
          "author_url": "",
          "post_date": "2021-06-03T08:10:59.360000",
          "content": "<p>This happens all the time and I believe someone summed it up best in another competition as \"people already using the trick will be unhappy while people new to it will be happy\".</p>\n<p>I personally think it's Kaggle etiquette not to share anything that might have a material impact on the LB late in the competition - the caution on high-scoring notebooks is another facet of it.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1334141,
          "author_name": "Gabriel Lindenmaier",
          "author_url": "",
          "post_date": "2021-06-03T09:41:52.520000",
          "content": "<p>With the recent changes on the leaderboard we at least know who already used the trick and who didn't.<br>\nQuite surprising for me to see so interrupting changes. But we will see if carries to the private board.</p>\n<p>I myself have decided to not upload a submission at all. I have done a lot of experiments to reproduce some papers and trying out stuff on myself. It was fun and I learned a lot. Meanwhile I never went for the highest-most score, just an efficient model (more on the encoder side). In the end I run out of time to train for long enough to be competitive (enough for my taste).</p>\n<p>I think I will make a post after the submission end, detailing my most interesting findings.</p>\n<p>Edit: Cheers to <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a> for keeping up with the teams in the gold-region so far.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1334958,
          "author_name": "Chun Ming Lee",
          "author_url": "",
          "post_date": "2021-06-03T23:37:02.507000",
          "content": "<p>I have to wonder how much of the shake-up in the gold zone was driven by this reveal (or maybe not at all).</p>\n<p>I would not be happy if I was in one of the teams being punished… </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1331111,
      "author_name": "shinewine",
      "author_url": "",
      "post_date": "2021-06-01T09:26:30.843000",
      "content": "<p>Wow! This idea is very very important! Thanks a lot for sharing! </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1331120,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-06-01T09:37:33.890000",
          "content": "<p>It was worth sharing, then! :) You're welcome!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1330928,
      "author_name": "Gabriel Lindenmaier",
      "author_url": "",
      "post_date": "2021-06-01T07:09:02.070000",
      "content": "<p>Well, then I hope that is enough for you to achieve the solo gold you wanted :)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1330936,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-06-01T07:14:44.117000",
          "content": "<p>Thank you! :)<br>\nI still have doubts, but now I have a little hope, at least.</p>\n<p>But I guess that a lot of teams left their submissions to the end, so… :/ :D</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1331031,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-06-01T08:27:02.570000",
          "content": "<p>i estimate gold level should be better 0.65~0.62</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1331136,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2021-06-01T09:47:54.717000",
          "content": "<p><a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> i think you were hiding your score for a month. I think you have a better model and will win gold I think.</p>\n<p>I estimate gold will be &lt;0.62</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1331142,
          "author_name": "nofreewill42",
          "author_url": "",
          "post_date": "2021-06-01T09:53:10.190000",
          "content": "<p>I also think gold will be less than 0.62<br>\nBut sadly, no, my best score I can achieve is 0.70 right now..<br>\nI don't think I'll have gold, but I'm trying my best :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1331244,
          "author_name": "Gabriel Lindenmaier",
          "author_url": "",
          "post_date": "2021-06-01T11:10:23.743000",
          "content": "<p>OT: Has here anybody enough experience to confidently say, if the public leader board compared to the private + public (complete) leaderboard may either be:</p>\n<ol>\n<li>A random sub-sample</li>\n<li>A stratified sub-sample</li>\n<li>A purposefully easier sub-sample.</li>\n</ol>\n<p>In other words: Is the public test-set usually representative of the private one or not? Or does this strictly vary with the competition?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1331607,
          "author_name": "Vikrant",
          "author_url": "",
          "post_date": "2021-06-01T15:17:26.760000",
          "content": "<p>Good Question.. Ideally, it should be a stratified sub sample but based on my bast experience most of the times it's random. That's why shakeup happens</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1331677,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2021-06-01T16:15:17.263000",
          "content": "<p><a href=\"https://www.kaggle.com/cepheidq\" target=\"_blank\">@cepheidq</a> sampling of the private set from whole test set can any of your mention and it varies with each competition.<br>\nIn this competition it might be random sampled or length wise sampling as <a href=\"https://www.kaggle.com/nofreewill\" target=\"_blank\">@nofreewill</a> did for his CV.<br>\npurposeful maybe but lower probability.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1333714,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-06-03T02:12:39.673000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1330914": "If you have multiple submissions, just modify the [normalization script](https://www.kaggle.com/nofreewill/normalize-your-predictions) so that you note what happened with each InChI (valid, modified, mol_is_none, exception, segfault) and choose valid/modified ones based on some strategy.\n\nYou can also mine the molecules for which you have no valid and/or modified predictions and predict only those with more resources, i.e. with more beams.\n\nI don't know how I missed this so far... it's all there throughout the discussions\n@hengck23 thank you for opening my eyes!",
    "1331028": "to be exact \"validation of inchi ~~rdkit~~ is more crucial\"\n\nrdkit validation doesn't means correct (although it is mostly correct). If you assume valid = correct, then :\n\nyour score is upper bounded by error of prediction that is valid. \n\n---\n\nyou may need to train a score predictor at the end of your decoder, or sum up the prob of each element of the seq.\n\ngiven a set of trained models, you are using heuristics to determine if the prediction is correct., e.g. valid inchi of rdkit. train a model to predict the LD score (given multiple predictions from different models and rdkit) may be better",
    "1332519": "Thanks. It did help me to decrease LB from 1.08 to 1.03. ",
    "1332477": "Thanks for sharing @nofreewill",
    "1331111": "Wow! This idea is very very important! Thanks a lot for sharing! ",
    "1330928": "Well, then I hope that is enough for you to achieve the solo gold you wanted :)",
    "1333714": ""
  }
}