{
  "id": 238475,
  "title": "Why is LB consistently higher than CV?",
  "url": "/competitions/bms-molecular-translation/discussion/238475",
  "author_name": "",
  "post_date": "2021-05-12T09:17:51.494396200Z",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>At first, I took this for granted. I mean, isn't that always the case? We choose the best performing CV model to evaluate then it's likely that this selection bias means the LB result regresses to the norm.</p>\n<p>BUT, for this competition, many of us are only submitting once in a while, and it takes 48 hrs to train a model… So it's not like there's much room for selection bias. So is there something fundamentally different about the test data distribution?</p>",
  "messages": [
    {
      "id": "1303836",
      "postDate": "05/12/2021 09:17:51",
      "content": "<p>At first, I took this for granted. I mean, isn't that always the case? We choose the best performing CV model to evaluate then it's likely that this selection bias means the LB result regresses to the norm.</p>\n<p>BUT, for this competition, many of us are only submitting once in a while, and it takes 48 hrs to train a model… So it's not like there's much room for selection bias. So is there something fundamentally different about the test data distribution?</p>",
      "rawMarkdown": "At first, I took this for granted. I mean, isn't that always the case? We choose the best performing CV model to evaluate then it's likely that this selection bias means the LB result regresses to the norm.\n\nBUT, for this competition, many of us are only submitting once in a while, and it takes 48 hrs to train a model... So it's not like there's much room for selection bias. So is there something fundamentally different about the test data distribution?",
      "votes": null
    },
    {
      "id": "1303881",
      "postDate": "05/12/2021 10:03:08",
      "content": "<p>You mean <strong>higher</strong> LB than CV, right?</p>",
      "rawMarkdown": "You mean **higher** LB than CV, right?",
      "votes": null
    },
    {
      "id": "1303930",
      "postDate": "05/12/2021 10:45:17",
      "content": "<p>Yes, thanks</p>",
      "rawMarkdown": "Yes, thanks",
      "votes": null
    },
    {
      "id": "1304949",
      "postDate": "05/13/2021 02:52:35",
      "content": "<p>I think you answered your question at the end there: the test set must be drawn from a slightly different distribution than the train set. </p>\n<p>It could be related to the features of the molecules, or it could be from the algorithm used to add noise. </p>\n<p>The difference between CV and LB scores is not large, but it is there, and it is consistent. Any team which identifies how to address this will have an advantage (some may already have done so). </p>",
      "rawMarkdown": "I think you answered your question at the end there: the test set must be drawn from a slightly different distribution than the train set. \n\nIt could be related to the features of the molecules, or it could be from the algorithm used to add noise. \n\nThe difference between CV and LB scores is not large, but it is there, and it is consistent. Any team which identifies how to address this will have an advantage (some may already have done so).",
      "votes": null
    },
    {
      "id": "1305088",
      "postDate": "05/13/2021 05:06:15",
      "content": "<p>I just remembered another comment by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, regarding a tokeniser missing certain keys. E.g. if train data does not have a molecule numbered 164, then it will never be predicted in the test set. I'm not sure this would explain a difference of this magnitude though. </p>",
      "rawMarkdown": "I just remembered another comment by @hengck23, regarding a tokeniser missing certain keys. E.g. if train data does not have a molecule numbered 164, then it will never be predicted in the test set. I'm not sure this would explain a difference of this magnitude though.",
      "votes": null
    },
    {
      "id": "1305427",
      "postDate": "05/13/2021 09:09:31",
      "content": "<p>Good point, but I agree with your last bit. The InChI lengths form a long tail</p>",
      "rawMarkdown": "Good point, but I agree with your last bit. The InChI lengths form a long tail",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1303881,
      "author_name": "nofreewill",
      "author_url": "",
      "post_date": "05/12/2021 10:03:08",
      "content": "<p>You mean <strong>higher</strong> LB than CV, right?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1303930,
          "author_name": "alexandersoare",
          "author_url": "",
          "post_date": "05/12/2021 10:45:17",
          "content": "<p>Yes, thanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1304949,
      "author_name": "talktocharles",
      "author_url": "",
      "post_date": "05/13/2021 02:52:35",
      "content": "<p>I think you answered your question at the end there: the test set must be drawn from a slightly different distribution than the train set. </p>\n<p>It could be related to the features of the molecules, or it could be from the algorithm used to add noise. </p>\n<p>The difference between CV and LB scores is not large, but it is there, and it is consistent. Any team which identifies how to address this will have an advantage (some may already have done so). </p>",
      "votes": null,
      "replies": [
        {
          "id": 1305088,
          "author_name": "talktocharles",
          "author_url": "",
          "post_date": "05/13/2021 05:06:15",
          "content": "<p>I just remembered another comment by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, regarding a tokeniser missing certain keys. E.g. if train data does not have a molecule numbered 164, then it will never be predicted in the test set. I'm not sure this would explain a difference of this magnitude though. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1305427,
          "author_name": "alexandersoare",
          "author_url": "",
          "post_date": "05/13/2021 09:09:31",
          "content": "<p>Good point, but I agree with your last bit. The InChI lengths form a long tail</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1303836": "At first, I took this for granted. I mean, isn't that always the case? We choose the best performing CV model to evaluate then it's likely that this selection bias means the LB result regresses to the norm.\n\nBUT, for this competition, many of us are only submitting once in a while, and it takes 48 hrs to train a model... So it's not like there's much room for selection bias. So is there something fundamentally different about the test data distribution?",
    "1303881": "You mean **higher** LB than CV, right?",
    "1303930": "Yes, thanks",
    "1304949": "I think you answered your question at the end there: the test set must be drawn from a slightly different distribution than the train set. \n\nIt could be related to the features of the molecules, or it could be from the algorithm used to add noise. \n\nThe difference between CV and LB scores is not large, but it is there, and it is consistent. Any team which identifies how to address this will have an advantage (some may already have done so).",
    "1305088": "I just remembered another comment by @hengck23, regarding a tokeniser missing certain keys. E.g. if train data does not have a molecule numbered 164, then it will never be predicted in the test set. I'm not sure this would explain a difference of this magnitude though.",
    "1305427": "Good point, but I agree with your last bit. The InChI lengths form a long tail"
  },
  "source": "meta"
}