{
  "id": 239558,
  "title": "CV LB Gap?",
  "url": "/competitions/bms-molecular-translation/discussion/239558",
  "author_name": "",
  "post_date": "2021-05-16T17:40:46.117211800Z",
  "votes": 6,
  "comment_count": 16,
  "views": 0,
  "content": "<p>I recently made a submission for 4.64 LB but had a CV of 2.xx. Has anyone else experience these kinds of gaps and if so, how did you mitigate it(ex. splitting strategy). Thanks!</p>",
  "messages": [
    {
      "id": "1310497",
      "postDate": "05/16/2021 17:40:46",
      "content": "<p>I recently made a submission for 4.64 LB but had a CV of 2.xx. Has anyone else experience these kinds of gaps and if so, how did you mitigate it(ex. splitting strategy). Thanks!</p>",
      "rawMarkdown": "I recently made a submission for 4.64 LB but had a CV of 2.xx. Has anyone else experience these kinds of gaps and if so, how did you mitigate it(ex. splitting strategy). Thanks!",
      "votes": null
    },
    {
      "id": "1310612",
      "postDate": "05/16/2021 18:57:33",
      "content": "<p>How many molecules do you use, and how did you choose them?<br>\nI sorted training data by length of inchi, and grabbed every 50th -&gt; ~50k images<br>\nI have had a pretty small gap so far.</p>",
      "rawMarkdown": "How many molecules do you use, and how did you choose them?\nI sorted training data by length of inchi, and grabbed every 50th -> ~50k images\nI have had a pretty small gap so far.",
      "votes": null
    },
    {
      "id": "1310613",
      "postDate": "05/16/2021 18:59:15",
      "content": "<p>If the gap doesn't go away, then that most probably is caused by differences in the nature of training and test data.</p>",
      "rawMarkdown": "If the gap doesn't go away, then that most probably is caused by differences in the nature of training and test data.",
      "votes": null
    },
    {
      "id": "1310622",
      "postDate": "05/16/2021 19:10:05",
      "content": "<p>I'm not sure what you mean by how many molecules. If it means how much data, I'm using all of it(2.3 mil)</p>",
      "rawMarkdown": "I'm not sure what you mean by how many molecules. If it means how much data, I'm using all of it(2.3 mil)",
      "votes": null
    },
    {
      "id": "1310624",
      "postDate": "05/16/2021 19:10:24",
      "content": "<p>I'll try doing this splitting strategy. Thanks for the help!</p>",
      "rawMarkdown": "I'll try doing this splitting strategy. Thanks for the help!",
      "votes": null
    },
    {
      "id": "1310642",
      "postDate": "05/16/2021 19:27:50",
      "content": "<p>Oh, I just meant, how many molecules do you have in your validation after you split the data :)</p>",
      "rawMarkdown": "Oh, I just meant, how many molecules do you have in your validation after you split the data :)",
      "votes": null
    },
    {
      "id": "1310647",
      "postDate": "05/16/2021 19:33:27",
      "content": "<p>Oh I see :). I did 80000.</p>",
      "rawMarkdown": "Oh I see :). I did 80000.",
      "votes": null
    },
    {
      "id": "1310700",
      "postDate": "05/16/2021 21:16:33",
      "content": "<p>Assuming you've created your validation sample in a relatively smart way (sampling by length/number-of-tokens), I think the difference is due to the more frequent rotation of images found in the test dataset. If you can identify and fix this prior to testing, I assume the gap will decrease to a negligible size.</p>",
      "rawMarkdown": "Assuming you've created your validation sample in a relatively smart way (sampling by length/number-of-tokens), I think the difference is due to the more frequent rotation of images found in the test dataset. If you can identify and fix this prior to testing, I assume the gap will decrease to a negligible size.",
      "votes": null
    },
    {
      "id": "1310714",
      "postDate": "05/16/2021 21:38:08",
      "content": "<p>Does this mean like train a classifier or do the w &gt; h trick?</p>",
      "rawMarkdown": "Does this mean like train a classifier or do the w > h trick?",
      "votes": null
    },
    {
      "id": "1310733",
      "postDate": "05/16/2021 22:28:28",
      "content": "<p>That's the direction I will be moving in probably. if you take all the images with an AR of 1 in the test dataset and plot some… you'll notice that there are some that are rotated and some that aren't… as such… we can't just use a heuristic based on width and height.</p>\n<p>This is a pretty easy way to improve if you can solve it. </p>\n<p>A simple model to detect orientation makes sense to me.</p>",
      "rawMarkdown": "That's the direction I will be moving in probably. if you take all the images with an AR of 1 in the test dataset and plot some... you'll notice that there are some that are rotated and some that aren't... as such... we can't just use a heuristic based on width and height.\n\nThis is a pretty easy way to improve if you can solve it. \n\nA simple model to detect orientation makes sense to me.",
      "votes": null
    },
    {
      "id": "1319442",
      "postDate": "05/23/2021 07:42:25",
      "content": "<p>I've said that I sort molecules by inchi length before I split. I did so previously, but I see now that I've removed that part. Sorry for giving you a misinformation!</p>",
      "rawMarkdown": "I've said that I sort molecules by inchi length before I split. I did so previously, but I see now that I've removed that part. Sorry for giving you a misinformation!",
      "votes": null
    },
    {
      "id": "1319561",
      "postDate": "05/23/2021 10:17:45",
      "content": "<p>I think you mentioned in another post that you sample by \"InChI\" rarity. Could I ask how you numerically determine how rare a molecule is? Thanks!</p>",
      "rawMarkdown": "I think you mentioned in another post that you sample by \"InChI\" rarity. Could I ask how you numerically determine how rare a molecule is? Thanks!",
      "votes": null
    },
    {
      "id": "1319602",
      "postDate": "05/23/2021 11:13:07",
      "content": "<p>I'll share the details after the competition :)</p>",
      "rawMarkdown": "I'll share the details after the competition :)",
      "votes": null
    },
    {
      "id": "1319609",
      "postDate": "05/23/2021 11:18:17",
      "content": "<p>I understand why you wouldn't want to share all details about your sampling strategy. But could I ask whether or not you sampled in this way(by rarity). You don't have to go any further than a yes or no. Regards, Andrew.</p>",
      "rawMarkdown": "I understand why you wouldn't want to share all details about your sampling strategy. But could I ask whether or not you sampled in this way(by rarity). You don't have to go any further than a yes or no. Regards, Andrew.",
      "votes": null
    },
    {
      "id": "1319618",
      "postDate": "05/23/2021 11:28:24",
      "content": "<p>By rarity, and also by (kind of) complexity.</p>",
      "rawMarkdown": "By rarity, and also by (kind of) complexity.",
      "votes": null
    },
    {
      "id": "1319623",
      "postDate": "05/23/2021 11:35:08",
      "content": "<p>Ok Thanks so much!</p>",
      "rawMarkdown": "Ok Thanks so much!",
      "votes": null
    },
    {
      "id": "1319625",
      "postDate": "05/23/2021 11:36:51",
      "content": "<p>You're welcome! :)</p>",
      "rawMarkdown": "You're welcome! :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1310612,
      "author_name": "nofreewill",
      "author_url": "",
      "post_date": "05/16/2021 18:57:33",
      "content": "<p>How many molecules do you use, and how did you choose them?<br>\nI sorted training data by length of inchi, and grabbed every 50th -&gt; ~50k images<br>\nI have had a pretty small gap so far.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1310613,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "05/16/2021 18:59:15",
          "content": "<p>If the gap doesn't go away, then that most probably is caused by differences in the nature of training and test data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1310622,
          "author_name": "andrewshao05",
          "author_url": "",
          "post_date": "05/16/2021 19:10:05",
          "content": "<p>I'm not sure what you mean by how many molecules. If it means how much data, I'm using all of it(2.3 mil)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1310624,
          "author_name": "andrewshao05",
          "author_url": "",
          "post_date": "05/16/2021 19:10:24",
          "content": "<p>I'll try doing this splitting strategy. Thanks for the help!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1310642,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "05/16/2021 19:27:50",
          "content": "<p>Oh, I just meant, how many molecules do you have in your validation after you split the data :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1310647,
          "author_name": "andrewshao05",
          "author_url": "",
          "post_date": "05/16/2021 19:33:27",
          "content": "<p>Oh I see :). I did 80000.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1319442,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "05/23/2021 07:42:25",
          "content": "<p>I've said that I sort molecules by inchi length before I split. I did so previously, but I see now that I've removed that part. Sorry for giving you a misinformation!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1319561,
          "author_name": "andrewshao05",
          "author_url": "",
          "post_date": "05/23/2021 10:17:45",
          "content": "<p>I think you mentioned in another post that you sample by \"InChI\" rarity. Could I ask how you numerically determine how rare a molecule is? Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1319602,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "05/23/2021 11:13:07",
          "content": "<p>I'll share the details after the competition :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1319609,
          "author_name": "andrewshao05",
          "author_url": "",
          "post_date": "05/23/2021 11:18:17",
          "content": "<p>I understand why you wouldn't want to share all details about your sampling strategy. But could I ask whether or not you sampled in this way(by rarity). You don't have to go any further than a yes or no. Regards, Andrew.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1319618,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "05/23/2021 11:28:24",
          "content": "<p>By rarity, and also by (kind of) complexity.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1319623,
          "author_name": "andrewshao05",
          "author_url": "",
          "post_date": "05/23/2021 11:35:08",
          "content": "<p>Ok Thanks so much!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1319625,
          "author_name": "nofreewill",
          "author_url": "",
          "post_date": "05/23/2021 11:36:51",
          "content": "<p>You're welcome! :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1310700,
      "author_name": "dschettler8845",
      "author_url": "",
      "post_date": "05/16/2021 21:16:33",
      "content": "<p>Assuming you've created your validation sample in a relatively smart way (sampling by length/number-of-tokens), I think the difference is due to the more frequent rotation of images found in the test dataset. If you can identify and fix this prior to testing, I assume the gap will decrease to a negligible size.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1310714,
          "author_name": "andrewshao05",
          "author_url": "",
          "post_date": "05/16/2021 21:38:08",
          "content": "<p>Does this mean like train a classifier or do the w &gt; h trick?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1310733,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "05/16/2021 22:28:28",
          "content": "<p>That's the direction I will be moving in probably. if you take all the images with an AR of 1 in the test dataset and plot some… you'll notice that there are some that are rotated and some that aren't… as such… we can't just use a heuristic based on width and height.</p>\n<p>This is a pretty easy way to improve if you can solve it. </p>\n<p>A simple model to detect orientation makes sense to me.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1310497": "I recently made a submission for 4.64 LB but had a CV of 2.xx. Has anyone else experience these kinds of gaps and if so, how did you mitigate it(ex. splitting strategy). Thanks!",
    "1310612": "How many molecules do you use, and how did you choose them?\nI sorted training data by length of inchi, and grabbed every 50th -> ~50k images\nI have had a pretty small gap so far.",
    "1310613": "If the gap doesn't go away, then that most probably is caused by differences in the nature of training and test data.",
    "1310622": "I'm not sure what you mean by how many molecules. If it means how much data, I'm using all of it(2.3 mil)",
    "1310624": "I'll try doing this splitting strategy. Thanks for the help!",
    "1310642": "Oh, I just meant, how many molecules do you have in your validation after you split the data :)",
    "1310647": "Oh I see :). I did 80000.",
    "1310700": "Assuming you've created your validation sample in a relatively smart way (sampling by length/number-of-tokens), I think the difference is due to the more frequent rotation of images found in the test dataset. If you can identify and fix this prior to testing, I assume the gap will decrease to a negligible size.",
    "1310714": "Does this mean like train a classifier or do the w > h trick?",
    "1310733": "That's the direction I will be moving in probably. if you take all the images with an AR of 1 in the test dataset and plot some... you'll notice that there are some that are rotated and some that aren't... as such... we can't just use a heuristic based on width and height.\n\nThis is a pretty easy way to improve if you can solve it. \n\nA simple model to detect orientation makes sense to me.",
    "1319442": "I've said that I sort molecules by inchi length before I split. I did so previously, but I see now that I've removed that part. Sorry for giving you a misinformation!",
    "1319561": "I think you mentioned in another post that you sample by \"InChI\" rarity. Could I ask how you numerically determine how rare a molecule is? Thanks!",
    "1319602": "I'll share the details after the competition :)",
    "1319609": "I understand why you wouldn't want to share all details about your sampling strategy. But could I ask whether or not you sampled in this way(by rarity). You don't have to go any further than a yes or no. Regards, Andrew.",
    "1319618": "By rarity, and also by (kind of) complexity.",
    "1319623": "Ok Thanks so much!",
    "1319625": "You're welcome! :)"
  },
  "source": "meta"
}