{
  "id": 131694,
  "title": "Large gap between LB/CV",
  "url": "/competitions/bengaliai-cv19/discussion/131694",
  "author_name": "",
  "post_date": "2020-02-21T04:48:57.791108500Z",
  "votes": 2,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Hi, all, I tested two models which validation is around 0.98, and I expect the LB should be around 0.97 from other people's experience. However, I only got 0.96. I am wondering if someone has met this before, do you have any suggestion for the possible problem?</p>",
  "messages": [
    {
      "id": "752502",
      "postDate": "02/21/2020 04:48:57",
      "content": "<p>Hi, all, I tested two models which validation is around 0.98, and I expect the LB should be around 0.97 from other people's experience. However, I only got 0.96. I am wondering if someone has met this before, do you have any suggestion for the possible problem?</p>",
      "rawMarkdown": "Hi, all, I tested two models which validation is around 0.98, and I expect the LB should be around 0.97 from other people's experience. However, I only got 0.96. I am wondering if someone has met this before, do you have any suggestion for the possible problem?",
      "votes": null
    },
    {
      "id": "752551",
      "postDate": "02/21/2020 06:33:05",
      "content": "<p>The distribution of words (graphemes), grapheme roots, vowel diacritics, and consonant diacritics in the test data is unknown, so it isn't possible to make a validation set just like the test dataset. </p>\n\n<p>The training data has 1292 words, 168 grapheme roots, 11 vowel diacritics, and 7 consonant diacritics. When you create your validation set, you can try to stratify with relation to all of these.</p>",
      "rawMarkdown": "The distribution of words (graphemes), grapheme roots, vowel diacritics, and consonant diacritics in the test data is unknown, so it isn't possible to make a validation set just like the test dataset. \n\nThe training data has 1292 words, 168 grapheme roots, 11 vowel diacritics, and 7 consonant diacritics. When you create your validation set, you can try to stratify with relation to all of these.",
      "votes": null
    },
    {
      "id": "752558",
      "postDate": "02/21/2020 06:44:42",
      "content": "<p>I do use the stratified method that Venn talked. So I suspect there is some other thing I am missing</p>",
      "rawMarkdown": "I do use the stratified method that Venn talked. So I suspect there is some other thing I am missing",
      "votes": null
    },
    {
      "id": "752561",
      "postDate": "02/21/2020 06:45:59",
      "content": "<p>check the img_size during train and test. It was the case for me once in the past when CV and LB didn't correlate</p>",
      "rawMarkdown": "check the img_size during train and test. It was the case for me once in the past when CV and LB didn't correlate",
      "votes": null
    },
    {
      "id": "752569",
      "postDate": "02/21/2020 07:04:20",
      "content": "<p>Are you weighting the macro recall of root as 2 whereas vowel and consonant macro recall has weight 1? That could inflate your CV.</p>",
      "rawMarkdown": "Are you weighting the macro recall of root as 2 whereas vowel and consonant macro recall has weight 1? That could inflate your CV.",
      "votes": null
    },
    {
      "id": "753865",
      "postDate": "02/22/2020 18:44:47",
      "content": "<p>I checked the img size and they are consistent</p>",
      "rawMarkdown": "I checked the img size and they are consistent",
      "votes": null
    },
    {
      "id": "754017",
      "postDate": "02/22/2020 23:35:54",
      "content": "<p>Hey, me too im having this problem.. I'm also using de 80/20 stratified method and im getting 98.1% val and only 96.56% lb (on that particular test). The only way I could reduce that gap was by adding weight tensor in loss function but then my cv is much lower and i can only get about 97% lb! Please someone help haha</p>",
      "rawMarkdown": "Hey, me too im having this problem.. I'm also using de 80/20 stratified method and im getting 98.1% val and only 96.56% lb (on that particular test). The only way I could reduce that gap was by adding weight tensor in loss function but then my cv is much lower and i can only get about 97% lb! Please someone help haha",
      "votes": null
    },
    {
      "id": "754028",
      "postDate": "02/23/2020 00:20:52",
      "content": "<p>how does that inflate cv? and isn't that the competition evaluation metric we are supposed to use i.e. 2, 1, 1 weighting on the macro recall?</p>",
      "rawMarkdown": "how does that inflate cv? and isn't that the competition evaluation metric we are supposed to use i.e. 2, 1, 1 weighting on the macro recall?",
      "votes": null
    },
    {
      "id": "754495",
      "postDate": "02/23/2020 16:24:49",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> what do you mean by stratify? If I'm not wrong, is it something like this:</p>\n\n<p>```</p>\n\n<h1>for way one - data generator</h1>\n\n<p>train_labels, val_labels = train_test_split(train, test_size = 0.20, \nrandom_state = SEED,\nstratify = train[['grapheme_root', 'vowel_diacritic', 'consonant_diacritic']])\n```</p>\n\n<p>My validation score on grapheme_root is very low. Is there anything I need to consider? Thanks.</p>",
      "rawMarkdown": "cdeotte what do you mean by stratify? If I'm not wrong, is it something like this:\n\n```\n# for way one - data generator\ntrain_labels, val_labels = train_test_split(train, test_size = 0.20, \nrandom_state = SEED,\nstratify = train[['grapheme_root', 'vowel_diacritic', 'consonant_diacritic']])\n```\n\nMy validation score on grapheme_root is very low. Is there anything I need to consider? Thanks.",
      "votes": null
    },
    {
      "id": "754523",
      "postDate": "02/23/2020 17:25:22",
      "content": "<p>the kaggle metric is average class recall. you may want to refer to the official kernel code and make sure that you have not mistaken as \"recall averaged all samples without regard to class\".</p>",
      "rawMarkdown": "the kaggle metric is average class recall. you may want to refer to the official kernel code and make sure that you have not mistaken as \"recall averaged all samples without regard to class\".",
      "votes": null
    },
    {
      "id": "755714",
      "postDate": "02/25/2020 03:51:05",
      "content": "<p>I basically follow the code in one of most vote kernel, and the recall is calculated by class then weighted average to 2 1 1.</p>",
      "rawMarkdown": "I basically follow the code in one of most vote kernel, and the recall is calculated by class then weighted average to 2 1 1.",
      "votes": null
    },
    {
      "id": "755716",
      "postDate": "02/25/2020 03:52:21",
      "content": "<p><a href=\"/yannmajewski\">@yannmajewski</a> I think 0.97 is a decent score with validation of 0.98, can you explain what you did to mitigate the issue?</p>",
      "rawMarkdown": "yannmajewski I think 0.97 is a decent score with validation of 0.98, can you explain what you did to mitigate the issue?",
      "votes": null
    },
    {
      "id": "757363",
      "postDate": "02/26/2020 17:19:02",
      "content": "<p>I adjusted the loss weight, but not helps</p>",
      "rawMarkdown": "I adjusted the loss weight, but not helps",
      "votes": null
    },
    {
      "id": "760801",
      "postDate": "03/01/2020 18:28:56",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "760810",
      "postDate": "03/01/2020 18:38:26",
      "content": "<p>Such a drop cannot be caused by using apex imho, sounds rather like randomness.</p>",
      "rawMarkdown": "Such a drop cannot be caused by using apex imho, sounds rather like randomness.",
      "votes": null
    },
    {
      "id": "760856",
      "postDate": "03/01/2020 20:00:37",
      "content": "<p>But after I removed apex code, the drop is gone. There is possible compatibility issue and bug for Apex with Pytorch 1.4. I recently upgraded to 1.4 and before this competition there is no such drop from apex.</p>",
      "rawMarkdown": "But after I removed apex code, the drop is gone. There is possible compatibility issue and bug for Apex with Pytorch 1.4. I recently upgraded to 1.4 and before this competition there is no such drop from apex.",
      "votes": null
    },
    {
      "id": "760868",
      "postDate": "03/01/2020 20:13:35",
      "content": "<p>But you retrained after removing apex? So it is a different fit.</p>",
      "rawMarkdown": "But you retrained after removing apex? So it is a different fit.",
      "votes": null
    },
    {
      "id": "761093",
      "postDate": "03/02/2020 06:04:57",
      "content": "<p>Let me explain what I found. Model A was trained by Apex, when directly load the model for evaluation on the validation dataset, there is 0.01 drop compared to the validation score in the training process. If I initialize a model B without Apex, when I load the model and reevaluate again, the validation score is consistent with the validation score in the training process</p>",
      "rawMarkdown": "Let me explain what I found. Model A was trained by Apex, when directly load the model for evaluation on the validation dataset, there is 0.01 drop compared to the validation score in the training process. If I initialize a model B without Apex, when I load the model and reevaluate again, the validation score is consistent with the validation score in the training process",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 752551,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/21/2020 06:33:05",
      "content": "<p>The distribution of words (graphemes), grapheme roots, vowel diacritics, and consonant diacritics in the test data is unknown, so it isn't possible to make a validation set just like the test dataset. </p>\n\n<p>The training data has 1292 words, 168 grapheme roots, 11 vowel diacritics, and 7 consonant diacritics. When you create your validation set, you can try to stratify with relation to all of these.</p>",
      "votes": null,
      "replies": [
        {
          "id": 752558,
          "author_name": "strideradu",
          "author_url": "",
          "post_date": "02/21/2020 06:44:42",
          "content": "<p>I do use the stratified method that Venn talked. So I suspect there is some other thing I am missing</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 752561,
          "author_name": "bibek777",
          "author_url": "",
          "post_date": "02/21/2020 06:45:59",
          "content": "<p>check the img_size during train and test. It was the case for me once in the past when CV and LB didn't correlate</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 752569,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/21/2020 07:04:20",
          "content": "<p>Are you weighting the macro recall of root as 2 whereas vowel and consonant macro recall has weight 1? That could inflate your CV.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 753865,
          "author_name": "strideradu",
          "author_url": "",
          "post_date": "02/22/2020 18:44:47",
          "content": "<p>I checked the img size and they are consistent</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 754017,
          "author_name": "yannmajewski",
          "author_url": "",
          "post_date": "02/22/2020 23:35:54",
          "content": "<p>Hey, me too im having this problem.. I'm also using de 80/20 stratified method and im getting 98.1% val and only 96.56% lb (on that particular test). The only way I could reduce that gap was by adding weight tensor in loss function but then my cv is much lower and i can only get about 97% lb! Please someone help haha</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 754028,
          "author_name": "samshipengs",
          "author_url": "",
          "post_date": "02/23/2020 00:20:52",
          "content": "<p>how does that inflate cv? and isn't that the competition evaluation metric we are supposed to use i.e. 2, 1, 1 weighting on the macro recall?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 754495,
          "author_name": "akashshingha850",
          "author_url": "",
          "post_date": "02/23/2020 16:24:49",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> what do you mean by stratify? If I'm not wrong, is it something like this:</p>\n\n<p>```</p>\n\n<h1>for way one - data generator</h1>\n\n<p>train_labels, val_labels = train_test_split(train, test_size = 0.20, \nrandom_state = SEED,\nstratify = train[['grapheme_root', 'vowel_diacritic', 'consonant_diacritic']])\n```</p>\n\n<p>My validation score on grapheme_root is very low. Is there anything I need to consider? Thanks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 755716,
          "author_name": "strideradu",
          "author_url": "",
          "post_date": "02/25/2020 03:52:21",
          "content": "<p><a href=\"/yannmajewski\">@yannmajewski</a> I think 0.97 is a decent score with validation of 0.98, can you explain what you did to mitigate the issue?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 754523,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/23/2020 17:25:22",
      "content": "<p>the kaggle metric is average class recall. you may want to refer to the official kernel code and make sure that you have not mistaken as \"recall averaged all samples without regard to class\".</p>",
      "votes": null,
      "replies": [
        {
          "id": 755714,
          "author_name": "strideradu",
          "author_url": "",
          "post_date": "02/25/2020 03:51:05",
          "content": "<p>I basically follow the code in one of most vote kernel, and the recall is calculated by class then weighted average to 2 1 1.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 757363,
      "author_name": "strideradu",
      "author_url": "",
      "post_date": "02/26/2020 17:19:02",
      "content": "<p>I adjusted the loss weight, but not helps</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 760801,
      "author_name": "strideradu",
      "author_url": "",
      "post_date": "03/01/2020 18:28:56",
      "content": "",
      "votes": null,
      "replies": [
        {
          "id": 760810,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "03/01/2020 18:38:26",
          "content": "<p>Such a drop cannot be caused by using apex imho, sounds rather like randomness.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 760856,
          "author_name": "strideradu",
          "author_url": "",
          "post_date": "03/01/2020 20:00:37",
          "content": "<p>But after I removed apex code, the drop is gone. There is possible compatibility issue and bug for Apex with Pytorch 1.4. I recently upgraded to 1.4 and before this competition there is no such drop from apex.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 760868,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "03/01/2020 20:13:35",
          "content": "<p>But you retrained after removing apex? So it is a different fit.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 761093,
          "author_name": "strideradu",
          "author_url": "",
          "post_date": "03/02/2020 06:04:57",
          "content": "<p>Let me explain what I found. Model A was trained by Apex, when directly load the model for evaluation on the validation dataset, there is 0.01 drop compared to the validation score in the training process. If I initialize a model B without Apex, when I load the model and reevaluate again, the validation score is consistent with the validation score in the training process</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "752502": "Hi, all, I tested two models which validation is around 0.98, and I expect the LB should be around 0.97 from other people's experience. However, I only got 0.96. I am wondering if someone has met this before, do you have any suggestion for the possible problem?",
    "752551": "The distribution of words (graphemes), grapheme roots, vowel diacritics, and consonant diacritics in the test data is unknown, so it isn't possible to make a validation set just like the test dataset. \n\nThe training data has 1292 words, 168 grapheme roots, 11 vowel diacritics, and 7 consonant diacritics. When you create your validation set, you can try to stratify with relation to all of these.",
    "752558": "I do use the stratified method that Venn talked. So I suspect there is some other thing I am missing",
    "752561": "check the img_size during train and test. It was the case for me once in the past when CV and LB didn't correlate",
    "752569": "Are you weighting the macro recall of root as 2 whereas vowel and consonant macro recall has weight 1? That could inflate your CV.",
    "753865": "I checked the img size and they are consistent",
    "754017": "Hey, me too im having this problem.. I'm also using de 80/20 stratified method and im getting 98.1% val and only 96.56% lb (on that particular test). The only way I could reduce that gap was by adding weight tensor in loss function but then my cv is much lower and i can only get about 97% lb! Please someone help haha",
    "754028": "how does that inflate cv? and isn't that the competition evaluation metric we are supposed to use i.e. 2, 1, 1 weighting on the macro recall?",
    "754495": "cdeotte what do you mean by stratify? If I'm not wrong, is it something like this:\n\n```\n# for way one - data generator\ntrain_labels, val_labels = train_test_split(train, test_size = 0.20, \nrandom_state = SEED,\nstratify = train[['grapheme_root', 'vowel_diacritic', 'consonant_diacritic']])\n```\n\nMy validation score on grapheme_root is very low. Is there anything I need to consider? Thanks.",
    "754523": "the kaggle metric is average class recall. you may want to refer to the official kernel code and make sure that you have not mistaken as \"recall averaged all samples without regard to class\".",
    "755714": "I basically follow the code in one of most vote kernel, and the recall is calculated by class then weighted average to 2 1 1.",
    "755716": "yannmajewski I think 0.97 is a decent score with validation of 0.98, can you explain what you did to mitigate the issue?",
    "757363": "I adjusted the loss weight, but not helps",
    "760801": "",
    "760810": "Such a drop cannot be caused by using apex imho, sounds rather like randomness.",
    "760856": "But after I removed apex code, the drop is gone. There is possible compatibility issue and bug for Apex with Pytorch 1.4. I recently upgraded to 1.4 and before this competition there is no such drop from apex.",
    "760868": "But you retrained after removing apex? So it is a different fit.",
    "761093": "Let me explain what I found. Model A was trained by Apex, when directly load the model for evaluation on the validation dataset, there is 0.01 drop compared to the validation score in the training process. If I initialize a model B without Apex, when I load the model and reevaluate again, the validation score is consistent with the validation score in the training process"
  },
  "source": "meta"
}