{
  "id": 93964,
  "title": "why my LB score is much smaller than my CV score?",
  "url": "/competitions/imet-2019-fgvc6/discussion/93964",
  "author_name": "",
  "post_date": "2019-05-31T12:55:10.575753300Z",
  "votes": null,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I use seresnext101 model in pytorch. I trained the model for 7 epochs,with the image size 300x300. Then I changed the image size to 388x388, trained the model for 5 epochs. </p>\n\n<p>The loss function is BCE. \nThe augmentation included randomsizecrop, hflid and color jitter. \nValidation and prediction load the image in the same way. \nTTA is 8. My CV is 0.676 and the public score is 0.45. </p>\n\n<p>What I should do to avoid this problem.</p>",
  "messages": [
    {
      "id": "540393",
      "postDate": "05/31/2019 12:55:10",
      "content": "<p>I use seresnext101 model in pytorch. I trained the model for 7 epochs,with the image size 300x300. Then I changed the image size to 388x388, trained the model for 5 epochs. </p>\n\n<p>The loss function is BCE. \nThe augmentation included randomsizecrop, hflid and color jitter. \nValidation and prediction load the image in the same way. \nTTA is 8. My CV is 0.676 and the public score is 0.45. </p>\n\n<p>What I should do to avoid this problem.</p>",
      "rawMarkdown": "I use seresnext101 model in pytorch. I trained the model for 7 epochs,with the image size 300x300. Then I changed the image size to 388x388, trained the model for 5 epochs. \n\nThe loss function is BCE. \nThe augmentation included randomsizecrop, hflid and color jitter. \nValidation and prediction load the image in the same way. \nTTA is 8. My CV is 0.676 and the public score is 0.45. \n\nWhat I should do to avoid this problem.",
      "votes": null
    },
    {
      "id": "540421",
      "postDate": "05/31/2019 13:35:13",
      "content": "<p>if the public score is 0.45 there is definitely some bug in the inference, so we can't help you. check how you pick thresholds, try no tta, see the images from your test data loader if they are okay. also be sure your resize/crop train image schema is meaningful for test inference</p>",
      "rawMarkdown": "if the public score is 0.45 there is definitely some bug in the inference, so we can't help you. check how you pick thresholds, try no tta, see the images from your test data loader if they are okay. also be sure your resize/crop train image schema is meaningful for test inference",
      "votes": null
    },
    {
      "id": "540424",
      "postDate": "05/31/2019 13:37:19",
      "content": "<p>I think 0.45 is the sign of some serious error in your code, so you should debug and check everything.</p>",
      "rawMarkdown": "I think 0.45 is the sign of some serious error in your code, so you should debug and check everything.",
      "votes": null
    },
    {
      "id": "540433",
      "postDate": "05/31/2019 14:07:22",
      "content": "<p>Thank you</p>",
      "rawMarkdown": "Thank you",
      "votes": null
    },
    {
      "id": "540434",
      "postDate": "05/31/2019 14:07:34",
      "content": "<p>Thank you</p>",
      "rawMarkdown": "Thank you",
      "votes": null
    },
    {
      "id": "540472",
      "postDate": "05/31/2019 15:02:43",
      "content": "<p>How did you make your validation data? I think there is leak.\nI think you 1st-trained the model with your train data(1st train data) and CV with validation data(1st validation data).\nIf you 2nd-trained the model with your train data(2nd train data), and CV with validation data(2nd validation data) which include 1st train data, it occurs leak.\nMake sure 2nd validation data don't include 1st train data.</p>",
      "rawMarkdown": "How did you make your validation data? I think there is leak.\nI think you 1st-trained the model with your train data(1st train data) and CV with validation data(1st validation data).\nIf you 2nd-trained the model with your train data(2nd train data), and CV with validation data(2nd validation data) which include 1st train data, it occurs leak.\nMake sure 2nd validation data don't include 1st train data.",
      "votes": null
    },
    {
      "id": "541161",
      "postDate": "06/01/2019 23:25:50",
      "content": "<p>The 0.676 CV is obviously too high</p>",
      "rawMarkdown": "The 0.676 CV is obviously too high",
      "votes": null
    },
    {
      "id": "541244",
      "postDate": "06/02/2019 04:54:28",
      "content": "<p>Hi, what method do you use to divide the data set?</p>",
      "rawMarkdown": "Hi, what method do you use to divide the data set?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 540421,
      "author_name": "yaroshevskiy",
      "author_url": "",
      "post_date": "05/31/2019 13:35:13",
      "content": "<p>if the public score is 0.45 there is definitely some bug in the inference, so we can't help you. check how you pick thresholds, try no tta, see the images from your test data loader if they are okay. also be sure your resize/crop train image schema is meaningful for test inference</p>",
      "votes": null,
      "replies": [
        {
          "id": 540434,
          "author_name": "saladjay",
          "author_url": "",
          "post_date": "05/31/2019 14:07:34",
          "content": "<p>Thank you</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 540424,
      "author_name": "demonplus",
      "author_url": "",
      "post_date": "05/31/2019 13:37:19",
      "content": "<p>I think 0.45 is the sign of some serious error in your code, so you should debug and check everything.</p>",
      "votes": null,
      "replies": [
        {
          "id": 540433,
          "author_name": "saladjay",
          "author_url": "",
          "post_date": "05/31/2019 14:07:22",
          "content": "<p>Thank you</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 540472,
      "author_name": "yasufuminakama",
      "author_url": "",
      "post_date": "05/31/2019 15:02:43",
      "content": "<p>How did you make your validation data? I think there is leak.\nI think you 1st-trained the model with your train data(1st train data) and CV with validation data(1st validation data).\nIf you 2nd-trained the model with your train data(2nd train data), and CV with validation data(2nd validation data) which include 1st train data, it occurs leak.\nMake sure 2nd validation data don't include 1st train data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 541161,
      "author_name": "strideradu",
      "author_url": "",
      "post_date": "06/01/2019 23:25:50",
      "content": "<p>The 0.676 CV is obviously too high</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 541244,
      "author_name": "guitar123",
      "author_url": "",
      "post_date": "06/02/2019 04:54:28",
      "content": "<p>Hi, what method do you use to divide the data set?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "540393": "I use seresnext101 model in pytorch. I trained the model for 7 epochs,with the image size 300x300. Then I changed the image size to 388x388, trained the model for 5 epochs. \n\nThe loss function is BCE. \nThe augmentation included randomsizecrop, hflid and color jitter. \nValidation and prediction load the image in the same way. \nTTA is 8. My CV is 0.676 and the public score is 0.45. \n\nWhat I should do to avoid this problem.",
    "540421": "if the public score is 0.45 there is definitely some bug in the inference, so we can't help you. check how you pick thresholds, try no tta, see the images from your test data loader if they are okay. also be sure your resize/crop train image schema is meaningful for test inference",
    "540424": "I think 0.45 is the sign of some serious error in your code, so you should debug and check everything.",
    "540433": "Thank you",
    "540434": "Thank you",
    "540472": "How did you make your validation data? I think there is leak.\nI think you 1st-trained the model with your train data(1st train data) and CV with validation data(1st validation data).\nIf you 2nd-trained the model with your train data(2nd train data), and CV with validation data(2nd validation data) which include 1st train data, it occurs leak.\nMake sure 2nd validation data don't include 1st train data.",
    "541161": "The 0.676 CV is obviously too high",
    "541244": "Hi, what method do you use to divide the data set?"
  },
  "source": "meta"
}