{
  "id": 319983,
  "title": "mismatch between test accuracy and validation accuracy",
  "url": "/competitions/sorghum-id-fgvc-9/discussion/319983",
  "author_name": "",
  "post_date": "2022-04-19T16:40:04.088808100Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi, I'm having a large gap between my validation accuracy (around 95%) and my test accuracy( 50%). I know there is a natural split in the data so it is probably not exactly iid, yet I find the gap to be very large.. Has someone else also experienced this or am I doing something wrong here?</p>\n<p>I used a stratified Split (10%) to create my validation set and am using a pretrained efficientnet-b2 model with mild data augmentation. I'm also using a jpeg encoded version of the dataset at 256x256 px.</p>",
  "messages": [
    {
      "id": "1761026",
      "postDate": "04/19/2022 16:40:04",
      "content": "<p>Hi, I'm having a large gap between my validation accuracy (around 95%) and my test accuracy( 50%). I know there is a natural split in the data so it is probably not exactly iid, yet I find the gap to be very large.. Has someone else also experienced this or am I doing something wrong here?</p>\n<p>I used a stratified Split (10%) to create my validation set and am using a pretrained efficientnet-b2 model with mild data augmentation. I'm also using a jpeg encoded version of the dataset at 256x256 px.</p>",
      "rawMarkdown": "Hi, I'm having a large gap between my validation accuracy (around 95%) and my test accuracy( 50%). I know there is a natural split in the data so it is probably not exactly iid, yet I find the gap to be very large.. Has someone else also experienced this or am I doing something wrong here?\n\nI used a stratified Split (10%) to create my validation set and am using a pretrained efficientnet-b2 model with mild data augmentation. I'm also using a jpeg encoded version of the dataset at 256x256 px.",
      "votes": null
    },
    {
      "id": "1762544",
      "postDate": "04/20/2022 19:20:09",
      "content": "<p>No, you are not doing anything wrong! My numbers are in the same ballpark. It is just an effect of the weird data split, as you observed. Makes this challenge more interesting, in my opinion…</p>\n<p>What we need is a way to penalize networks that lock on to the wrong features (e.g. the color of the dirt). If the plants can be segmented out before training the network, that should help too.</p>\n<p>If the training data included field IDs, that would have helped us construct a validation set that mirrors the train-test split. Maybe the timestamp can be used as a loose proxy for field location.</p>",
      "rawMarkdown": "No, you are not doing anything wrong! My numbers are in the same ballpark. It is just an effect of the weird data split, as you observed. Makes this challenge more interesting, in my opinion...\n\nWhat we need is a way to penalize networks that lock on to the wrong features (e.g. the color of the dirt). If the plants can be segmented out before training the network, that should help too.\n\nIf the training data included field IDs, that would have helped us construct a validation set that mirrors the train-test split. Maybe the timestamp can be used as a loose proxy for field location.",
      "votes": null
    },
    {
      "id": "1764185",
      "postDate": "04/22/2022 08:28:18",
      "content": "<p>Thanks for your reply! Glad to know it's due to the split and not a bad validation split or something.</p>\n<p>Yes it is indeed more interesting. I agree that getting rid of the dirt might be interesting to help with generalization to the other parcels.</p>",
      "rawMarkdown": "Thanks for your reply! Glad to know it's due to the split and not a bad validation split or something.\n\nYes it is indeed more interesting. I agree that getting rid of the dirt might be interesting to help with generalization to the other parcels.",
      "votes": null
    },
    {
      "id": "1767751",
      "postDate": "04/25/2022 15:54:45",
      "content": "<p>I guess, if your image is augmented in train set and not in valid set, it can make it hard to predict for train set. So it makes sense.<br>\nbecause train images has randomly transformed, validation loss can be lower than training loss.</p>",
      "rawMarkdown": "I guess, if your image is augmented in train set and not in valid set, it can make it hard to predict for train set. So it makes sense.\nbecause train images has randomly transformed, validation loss can be lower than training loss.",
      "votes": null
    },
    {
      "id": "1790953",
      "postDate": "05/15/2022 13:52:34",
      "content": "<p>HI , I too am not able to achieve val accuracy &gt;0.5 with effnetb2. I'm not sure what is the reason for this high bias. Initially i tried freezing all the layers except the top but was not able to achieve train accuracy&gt;0.45. which means fine tuning too will not have any significant improvements. Then I unfreeze all layers and got the train and val accuracy&gt;0.9 but test acc=0.5.  I'm stuck with this problem. Let me know if you find any solutions and reasons for high bias. In my opinion , it is due to overfitting since I am training all layers on train data which may cause overfitting very quickly, though I'm not sure.If that was the case why would i get val accuracy &gt;0.9.</p>",
      "rawMarkdown": "HI , I too am not able to achieve val accuracy >0.5 with effnetb2. I'm not sure what is the reason for this high bias. Initially i tried freezing all the layers except the top but was not able to achieve train accuracy>0.45. which means fine tuning too will not have any significant improvements. Then I unfreeze all layers and got the train and val accuracy>0.9 but test acc=0.5.  I'm stuck with this problem. Let me know if you find any solutions and reasons for high bias. In my opinion , it is due to overfitting since I am training all layers on train data which may cause overfitting very quickly, though I'm not sure.If that was the case why would i get val accuracy >0.9.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1762544,
      "author_name": "anlthms",
      "author_url": "",
      "post_date": "04/20/2022 19:20:09",
      "content": "<p>No, you are not doing anything wrong! My numbers are in the same ballpark. It is just an effect of the weird data split, as you observed. Makes this challenge more interesting, in my opinion…</p>\n<p>What we need is a way to penalize networks that lock on to the wrong features (e.g. the color of the dirt). If the plants can be segmented out before training the network, that should help too.</p>\n<p>If the training data included field IDs, that would have helped us construct a validation set that mirrors the train-test split. Maybe the timestamp can be used as a loose proxy for field location.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1764185,
      "author_name": "tlipss",
      "author_url": "",
      "post_date": "04/22/2022 08:28:18",
      "content": "<p>Thanks for your reply! Glad to know it's due to the split and not a bad validation split or something.</p>\n<p>Yes it is indeed more interesting. I agree that getting rid of the dirt might be interesting to help with generalization to the other parcels.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1767751,
      "author_name": "joonasyoon",
      "author_url": "",
      "post_date": "04/25/2022 15:54:45",
      "content": "<p>I guess, if your image is augmented in train set and not in valid set, it can make it hard to predict for train set. So it makes sense.<br>\nbecause train images has randomly transformed, validation loss can be lower than training loss.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1790953,
      "author_name": "aayushdeswal",
      "author_url": "",
      "post_date": "05/15/2022 13:52:34",
      "content": "<p>HI , I too am not able to achieve val accuracy &gt;0.5 with effnetb2. I'm not sure what is the reason for this high bias. Initially i tried freezing all the layers except the top but was not able to achieve train accuracy&gt;0.45. which means fine tuning too will not have any significant improvements. Then I unfreeze all layers and got the train and val accuracy&gt;0.9 but test acc=0.5.  I'm stuck with this problem. Let me know if you find any solutions and reasons for high bias. In my opinion , it is due to overfitting since I am training all layers on train data which may cause overfitting very quickly, though I'm not sure.If that was the case why would i get val accuracy &gt;0.9.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1761026": "Hi, I'm having a large gap between my validation accuracy (around 95%) and my test accuracy( 50%). I know there is a natural split in the data so it is probably not exactly iid, yet I find the gap to be very large.. Has someone else also experienced this or am I doing something wrong here?\n\nI used a stratified Split (10%) to create my validation set and am using a pretrained efficientnet-b2 model with mild data augmentation. I'm also using a jpeg encoded version of the dataset at 256x256 px.",
    "1762544": "No, you are not doing anything wrong! My numbers are in the same ballpark. It is just an effect of the weird data split, as you observed. Makes this challenge more interesting, in my opinion...\n\nWhat we need is a way to penalize networks that lock on to the wrong features (e.g. the color of the dirt). If the plants can be segmented out before training the network, that should help too.\n\nIf the training data included field IDs, that would have helped us construct a validation set that mirrors the train-test split. Maybe the timestamp can be used as a loose proxy for field location.",
    "1764185": "Thanks for your reply! Glad to know it's due to the split and not a bad validation split or something.\n\nYes it is indeed more interesting. I agree that getting rid of the dirt might be interesting to help with generalization to the other parcels.",
    "1767751": "I guess, if your image is augmented in train set and not in valid set, it can make it hard to predict for train set. So it makes sense.\nbecause train images has randomly transformed, validation loss can be lower than training loss.",
    "1790953": "HI , I too am not able to achieve val accuracy >0.5 with effnetb2. I'm not sure what is the reason for this high bias. Initially i tried freezing all the layers except the top but was not able to achieve train accuracy>0.45. which means fine tuning too will not have any significant improvements. Then I unfreeze all layers and got the train and val accuracy>0.9 but test acc=0.5.  I'm stuck with this problem. Let me know if you find any solutions and reasons for high bias. In my opinion , it is due to overfitting since I am training all layers on train data which may cause overfitting very quickly, though I'm not sure.If that was the case why would i get val accuracy >0.9."
  },
  "source": "meta"
}