{
  "id": 303532,
  "title": "Augmentations with YOLOV5",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/303532",
  "author_name": "",
  "post_date": "2022-01-28T05:06:03.142976300Z",
  "votes": 4,
  "comment_count": 13,
  "views": 0,
  "content": "<p>I have been trying to train a YOLOV5-M model and am wondering why adding augmentations is leading to a lower leaderboard score. With the default hyperparameters and 15 epochs for training, I achieve a leaderboard score of 0.508 but when changing the hyperparameter file to increase the \"shear\" and \"degrees\" to 2.0 and 3.0, respectively, my score drops to 0.403. However, the F2 score on the validation data increases by about 0.2 with this increased augmentation while mAP decreases by about 0.002.</p>\n<p>I also noticed that when I increase image inference size to 3600, the score drops significantly to 0.416 compared to 0.508.</p>\n<p>Does anyone have any suggestions or ideas as to why this may be occurring?</p>",
  "messages": [
    {
      "id": "1666824",
      "postDate": "01/28/2022 05:06:03",
      "content": "<p>I have been trying to train a YOLOV5-M model and am wondering why adding augmentations is leading to a lower leaderboard score. With the default hyperparameters and 15 epochs for training, I achieve a leaderboard score of 0.508 but when changing the hyperparameter file to increase the \"shear\" and \"degrees\" to 2.0 and 3.0, respectively, my score drops to 0.403. However, the F2 score on the validation data increases by about 0.2 with this increased augmentation while mAP decreases by about 0.002.</p>\n<p>I also noticed that when I increase image inference size to 3600, the score drops significantly to 0.416 compared to 0.508.</p>\n<p>Does anyone have any suggestions or ideas as to why this may be occurring?</p>",
      "rawMarkdown": "I have been trying to train a YOLOV5-M model and am wondering why adding augmentations is leading to a lower leaderboard score. With the default hyperparameters and 15 epochs for training, I achieve a leaderboard score of 0.508 but when changing the hyperparameter file to increase the \"shear\" and \"degrees\" to 2.0 and 3.0, respectively, my score drops to 0.403. However, the F2 score on the validation data increases by about 0.2 with this increased augmentation while mAP decreases by about 0.002.\n\nI also noticed that when I increase image inference size to 3600, the score drops significantly to 0.416 compared to 0.508.\n\nDoes anyone have any suggestions or ideas as to why this may be occurring?",
      "votes": null
    },
    {
      "id": "1666866",
      "postDate": "01/28/2022 05:58:17",
      "content": "<p>shear of 2 and degrees of 3 isn't very much.</p>\n<p>have you had a look at the train_batch jpgs of each run to see what the differences are?</p>\n<p>which defaults are you using? scratch or finetune?</p>",
      "rawMarkdown": "shear of 2 and degrees of 3 isn't very much.\n\nhave you had a look at the train_batch jpgs of each run to see what the differences are?\n\nwhich defaults are you using? scratch or finetune?",
      "votes": null
    },
    {
      "id": "1666871",
      "postDate": "01/28/2022 06:02:55",
      "content": "<p>The shear is definitely noticeable in the train_batch jpgs and I am using the scratch defaults. Does it make sense though for validation to improve while there is such a large decrease in leaderboard score?</p>",
      "rawMarkdown": "The shear is definitely noticeable in the train_batch jpgs and I am using the scratch defaults. Does it make sense though for validation to improve while there is such a large decrease in leaderboard score?",
      "votes": null
    },
    {
      "id": "1666877",
      "postDate": "01/28/2022 06:20:14",
      "content": "<p>Imho, I think maybe increasing resolution significantly may detect more potential objects, which may increase the FP rate. Moreover, the annotations are not perfect, so it is likely that there are missing positive labels in some images with small COTS. While your algorithm detects those small objects correctly, there are missing annotations for those objects, and the evaluator considers your detections as \"False positive\".</p>\n<p>You may have to inspect the detection differences closely to see whether it is true and exactly what is happening.</p>",
      "rawMarkdown": "Imho, I think maybe increasing resolution significantly may detect more potential objects, which may increase the FP rate. Moreover, the annotations are not perfect, so it is likely that there are missing positive labels in some images with small COTS. While your algorithm detects those small objects correctly, there are missing annotations for those objects, and the evaluator considers your detections as \"False positive\".\n\nYou may have to inspect the detection differences closely to see whether it is true and exactly what is happening.",
      "votes": null
    },
    {
      "id": "1667091",
      "postDate": "01/28/2022 09:58:40",
      "content": "<p>I've certainly seen some of my own models perform well on F2 of validation, but have a drop in the LB.</p>\n<p>The only thing we can conclude is that they are different sets of data.</p>\n<p>Also, the different models may have different CONF/F2 peaks. Use <a href=\"https://www.kaggle.com/locbaop/systematic-evaluate-f2-yolov5\" target=\"_blank\">https://www.kaggle.com/locbaop/systematic-evaluate-f2-yolov5</a> to determine optimal CONF score for LB</p>",
      "rawMarkdown": "I've certainly seen some of my own models perform well on F2 of validation, but have a drop in the LB.\n\nThe only thing we can conclude is that they are different sets of data.\n\nAlso, the different models may have different CONF/F2 peaks. Use https://www.kaggle.com/locbaop/systematic-evaluate-f2-yolov5 to determine optimal CONF score for LB",
      "votes": null
    },
    {
      "id": "1667254",
      "postDate": "01/28/2022 12:33:18",
      "content": "<p>I see but the strange thing is when I run my inferencing notebook with 3600 size images, the model struggles in finding many of the COTS as compared to the 1280 size images. While it seems a lower score makes sense, why would a larger inference size cause this? From the other discussions, it seems like larger inference size has been helping with scores. </p>",
      "rawMarkdown": "I see but the strange thing is when I run my inferencing notebook with 3600 size images, the model struggles in finding many of the COTS as compared to the 1280 size images. While it seems a lower score makes sense, why would a larger inference size cause this? From the other discussions, it seems like larger inference size has been helping with scores.",
      "votes": null
    },
    {
      "id": "1667460",
      "postDate": "01/28/2022 16:11:43",
      "content": "<p>Do you have any other ideas behind why these augmentations cause such a drop of 0.1 mAP? It seems unreasonable for a relatively small change and better CV scores lead to such a large drop. </p>",
      "rawMarkdown": "Do you have any other ideas behind why these augmentations cause such a drop of 0.1 mAP? It seems unreasonable for a relatively small change and better CV scores lead to such a large drop.",
      "votes": null
    },
    {
      "id": "1667639",
      "postDate": "01/28/2022 19:29:35",
      "content": "<p>What you observed is called CV-LB gap and this is one of the biggest problem (to solve) in this competition. Assuming your evaluation code is working well and you have no data leak between your local folds, this gap is usually related to a shift in data distribution between training and LB.</p>",
      "rawMarkdown": "What you observed is called CV-LB gap and this is one of the biggest problem (to solve) in this competition. Assuming your evaluation code is working well and you have no data leak between your local folds, this gap is usually related to a shift in data distribution between training and LB.",
      "votes": null
    },
    {
      "id": "1667692",
      "postDate": "01/28/2022 20:22:47",
      "content": "<p>I am using the 5 fold subsequences dataset, with fold 1 as the validation. Shouldn’t CV-LB gap be consistent with the data used? The same data trained with slightly increased augmentations leads to a significantly lower score. </p>",
      "rawMarkdown": "I am using the 5 fold subsequences dataset, with fold 1 as the validation. Shouldn’t CV-LB gap be consistent with the data used? The same data trained with slightly increased augmentations leads to a significantly lower score.",
      "votes": null
    },
    {
      "id": "1667738",
      "postDate": "01/28/2022 21:46:06",
      "content": "<p>The competition LB metric is F2, not mAP.</p>\n<p>Choosing the wrong CONF score itself would explain this, especially if it is close to 0.0 or 1.0.</p>\n<p>I would run both models through the \"evaluate F2\" notebook and look at the F2 curves to confirm whether your augmented model is worse than the baseline.</p>",
      "rawMarkdown": "The competition LB metric is F2, not mAP.\n\nChoosing the wrong CONF score itself would explain this, especially if it is close to 0.0 or 1.0.\n\nI would run both models through the \"evaluate F2\" notebook and look at the F2 curves to confirm whether your augmented model is worse than the baseline.",
      "votes": null
    },
    {
      "id": "1667773",
      "postDate": "01/28/2022 23:22:48",
      "content": "<p>Sorry, I meant drop of 0.1 in F2 score on public leaderboard.</p>\n<p>I just ran both models through the notebook and made sure to change the data being used to fold 1 from the 5-fold subsequence dataset (which was my validation set for training) and the model with lighter augmentations gave an F2 score of 0.5995 at a confidence threshold of 0.44 while the model with more augmentations gave an F2 score of 0.600 at a confidence threshold of 0.17. These were around the same thresholds I used when making the submissions stated above.</p>",
      "rawMarkdown": "Sorry, I meant drop of 0.1 in F2 score on public leaderboard.\n\nI just ran both models through the notebook and made sure to change the data being used to fold 1 from the 5-fold subsequence dataset (which was my validation set for training) and the model with lighter augmentations gave an F2 score of 0.5995 at a confidence threshold of 0.44 while the model with more augmentations gave an F2 score of 0.600 at a confidence threshold of 0.17. These were around the same thresholds I used when making the submissions stated above.",
      "votes": null
    },
    {
      "id": "1667785",
      "postDate": "01/29/2022 00:05:14",
      "content": "<p>Yes that's informative.</p>\n<p>I've observed with my own models that the one that gave a lower validation F2 sometimes have higher LB score as well. It is hard for me to explain why. Which is why I say that the hidden dataset is somewhat different to the training dataset.</p>\n<p>Higher optimal CONF usually means they have higher false positive rates, due to their willingness to call uncertainties at higher confidences. So perhaps your lighter model has higher recall because of this?</p>\n<p>Are your two models trained at the same number of epochs? It may be your higher augmentation model was trained longer and is overfitting.</p>",
      "rawMarkdown": "Yes that's informative.\n\nI've observed with my own models that the one that gave a lower validation F2 sometimes have higher LB score as well. It is hard for me to explain why. Which is why I say that the hidden dataset is somewhat different to the training dataset.\n\nHigher optimal CONF usually means they have higher false positive rates, due to their willingness to call uncertainties at higher confidences. So perhaps your lighter model has higher recall because of this?\n\nAre your two models trained at the same number of epochs? It may be your higher augmentation model was trained longer and is overfitting.",
      "votes": null
    },
    {
      "id": "1667788",
      "postDate": "01/29/2022 00:11:35",
      "content": "<p>The lighter model did have a slightly higher recall. Yes, both models are trained for 15 epochs.</p>",
      "rawMarkdown": "The lighter model did have a slightly higher recall. Yes, both models are trained for 15 epochs.",
      "votes": null
    },
    {
      "id": "1669902",
      "postDate": "01/31/2022 03:51:58",
      "content": "<p>I recently tried training a model with only 5 degrees of rotation and the same model without any rotation. Although CV scores were similar, leaderboard scores differed greatly with 0.438 for the rotation model and 0.513 for no rotation. Why is this?</p>",
      "rawMarkdown": "I recently tried training a model with only 5 degrees of rotation and the same model without any rotation. Although CV scores were similar, leaderboard scores differed greatly with 0.438 for the rotation model and 0.513 for no rotation. Why is this?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1666866,
      "author_name": "alexchwong",
      "author_url": "",
      "post_date": "01/28/2022 05:58:17",
      "content": "<p>shear of 2 and degrees of 3 isn't very much.</p>\n<p>have you had a look at the train_batch jpgs of each run to see what the differences are?</p>\n<p>which defaults are you using? scratch or finetune?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1666871,
          "author_name": "ayu055",
          "author_url": "",
          "post_date": "01/28/2022 06:02:55",
          "content": "<p>The shear is definitely noticeable in the train_batch jpgs and I am using the scratch defaults. Does it make sense though for validation to improve while there is such a large decrease in leaderboard score?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1667091,
          "author_name": "alexchwong",
          "author_url": "",
          "post_date": "01/28/2022 09:58:40",
          "content": "<p>I've certainly seen some of my own models perform well on F2 of validation, but have a drop in the LB.</p>\n<p>The only thing we can conclude is that they are different sets of data.</p>\n<p>Also, the different models may have different CONF/F2 peaks. Use <a href=\"https://www.kaggle.com/locbaop/systematic-evaluate-f2-yolov5\" target=\"_blank\">https://www.kaggle.com/locbaop/systematic-evaluate-f2-yolov5</a> to determine optimal CONF score for LB</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1667460,
          "author_name": "ayu055",
          "author_url": "",
          "post_date": "01/28/2022 16:11:43",
          "content": "<p>Do you have any other ideas behind why these augmentations cause such a drop of 0.1 mAP? It seems unreasonable for a relatively small change and better CV scores lead to such a large drop. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1667738,
          "author_name": "alexchwong",
          "author_url": "",
          "post_date": "01/28/2022 21:46:06",
          "content": "<p>The competition LB metric is F2, not mAP.</p>\n<p>Choosing the wrong CONF score itself would explain this, especially if it is close to 0.0 or 1.0.</p>\n<p>I would run both models through the \"evaluate F2\" notebook and look at the F2 curves to confirm whether your augmented model is worse than the baseline.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1667773,
          "author_name": "ayu055",
          "author_url": "",
          "post_date": "01/28/2022 23:22:48",
          "content": "<p>Sorry, I meant drop of 0.1 in F2 score on public leaderboard.</p>\n<p>I just ran both models through the notebook and made sure to change the data being used to fold 1 from the 5-fold subsequence dataset (which was my validation set for training) and the model with lighter augmentations gave an F2 score of 0.5995 at a confidence threshold of 0.44 while the model with more augmentations gave an F2 score of 0.600 at a confidence threshold of 0.17. These were around the same thresholds I used when making the submissions stated above.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1667785,
          "author_name": "alexchwong",
          "author_url": "",
          "post_date": "01/29/2022 00:05:14",
          "content": "<p>Yes that's informative.</p>\n<p>I've observed with my own models that the one that gave a lower validation F2 sometimes have higher LB score as well. It is hard for me to explain why. Which is why I say that the hidden dataset is somewhat different to the training dataset.</p>\n<p>Higher optimal CONF usually means they have higher false positive rates, due to their willingness to call uncertainties at higher confidences. So perhaps your lighter model has higher recall because of this?</p>\n<p>Are your two models trained at the same number of epochs? It may be your higher augmentation model was trained longer and is overfitting.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1667788,
          "author_name": "ayu055",
          "author_url": "",
          "post_date": "01/29/2022 00:11:35",
          "content": "<p>The lighter model did have a slightly higher recall. Yes, both models are trained for 15 epochs.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1666877,
      "author_name": "anhnhunhat",
      "author_url": "",
      "post_date": "01/28/2022 06:20:14",
      "content": "<p>Imho, I think maybe increasing resolution significantly may detect more potential objects, which may increase the FP rate. Moreover, the annotations are not perfect, so it is likely that there are missing positive labels in some images with small COTS. While your algorithm detects those small objects correctly, there are missing annotations for those objects, and the evaluator considers your detections as \"False positive\".</p>\n<p>You may have to inspect the detection differences closely to see whether it is true and exactly what is happening.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1667254,
          "author_name": "ayu055",
          "author_url": "",
          "post_date": "01/28/2022 12:33:18",
          "content": "<p>I see but the strange thing is when I run my inferencing notebook with 3600 size images, the model struggles in finding many of the COTS as compared to the 1280 size images. While it seems a lower score makes sense, why would a larger inference size cause this? From the other discussions, it seems like larger inference size has been helping with scores. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1667639,
      "author_name": "alexandrecc",
      "author_url": "",
      "post_date": "01/28/2022 19:29:35",
      "content": "<p>What you observed is called CV-LB gap and this is one of the biggest problem (to solve) in this competition. Assuming your evaluation code is working well and you have no data leak between your local folds, this gap is usually related to a shift in data distribution between training and LB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1667692,
          "author_name": "ayu055",
          "author_url": "",
          "post_date": "01/28/2022 20:22:47",
          "content": "<p>I am using the 5 fold subsequences dataset, with fold 1 as the validation. Shouldn’t CV-LB gap be consistent with the data used? The same data trained with slightly increased augmentations leads to a significantly lower score. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1669902,
      "author_name": "ayu055",
      "author_url": "",
      "post_date": "01/31/2022 03:51:58",
      "content": "<p>I recently tried training a model with only 5 degrees of rotation and the same model without any rotation. Although CV scores were similar, leaderboard scores differed greatly with 0.438 for the rotation model and 0.513 for no rotation. Why is this?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1666824": "I have been trying to train a YOLOV5-M model and am wondering why adding augmentations is leading to a lower leaderboard score. With the default hyperparameters and 15 epochs for training, I achieve a leaderboard score of 0.508 but when changing the hyperparameter file to increase the \"shear\" and \"degrees\" to 2.0 and 3.0, respectively, my score drops to 0.403. However, the F2 score on the validation data increases by about 0.2 with this increased augmentation while mAP decreases by about 0.002.\n\nI also noticed that when I increase image inference size to 3600, the score drops significantly to 0.416 compared to 0.508.\n\nDoes anyone have any suggestions or ideas as to why this may be occurring?",
    "1666866": "shear of 2 and degrees of 3 isn't very much.\n\nhave you had a look at the train_batch jpgs of each run to see what the differences are?\n\nwhich defaults are you using? scratch or finetune?",
    "1666871": "The shear is definitely noticeable in the train_batch jpgs and I am using the scratch defaults. Does it make sense though for validation to improve while there is such a large decrease in leaderboard score?",
    "1666877": "Imho, I think maybe increasing resolution significantly may detect more potential objects, which may increase the FP rate. Moreover, the annotations are not perfect, so it is likely that there are missing positive labels in some images with small COTS. While your algorithm detects those small objects correctly, there are missing annotations for those objects, and the evaluator considers your detections as \"False positive\".\n\nYou may have to inspect the detection differences closely to see whether it is true and exactly what is happening.",
    "1667091": "I've certainly seen some of my own models perform well on F2 of validation, but have a drop in the LB.\n\nThe only thing we can conclude is that they are different sets of data.\n\nAlso, the different models may have different CONF/F2 peaks. Use https://www.kaggle.com/locbaop/systematic-evaluate-f2-yolov5 to determine optimal CONF score for LB",
    "1667254": "I see but the strange thing is when I run my inferencing notebook with 3600 size images, the model struggles in finding many of the COTS as compared to the 1280 size images. While it seems a lower score makes sense, why would a larger inference size cause this? From the other discussions, it seems like larger inference size has been helping with scores.",
    "1667460": "Do you have any other ideas behind why these augmentations cause such a drop of 0.1 mAP? It seems unreasonable for a relatively small change and better CV scores lead to such a large drop.",
    "1667639": "What you observed is called CV-LB gap and this is one of the biggest problem (to solve) in this competition. Assuming your evaluation code is working well and you have no data leak between your local folds, this gap is usually related to a shift in data distribution between training and LB.",
    "1667692": "I am using the 5 fold subsequences dataset, with fold 1 as the validation. Shouldn’t CV-LB gap be consistent with the data used? The same data trained with slightly increased augmentations leads to a significantly lower score.",
    "1667738": "The competition LB metric is F2, not mAP.\n\nChoosing the wrong CONF score itself would explain this, especially if it is close to 0.0 or 1.0.\n\nI would run both models through the \"evaluate F2\" notebook and look at the F2 curves to confirm whether your augmented model is worse than the baseline.",
    "1667773": "Sorry, I meant drop of 0.1 in F2 score on public leaderboard.\n\nI just ran both models through the notebook and made sure to change the data being used to fold 1 from the 5-fold subsequence dataset (which was my validation set for training) and the model with lighter augmentations gave an F2 score of 0.5995 at a confidence threshold of 0.44 while the model with more augmentations gave an F2 score of 0.600 at a confidence threshold of 0.17. These were around the same thresholds I used when making the submissions stated above.",
    "1667785": "Yes that's informative.\n\nI've observed with my own models that the one that gave a lower validation F2 sometimes have higher LB score as well. It is hard for me to explain why. Which is why I say that the hidden dataset is somewhat different to the training dataset.\n\nHigher optimal CONF usually means they have higher false positive rates, due to their willingness to call uncertainties at higher confidences. So perhaps your lighter model has higher recall because of this?\n\nAre your two models trained at the same number of epochs? It may be your higher augmentation model was trained longer and is overfitting.",
    "1667788": "The lighter model did have a slightly higher recall. Yes, both models are trained for 15 epochs.",
    "1669902": "I recently tried training a model with only 5 degrees of rotation and the same model without any rotation. Although CV scores were similar, leaderboard scores differed greatly with 0.438 for the rotation model and 0.513 for no rotation. Why is this?"
  },
  "source": "meta"
}