{
  "id": 208402,
  "title": "Important points to boost the LB score",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/208402",
  "author_name": "",
  "post_date": "2021-01-03T10:11:48.707182400Z",
  "votes": 48,
  "comment_count": 36,
  "views": 0,
  "content": "<p>I thought it would be great to have a list of <strong>consolidated points</strong> which have/are proved to be useful to push the LB score (no matter how small).</p>\n<p>The following are some points that I have tried, have read from the discussion forum and from the notebooks published:</p>\n<ol>\n<li>Using Label smoothing (Both 1st and 2nd because of the presence of noisy labels in the dataset)</li>\n<li>Bi-tempered-logistic-loss (Description found <a href=\"https://ai.googleblog.com/2019/08/bi-tempered-logistic-loss-for-training.html\" target=\"_blank\">here</a>).</li>\n<li>Keeping the batch normalization layers frozen.</li>\n<li>Training with early stopping criterion.</li>\n<li>Variants (mainly B3, B4) of EfficientNet with 'imagenet' weights.</li>\n<li>Employ Image augmentation techniques on training data. Care has to be taken while applying augmentation so as not to induce train-test contamination.</li>\n<li>N-Fold CV</li>\n</ol>\n<p>I am sure this list is not even close to being completed and I really think this list would be a very good learning source for everyone. There will be tons of other things that many people would have tried and worked for them. This list can only be completed by everyone's contribution. So, please feel free to comment on the things you may have tried and proved useful or point out something that I have written is wrong.</p>",
  "messages": [
    {
      "id": "1136668",
      "postDate": "01/03/2021 10:11:48",
      "content": "<p>I thought it would be great to have a list of <strong>consolidated points</strong> which have/are proved to be useful to push the LB score (no matter how small).</p>\n<p>The following are some points that I have tried, have read from the discussion forum and from the notebooks published:</p>\n<ol>\n<li>Using Label smoothing (Both 1st and 2nd because of the presence of noisy labels in the dataset)</li>\n<li>Bi-tempered-logistic-loss (Description found <a href=\"https://ai.googleblog.com/2019/08/bi-tempered-logistic-loss-for-training.html\" target=\"_blank\">here</a>).</li>\n<li>Keeping the batch normalization layers frozen.</li>\n<li>Training with early stopping criterion.</li>\n<li>Variants (mainly B3, B4) of EfficientNet with 'imagenet' weights.</li>\n<li>Employ Image augmentation techniques on training data. Care has to be taken while applying augmentation so as not to induce train-test contamination.</li>\n<li>N-Fold CV</li>\n</ol>\n<p>I am sure this list is not even close to being completed and I really think this list would be a very good learning source for everyone. There will be tons of other things that many people would have tried and worked for them. This list can only be completed by everyone's contribution. So, please feel free to comment on the things you may have tried and proved useful or point out something that I have written is wrong.</p>",
      "rawMarkdown": "I thought it would be great to have a list of **consolidated points** which have/are proved to be useful to push the LB score (no matter how small).\n\nThe following are some points that I have tried, have read from the discussion forum and from the notebooks published:\n\n1. Using Label smoothing (Both 1st and 2nd because of the presence of noisy labels in the dataset)\n2. Bi-tempered-logistic-loss (Description found [here](https://ai.googleblog.com/2019/08/bi-tempered-logistic-loss-for-training.html)).\n3. Keeping the batch normalization layers frozen.\n4. Training with early stopping criterion.\n5. Variants (mainly B3, B4) of EfficientNet with 'imagenet' weights.\n6. Employ Image augmentation techniques on training data. Care has to be taken while applying augmentation so as not to induce train-test contamination.\n7. N-Fold CV\n\nI am sure this list is not even close to being completed and I really think this list would be a very good learning source for everyone. There will be tons of other things that many people would have tried and worked for them. This list can only be completed by everyone's contribution. So, please feel free to comment on the things you may have tried and proved useful or point out something that I have written is wrong.",
      "votes": null
    },
    {
      "id": "1138687",
      "postDate": "01/04/2021 21:34:26",
      "content": "<p>Calculating the <code>mean</code> and <code>std</code> for a certain image size on that dataset also increased LB. It peformed better than using ImageNets <code>mean</code> and <code>std</code> to normalize the images.</p>",
      "rawMarkdown": "Calculating the `mean` and `std` for a certain image size on that dataset also increased LB. It peformed better than using ImageNets `mean` and `std` to normalize the images.",
      "votes": null
    },
    {
      "id": "1139163",
      "postDate": "01/05/2021 08:04:05",
      "content": "<p>Do you mean that you created <strong>buckets</strong> of <strong>mean</strong> and <strong>std</strong> based on <strong>image size</strong> and then normalized the image rather than using Imagenet's common mean and std to normalize the images?</p>",
      "rawMarkdown": "Do you mean that you created **buckets** of **mean** and **std** based on **image size** and then normalized the image rather than using Imagenet's common mean and std to normalize the images?",
      "votes": null
    },
    {
      "id": "1139231",
      "postDate": "01/05/2021 09:02:47",
      "content": "<ul>\n<li>5. For us, noisy student weight is better than the imagenet weights for EfficientNet</li>\n<li>7. N=5 for most of machine learning problem</li>\n</ul>",
      "rawMarkdown": "5. For us, noisy student weight is better than the imagenet weights for EfficientNet\n- 7. N=5 for most of machine learning problem",
      "votes": null
    },
    {
      "id": "1139293",
      "postDate": "01/05/2021 09:53:29",
      "content": "<p>Did you find that bi-tempered logistic loss gave you an improved CV over other loss functions?</p>",
      "rawMarkdown": "Did you find that bi-tempered logistic loss gave you an improved CV over other loss functions?",
      "votes": null
    },
    {
      "id": "1139341",
      "postDate": "01/05/2021 10:28:22",
      "content": "<p><a href=\"https://www.kaggle.com/param1\" target=\"_blank\">@param1</a> Yes it did indeed</p>",
      "rawMarkdown": "param1 Yes it did indeed",
      "votes": null
    },
    {
      "id": "1139345",
      "postDate": "01/05/2021 10:29:47",
      "content": "<p>I meant to try out the noisy-student weight but haven't done it yet. <br>\nThanks a lot for this information !!</p>",
      "rawMarkdown": "I meant to try out the noisy-student weight but haven't done it yet. \nThanks a lot for this information !!",
      "votes": null
    },
    {
      "id": "1140509",
      "postDate": "01/06/2021 04:26:00",
      "content": "<p>How can I implement noisy student weights? the GitHub tensorflow version is not working for me </p>",
      "rawMarkdown": "How can I implement noisy student weights? the GitHub tensorflow version is not working for me",
      "votes": null
    },
    {
      "id": "1140539",
      "postDate": "01/06/2021 05:00:47",
      "content": "<p>Well, you can train using the noisy student training method from the original EfficientNet github repo to get the noisy student…<br>\nBut, what I meant here is to use the pretrained weight from the noisy student checkpoint instead of the imagenet checkpoint. It is just a pretrained weight. Nothing to implement.<br>\nIf you are using keras, you can get the noisy student pretrained weight here (last section)<br>\n<a href=\"https://keras.io/examples/vision/image_classification_efficientnet_fine_tuning/\" target=\"_blank\">https://keras.io/examples/vision/image_classification_efficientnet_fine_tuning/</a></p>",
      "rawMarkdown": "Well, you can train using the noisy student training method from the original EfficientNet github repo to get the noisy student...\nBut, what I meant here is to use the pretrained weight from the noisy student checkpoint instead of the imagenet checkpoint. It is just a pretrained weight. Nothing to implement.\nIf you are using keras, you can get the noisy student pretrained weight here (last section)\nhttps://keras.io/examples/vision/image_classification_efficientnet_fine_tuning/",
      "votes": null
    },
    {
      "id": "1140613",
      "postDate": "01/06/2021 06:05:48",
      "content": "<p>Thanks for that</p>",
      "rawMarkdown": "Thanks for that",
      "votes": null
    },
    {
      "id": "1141330",
      "postDate": "01/06/2021 16:29:44",
      "content": "<p>I mean that you take each image from your dataset and calculate their <code>mean</code> and <code>std</code> afterwards you build an average for all these calculated values and you get an <code>average_mean</code> and <code>average_std</code> which you can use to normalize the images. It works better for me both CV and LB using my own <code>mean</code> and <code>std</code> instead of using the ones from ImageNet. It is also recommended because you are extracting features on your dataset basis and not on ImageNet basis.</p>\n<p>OpenCV comes with a handy function for this use-case: <code>mean, std = cv2.meanStdDev(img)</code></p>",
      "rawMarkdown": "I mean that you take each image from your dataset and calculate their `mean` and `std` afterwards you build an average for all these calculated values and you get an `average_mean` and `average_std` which you can use to normalize the images. It works better for me both CV and LB using my own `mean` and `std` instead of using the ones from ImageNet. It is also recommended because you are extracting features on your dataset basis and not on ImageNet basis.\n\nOpenCV comes with a handy function for this use-case: `mean, std = cv2.meanStdDev(img)`",
      "votes": null
    },
    {
      "id": "1142660",
      "postDate": "01/07/2021 14:23:35",
      "content": "<p>Okay, so you are basically calculating your own <strong>mean</strong> and <strong>std</strong> for the data, rather than relying on Imagenet's mean and std. Sounds simple and effective. Thanks a lot for this.</p>",
      "rawMarkdown": "Okay, so you are basically calculating your own **mean** and **std** for the data, rather than relying on Imagenet's mean and std. Sounds simple and effective. Thanks a lot for this.",
      "votes": null
    },
    {
      "id": "1144147",
      "postDate": "01/08/2021 09:06:11",
      "content": "<p>I found this just now and got 2 new points to try. Thanks for the post, I am getting 0.885 now and in a slump after trying so many techniques and hacks, I have exhausted my GPU limit and now training on my laptop. Please update if you find anything else.</p>",
      "rawMarkdown": "I found this just now and got 2 new points to try. Thanks for the post, I am getting 0.885 now and in a slump after trying so many techniques and hacks, I have exhausted my GPU limit and now training on my laptop. Please update if you find anything else.",
      "votes": null
    },
    {
      "id": "1144490",
      "postDate": "01/08/2021 13:39:16",
      "content": "<p>For sure I will. For exhausting your GPU limit, what batch size are you using?</p>",
      "rawMarkdown": "For sure I will. For exhausting your GPU limit, what batch size are you using?",
      "votes": null
    },
    {
      "id": "1144563",
      "postDate": "01/08/2021 14:33:29",
      "content": "<p>Can somebody help me implement label smoothing? I am getting this error <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/209093\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/209093</a></p>",
      "rawMarkdown": "Can somebody help me implement label smoothing? I am getting this error https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/209093",
      "votes": null
    },
    {
      "id": "1144636",
      "postDate": "01/08/2021 15:27:23",
      "content": "<p>Sorry for not being clear, I meant GPU quota. BTW I am using 32 as my batch size. Anymore than that my system is throwing an out of memory error.</p>",
      "rawMarkdown": "Sorry for not being clear, I meant GPU quota. BTW I am using 32 as my batch size. Anymore than that my system is throwing an out of memory error.",
      "votes": null
    },
    {
      "id": "1144667",
      "postDate": "01/08/2021 15:47:37",
      "content": "<p><a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a> - Can you post the mean and std dev of the dataset here please. I tried calculating it but ran out of memory. </p>",
      "rawMarkdown": "aliabdin1 - Can you post the mean and std dev of the dataset here please. I tried calculating it but ran out of memory.",
      "votes": null
    },
    {
      "id": "1144818",
      "postDate": "01/08/2021 17:30:55",
      "content": "<p>Its depending on the image size, which one are you using?</p>",
      "rawMarkdown": "Its depending on the image size, which one are you using?",
      "votes": null
    },
    {
      "id": "1144822",
      "postDate": "01/08/2021 17:35:22",
      "content": "<p>I can't help you with your certain tf problem but you can implement label smoothing like following:</p>\n<pre><code>if self.label_smoothing &gt; 0:\n      labels = (1-self.label_smoothing) * labels + (self.label_smoothing/self.num_classes)\n</code></pre>\n<p>This is a part of my rainforest dataset. </p>\n<ul>\n<li>labels have to be one-hot encoded (in my case its just a PyTorch tensor)</li>\n<li>label_smoothing adjusts how much smoothing you want</li>\n<li>num_classes are the amount of classes in your dataset</li>\n</ul>",
      "rawMarkdown": "I can't help you with your certain tf problem but you can implement label smoothing like following:\n\n```\nif self.label_smoothing > 0:\n      labels = (1-self.label_smoothing) * labels + (self.label_smoothing/self.num_classes)\n```\n\nThis is a part of my rainforest dataset. \n\n- labels have to be one-hot encoded (in my case its just a PyTorch tensor)\n- label_smoothing adjusts how much smoothing you want\n- num_classes are the amount of classes in your dataset",
      "votes": null
    },
    {
      "id": "1144978",
      "postDate": "01/08/2021 19:42:52",
      "content": "<p>I am mostly using 384, 512, and 576.</p>",
      "rawMarkdown": "I am mostly using 384, 512, and 576.",
      "votes": null
    },
    {
      "id": "1144994",
      "postDate": "01/08/2021 19:53:15",
      "content": "<p>For 512x512 I calculated:</p>\n<p><code>mean=(0.42984136, 0.49624753, 0.3129598)</code> and <code>std=(0.21417203, 0.21910103, 0.19542212)</code></p>",
      "rawMarkdown": "For 512x512 I calculated:\n\n`mean=(0.42984136, 0.49624753, 0.3129598)` and `std=(0.21417203, 0.21910103, 0.19542212)`",
      "votes": null
    },
    {
      "id": "1146423",
      "postDate": "01/09/2021 19:01:20",
      "content": "<p>if you use Tensorflow, you can use Crossentropy Categorical loss which provide label smoothing very easily. <a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/losses/categorical_crossentropy\" target=\"_blank\">https://www.tensorflow.org/api_docs/python/tf/keras/losses/categorical_crossentropy</a></p>",
      "rawMarkdown": "if you use Tensorflow, you can use Crossentropy Categorical loss which provide label smoothing very easily. https://www.tensorflow.org/api_docs/python/tf/keras/losses/categorical_crossentropy",
      "votes": null
    },
    {
      "id": "1146496",
      "postDate": "01/09/2021 19:58:01",
      "content": "<p>According to his opened discussion: <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/209093\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/209093</a> the already implemented CE with label smoothing is causing an error</p>",
      "rawMarkdown": "According to his opened discussion: https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/209093 the already implemented CE with label smoothing is causing an error",
      "votes": null
    },
    {
      "id": "1146539",
      "postDate": "01/09/2021 20:46:01",
      "content": "<p>Do you think applying GANs for Data Augmentation will be a good step? Because I have read some of the discussion forums and they have said that there are mislabeled images. So, can you comment on this? </p>",
      "rawMarkdown": "Do you think applying GANs for Data Augmentation will be a good step? Because I have read some of the discussion forums and they have said that there are mislabeled images. So, can you comment on this?",
      "votes": null
    },
    {
      "id": "1146752",
      "postDate": "01/10/2021 02:49:56",
      "content": "<p>Yes that is the problem </p>",
      "rawMarkdown": "Yes that is the problem",
      "votes": null
    },
    {
      "id": "1148433",
      "postDate": "01/11/2021 06:04:02",
      "content": "<p>Also, remember to be very light with TTA. A little rotation or simple TTA augmentation (like normalization) is enough.</p>",
      "rawMarkdown": "Also, remember to be very light with TTA. A little rotation or simple TTA augmentation (like normalization) is enough.",
      "votes": null
    },
    {
      "id": "1149211",
      "postDate": "01/11/2021 17:07:58",
      "content": "<p>You could just do the label-smoothing by yourself and feed the transformed tensors into your loss-function. You can use my implementation from above.</p>",
      "rawMarkdown": "You could just do the label-smoothing by yourself and feed the transformed tensors into your loss-function. You can use my implementation from above.",
      "votes": null
    },
    {
      "id": "1149417",
      "postDate": "01/11/2021 20:25:16",
      "content": "<p><a href=\"https://www.kaggle.com/prvnkmr\" target=\"_blank\">@prvnkmr</a> Great work man!</p>",
      "rawMarkdown": "prvnkmr Great work man!",
      "votes": null
    },
    {
      "id": "1151597",
      "postDate": "01/13/2021 12:43:41",
      "content": "<p>Hi, Ali! Thanks a lot for the insight! Could you please answer my question: do you calculate mean and std for the whole dataset or for the train and test part separately?</p>",
      "rawMarkdown": "Hi, Ali! Thanks a lot for the insight! Could you please answer my question: do you calculate mean and std for the whole dataset or for the train and test part separately?",
      "votes": null
    },
    {
      "id": "1151611",
      "postDate": "01/13/2021 12:57:18",
      "content": "<p>Only for the trainset to get the normalization as good as possible during training and then I apply the same calculated <code>mean</code> and <code>std</code> during inference</p>",
      "rawMarkdown": "Only for the trainset to get the normalization as good as possible during training and then I apply the same calculated `mean` and `std` during inference",
      "votes": null
    },
    {
      "id": "1151628",
      "postDate": "01/13/2021 13:12:44",
      "content": "<p>Thanks! One more question: why did you ask for input size to calculate mean and std in one of your previous answers? Do you first resize corresponding to selected input size and then calculate mean and std? If yes, do you do that with opencv as well? </p>",
      "rawMarkdown": "Thanks! One more question: why did you ask for input size to calculate mean and std in one of your previous answers? Do you first resize corresponding to selected input size and then calculate mean and std? If yes, do you do that with opencv as well?",
      "votes": null
    },
    {
      "id": "1152072",
      "postDate": "01/13/2021 19:14:06",
      "content": "<p>Yes I do everything with OpenCV. You have to resize the images accordingly to get their right corresponding <code>mean</code> and <code>std</code></p>",
      "rawMarkdown": "Yes I do everything with OpenCV. You have to resize the images accordingly to get their right corresponding `mean` and `std`",
      "votes": null
    },
    {
      "id": "1152174",
      "postDate": "01/13/2021 22:07:44",
      "content": "<p>Hmm, I tried to implement calculation using opencv as you suggested, but the output mean and std are very low..Could you please share the code that you use to calculate that numbers? Would be very helpful</p>",
      "rawMarkdown": "Hmm, I tried to implement calculation using opencv as you suggested, but the output mean and std are very low..Could you please share the code that you use to calculate that numbers? Would be very helpful",
      "votes": null
    },
    {
      "id": "1152180",
      "postDate": "01/13/2021 22:15:55",
      "content": "<p>Piece of code that I wrote to calculate mean and std:</p>\n<p>`def get_mean_std_opencv(train_df):</p>\n<pre><code>sum_mean = 0\nsum_std = 0\nfile_name_list = train_df[\"image_id\"].values\nfor file_name in file_name_list:\n    print(file_name)\n    file_path = f\"{CFG.TRAIN_PATH}/{file_name}\"\n    image = cv2.imread(file_path)\n    image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n    resized_image = cv2.resize(image, (CFG.size, CFG.size))\n    mean, std = cv2.meanStdDev(resized_image)\n    sum_mean += mean\n    sum_std += std\n    avg_mean = sum_mean/len(file_name_list)\n    avg_std = sum_std/len(file_name_list)\n    return avg_mean, avg_std`\n</code></pre>",
      "rawMarkdown": "Piece of code that I wrote to calculate mean and std:\n\n`def get_mean_std_opencv(train_df):\n\n    sum_mean = 0\n    sum_std = 0\n    file_name_list = train_df[\"image_id\"].values\n    for file_name in file_name_list:\n        print(file_name)\n        file_path = f\"{CFG.TRAIN_PATH}/{file_name}\"\n        image = cv2.imread(file_path)\n        image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n        resized_image = cv2.resize(image, (CFG.size, CFG.size))\n        mean, std = cv2.meanStdDev(resized_image)\n        sum_mean += mean\n        sum_std += std\n        avg_mean = sum_mean/len(file_name_list)\n        avg_std = sum_std/len(file_name_list)\n        return avg_mean, avg_std`",
      "votes": null
    },
    {
      "id": "1154758",
      "postDate": "01/15/2021 21:41:39",
      "content": "<p>I made my notebook public on how to calculate std and mean of images. Comments within the notebook will come soon:</p>\n<p><a href=\"https://www.kaggle.com/aliabdin1/calculate-mean-std-of-images\" target=\"_blank\">https://www.kaggle.com/aliabdin1/calculate-mean-std-of-images</a></p>",
      "rawMarkdown": "I made my notebook public on how to calculate std and mean of images. Comments within the notebook will come soon:\n\nhttps://www.kaggle.com/aliabdin1/calculate-mean-std-of-images",
      "votes": null
    },
    {
      "id": "1158903",
      "postDate": "01/18/2021 20:57:07",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/prvnkmr\" target=\"_blank\">@prvnkmr</a> ! Thanks a lot for the insights! Could you please clarify what you mean by freezing the batch norm layers? Which layers exactly?</p>",
      "rawMarkdown": "Hi, @prvnkmr ! Thanks a lot for the insights! Could you please clarify what you mean by freezing the batch norm layers? Which layers exactly?",
      "votes": null
    },
    {
      "id": "1197122",
      "postDate": "02/12/2021 00:06:59",
      "content": "<p>Have anyone of you experienced that just a simple rescaling of the pixel values (i.e. division by 255) gives better results compared when using image net stats, or cassava dataset stats? <br>\nCheers.</p>",
      "rawMarkdown": "Have anyone of you experienced that just a simple rescaling of the pixel values (i.e. division by 255) gives better results compared when using image net stats, or cassava dataset stats? \nCheers.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1138687,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "01/04/2021 21:34:26",
      "content": "<p>Calculating the <code>mean</code> and <code>std</code> for a certain image size on that dataset also increased LB. It peformed better than using ImageNets <code>mean</code> and <code>std</code> to normalize the images.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1139163,
          "author_name": "prvnkmr",
          "author_url": "",
          "post_date": "01/05/2021 08:04:05",
          "content": "<p>Do you mean that you created <strong>buckets</strong> of <strong>mean</strong> and <strong>std</strong> based on <strong>image size</strong> and then normalized the image rather than using Imagenet's common mean and std to normalize the images?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1141330,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "01/06/2021 16:29:44",
          "content": "<p>I mean that you take each image from your dataset and calculate their <code>mean</code> and <code>std</code> afterwards you build an average for all these calculated values and you get an <code>average_mean</code> and <code>average_std</code> which you can use to normalize the images. It works better for me both CV and LB using my own <code>mean</code> and <code>std</code> instead of using the ones from ImageNet. It is also recommended because you are extracting features on your dataset basis and not on ImageNet basis.</p>\n<p>OpenCV comes with a handy function for this use-case: <code>mean, std = cv2.meanStdDev(img)</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1142660,
          "author_name": "prvnkmr",
          "author_url": "",
          "post_date": "01/07/2021 14:23:35",
          "content": "<p>Okay, so you are basically calculating your own <strong>mean</strong> and <strong>std</strong> for the data, rather than relying on Imagenet's mean and std. Sounds simple and effective. Thanks a lot for this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1144667,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "01/08/2021 15:47:37",
          "content": "<p><a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a> - Can you post the mean and std dev of the dataset here please. I tried calculating it but ran out of memory. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1144818,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "01/08/2021 17:30:55",
          "content": "<p>Its depending on the image size, which one are you using?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1144978,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "01/08/2021 19:42:52",
          "content": "<p>I am mostly using 384, 512, and 576.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1144994,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "01/08/2021 19:53:15",
          "content": "<p>For 512x512 I calculated:</p>\n<p><code>mean=(0.42984136, 0.49624753, 0.3129598)</code> and <code>std=(0.21417203, 0.21910103, 0.19542212)</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1151597,
          "author_name": "etagiev",
          "author_url": "",
          "post_date": "01/13/2021 12:43:41",
          "content": "<p>Hi, Ali! Thanks a lot for the insight! Could you please answer my question: do you calculate mean and std for the whole dataset or for the train and test part separately?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1151611,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "01/13/2021 12:57:18",
          "content": "<p>Only for the trainset to get the normalization as good as possible during training and then I apply the same calculated <code>mean</code> and <code>std</code> during inference</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1151628,
          "author_name": "etagiev",
          "author_url": "",
          "post_date": "01/13/2021 13:12:44",
          "content": "<p>Thanks! One more question: why did you ask for input size to calculate mean and std in one of your previous answers? Do you first resize corresponding to selected input size and then calculate mean and std? If yes, do you do that with opencv as well? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1152072,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "01/13/2021 19:14:06",
          "content": "<p>Yes I do everything with OpenCV. You have to resize the images accordingly to get their right corresponding <code>mean</code> and <code>std</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1152174,
          "author_name": "etagiev",
          "author_url": "",
          "post_date": "01/13/2021 22:07:44",
          "content": "<p>Hmm, I tried to implement calculation using opencv as you suggested, but the output mean and std are very low..Could you please share the code that you use to calculate that numbers? Would be very helpful</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1152180,
          "author_name": "etagiev",
          "author_url": "",
          "post_date": "01/13/2021 22:15:55",
          "content": "<p>Piece of code that I wrote to calculate mean and std:</p>\n<p>`def get_mean_std_opencv(train_df):</p>\n<pre><code>sum_mean = 0\nsum_std = 0\nfile_name_list = train_df[\"image_id\"].values\nfor file_name in file_name_list:\n    print(file_name)\n    file_path = f\"{CFG.TRAIN_PATH}/{file_name}\"\n    image = cv2.imread(file_path)\n    image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n    resized_image = cv2.resize(image, (CFG.size, CFG.size))\n    mean, std = cv2.meanStdDev(resized_image)\n    sum_mean += mean\n    sum_std += std\n    avg_mean = sum_mean/len(file_name_list)\n    avg_std = sum_std/len(file_name_list)\n    return avg_mean, avg_std`\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1154758,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "01/15/2021 21:41:39",
          "content": "<p>I made my notebook public on how to calculate std and mean of images. Comments within the notebook will come soon:</p>\n<p><a href=\"https://www.kaggle.com/aliabdin1/calculate-mean-std-of-images\" target=\"_blank\">https://www.kaggle.com/aliabdin1/calculate-mean-std-of-images</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1139231,
      "author_name": "louis925",
      "author_url": "",
      "post_date": "01/05/2021 09:02:47",
      "content": "<ul>\n<li>5. For us, noisy student weight is better than the imagenet weights for EfficientNet</li>\n<li>7. N=5 for most of machine learning problem</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1139345,
          "author_name": "prvnkmr",
          "author_url": "",
          "post_date": "01/05/2021 10:29:47",
          "content": "<p>I meant to try out the noisy-student weight but haven't done it yet. <br>\nThanks a lot for this information !!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1140509,
          "author_name": "mithilsalunkhe",
          "author_url": "",
          "post_date": "01/06/2021 04:26:00",
          "content": "<p>How can I implement noisy student weights? the GitHub tensorflow version is not working for me </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1140539,
          "author_name": "louis925",
          "author_url": "",
          "post_date": "01/06/2021 05:00:47",
          "content": "<p>Well, you can train using the noisy student training method from the original EfficientNet github repo to get the noisy student…<br>\nBut, what I meant here is to use the pretrained weight from the noisy student checkpoint instead of the imagenet checkpoint. It is just a pretrained weight. Nothing to implement.<br>\nIf you are using keras, you can get the noisy student pretrained weight here (last section)<br>\n<a href=\"https://keras.io/examples/vision/image_classification_efficientnet_fine_tuning/\" target=\"_blank\">https://keras.io/examples/vision/image_classification_efficientnet_fine_tuning/</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1140613,
          "author_name": "mithilsalunkhe",
          "author_url": "",
          "post_date": "01/06/2021 06:05:48",
          "content": "<p>Thanks for that</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1139293,
      "author_name": "param1",
      "author_url": "",
      "post_date": "01/05/2021 09:53:29",
      "content": "<p>Did you find that bi-tempered logistic loss gave you an improved CV over other loss functions?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1139341,
          "author_name": "prvnkmr",
          "author_url": "",
          "post_date": "01/05/2021 10:28:22",
          "content": "<p><a href=\"https://www.kaggle.com/param1\" target=\"_blank\">@param1</a> Yes it did indeed</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1144147,
      "author_name": "mohneesh7",
      "author_url": "",
      "post_date": "01/08/2021 09:06:11",
      "content": "<p>I found this just now and got 2 new points to try. Thanks for the post, I am getting 0.885 now and in a slump after trying so many techniques and hacks, I have exhausted my GPU limit and now training on my laptop. Please update if you find anything else.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1144490,
          "author_name": "prvnkmr",
          "author_url": "",
          "post_date": "01/08/2021 13:39:16",
          "content": "<p>For sure I will. For exhausting your GPU limit, what batch size are you using?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1144636,
          "author_name": "mohneesh7",
          "author_url": "",
          "post_date": "01/08/2021 15:27:23",
          "content": "<p>Sorry for not being clear, I meant GPU quota. BTW I am using 32 as my batch size. Anymore than that my system is throwing an out of memory error.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1144563,
      "author_name": "mithilsalunkhe",
      "author_url": "",
      "post_date": "01/08/2021 14:33:29",
      "content": "<p>Can somebody help me implement label smoothing? I am getting this error <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/209093\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/209093</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1144822,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "01/08/2021 17:35:22",
          "content": "<p>I can't help you with your certain tf problem but you can implement label smoothing like following:</p>\n<pre><code>if self.label_smoothing &gt; 0:\n      labels = (1-self.label_smoothing) * labels + (self.label_smoothing/self.num_classes)\n</code></pre>\n<p>This is a part of my rainforest dataset. </p>\n<ul>\n<li>labels have to be one-hot encoded (in my case its just a PyTorch tensor)</li>\n<li>label_smoothing adjusts how much smoothing you want</li>\n<li>num_classes are the amount of classes in your dataset</li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1146423,
          "author_name": "durbin164",
          "author_url": "",
          "post_date": "01/09/2021 19:01:20",
          "content": "<p>if you use Tensorflow, you can use Crossentropy Categorical loss which provide label smoothing very easily. <a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/losses/categorical_crossentropy\" target=\"_blank\">https://www.tensorflow.org/api_docs/python/tf/keras/losses/categorical_crossentropy</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1146496,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "01/09/2021 19:58:01",
          "content": "<p>According to his opened discussion: <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/209093\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/209093</a> the already implemented CE with label smoothing is causing an error</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1146752,
          "author_name": "mithilsalunkhe",
          "author_url": "",
          "post_date": "01/10/2021 02:49:56",
          "content": "<p>Yes that is the problem </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1149211,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "01/11/2021 17:07:58",
          "content": "<p>You could just do the label-smoothing by yourself and feed the transformed tensors into your loss-function. You can use my implementation from above.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1146539,
      "author_name": "hotsonhonet",
      "author_url": "",
      "post_date": "01/09/2021 20:46:01",
      "content": "<p>Do you think applying GANs for Data Augmentation will be a good step? Because I have read some of the discussion forums and they have said that there are mislabeled images. So, can you comment on this? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1148433,
      "author_name": "prvnkmr",
      "author_url": "",
      "post_date": "01/11/2021 06:04:02",
      "content": "<p>Also, remember to be very light with TTA. A little rotation or simple TTA augmentation (like normalization) is enough.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1149417,
      "author_name": "jonbjones",
      "author_url": "",
      "post_date": "01/11/2021 20:25:16",
      "content": "<p><a href=\"https://www.kaggle.com/prvnkmr\" target=\"_blank\">@prvnkmr</a> Great work man!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1158903,
      "author_name": "etagiev",
      "author_url": "",
      "post_date": "01/18/2021 20:57:07",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/prvnkmr\" target=\"_blank\">@prvnkmr</a> ! Thanks a lot for the insights! Could you please clarify what you mean by freezing the batch norm layers? Which layers exactly?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1197122,
      "author_name": "marjan1111",
      "author_url": "",
      "post_date": "02/12/2021 00:06:59",
      "content": "<p>Have anyone of you experienced that just a simple rescaling of the pixel values (i.e. division by 255) gives better results compared when using image net stats, or cassava dataset stats? <br>\nCheers.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1136668": "I thought it would be great to have a list of **consolidated points** which have/are proved to be useful to push the LB score (no matter how small).\n\nThe following are some points that I have tried, have read from the discussion forum and from the notebooks published:\n\n1. Using Label smoothing (Both 1st and 2nd because of the presence of noisy labels in the dataset)\n2. Bi-tempered-logistic-loss (Description found [here](https://ai.googleblog.com/2019/08/bi-tempered-logistic-loss-for-training.html)).\n3. Keeping the batch normalization layers frozen.\n4. Training with early stopping criterion.\n5. Variants (mainly B3, B4) of EfficientNet with 'imagenet' weights.\n6. Employ Image augmentation techniques on training data. Care has to be taken while applying augmentation so as not to induce train-test contamination.\n7. N-Fold CV\n\nI am sure this list is not even close to being completed and I really think this list would be a very good learning source for everyone. There will be tons of other things that many people would have tried and worked for them. This list can only be completed by everyone's contribution. So, please feel free to comment on the things you may have tried and proved useful or point out something that I have written is wrong.",
    "1138687": "Calculating the `mean` and `std` for a certain image size on that dataset also increased LB. It peformed better than using ImageNets `mean` and `std` to normalize the images.",
    "1139163": "Do you mean that you created **buckets** of **mean** and **std** based on **image size** and then normalized the image rather than using Imagenet's common mean and std to normalize the images?",
    "1139231": "5. For us, noisy student weight is better than the imagenet weights for EfficientNet\n- 7. N=5 for most of machine learning problem",
    "1139293": "Did you find that bi-tempered logistic loss gave you an improved CV over other loss functions?",
    "1139341": "param1 Yes it did indeed",
    "1139345": "I meant to try out the noisy-student weight but haven't done it yet. \nThanks a lot for this information !!",
    "1140509": "How can I implement noisy student weights? the GitHub tensorflow version is not working for me",
    "1140539": "Well, you can train using the noisy student training method from the original EfficientNet github repo to get the noisy student...\nBut, what I meant here is to use the pretrained weight from the noisy student checkpoint instead of the imagenet checkpoint. It is just a pretrained weight. Nothing to implement.\nIf you are using keras, you can get the noisy student pretrained weight here (last section)\nhttps://keras.io/examples/vision/image_classification_efficientnet_fine_tuning/",
    "1140613": "Thanks for that",
    "1141330": "I mean that you take each image from your dataset and calculate their `mean` and `std` afterwards you build an average for all these calculated values and you get an `average_mean` and `average_std` which you can use to normalize the images. It works better for me both CV and LB using my own `mean` and `std` instead of using the ones from ImageNet. It is also recommended because you are extracting features on your dataset basis and not on ImageNet basis.\n\nOpenCV comes with a handy function for this use-case: `mean, std = cv2.meanStdDev(img)`",
    "1142660": "Okay, so you are basically calculating your own **mean** and **std** for the data, rather than relying on Imagenet's mean and std. Sounds simple and effective. Thanks a lot for this.",
    "1144147": "I found this just now and got 2 new points to try. Thanks for the post, I am getting 0.885 now and in a slump after trying so many techniques and hacks, I have exhausted my GPU limit and now training on my laptop. Please update if you find anything else.",
    "1144490": "For sure I will. For exhausting your GPU limit, what batch size are you using?",
    "1144563": "Can somebody help me implement label smoothing? I am getting this error https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/209093",
    "1144636": "Sorry for not being clear, I meant GPU quota. BTW I am using 32 as my batch size. Anymore than that my system is throwing an out of memory error.",
    "1144667": "aliabdin1 - Can you post the mean and std dev of the dataset here please. I tried calculating it but ran out of memory.",
    "1144818": "Its depending on the image size, which one are you using?",
    "1144822": "I can't help you with your certain tf problem but you can implement label smoothing like following:\n\n```\nif self.label_smoothing > 0:\n      labels = (1-self.label_smoothing) * labels + (self.label_smoothing/self.num_classes)\n```\n\nThis is a part of my rainforest dataset. \n\n- labels have to be one-hot encoded (in my case its just a PyTorch tensor)\n- label_smoothing adjusts how much smoothing you want\n- num_classes are the amount of classes in your dataset",
    "1144978": "I am mostly using 384, 512, and 576.",
    "1144994": "For 512x512 I calculated:\n\n`mean=(0.42984136, 0.49624753, 0.3129598)` and `std=(0.21417203, 0.21910103, 0.19542212)`",
    "1146423": "if you use Tensorflow, you can use Crossentropy Categorical loss which provide label smoothing very easily. https://www.tensorflow.org/api_docs/python/tf/keras/losses/categorical_crossentropy",
    "1146496": "According to his opened discussion: https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/209093 the already implemented CE with label smoothing is causing an error",
    "1146539": "Do you think applying GANs for Data Augmentation will be a good step? Because I have read some of the discussion forums and they have said that there are mislabeled images. So, can you comment on this?",
    "1146752": "Yes that is the problem",
    "1148433": "Also, remember to be very light with TTA. A little rotation or simple TTA augmentation (like normalization) is enough.",
    "1149211": "You could just do the label-smoothing by yourself and feed the transformed tensors into your loss-function. You can use my implementation from above.",
    "1149417": "prvnkmr Great work man!",
    "1151597": "Hi, Ali! Thanks a lot for the insight! Could you please answer my question: do you calculate mean and std for the whole dataset or for the train and test part separately?",
    "1151611": "Only for the trainset to get the normalization as good as possible during training and then I apply the same calculated `mean` and `std` during inference",
    "1151628": "Thanks! One more question: why did you ask for input size to calculate mean and std in one of your previous answers? Do you first resize corresponding to selected input size and then calculate mean and std? If yes, do you do that with opencv as well?",
    "1152072": "Yes I do everything with OpenCV. You have to resize the images accordingly to get their right corresponding `mean` and `std`",
    "1152174": "Hmm, I tried to implement calculation using opencv as you suggested, but the output mean and std are very low..Could you please share the code that you use to calculate that numbers? Would be very helpful",
    "1152180": "Piece of code that I wrote to calculate mean and std:\n\n`def get_mean_std_opencv(train_df):\n\n    sum_mean = 0\n    sum_std = 0\n    file_name_list = train_df[\"image_id\"].values\n    for file_name in file_name_list:\n        print(file_name)\n        file_path = f\"{CFG.TRAIN_PATH}/{file_name}\"\n        image = cv2.imread(file_path)\n        image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n        resized_image = cv2.resize(image, (CFG.size, CFG.size))\n        mean, std = cv2.meanStdDev(resized_image)\n        sum_mean += mean\n        sum_std += std\n        avg_mean = sum_mean/len(file_name_list)\n        avg_std = sum_std/len(file_name_list)\n        return avg_mean, avg_std`",
    "1154758": "I made my notebook public on how to calculate std and mean of images. Comments within the notebook will come soon:\n\nhttps://www.kaggle.com/aliabdin1/calculate-mean-std-of-images",
    "1158903": "Hi, @prvnkmr ! Thanks a lot for the insights! Could you please clarify what you mean by freezing the batch norm layers? Which layers exactly?",
    "1197122": "Have anyone of you experienced that just a simple rescaling of the pixel values (i.e. division by 255) gives better results compared when using image net stats, or cassava dataset stats? \nCheers."
  },
  "source": "meta"
}