{
  "id": 225843,
  "title": "Any tips for training Efficientnets",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/225843",
  "author_name": "",
  "post_date": "2021-03-14T09:06:35.611109200Z",
  "votes": 4,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Hi guys, just wondering if any of you faces trouble getting a stable pipeline for say, efficientnet b5 noisy student? I used <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> heavy aug pipeline as well as the warmup scheduled. I also used <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> multihead pipeline on top of it. However, the results are suboptimal and tend to “overfit” because loss increases over the epochs. </p>",
  "messages": [
    {
      "id": "1237563",
      "postDate": "03/14/2021 09:06:35",
      "content": "<p>Hi guys, just wondering if any of you faces trouble getting a stable pipeline for say, efficientnet b5 noisy student? I used <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> heavy aug pipeline as well as the warmup scheduled. I also used <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> multihead pipeline on top of it. However, the results are suboptimal and tend to “overfit” because loss increases over the epochs. </p>",
      "rawMarkdown": "Hi guys, just wondering if any of you faces trouble getting a stable pipeline for say, efficientnet b5 noisy student? I used @underwearfitting heavy aug pipeline as well as the warmup scheduled. I also used @ttahara multihead pipeline on top of it. However, the results are suboptimal and tend to “overfit” because loss increases over the epochs.",
      "votes": null
    },
    {
      "id": "1237732",
      "postDate": "03/14/2021 12:07:42",
      "content": "<p>I tried effnets only on tpu with tensorflow/keras to experiment some stuff. I had similar problems, it was overfitting pretty fast, but I didn't dig deeper into reasons behind it, so cannot say it's common for effnets on this competition…</p>",
      "rawMarkdown": "I tried effnets only on tpu with tensorflow/keras to experiment some stuff. I had similar problems, it was overfitting pretty fast, but I didn't dig deeper into reasons behind it, so cannot say it's common for effnets on this competition...",
      "votes": null
    },
    {
      "id": "1238189",
      "postDate": "03/14/2021 18:19:43",
      "content": "<p>Freezing batch norm really helped me.<br>\ncv 0.961, lb 0.963, difference in val and train loss practically the same as for resnet200.<br>\nwithout freezing - 0.958/0.959 and val loss barely decreasing<br>\nBut for other models frozen batch norm didnt help at all, dont know why so xD</p>",
      "rawMarkdown": "Freezing batch norm really helped me.\ncv 0.961, lb 0.963, difference in val and train loss practically the same as for resnet200.\nwithout freezing - 0.958/0.959 and val loss barely decreasing\nBut for other models frozen batch norm didnt help at all, dont know why so xD",
      "votes": null
    },
    {
      "id": "1238474",
      "postDate": "03/15/2021 02:07:37",
      "content": "<p>Try running atleast 2 epochs without augumentations (only horizontal flip, resize and normalize), it gave me booost from cv 96.02 to cv 96.27!!, try increasing you image size and also try freezing batch norm, it gives some boost especially to effnets.</p>",
      "rawMarkdown": "Try running atleast 2 epochs without augumentations (only horizontal flip, resize and normalize), it gave me booost from cv 96.02 to cv 96.27!!, try increasing you image size and also try freezing batch norm, it gives some boost especially to effnets.",
      "votes": null
    },
    {
      "id": "1238516",
      "postDate": "03/15/2021 04:14:05",
      "content": "<p>I’m still confused on the proper code to freeze batch norm layers. It seems that the code I got from previous discussion didn’t set model to evaluation mode. Do I need to?</p>",
      "rawMarkdown": "I’m still confused on the proper code to freeze batch norm layers. It seems that the code I got from previous discussion didn’t set model to evaluation mode. Do I need to?",
      "votes": null
    },
    {
      "id": "1238694",
      "postDate": "03/15/2021 08:09:58",
      "content": "<p>Here is what i use:</p>\n<pre><code>def set_batchnorm_eval(m): \n    classname = m.__class__.__name__\n    if classname.find('BatchNorm') != -1:\n        m.eval()\n</code></pre>\n<p>And then applying this function to model after i switch model to train mode:<br>\n<code>model.apply(set_batchnorm_eval)</code></p>",
      "rawMarkdown": "Here is what i use:\n```\ndef set_batchnorm_eval(m): \n    classname = m.__class__.__name__\n    if classname.find('BatchNorm') != -1:\n        m.eval()\n```\n\nAnd then applying this function to model after i switch model to train mode:\n`model.apply(set_batchnorm_eval)`",
      "votes": null
    },
    {
      "id": "1238878",
      "postDate": "03/15/2021 11:22:18",
      "content": "<p>Thanks a lot! Did u use the image net weights or pretrained ones from this dataset?</p>",
      "rawMarkdown": "Thanks a lot! Did u use the image net weights or pretrained ones from this dataset?",
      "votes": null
    },
    {
      "id": "1239002",
      "postDate": "03/15/2021 12:15:23",
      "content": "<p>Glad i could help)<br>\nThe ones from the dataset</p>",
      "rawMarkdown": "Glad i could help)\nThe ones from the dataset",
      "votes": null
    },
    {
      "id": "1239139",
      "postDate": "03/15/2021 13:36:25",
      "content": "<p><a href=\"https://www.kaggle.com/edyanakov\" target=\"_blank\">@edyanakov</a> Glad that you helped. However, even after freezing batchnorm, I get insanely high cv score in the first 1-2 epochs (0.962-0.965), and then afterwards become super low like 0.95. However, I am using the <code>multihead</code> approach - meaning using 4 different classifier layers. I guess simplicity should prevail…?</p>",
      "rawMarkdown": "edyanakov Glad that you helped. However, even after freezing batchnorm, I get insanely high cv score in the first 1-2 epochs (0.962-0.965), and then afterwards become super low like 0.95. However, I am using the `multihead` approach - meaning using 4 different classifier layers. I guess simplicity should prevail...?",
      "votes": null
    },
    {
      "id": "1239252",
      "postDate": "03/15/2021 15:09:20",
      "content": "<p>I tried  <a href=\"https://github.comqubvel/efficientnet\" target=\"_blank\">qubvel/efficientnet</a> and found it was much better than <a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/applications/efficientnet\" target=\"_blank\">tensorflow efficientnet</a>. I don't know what is the difference between the two versions.</p>",
      "rawMarkdown": "I tried  [qubvel/efficientnet](https://github.comqubvel/efficientnet) and found it was much better than [tensorflow efficientnet](https://www.tensorflow.org/api_docs/python/tf/keras/applications/efficientnet). I don't know what is the difference between the two versions.",
      "votes": null
    },
    {
      "id": "1239388",
      "postDate": "03/15/2021 17:15:21",
      "content": "<p>Two differences that i am aware of. (1) They expect different preprocessing for the input images. (2) They use different pretrained imagenet weights.</p>",
      "rawMarkdown": "Two differences that i am aware of. (1) They expect different preprocessing for the input images. (2) They use different pretrained imagenet weights.",
      "votes": null
    },
    {
      "id": "1239396",
      "postDate": "03/15/2021 17:19:55",
      "content": "<p>I noticed they use a different dropout layer. maybe there are other different layers</p>",
      "rawMarkdown": "I noticed they use a different dropout layer. maybe there are other different layers",
      "votes": null
    },
    {
      "id": "1239558",
      "postDate": "03/15/2021 20:01:05",
      "content": "<p>Yeah, seems like.</p>",
      "rawMarkdown": "Yeah, seems like.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1237732,
      "author_name": "datafan07",
      "author_url": "",
      "post_date": "03/14/2021 12:07:42",
      "content": "<p>I tried effnets only on tpu with tensorflow/keras to experiment some stuff. I had similar problems, it was overfitting pretty fast, but I didn't dig deeper into reasons behind it, so cannot say it's common for effnets on this competition…</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1238189,
      "author_name": "edyanakov",
      "author_url": "",
      "post_date": "03/14/2021 18:19:43",
      "content": "<p>Freezing batch norm really helped me.<br>\ncv 0.961, lb 0.963, difference in val and train loss practically the same as for resnet200.<br>\nwithout freezing - 0.958/0.959 and val loss barely decreasing<br>\nBut for other models frozen batch norm didnt help at all, dont know why so xD</p>",
      "votes": null,
      "replies": [
        {
          "id": 1238516,
          "author_name": "reighns",
          "author_url": "",
          "post_date": "03/15/2021 04:14:05",
          "content": "<p>I’m still confused on the proper code to freeze batch norm layers. It seems that the code I got from previous discussion didn’t set model to evaluation mode. Do I need to?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1238694,
          "author_name": "edyanakov",
          "author_url": "",
          "post_date": "03/15/2021 08:09:58",
          "content": "<p>Here is what i use:</p>\n<pre><code>def set_batchnorm_eval(m): \n    classname = m.__class__.__name__\n    if classname.find('BatchNorm') != -1:\n        m.eval()\n</code></pre>\n<p>And then applying this function to model after i switch model to train mode:<br>\n<code>model.apply(set_batchnorm_eval)</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1238878,
          "author_name": "reighns",
          "author_url": "",
          "post_date": "03/15/2021 11:22:18",
          "content": "<p>Thanks a lot! Did u use the image net weights or pretrained ones from this dataset?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1239002,
          "author_name": "edyanakov",
          "author_url": "",
          "post_date": "03/15/2021 12:15:23",
          "content": "<p>Glad i could help)<br>\nThe ones from the dataset</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1239139,
          "author_name": "reighns",
          "author_url": "",
          "post_date": "03/15/2021 13:36:25",
          "content": "<p><a href=\"https://www.kaggle.com/edyanakov\" target=\"_blank\">@edyanakov</a> Glad that you helped. However, even after freezing batchnorm, I get insanely high cv score in the first 1-2 epochs (0.962-0.965), and then afterwards become super low like 0.95. However, I am using the <code>multihead</code> approach - meaning using 4 different classifier layers. I guess simplicity should prevail…?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1239558,
          "author_name": "edyanakov",
          "author_url": "",
          "post_date": "03/15/2021 20:01:05",
          "content": "<p>Yeah, seems like.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1238474,
      "author_name": "anku5hk",
      "author_url": "",
      "post_date": "03/15/2021 02:07:37",
      "content": "<p>Try running atleast 2 epochs without augumentations (only horizontal flip, resize and normalize), it gave me booost from cv 96.02 to cv 96.27!!, try increasing you image size and also try freezing batch norm, it gives some boost especially to effnets.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1239252,
      "author_name": "mariammohamed",
      "author_url": "",
      "post_date": "03/15/2021 15:09:20",
      "content": "<p>I tried  <a href=\"https://github.comqubvel/efficientnet\" target=\"_blank\">qubvel/efficientnet</a> and found it was much better than <a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/applications/efficientnet\" target=\"_blank\">tensorflow efficientnet</a>. I don't know what is the difference between the two versions.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1239388,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "03/15/2021 17:15:21",
          "content": "<p>Two differences that i am aware of. (1) They expect different preprocessing for the input images. (2) They use different pretrained imagenet weights.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1239396,
          "author_name": "mariammohamed",
          "author_url": "",
          "post_date": "03/15/2021 17:19:55",
          "content": "<p>I noticed they use a different dropout layer. maybe there are other different layers</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1237563": "Hi guys, just wondering if any of you faces trouble getting a stable pipeline for say, efficientnet b5 noisy student? I used @underwearfitting heavy aug pipeline as well as the warmup scheduled. I also used @ttahara multihead pipeline on top of it. However, the results are suboptimal and tend to “overfit” because loss increases over the epochs.",
    "1237732": "I tried effnets only on tpu with tensorflow/keras to experiment some stuff. I had similar problems, it was overfitting pretty fast, but I didn't dig deeper into reasons behind it, so cannot say it's common for effnets on this competition...",
    "1238189": "Freezing batch norm really helped me.\ncv 0.961, lb 0.963, difference in val and train loss practically the same as for resnet200.\nwithout freezing - 0.958/0.959 and val loss barely decreasing\nBut for other models frozen batch norm didnt help at all, dont know why so xD",
    "1238474": "Try running atleast 2 epochs without augumentations (only horizontal flip, resize and normalize), it gave me booost from cv 96.02 to cv 96.27!!, try increasing you image size and also try freezing batch norm, it gives some boost especially to effnets.",
    "1238516": "I’m still confused on the proper code to freeze batch norm layers. It seems that the code I got from previous discussion didn’t set model to evaluation mode. Do I need to?",
    "1238694": "Here is what i use:\n```\ndef set_batchnorm_eval(m): \n    classname = m.__class__.__name__\n    if classname.find('BatchNorm') != -1:\n        m.eval()\n```\n\nAnd then applying this function to model after i switch model to train mode:\n`model.apply(set_batchnorm_eval)`",
    "1238878": "Thanks a lot! Did u use the image net weights or pretrained ones from this dataset?",
    "1239002": "Glad i could help)\nThe ones from the dataset",
    "1239139": "edyanakov Glad that you helped. However, even after freezing batchnorm, I get insanely high cv score in the first 1-2 epochs (0.962-0.965), and then afterwards become super low like 0.95. However, I am using the `multihead` approach - meaning using 4 different classifier layers. I guess simplicity should prevail...?",
    "1239252": "I tried  [qubvel/efficientnet](https://github.comqubvel/efficientnet) and found it was much better than [tensorflow efficientnet](https://www.tensorflow.org/api_docs/python/tf/keras/applications/efficientnet). I don't know what is the difference between the two versions.",
    "1239388": "Two differences that i am aware of. (1) They expect different preprocessing for the input images. (2) They use different pretrained imagenet weights.",
    "1239396": "I noticed they use a different dropout layer. maybe there are other different layers",
    "1239558": "Yeah, seems like."
  },
  "source": "meta"
}