{
  "id": 169641,
  "title": "Efficient net training behaving weirdly",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/169641",
  "author_name": "",
  "post_date": "2020-07-24T15:49:32.433397800Z",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I see many efficient nets getting ~0.95 auc on training. My implementation can only reach 0.85 LB score. And during training it never climbs over 0.65. Validation loss stagnates and the model early stops.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1821809%2F215628a5d628946008cbfd3a4768c2d9%2F__results___26_0.png?generation=1595605504331009&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1821809%2Fb243b5396d63f825273cffe570d598a7%2F__results___26_1.png?generation=1595605593703559&amp;alt=media\" alt=\"\"></p>\n\n<p>I have no idea what I'm doing wrong I have an inception V3 version which behaves fine.</p>\n\n<p>Kernel is this if you want to check it out\n<a href=\"https://www.kaggle.com/amneves/tensorflow-efficientnetbx-transfer-learning\">https://www.kaggle.com/amneves/tensorflow-efficientnetbx-transfer-learning</a></p>",
  "messages": [
    {
      "id": "943820",
      "postDate": "07/24/2020 15:49:32",
      "content": "<p>I see many efficient nets getting ~0.95 auc on training. My implementation can only reach 0.85 LB score. And during training it never climbs over 0.65. Validation loss stagnates and the model early stops.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1821809%2F215628a5d628946008cbfd3a4768c2d9%2F__results___26_0.png?generation=1595605504331009&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1821809%2Fb243b5396d63f825273cffe570d598a7%2F__results___26_1.png?generation=1595605593703559&amp;alt=media\" alt=\"\"></p>\n\n<p>I have no idea what I'm doing wrong I have an inception V3 version which behaves fine.</p>\n\n<p>Kernel is this if you want to check it out\n<a href=\"https://www.kaggle.com/amneves/tensorflow-efficientnetbx-transfer-learning\">https://www.kaggle.com/amneves/tensorflow-efficientnetbx-transfer-learning</a></p>",
      "rawMarkdown": "I see many efficient nets getting ~0.95 auc on training. My implementation can only reach 0.85 LB score. And during training it never climbs over 0.65. Validation loss stagnates and the model early stops.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1821809%2F215628a5d628946008cbfd3a4768c2d9%2F__results___26_0.png?generation=1595605504331009&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1821809%2Fb243b5396d63f825273cffe570d598a7%2F__results___26_1.png?generation=1595605593703559&amp;alt=media)\n\nI have no idea what I'm doing wrong I have an inception V3 version which behaves fine.\n\nKernel is this if you want to check it out\n[https://www.kaggle.com/amneves/tensorflow-efficientnetbx-transfer-learning](https://www.kaggle.com/amneves/tensorflow-efficientnetbx-transfer-learning)",
      "votes": null
    },
    {
      "id": "943914",
      "postDate": "07/24/2020 17:05:27",
      "content": "<p>lr</p>",
      "rawMarkdown": "lr",
      "votes": null
    },
    {
      "id": "943915",
      "postDate": "07/24/2020 17:07:17",
      "content": "<p>That's what I thought. Any tips?</p>",
      "rawMarkdown": "That's what I thought. Any tips?",
      "votes": null
    },
    {
      "id": "943919",
      "postDate": "07/24/2020 17:11:19",
      "content": "<p>On Epoch 6/16  , lr: 0.0205.I think lr is too large.</p>",
      "rawMarkdown": "On Epoch 6/16  , lr: 0.0205.I think lr is too large.",
      "votes": null
    },
    {
      "id": "943958",
      "postDate": "07/24/2020 17:59:12",
      "content": "<p>As commented, it looks like your Learning Rate is too large. You have the code:</p>\n\n<p>lr_max     = 0.000020 * REPLICAS * batch_size</p>\n\n<p>I don't think you want to multiply by batch_size.</p>\n\n<p>-Rich</p>",
      "rawMarkdown": "As commented, it looks like your Learning Rate is too large. You have the code:\n\n lr_max     = 0.000020 * REPLICAS * batch_size\n\nI don't think you want to multiply by batch_size.\n\n-Rich",
      "votes": null
    },
    {
      "id": "944526",
      "postDate": "07/25/2020 07:07:33",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1690820%2F555ba305717d78a2d6a031f190de9e3f%2FScreenshot%20from%202020-07-25%2010-02-53.png?generation=1595660606549263&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1690820%2F9724b20f1b7d0aac873252dedf379cd5%2FScreenshot%20from%202020-07-25%2010-03-05.png?generation=1595660621477968&amp;alt=media\" alt=\"\"></p>\n\n<p>Red - train log, Blue - valid log\nAs you can see have the same problem - models starts to overfit after a few epochs.\nSo possible solutions are:\n- Use best model from metric and use ansambling to make it stable\n- Use lower LR\n- Use models that have fewer number of params but overall better accuracy (for example effnets)</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1690820%2F555ba305717d78a2d6a031f190de9e3f%2FScreenshot%20from%202020-07-25%2010-02-53.png?generation=1595660606549263&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1690820%2F9724b20f1b7d0aac873252dedf379cd5%2FScreenshot%20from%202020-07-25%2010-03-05.png?generation=1595660621477968&amp;alt=media)\n\nRed - train log, Blue - valid log\nAs you can see have the same problem - models starts to overfit after a few epochs.\nSo possible solutions are:\n- Use best model from metric and use ansambling to make it stable\n- Use lower LR\n- Use models that have fewer number of params but overall better accuracy (for example effnets)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 943914,
      "author_name": "chihantsai",
      "author_url": "",
      "post_date": "07/24/2020 17:05:27",
      "content": "<p>lr</p>",
      "votes": null,
      "replies": [
        {
          "id": 943915,
          "author_name": "amneves",
          "author_url": "",
          "post_date": "07/24/2020 17:07:17",
          "content": "<p>That's what I thought. Any tips?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 943919,
          "author_name": "chihantsai",
          "author_url": "",
          "post_date": "07/24/2020 17:11:19",
          "content": "<p>On Epoch 6/16  , lr: 0.0205.I think lr is too large.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 943958,
      "author_name": "richardepstein",
      "author_url": "",
      "post_date": "07/24/2020 17:59:12",
      "content": "<p>As commented, it looks like your Learning Rate is too large. You have the code:</p>\n\n<p>lr_max     = 0.000020 * REPLICAS * batch_size</p>\n\n<p>I don't think you want to multiply by batch_size.</p>\n\n<p>-Rich</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 944526,
      "author_name": "vladimirsydor",
      "author_url": "",
      "post_date": "07/25/2020 07:07:33",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1690820%2F555ba305717d78a2d6a031f190de9e3f%2FScreenshot%20from%202020-07-25%2010-02-53.png?generation=1595660606549263&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1690820%2F9724b20f1b7d0aac873252dedf379cd5%2FScreenshot%20from%202020-07-25%2010-03-05.png?generation=1595660621477968&amp;alt=media\" alt=\"\"></p>\n\n<p>Red - train log, Blue - valid log\nAs you can see have the same problem - models starts to overfit after a few epochs.\nSo possible solutions are:\n- Use best model from metric and use ansambling to make it stable\n- Use lower LR\n- Use models that have fewer number of params but overall better accuracy (for example effnets)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "943820": "I see many efficient nets getting ~0.95 auc on training. My implementation can only reach 0.85 LB score. And during training it never climbs over 0.65. Validation loss stagnates and the model early stops.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1821809%2F215628a5d628946008cbfd3a4768c2d9%2F__results___26_0.png?generation=1595605504331009&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1821809%2Fb243b5396d63f825273cffe570d598a7%2F__results___26_1.png?generation=1595605593703559&amp;alt=media)\n\nI have no idea what I'm doing wrong I have an inception V3 version which behaves fine.\n\nKernel is this if you want to check it out\n[https://www.kaggle.com/amneves/tensorflow-efficientnetbx-transfer-learning](https://www.kaggle.com/amneves/tensorflow-efficientnetbx-transfer-learning)",
    "943914": "lr",
    "943915": "That's what I thought. Any tips?",
    "943919": "On Epoch 6/16  , lr: 0.0205.I think lr is too large.",
    "943958": "As commented, it looks like your Learning Rate is too large. You have the code:\n\n lr_max     = 0.000020 * REPLICAS * batch_size\n\nI don't think you want to multiply by batch_size.\n\n-Rich",
    "944526": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1690820%2F555ba305717d78a2d6a031f190de9e3f%2FScreenshot%20from%202020-07-25%2010-02-53.png?generation=1595660606549263&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1690820%2F9724b20f1b7d0aac873252dedf379cd5%2FScreenshot%20from%202020-07-25%2010-03-05.png?generation=1595660621477968&amp;alt=media)\n\nRed - train log, Blue - valid log\nAs you can see have the same problem - models starts to overfit after a few epochs.\nSo possible solutions are:\n- Use best model from metric and use ansambling to make it stable\n- Use lower LR\n- Use models that have fewer number of params but overall better accuracy (for example effnets)"
  },
  "source": "meta"
}