{
  "id": 310206,
  "title": "how to adjust learning rate?",
  "url": "/competitions/happy-whale-and-dolphin/discussion/310206",
  "author_name": "",
  "post_date": "2022-02-28T06:07:53.686576200Z",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I change the model to big one in my forked pytorch-arcface-gempooling-tpu-train notebook.</p>\n<p>however, the loss goes to plateau without decreasing.</p>\n<p>Any suggestion </p>\n<p>%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%<br>\nLaunching a training on 8 TPU cores.<br>\nValidation Loss: 17.33965392112732<br>\nModel Saved</p>\n<p>Validation Loss: 16.07572067975998<br>\nModel Saved</p>\n<p>Validation Loss: 15.178503835201264<br>\nModel Saved</p>\n<p>Validation Loss: 14.623040127754212<br>\nModel Saved</p>\n<p>Validation Loss: 14.062633389234543<br>\nModel Saved</p>\n<p>Validation Loss: 14.027782171964645<br>\nModel Saved</p>\n<p>Validation Loss: 13.651807701587677<br>\nModel Saved</p>\n<p>Validation Loss: 13.665936523675919<br>\nModel Saved</p>\n<p>Validation Loss: 13.792583948373794<br>\nModel Saved</p>\n<p>Validation Loss: 13.604751998186112<br>\nModel Saved</p>\n<p>Validation Loss: 13.649900329113006<br>\nModel Saved</p>\n<p>Validation Loss: 13.736776286363602<br>\nModel Saved</p>\n<p>Validation Loss: 13.664650583267212<br>\nModel Saved</p>\n<p>Validation Loss: 13.632439285516739<br>\nModel Saved</p>\n<p>Validation Loss: 13.798904740810395<br>\nModel Saved</p>\n<p>Validation Loss: 13.680295634269715<br>\nModel Saved</p>\n<p>Validation Loss: 13.6862082362175<br>\nModel Saved</p>\n<p>Validation Loss: 13.732640331983566<br>\nModel Saved</p>\n<p>%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%</p>",
  "messages": [
    {
      "id": "1707099",
      "postDate": "02/28/2022 06:07:53",
      "content": "<p>I change the model to big one in my forked pytorch-arcface-gempooling-tpu-train notebook.</p>\n<p>however, the loss goes to plateau without decreasing.</p>\n<p>Any suggestion </p>\n<p>%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%<br>\nLaunching a training on 8 TPU cores.<br>\nValidation Loss: 17.33965392112732<br>\nModel Saved</p>\n<p>Validation Loss: 16.07572067975998<br>\nModel Saved</p>\n<p>Validation Loss: 15.178503835201264<br>\nModel Saved</p>\n<p>Validation Loss: 14.623040127754212<br>\nModel Saved</p>\n<p>Validation Loss: 14.062633389234543<br>\nModel Saved</p>\n<p>Validation Loss: 14.027782171964645<br>\nModel Saved</p>\n<p>Validation Loss: 13.651807701587677<br>\nModel Saved</p>\n<p>Validation Loss: 13.665936523675919<br>\nModel Saved</p>\n<p>Validation Loss: 13.792583948373794<br>\nModel Saved</p>\n<p>Validation Loss: 13.604751998186112<br>\nModel Saved</p>\n<p>Validation Loss: 13.649900329113006<br>\nModel Saved</p>\n<p>Validation Loss: 13.736776286363602<br>\nModel Saved</p>\n<p>Validation Loss: 13.664650583267212<br>\nModel Saved</p>\n<p>Validation Loss: 13.632439285516739<br>\nModel Saved</p>\n<p>Validation Loss: 13.798904740810395<br>\nModel Saved</p>\n<p>Validation Loss: 13.680295634269715<br>\nModel Saved</p>\n<p>Validation Loss: 13.6862082362175<br>\nModel Saved</p>\n<p>Validation Loss: 13.732640331983566<br>\nModel Saved</p>\n<p>%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%</p>",
      "rawMarkdown": "I change the model to big one in my forked pytorch-arcface-gempooling-tpu-train notebook.\n\nhowever, the loss goes to plateau without decreasing.\n\nAny suggestion \n\n\n%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%\nLaunching a training on 8 TPU cores.\nValidation Loss: 17.33965392112732\nModel Saved\n\nValidation Loss: 16.07572067975998\nModel Saved\n\nValidation Loss: 15.178503835201264\nModel Saved\n\nValidation Loss: 14.623040127754212\nModel Saved\n\nValidation Loss: 14.062633389234543\nModel Saved\n\nValidation Loss: 14.027782171964645\nModel Saved\n\nValidation Loss: 13.651807701587677\nModel Saved\n\nValidation Loss: 13.665936523675919\nModel Saved\n\nValidation Loss: 13.792583948373794\nModel Saved\n\nValidation Loss: 13.604751998186112\nModel Saved\n\nValidation Loss: 13.649900329113006\nModel Saved\n\nValidation Loss: 13.736776286363602\nModel Saved\n\nValidation Loss: 13.664650583267212\nModel Saved\n\nValidation Loss: 13.632439285516739\nModel Saved\n\nValidation Loss: 13.798904740810395\nModel Saved\n\nValidation Loss: 13.680295634269715\nModel Saved\n\nValidation Loss: 13.6862082362175\nModel Saved\n\nValidation Loss: 13.732640331983566\nModel Saved\n\n%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%",
      "votes": null
    },
    {
      "id": "1707160",
      "postDate": "02/28/2022 07:11:13",
      "content": "<p>I'd recommend to try some Find LR technic, like <a href=\"https://gist.github.com/karanchahal/dc0575bf21b976ea633dea8ceffaf9dc\" target=\"_blank\">this</a> one</p>",
      "rawMarkdown": "I'd recommend to try some Find LR technic, like [this](https://gist.github.com/karanchahal/dc0575bf21b976ea633dea8ceffaf9dc) one",
      "votes": null
    },
    {
      "id": "1707180",
      "postDate": "02/28/2022 07:35:50",
      "content": "<p>thank you!  I know that fast.ai has one.  I will have a look at the you post.</p>",
      "rawMarkdown": "thank you!  I know that fast.ai has one.  I will have a look at the you post.",
      "votes": null
    },
    {
      "id": "1707187",
      "postDate": "02/28/2022 07:40:49",
      "content": "<p>simple version is just change your scheduler to ExponentialLR scheduler and see by eyes :)</p>",
      "rawMarkdown": "simple version is just change your scheduler to ExponentialLR scheduler and see by eyes :)",
      "votes": null
    },
    {
      "id": "1707192",
      "postDate": "02/28/2022 07:46:06",
      "content": "<p>Also remember The Linear Scaling Rule: When the minibatch size is multiplied by k, multiply the learning rate by k. I guess when you changed model to bigger you decrease your batch size, do the same with LR. Batch size 8 -&gt; 4, LR 0.01 -&gt; 0.005 for example</p>",
      "rawMarkdown": "Also remember The Linear Scaling Rule: When the minibatch size is multiplied by k, multiply the learning rate by k. I guess when you changed model to bigger you decrease your batch size, do the same with LR. Batch size 8 -> 4, LR 0.01 -> 0.005 for example",
      "votes": null
    },
    {
      "id": "1707493",
      "postDate": "02/28/2022 14:06:36",
      "content": "<p>Thanks for your reply.   I found one notebook on kaggle  with different scheduler demos.</p>\n<p>However, I decide to give up replicating pytorch notebook,  since TF models are good enough.</p>",
      "rawMarkdown": "Thanks for your reply.   I found one notebook on kaggle  with different scheduler demos.\n\nHowever, I decide to give up replicating pytorch notebook,  since TF models are good enough.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1707160,
      "author_name": "kwentar",
      "author_url": "",
      "post_date": "02/28/2022 07:11:13",
      "content": "<p>I'd recommend to try some Find LR technic, like <a href=\"https://gist.github.com/karanchahal/dc0575bf21b976ea633dea8ceffaf9dc\" target=\"_blank\">this</a> one</p>",
      "votes": null,
      "replies": [
        {
          "id": 1707180,
          "author_name": "dragonzhang",
          "author_url": "",
          "post_date": "02/28/2022 07:35:50",
          "content": "<p>thank you!  I know that fast.ai has one.  I will have a look at the you post.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1707187,
          "author_name": "kwentar",
          "author_url": "",
          "post_date": "02/28/2022 07:40:49",
          "content": "<p>simple version is just change your scheduler to ExponentialLR scheduler and see by eyes :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1707192,
      "author_name": "kwentar",
      "author_url": "",
      "post_date": "02/28/2022 07:46:06",
      "content": "<p>Also remember The Linear Scaling Rule: When the minibatch size is multiplied by k, multiply the learning rate by k. I guess when you changed model to bigger you decrease your batch size, do the same with LR. Batch size 8 -&gt; 4, LR 0.01 -&gt; 0.005 for example</p>",
      "votes": null,
      "replies": [
        {
          "id": 1707493,
          "author_name": "dragonzhang",
          "author_url": "",
          "post_date": "02/28/2022 14:06:36",
          "content": "<p>Thanks for your reply.   I found one notebook on kaggle  with different scheduler demos.</p>\n<p>However, I decide to give up replicating pytorch notebook,  since TF models are good enough.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1707099": "I change the model to big one in my forked pytorch-arcface-gempooling-tpu-train notebook.\n\nhowever, the loss goes to plateau without decreasing.\n\nAny suggestion \n\n\n%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%\nLaunching a training on 8 TPU cores.\nValidation Loss: 17.33965392112732\nModel Saved\n\nValidation Loss: 16.07572067975998\nModel Saved\n\nValidation Loss: 15.178503835201264\nModel Saved\n\nValidation Loss: 14.623040127754212\nModel Saved\n\nValidation Loss: 14.062633389234543\nModel Saved\n\nValidation Loss: 14.027782171964645\nModel Saved\n\nValidation Loss: 13.651807701587677\nModel Saved\n\nValidation Loss: 13.665936523675919\nModel Saved\n\nValidation Loss: 13.792583948373794\nModel Saved\n\nValidation Loss: 13.604751998186112\nModel Saved\n\nValidation Loss: 13.649900329113006\nModel Saved\n\nValidation Loss: 13.736776286363602\nModel Saved\n\nValidation Loss: 13.664650583267212\nModel Saved\n\nValidation Loss: 13.632439285516739\nModel Saved\n\nValidation Loss: 13.798904740810395\nModel Saved\n\nValidation Loss: 13.680295634269715\nModel Saved\n\nValidation Loss: 13.6862082362175\nModel Saved\n\nValidation Loss: 13.732640331983566\nModel Saved\n\n%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%%",
    "1707160": "I'd recommend to try some Find LR technic, like [this](https://gist.github.com/karanchahal/dc0575bf21b976ea633dea8ceffaf9dc) one",
    "1707180": "thank you!  I know that fast.ai has one.  I will have a look at the you post.",
    "1707187": "simple version is just change your scheduler to ExponentialLR scheduler and see by eyes :)",
    "1707192": "Also remember The Linear Scaling Rule: When the minibatch size is multiplied by k, multiply the learning rate by k. I guess when you changed model to bigger you decrease your batch size, do the same with LR. Batch size 8 -> 4, LR 0.01 -> 0.005 for example",
    "1707493": "Thanks for your reply.   I found one notebook on kaggle  with different scheduler demos.\n\nHowever, I decide to give up replicating pytorch notebook,  since TF models are good enough."
  },
  "source": "meta"
}