{
  "id": 307653,
  "title": "resume training problem.",
  "url": "/competitions/happy-whale-and-dolphin/discussion/307653",
  "author_name": "",
  "post_date": "2022-02-15T05:53:42.454959300Z",
  "votes": 6,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I forked \"HappyWhale ArcFace Baseline (TPU)\", and use a bigger model to train on Colab.  however It is pretty slow.  so I tried to resume  training.</p>\n<p>however, since the new learning schedule is not same as before,  I got this warning message:</p>\n<p>WARNING:tensorflow:Callback method <code>on_train_batch_end</code> is slow compared to the batch time (batch time: 0.0185s vs <code>on_train_batch_end</code> time: 27.2315s). Check your callbacks.</p>\n<p>how to solve this problem?</p>",
  "messages": [
    {
      "id": "1690832",
      "postDate": "02/15/2022 05:53:42",
      "content": "<p>I forked \"HappyWhale ArcFace Baseline (TPU)\", and use a bigger model to train on Colab.  however It is pretty slow.  so I tried to resume  training.</p>\n<p>however, since the new learning schedule is not same as before,  I got this warning message:</p>\n<p>WARNING:tensorflow:Callback method <code>on_train_batch_end</code> is slow compared to the batch time (batch time: 0.0185s vs <code>on_train_batch_end</code> time: 27.2315s). Check your callbacks.</p>\n<p>how to solve this problem?</p>",
      "rawMarkdown": "I forked \"HappyWhale ArcFace Baseline (TPU)\", and use a bigger model to train on Colab.  however It is pretty slow.  so I tried to resume  training.\n\nhowever, since the new learning schedule is not same as before,  I got this warning message:\n\nWARNING:tensorflow:Callback method `on_train_batch_end` is slow compared to the batch time (batch time: 0.0185s vs `on_train_batch_end` time: 27.2315s). Check your callbacks.\n\nhow to solve this problem?",
      "votes": null
    },
    {
      "id": "1691486",
      "postDate": "02/15/2022 12:40:46",
      "content": "<p><a href=\"https://www.kaggle.com/dragonzhang\" target=\"_blank\">@dragonzhang</a> Have you tried increasing the batch size? I have been through a similar problem in the past.</p>",
      "rawMarkdown": "dragonzhang Have you tried increasing the batch size? I have been through a similar problem in the past.",
      "votes": null
    },
    {
      "id": "1691503",
      "postDate": "02/15/2022 12:48:08",
      "content": "<p>No,  I reduce the batch size. </p>",
      "rawMarkdown": "No,  I reduce the batch size.",
      "votes": null
    },
    {
      "id": "1691960",
      "postDate": "02/15/2022 17:49:17",
      "content": "<p>Probably Snapshot callback is slow</p>",
      "rawMarkdown": "Probably Snapshot callback is slow",
      "votes": null
    },
    {
      "id": "1693554",
      "postDate": "02/16/2022 18:25:18",
      "content": "<p>by experiments, I found that the lr_max  set by config.BATCH_SIZE is not suitable, because I have changed  it.</p>",
      "rawMarkdown": "by experiments, I found that the lr_max  set by config.BATCH_SIZE is not suitable, because I have changed  it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1691486,
      "author_name": "yerramvarun",
      "author_url": "",
      "post_date": "02/15/2022 12:40:46",
      "content": "<p><a href=\"https://www.kaggle.com/dragonzhang\" target=\"_blank\">@dragonzhang</a> Have you tried increasing the batch size? I have been through a similar problem in the past.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1691503,
          "author_name": "dragonzhang",
          "author_url": "",
          "post_date": "02/15/2022 12:48:08",
          "content": "<p>No,  I reduce the batch size. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1691960,
      "author_name": "atamazian",
      "author_url": "",
      "post_date": "02/15/2022 17:49:17",
      "content": "<p>Probably Snapshot callback is slow</p>",
      "votes": null,
      "replies": [
        {
          "id": 1693554,
          "author_name": "dragonzhang",
          "author_url": "",
          "post_date": "02/16/2022 18:25:18",
          "content": "<p>by experiments, I found that the lr_max  set by config.BATCH_SIZE is not suitable, because I have changed  it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1690832": "I forked \"HappyWhale ArcFace Baseline (TPU)\", and use a bigger model to train on Colab.  however It is pretty slow.  so I tried to resume  training.\n\nhowever, since the new learning schedule is not same as before,  I got this warning message:\n\nWARNING:tensorflow:Callback method `on_train_batch_end` is slow compared to the batch time (batch time: 0.0185s vs `on_train_batch_end` time: 27.2315s). Check your callbacks.\n\nhow to solve this problem?",
    "1691486": "dragonzhang Have you tried increasing the batch size? I have been through a similar problem in the past.",
    "1691503": "No,  I reduce the batch size.",
    "1691960": "Probably Snapshot callback is slow",
    "1693554": "by experiments, I found that the lr_max  set by config.BATCH_SIZE is not suitable, because I have changed  it."
  },
  "source": "meta"
}