{
  "id": 444850,
  "title": "Training Bug 🤷",
  "url": "/competitions/bengaliai-speech/discussion/444850",
  "author_name": "",
  "post_date": "2023-10-03T23:58:25.782792600Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>When I was training the model last night, I encounter a bug never seen before:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13889710%2Fcef7b4f07c348c98ee7a6164f5870c2a%2FScreenshot%202023-10-04%20at%207.28.22%20AM.png?generation=1696375826209112&amp;alt=media\" alt=\"\"></p>\n<p>There are other weirds signs as well in wandb:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13889710%2F17ec39e19c5c99ff0f8a218476140786%2FScreenshot%202023-10-04%20at%207.31.55%20AM.png?generation=1696375966626611&amp;alt=media\" alt=\"\"></p>\n<p>The global step and epoch decreases which normally only increases. </p>\n<p>My assumption: <br>\nThe <code>padding</code> is set to <code>True</code> in the Data Collator. The sequence length of each batch  thus varies, resulting in different cuda memory consumption. Considering that I have added several length audios, this may be the source of the problem.</p>\n<p>Few problems I can't explain:<br>\n1) If there is  the batch-size reduced to zero, it should be caused by setting <code>auto_find_batch_size=True</code>. In that case, how can few thousands steps be taken before crashing.<br>\n2) why does the global_step/epoch decrease?</p>\n<p>Since truncating my audio to a suitable size is lots of work, I plea for more insight before attempting to debug.🫡</p>",
  "messages": [
    {
      "id": "2466542",
      "postDate": "10/03/2023 23:58:25",
      "content": "<p>When I was training the model last night, I encounter a bug never seen before:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13889710%2Fcef7b4f07c348c98ee7a6164f5870c2a%2FScreenshot%202023-10-04%20at%207.28.22%20AM.png?generation=1696375826209112&amp;alt=media\" alt=\"\"></p>\n<p>There are other weirds signs as well in wandb:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13889710%2F17ec39e19c5c99ff0f8a218476140786%2FScreenshot%202023-10-04%20at%207.31.55%20AM.png?generation=1696375966626611&amp;alt=media\" alt=\"\"></p>\n<p>The global step and epoch decreases which normally only increases. </p>\n<p>My assumption: <br>\nThe <code>padding</code> is set to <code>True</code> in the Data Collator. The sequence length of each batch  thus varies, resulting in different cuda memory consumption. Considering that I have added several length audios, this may be the source of the problem.</p>\n<p>Few problems I can't explain:<br>\n1) If there is  the batch-size reduced to zero, it should be caused by setting <code>auto_find_batch_size=True</code>. In that case, how can few thousands steps be taken before crashing.<br>\n2) why does the global_step/epoch decrease?</p>\n<p>Since truncating my audio to a suitable size is lots of work, I plea for more insight before attempting to debug.🫡</p>",
      "rawMarkdown": "When I was training the model last night, I encounter a bug never seen before:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13889710%2Fcef7b4f07c348c98ee7a6164f5870c2a%2FScreenshot%202023-10-04%20at%207.28.22%20AM.png?generation=1696375826209112&alt=media)\n\nThere are other weirds signs as well in wandb:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13889710%2F17ec39e19c5c99ff0f8a218476140786%2FScreenshot%202023-10-04%20at%207.31.55%20AM.png?generation=1696375966626611&alt=media)\n\nThe global step and epoch decreases which normally only increases. \n\nMy assumption: \nThe `padding` is set to `True` in the Data Collator. The sequence length of each batch  thus varies, resulting in different cuda memory consumption. Considering that I have added several length audios, this may be the source of the problem.\n\nFew problems I can't explain:\n1) If there is  the batch-size reduced to zero, it should be caused by setting `auto_find_batch_size=True`. In that case, how can few thousands steps be taken before crashing.\n2) why does the global_step/epoch decrease?\n\nSince truncating my audio to a suitable size is lots of work, I plea for more insight before attempting to debug.🫡",
      "votes": null
    },
    {
      "id": "2466543",
      "postDate": "10/04/2023 00:00:17",
      "content": "<p>I'm really kind of new in this field, forgive if I am wrong in some of my understanding😵‍💫</p>",
      "rawMarkdown": "I'm really kind of new in this field, forgive if I am wrong in some of my understanding😵‍💫",
      "votes": null
    },
    {
      "id": "2466639",
      "postDate": "10/04/2023 02:26:06",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13889710%2Fdd2273251e52eedbf1b99d1e263fca78%2FScreenshot%202023-10-04%20at%2010.21.53%20AM.png?generation=1696386227439904&amp;alt=media\" alt=\"\"></p>\n<p>This is not even a singular case. In a previous <strong>successful</strong> session, there is also this weird decrease in the global step and percentage of epoch ran. It never happens again afterwards</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13889710%2Fdd2273251e52eedbf1b99d1e263fca78%2FScreenshot%202023-10-04%20at%2010.21.53%20AM.png?generation=1696386227439904&alt=media)\n\nThis is not even a singular case. In a previous **successful** session, there is also this weird decrease in the global step and percentage of epoch ran. It never happens again afterwards",
      "votes": null
    },
    {
      "id": "2467488",
      "postDate": "10/04/2023 15:38:34",
      "content": "<p>Since the error is from accelerater pack, there is one function in train that needs accelerater: </p>\n<p>auto_find_batch_size</p>\n<p>That means your GPU was out of memory on every nonzero size.<br>\nmaybe you can set fp16=True, or disable adam scheduler.</p>",
      "rawMarkdown": "Since the error is from accelerater pack, there is one function in train that needs accelerater: \n\nauto_find_batch_size\n\nThat means your GPU was out of memory on every nonzero size.\nmaybe you can set fp16=True, or disable adam scheduler.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2466543,
      "author_name": "renyiwei",
      "author_url": "",
      "post_date": "10/04/2023 00:00:17",
      "content": "<p>I'm really kind of new in this field, forgive if I am wrong in some of my understanding😵‍💫</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2466639,
      "author_name": "renyiwei",
      "author_url": "",
      "post_date": "10/04/2023 02:26:06",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13889710%2Fdd2273251e52eedbf1b99d1e263fca78%2FScreenshot%202023-10-04%20at%2010.21.53%20AM.png?generation=1696386227439904&amp;alt=media\" alt=\"\"></p>\n<p>This is not even a singular case. In a previous <strong>successful</strong> session, there is also this weird decrease in the global step and percentage of epoch ran. It never happens again afterwards</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2467488,
      "author_name": "hongori",
      "author_url": "",
      "post_date": "10/04/2023 15:38:34",
      "content": "<p>Since the error is from accelerater pack, there is one function in train that needs accelerater: </p>\n<p>auto_find_batch_size</p>\n<p>That means your GPU was out of memory on every nonzero size.<br>\nmaybe you can set fp16=True, or disable adam scheduler.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2466542": "When I was training the model last night, I encounter a bug never seen before:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13889710%2Fcef7b4f07c348c98ee7a6164f5870c2a%2FScreenshot%202023-10-04%20at%207.28.22%20AM.png?generation=1696375826209112&alt=media)\n\nThere are other weirds signs as well in wandb:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13889710%2F17ec39e19c5c99ff0f8a218476140786%2FScreenshot%202023-10-04%20at%207.31.55%20AM.png?generation=1696375966626611&alt=media)\n\nThe global step and epoch decreases which normally only increases. \n\nMy assumption: \nThe `padding` is set to `True` in the Data Collator. The sequence length of each batch  thus varies, resulting in different cuda memory consumption. Considering that I have added several length audios, this may be the source of the problem.\n\nFew problems I can't explain:\n1) If there is  the batch-size reduced to zero, it should be caused by setting `auto_find_batch_size=True`. In that case, how can few thousands steps be taken before crashing.\n2) why does the global_step/epoch decrease?\n\nSince truncating my audio to a suitable size is lots of work, I plea for more insight before attempting to debug.🫡",
    "2466543": "I'm really kind of new in this field, forgive if I am wrong in some of my understanding😵‍💫",
    "2466639": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13889710%2Fdd2273251e52eedbf1b99d1e263fca78%2FScreenshot%202023-10-04%20at%2010.21.53%20AM.png?generation=1696386227439904&alt=media)\n\nThis is not even a singular case. In a previous **successful** session, there is also this weird decrease in the global step and percentage of epoch ran. It never happens again afterwards",
    "2467488": "Since the error is from accelerater pack, there is one function in train that needs accelerater: \n\nauto_find_batch_size\n\nThat means your GPU was out of memory on every nonzero size.\nmaybe you can set fp16=True, or disable adam scheduler."
  },
  "source": "meta"
}