{
  "id": 437993,
  "title": "Unstable Training ",
  "url": "/competitions/bengaliai-speech/discussion/437993",
  "author_name": "",
  "post_date": "2023-09-09T03:12:51.208351Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I am stuck at finetuning <a href=\"https://www.kaggle.com/code/nischaydnk/bengali-finetuning-baseline-wav2vec2-inference\" target=\"_blank\">Nischay Dhankhar's finetuned model</a>. I am currently using <a href=\"https://www.kaggle.com/code/heyytanay/pytorch-training-wav2vec2-for-bengaliai\" target=\"_blank\">Tanay Mehta's method</a>; however, by simply replacing \"model_name\" with the downloaded model would always result in Nan loss. </p>\n<p>I suspect that it's because some parameters were not initialized, which causes into instability. I will always receive the following message:</p>\n<blockquote>\n  <p>Some weights of Wav2Vec2ForCTC were not initialized from the model checkpoint at ../input/bengali_wav2vec2_finetuned and are newly initialized: ['wav2vec2.masked_spec_embed']<br>\n  You should probably TRAIN this model on a down-stream task to be able to use it for predictions and inference.</p>\n</blockquote>\n<p>How to resolve this issue? </p>",
  "messages": [
    {
      "id": "2430040",
      "postDate": "09/09/2023 03:12:51",
      "content": "<p>I am stuck at finetuning <a href=\"https://www.kaggle.com/code/nischaydnk/bengali-finetuning-baseline-wav2vec2-inference\" target=\"_blank\">Nischay Dhankhar's finetuned model</a>. I am currently using <a href=\"https://www.kaggle.com/code/heyytanay/pytorch-training-wav2vec2-for-bengaliai\" target=\"_blank\">Tanay Mehta's method</a>; however, by simply replacing \"model_name\" with the downloaded model would always result in Nan loss. </p>\n<p>I suspect that it's because some parameters were not initialized, which causes into instability. I will always receive the following message:</p>\n<blockquote>\n  <p>Some weights of Wav2Vec2ForCTC were not initialized from the model checkpoint at ../input/bengali_wav2vec2_finetuned and are newly initialized: ['wav2vec2.masked_spec_embed']<br>\n  You should probably TRAIN this model on a down-stream task to be able to use it for predictions and inference.</p>\n</blockquote>\n<p>How to resolve this issue? </p>",
      "rawMarkdown": "I am stuck at finetuning [Nischay Dhankhar's finetuned model](https://www.kaggle.com/code/nischaydnk/bengali-finetuning-baseline-wav2vec2-inference). I am currently using [Tanay Mehta's method](https://www.kaggle.com/code/heyytanay/pytorch-training-wav2vec2-for-bengaliai); however, by simply replacing \"model_name\" with the downloaded model would always result in Nan loss. \n\nI suspect that it's because some parameters were not initialized, which causes into instability. I will always receive the following message:\n\n>Some weights of Wav2Vec2ForCTC were not initialized from the model checkpoint at ../input/bengali_wav2vec2_finetuned and are newly initialized: ['wav2vec2.masked_spec_embed']\nYou should probably TRAIN this model on a down-stream task to be able to use it for predictions and inference.\n\nHow to resolve this issue?",
      "votes": null
    },
    {
      "id": "2431262",
      "postDate": "09/10/2023 01:02:37",
      "content": "<p>Did you try the solution of <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/436508\" target=\"_blank\">this discussion</a>?</p>\n<p>You may add the args in loading CTC like:</p>\n<pre><code>Wav2Vec2ForCTC.from_pretrained(\n            path,\n            ctc_zero_infinity=,\n            diversity_loss_weight=,\n            **your_kwargs\n)\n</code></pre>\n<p>Hope this works for you.</p>",
      "rawMarkdown": "Did you try the solution of [this discussion](https://www.kaggle.com/competitions/bengaliai-speech/discussion/436508)?\n\nYou may add the args in loading CTC like:\n```python\nWav2Vec2ForCTC.from_pretrained(\n            path,\n            ctc_zero_infinity=True,\n            diversity_loss_weight=100,\n            **your_kwargs\n)\n```\n\nHope this works for you.",
      "votes": null
    },
    {
      "id": "2431269",
      "postDate": "09/10/2023 01:43:14",
      "content": "<p>A lot of thanks, I don't see Nan or Inf anymore🫡. It also gives rises to a new issue: out of cuda memory after a minute or two.   needs further inspection😅</p>",
      "rawMarkdown": "A lot of thanks, I don't see Nan or Inf anymore🫡. It also gives rises to a new issue: out of cuda memory after a minute or two.   needs further inspection😅",
      "votes": null
    },
    {
      "id": "2431636",
      "postDate": "09/10/2023 09:15:49",
      "content": "<p>out of cuda memory is not due to above params.</p>\n<p>Kindly check the batch size parameters. Opt for fp16 training</p>",
      "rawMarkdown": "out of cuda memory is not due to above params.\n\nKindly check the batch size parameters. Opt for fp16 training",
      "votes": null
    },
    {
      "id": "2432140",
      "postDate": "09/10/2023 15:56:05",
      "content": "<p>Thanks. I've used mixed precision and shrunk the batch size. Btw, why does the speed get a lot faster when  bs = 4 in comparison to bs = 8? I finished fine-tuning in just two hours with bs = 4</p>",
      "rawMarkdown": "Thanks. I've used mixed precision and shrunk the batch size. Btw, why does the speed get a lot faster when  bs = 4 in comparison to bs = 8? I finished fine-tuning in just two hours with bs = 4",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2431262,
      "author_name": "tsumli",
      "author_url": "",
      "post_date": "09/10/2023 01:02:37",
      "content": "<p>Did you try the solution of <a href=\"https://www.kaggle.com/competitions/bengaliai-speech/discussion/436508\" target=\"_blank\">this discussion</a>?</p>\n<p>You may add the args in loading CTC like:</p>\n<pre><code>Wav2Vec2ForCTC.from_pretrained(\n            path,\n            ctc_zero_infinity=,\n            diversity_loss_weight=,\n            **your_kwargs\n)\n</code></pre>\n<p>Hope this works for you.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2431269,
          "author_name": "renyiwei",
          "author_url": "",
          "post_date": "09/10/2023 01:43:14",
          "content": "<p>A lot of thanks, I don't see Nan or Inf anymore🫡. It also gives rises to a new issue: out of cuda memory after a minute or two.   needs further inspection😅</p>",
          "votes": null,
          "replies": [
            {
              "id": 2431636,
              "author_name": "dhakshiin1601",
              "author_url": "",
              "post_date": "09/10/2023 09:15:49",
              "content": "<p>out of cuda memory is not due to above params.</p>\n<p>Kindly check the batch size parameters. Opt for fp16 training</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2432140,
                  "author_name": "renyiwei",
                  "author_url": "",
                  "post_date": "09/10/2023 15:56:05",
                  "content": "<p>Thanks. I've used mixed precision and shrunk the batch size. Btw, why does the speed get a lot faster when  bs = 4 in comparison to bs = 8? I finished fine-tuning in just two hours with bs = 4</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2430040": "I am stuck at finetuning [Nischay Dhankhar's finetuned model](https://www.kaggle.com/code/nischaydnk/bengali-finetuning-baseline-wav2vec2-inference). I am currently using [Tanay Mehta's method](https://www.kaggle.com/code/heyytanay/pytorch-training-wav2vec2-for-bengaliai); however, by simply replacing \"model_name\" with the downloaded model would always result in Nan loss. \n\nI suspect that it's because some parameters were not initialized, which causes into instability. I will always receive the following message:\n\n>Some weights of Wav2Vec2ForCTC were not initialized from the model checkpoint at ../input/bengali_wav2vec2_finetuned and are newly initialized: ['wav2vec2.masked_spec_embed']\nYou should probably TRAIN this model on a down-stream task to be able to use it for predictions and inference.\n\nHow to resolve this issue?",
    "2431262": "Did you try the solution of [this discussion](https://www.kaggle.com/competitions/bengaliai-speech/discussion/436508)?\n\nYou may add the args in loading CTC like:\n```python\nWav2Vec2ForCTC.from_pretrained(\n            path,\n            ctc_zero_infinity=True,\n            diversity_loss_weight=100,\n            **your_kwargs\n)\n```\n\nHope this works for you.",
    "2431269": "A lot of thanks, I don't see Nan or Inf anymore🫡. It also gives rises to a new issue: out of cuda memory after a minute or two.   needs further inspection😅",
    "2431636": "out of cuda memory is not due to above params.\n\nKindly check the batch size parameters. Opt for fp16 training",
    "2432140": "Thanks. I've used mixed precision and shrunk the batch size. Btw, why does the speed get a lot faster when  bs = 4 in comparison to bs = 8? I finished fine-tuning in just two hours with bs = 4"
  },
  "source": "meta"
}