{
  "id": 437296,
  "title": "Why I am getting \"Notebook out of memory\" error upon submission",
  "url": "/competitions/bengaliai-speech/discussion/437296",
  "author_name": "",
  "post_date": "2023-09-06T08:00:50.844506300Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Over the past 7-10 days I have been facing a peculiar problem. I am trying to make my submission to the competition, but am getting \"Notebook out of memory\" error again and again as you can see in the picture. </p>\n<p>I have tried a variety of things to get rid of the problem, like keeping bare minimum package installations &amp; imports, deleting variables at various junctures and using garbage collection using gc.collect() and even reducing my training dataset to just 50 records. Nothing is working. I am running a simple, straightforward model using \"facebook-wav2vec2largexlsr53\". </p>\n<p>Apart from this model as a dataset, I am installing \"Jiwer package with Rapidfuzz\" (see the bottom pic). When I SAVEALL or just RUN ALL the notebook, the kernel runs perfectly even with 10000 records and I get the result as expected, but the problem is only for submission (on pressing the SUBMIT button either from output or from the right panel of notebook). By the way I was able to make a submission around 2 weeks back. </p>\n<p>What am I doing wrong? How do get rid of this problem and make a submission? This has already caused enormous wastage of time and even Kaggle's valuable platform resources. If I am not doing anything wrong and the problem is with Kaggle platform, how do I take it to their attention? Appreciate inputs.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3761209%2F1c3858df5c1944df73d71468d09b38cf%2Fkaggle.jpg?generation=1693987462647153&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3761209%2F11b24c0886218c92d822ac9488369009%2Fkaggle.jpg?generation=1693987206767909&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3761209%2F54543f32bd23a02144aacc1c232691b1%2Finputs.jpg?generation=1693987222446951&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "2425837",
      "postDate": "09/06/2023 08:00:50",
      "content": "<p>Over the past 7-10 days I have been facing a peculiar problem. I am trying to make my submission to the competition, but am getting \"Notebook out of memory\" error again and again as you can see in the picture. </p>\n<p>I have tried a variety of things to get rid of the problem, like keeping bare minimum package installations &amp; imports, deleting variables at various junctures and using garbage collection using gc.collect() and even reducing my training dataset to just 50 records. Nothing is working. I am running a simple, straightforward model using \"facebook-wav2vec2largexlsr53\". </p>\n<p>Apart from this model as a dataset, I am installing \"Jiwer package with Rapidfuzz\" (see the bottom pic). When I SAVEALL or just RUN ALL the notebook, the kernel runs perfectly even with 10000 records and I get the result as expected, but the problem is only for submission (on pressing the SUBMIT button either from output or from the right panel of notebook). By the way I was able to make a submission around 2 weeks back. </p>\n<p>What am I doing wrong? How do get rid of this problem and make a submission? This has already caused enormous wastage of time and even Kaggle's valuable platform resources. If I am not doing anything wrong and the problem is with Kaggle platform, how do I take it to their attention? Appreciate inputs.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3761209%2F1c3858df5c1944df73d71468d09b38cf%2Fkaggle.jpg?generation=1693987462647153&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3761209%2F11b24c0886218c92d822ac9488369009%2Fkaggle.jpg?generation=1693987206767909&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3761209%2F54543f32bd23a02144aacc1c232691b1%2Finputs.jpg?generation=1693987222446951&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Over the past 7-10 days I have been facing a peculiar problem. I am trying to make my submission to the competition, but am getting \"Notebook out of memory\" error again and again as you can see in the picture. \n\nI have tried a variety of things to get rid of the problem, like keeping bare minimum package installations & imports, deleting variables at various junctures and using garbage collection using gc.collect() and even reducing my training dataset to just 50 records. Nothing is working. I am running a simple, straightforward model using \"facebook-wav2vec2largexlsr53\". \n\nApart from this model as a dataset, I am installing \"Jiwer package with Rapidfuzz\" (see the bottom pic). When I SAVEALL or just RUN ALL the notebook, the kernel runs perfectly even with 10000 records and I get the result as expected, but the problem is only for submission (on pressing the SUBMIT button either from output or from the right panel of notebook). By the way I was able to make a submission around 2 weeks back. \n\nWhat am I doing wrong? How do get rid of this problem and make a submission? This has already caused enormous wastage of time and even Kaggle's valuable platform resources. If I am not doing anything wrong and the problem is with Kaggle platform, how do I take it to their attention? Appreciate inputs.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3761209%2F1c3858df5c1944df73d71468d09b38cf%2Fkaggle.jpg?generation=1693987462647153&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3761209%2F11b24c0886218c92d822ac9488369009%2Fkaggle.jpg?generation=1693987206767909&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3761209%2F54543f32bd23a02144aacc1c232691b1%2Finputs.jpg?generation=1693987222446951&alt=media)",
      "votes": null
    },
    {
      "id": "2426259",
      "postDate": "09/06/2023 14:04:23",
      "content": "<p>By the way the same problem pops up even when I try the base version, 'facebook-wav2vec2-base'</p>",
      "rawMarkdown": "By the way the same problem pops up even when I try the base version, 'facebook-wav2vec2-base'",
      "votes": null
    },
    {
      "id": "2426287",
      "postDate": "09/06/2023 14:29:35",
      "content": "<p>Are you using num_worker argument if yes then removed or set to 1?<br>\nReduce the batch size..</p>",
      "rawMarkdown": "Are you using num_worker argument if yes then removed or set to 1?\nReduce the batch size..",
      "votes": null
    },
    {
      "id": "2427424",
      "postDate": "09/07/2023 07:57:25",
      "content": "<p><a href=\"https://www.kaggle.com/rajgothi\" target=\"_blank\">@rajgothi</a> Thnx for reverting. Batch size is 16. Isn't that small enough? Do you mean \"num_proc\"? I have that argument in various dataset processing steps. I do not have \"num_worker\". I did try out setting \"num_proc\" to 1, but problem continues.</p>",
      "rawMarkdown": "rajgothi Thnx for reverting. Batch size is 16. Isn't that small enough? Do you mean \"num_proc\"? I have that argument in various dataset processing steps. I do not have \"num_worker\". I did try out setting \"num_proc\" to 1, but problem continues.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2426259,
      "author_name": "valmetisrinivas",
      "author_url": "",
      "post_date": "09/06/2023 14:04:23",
      "content": "<p>By the way the same problem pops up even when I try the base version, 'facebook-wav2vec2-base'</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2426287,
      "author_name": "rajgothi",
      "author_url": "",
      "post_date": "09/06/2023 14:29:35",
      "content": "<p>Are you using num_worker argument if yes then removed or set to 1?<br>\nReduce the batch size..</p>",
      "votes": null,
      "replies": [
        {
          "id": 2427424,
          "author_name": "valmetisrinivas",
          "author_url": "",
          "post_date": "09/07/2023 07:57:25",
          "content": "<p><a href=\"https://www.kaggle.com/rajgothi\" target=\"_blank\">@rajgothi</a> Thnx for reverting. Batch size is 16. Isn't that small enough? Do you mean \"num_proc\"? I have that argument in various dataset processing steps. I do not have \"num_worker\". I did try out setting \"num_proc\" to 1, but problem continues.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2425837": "Over the past 7-10 days I have been facing a peculiar problem. I am trying to make my submission to the competition, but am getting \"Notebook out of memory\" error again and again as you can see in the picture. \n\nI have tried a variety of things to get rid of the problem, like keeping bare minimum package installations & imports, deleting variables at various junctures and using garbage collection using gc.collect() and even reducing my training dataset to just 50 records. Nothing is working. I am running a simple, straightforward model using \"facebook-wav2vec2largexlsr53\". \n\nApart from this model as a dataset, I am installing \"Jiwer package with Rapidfuzz\" (see the bottom pic). When I SAVEALL or just RUN ALL the notebook, the kernel runs perfectly even with 10000 records and I get the result as expected, but the problem is only for submission (on pressing the SUBMIT button either from output or from the right panel of notebook). By the way I was able to make a submission around 2 weeks back. \n\nWhat am I doing wrong? How do get rid of this problem and make a submission? This has already caused enormous wastage of time and even Kaggle's valuable platform resources. If I am not doing anything wrong and the problem is with Kaggle platform, how do I take it to their attention? Appreciate inputs.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3761209%2F1c3858df5c1944df73d71468d09b38cf%2Fkaggle.jpg?generation=1693987462647153&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3761209%2F11b24c0886218c92d822ac9488369009%2Fkaggle.jpg?generation=1693987206767909&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3761209%2F54543f32bd23a02144aacc1c232691b1%2Finputs.jpg?generation=1693987222446951&alt=media)",
    "2426259": "By the way the same problem pops up even when I try the base version, 'facebook-wav2vec2-base'",
    "2426287": "Are you using num_worker argument if yes then removed or set to 1?\nReduce the batch size..",
    "2427424": "rajgothi Thnx for reverting. Batch size is 16. Isn't that small enough? Do you mean \"num_proc\"? I have that argument in various dataset processing steps. I do not have \"num_worker\". I did try out setting \"num_proc\" to 1, but problem continues."
  },
  "source": "meta"
}