{
  "id": 398879,
  "title": "Submission out of memory error",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/398879",
  "author_name": "GianPaolob",
  "post_date": "2023-04-01T07:57:04.333000",
  "votes": 6,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Hi all..<br>\nI am getting an out of memory when submitting.<br>\nThe notebook successfully commit and properly generates the submission file , however after that I can get any score and an out of memory error is raised.<br>\nHere are the screenshot of submission file and raised error:<br>\n&lt;img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1332227%2Fa11755b351d244685d418f46701c28aa%2F31.03.2023_17.17.06_REC.png?generation=1680335803113798&amp;alt=media\" alt=\"!<a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1332227%2F1f05b4dfc6566ed19726d79904009a3c%2F31.03.2023_17.19.36_REC.png?generation=1680335792804837&amp;alt=media\" target=\"_blank\">\" /&gt;</a></p>\n<p>Has someone else experienced such kind of behaviour ?</p>",
  "messages": [
    {
      "id": 2205062,
      "postDate": "2023-04-01T07:57:04.333Z",
      "content": "<p>Hi all..<br>\nI am getting an out of memory when submitting.<br>\nThe notebook successfully commit and properly generates the submission file , however after that I can get any score and an out of memory error is raised.<br>\nHere are the screenshot of submission file and raised error:<br>\n&lt;img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1332227%2Fa11755b351d244685d418f46701c28aa%2F31.03.2023_17.17.06_REC.png?generation=1680335803113798&amp;alt=media\" alt=\"!<a href=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1332227%2F1f05b4dfc6566ed19726d79904009a3c%2F31.03.2023_17.19.36_REC.png?generation=1680335792804837&amp;alt=media\" target=\"_blank\">\" /&gt;</a></p>\n<p>Has someone else experienced such kind of behaviour ?</p>",
      "rawMarkdown": "Hi all..\nI am getting an out of memory when submitting.\nThe notebook successfully commit and properly generates the submission file , however after that I can get any score and an out of memory error is raised.\nHere are the screenshot of submission file and raised error:\n![![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1332227%2Fa11755b351d244685d418f46701c28aa%2F31.03.2023_17.17.06_REC.png?generation=1680335803113798&alt=media)](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1332227%2F1f05b4dfc6566ed19726d79904009a3c%2F31.03.2023_17.19.36_REC.png?generation=1680335792804837&alt=media)\n\nHas someone else experienced such kind of behaviour ?",
      "votes": 6
    },
    {
      "id": 2205317,
      "postDate": "2023-04-01T12:30:01.660Z",
      "content": "<p>Are you loading all the data in a big data frame and then do predictions? You can also do chunk predictions, in other words read a file, preprocess and extract all the features you need and then do predictions and store this prediction in memory, then read the next file and so on, this way you are predicting small dataframes. This way you can use more than 300 features </p>",
      "rawMarkdown": "Are you loading all the data in a big data frame and then do predictions? You can also do chunk predictions, in other words read a file, preprocess and extract all the features you need and then do predictions and store this prediction in memory, then read the next file and so on, this way you are predicting small dataframes. This way you can use more than 300 features ",
      "votes": 1
    },
    {
      "id": 2205173,
      "postDate": "2023-04-01T10:13:39.037Z",
      "content": "<p>Perhaps you could insert some delete variables and garbage collection before doing the final predictions.</p>",
      "rawMarkdown": "Perhaps you could insert some delete variables and garbage collection before doing the final predictions.",
      "votes": 1
    },
    {
      "id": 2273929,
      "postDate": "2023-05-25T14:07:11.857Z",
      "content": "<p>I ran into a similar problem.<br>\nIs there any memory savings by just reducing the amount of data you put on the model?<br>\nFor example, reducing the input data from 2 million to 1 million.</p>",
      "rawMarkdown": "I ran into a similar problem.\nIs there any memory savings by just reducing the amount of data you put on the model?\nFor example, reducing the input data from 2 million to 1 million."
    },
    {
      "id": 2229214,
      "postDate": "2023-04-21T06:50:28.310Z",
      "content": "<p>You can use numpy instead of pandas and you'll get a huge memory improvement</p>",
      "rawMarkdown": "You can use numpy instead of pandas and you'll get a huge memory improvement"
    },
    {
      "id": 2210693,
      "postDate": "2023-04-05T15:04:14.737Z",
      "content": "<p>For one of my notebooks (GPU), I was facing memory issues during submission. Turns out the problem was that my prediction step created a single large Numpy array which was much larger with the hidden data added to test during submission.</p>\n<p>A good test to do for finding the error is to see if trying to do the same prediction step for all train folders (defog, tdscfog, notype) leads to a \"out of memory\" error.</p>",
      "rawMarkdown": "For one of my notebooks (GPU), I was facing memory issues during submission. Turns out the problem was that my prediction step created a single large Numpy array which was much larger with the hidden data added to test during submission.\n\nA good test to do for finding the error is to see if trying to do the same prediction step for all train folders (defog, tdscfog, notype) leads to a \"out of memory\" error."
    },
    {
      "id": 2207203,
      "postDate": "2023-04-03T07:58:25.783Z",
      "content": "<p>Yes, I faced the same issue, notebook successfully committed, but during actual submission it ran out of memory.</p>\n<p>It could be due to the larger size of the test set, as mentioned in other comments.</p>\n<p>Also, Kaggle recently upgraded the notebook RAM to 30 GB (<a href=\"https://www.kaggle.com/product-feedback/361104)\" target=\"_blank\">https://www.kaggle.com/product-feedback/361104)</a>, but not sure if that also holds for the submission notebook.</p>",
      "rawMarkdown": "Yes, I faced the same issue, notebook successfully committed, but during actual submission it ran out of memory.\n\nIt could be due to the larger size of the test set, as mentioned in other comments.\n\nAlso, Kaggle recently upgraded the notebook RAM to 30 GB (https://www.kaggle.com/product-feedback/361104), but not sure if that also holds for the submission notebook."
    },
    {
      "id": 2205287,
      "postDate": "2023-04-01T12:06:51.230Z",
      "content": "<p>Hey! <br>\nPlease check if your data types are efficient and that you don't use too many features.<br>\nAlso - what model are you using? <br>\nMabey tweaking the hyperparameters for memory-efficient options can help.</p>\n<p>Good luck!</p>",
      "rawMarkdown": "Hey! \nPlease check if your data types are efficient and that you don't use too many features.\nAlso - what model are you using? \nMabey tweaking the hyperparameters for memory-efficient options can help.\n\nGood luck!"
    },
    {
      "id": 2205063,
      "postDate": "2023-04-01T07:57:45.050Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1332227%2Fbcd93bd8c0ce4f79d3f82530ec3ebc5a%2F31.03.2023_17.19.36_REC.png?generation=1680335849808774&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1332227%2Fbcd93bd8c0ce4f79d3f82530ec3ebc5a%2F31.03.2023_17.19.36_REC.png?generation=1680335849808774&alt=media)\n"
    },
    {
      "id": 2205105,
      "postDate": "2023-04-01T08:44:49.847Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 2205137,
          "postDate": "2023-04-01T09:29:49.587Z",
          "content": "<p>Hi ..<br>\nThe submission file contains 286370 lines , plus the header…</p>",
          "rawMarkdown": "Hi ..\nThe submission file contains 286370 lines , plus the header...",
          "replies": [
            {
              "id": 2205146,
              "postDate": "2023-04-01T09:46:00.023Z",
              "rawMarkdown": "",
              "votes": 2,
              "isDeleted": true
            },
            {
              "id": 2205172,
              "postDate": "2023-04-01T10:13:31.383Z",
              "content": "<p>Many thanks for you help .. I was thinking the same thing .. I'll for memory leaks </p>",
              "rawMarkdown": "Many thanks for you help .. I was thinking the same thing .. I'll for memory leaks "
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2205317,
      "author_name": "Martin Kovacevic Buvinic",
      "author_url": "",
      "post_date": "2023-04-01T12:30:01.660000",
      "content": "<p>Are you loading all the data in a big data frame and then do predictions? You can also do chunk predictions, in other words read a file, preprocess and extract all the features you need and then do predictions and store this prediction in memory, then read the next file and so on, this way you are predicting small dataframes. This way you can use more than 300 features </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2205173,
      "author_name": "PeterSorensen",
      "author_url": "",
      "post_date": "2023-04-01T10:13:39.037000",
      "content": "<p>Perhaps you could insert some delete variables and garbage collection before doing the final predictions.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2273929,
      "author_name": "CHRONO",
      "author_url": "",
      "post_date": "2023-05-25T14:07:11.857000",
      "content": "<p>I ran into a similar problem.<br>\nIs there any memory savings by just reducing the amount of data you put on the model?<br>\nFor example, reducing the input data from 2 million to 1 million.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2229214,
      "author_name": "Quim Quadrada",
      "author_url": "",
      "post_date": "2023-04-21T06:50:28.310000",
      "content": "<p>You can use numpy instead of pandas and you'll get a huge memory improvement</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2210693,
      "author_name": "coderRKJ",
      "author_url": "",
      "post_date": "2023-04-05T15:04:14.737000",
      "content": "<p>For one of my notebooks (GPU), I was facing memory issues during submission. Turns out the problem was that my prediction step created a single large Numpy array which was much larger with the hidden data added to test during submission.</p>\n<p>A good test to do for finding the error is to see if trying to do the same prediction step for all train folders (defog, tdscfog, notype) leads to a \"out of memory\" error.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2207203,
      "author_name": "William Wu Chengyuan",
      "author_url": "",
      "post_date": "2023-04-03T07:58:25.783000",
      "content": "<p>Yes, I faced the same issue, notebook successfully committed, but during actual submission it ran out of memory.</p>\n<p>It could be due to the larger size of the test set, as mentioned in other comments.</p>\n<p>Also, Kaggle recently upgraded the notebook RAM to 30 GB (<a href=\"https://www.kaggle.com/product-feedback/361104)\" target=\"_blank\">https://www.kaggle.com/product-feedback/361104)</a>, but not sure if that also holds for the submission notebook.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2205287,
      "author_name": "Aviv Levi",
      "author_url": "",
      "post_date": "2023-04-01T12:06:51.230000",
      "content": "<p>Hey! <br>\nPlease check if your data types are efficient and that you don't use too many features.<br>\nAlso - what model are you using? <br>\nMabey tweaking the hyperparameters for memory-efficient options can help.</p>\n<p>Good luck!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2205063,
      "author_name": "GianPaolob",
      "author_url": "",
      "post_date": "2023-04-01T07:57:45.050000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1332227%2Fbcd93bd8c0ce4f79d3f82530ec3ebc5a%2F31.03.2023_17.19.36_REC.png?generation=1680335849808774&amp;alt=media\" alt=\"\"></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2205105,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-04-01T08:44:49.847000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2205137,
          "author_name": "GianPaolob",
          "author_url": "",
          "post_date": "2023-04-01T09:29:49.587000",
          "content": "<p>Hi ..<br>\nThe submission file contains 286370 lines , plus the header…</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2205146,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-04-01T09:46:00.023000",
              "content": "",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2205172,
              "author_name": "GianPaolob",
              "author_url": "",
              "post_date": "2023-04-01T10:13:31.383000",
              "content": "<p>Many thanks for you help .. I was thinking the same thing .. I'll for memory leaks </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2205062": "Hi all..\nI am getting an out of memory when submitting.\nThe notebook successfully commit and properly generates the submission file , however after that I can get any score and an out of memory error is raised.\nHere are the screenshot of submission file and raised error:\n![![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1332227%2Fa11755b351d244685d418f46701c28aa%2F31.03.2023_17.17.06_REC.png?generation=1680335803113798&alt=media)](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1332227%2F1f05b4dfc6566ed19726d79904009a3c%2F31.03.2023_17.19.36_REC.png?generation=1680335792804837&alt=media)\n\nHas someone else experienced such kind of behaviour ?",
    "2205317": "Are you loading all the data in a big data frame and then do predictions? You can also do chunk predictions, in other words read a file, preprocess and extract all the features you need and then do predictions and store this prediction in memory, then read the next file and so on, this way you are predicting small dataframes. This way you can use more than 300 features ",
    "2205173": "Perhaps you could insert some delete variables and garbage collection before doing the final predictions.",
    "2273929": "I ran into a similar problem.\nIs there any memory savings by just reducing the amount of data you put on the model?\nFor example, reducing the input data from 2 million to 1 million.",
    "2229214": "You can use numpy instead of pandas and you'll get a huge memory improvement",
    "2210693": "For one of my notebooks (GPU), I was facing memory issues during submission. Turns out the problem was that my prediction step created a single large Numpy array which was much larger with the hidden data added to test during submission.\n\nA good test to do for finding the error is to see if trying to do the same prediction step for all train folders (defog, tdscfog, notype) leads to a \"out of memory\" error.",
    "2207203": "Yes, I faced the same issue, notebook successfully committed, but during actual submission it ran out of memory.\n\nIt could be due to the larger size of the test set, as mentioned in other comments.\n\nAlso, Kaggle recently upgraded the notebook RAM to 30 GB (https://www.kaggle.com/product-feedback/361104), but not sure if that also holds for the submission notebook.",
    "2205287": "Hey! \nPlease check if your data types are efficient and that you don't use too many features.\nAlso - what model are you using? \nMabey tweaking the hyperparameters for memory-efficient options can help.\n\nGood luck!",
    "2205063": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1332227%2Fbcd93bd8c0ce4f79d3f82530ec3ebc5a%2F31.03.2023_17.19.36_REC.png?generation=1680335849808774&alt=media)\n",
    "2205105": ""
  }
}