{
  "id": 204972,
  "title": "Run out of RAM - Large Dataset",
  "url": "/competitions/predict-volcanic-eruptions-ingv-oe/discussion/204972",
  "author_name": "",
  "post_date": "2020-12-17T19:31:27.935088900Z",
  "votes": 2,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Hello all,</p>\n<p>Due to the size of this dataset, I keep running out of RAM in the Kaggle notebooks every time I try to store the training data in a variable. Does anyone have any recommendations? </p>\n<p>Thanks!</p>",
  "messages": [
    {
      "id": "1117179",
      "postDate": "12/17/2020 19:31:27",
      "content": "<p>Hello all,</p>\n<p>Due to the size of this dataset, I keep running out of RAM in the Kaggle notebooks every time I try to store the training data in a variable. Does anyone have any recommendations? </p>\n<p>Thanks!</p>",
      "rawMarkdown": "Hello all,\n\nDue to the size of this dataset, I keep running out of RAM in the Kaggle notebooks every time I try to store the training data in a variable. Does anyone have any recommendations? \n\nThanks!",
      "votes": null
    },
    {
      "id": "1117288",
      "postDate": "12/17/2020 21:59:22",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/kyleberdy\" target=\"_blank\">@kyleberdy</a>,</p>\n<p>I did was take groups of samples and calculate features by sample and sensor. Then I have worked with the features of the data.</p>\n<p>I determined the number of observations per group based on memory.</p>",
      "rawMarkdown": "Hello @kyleberdy,\n\nI did was take groups of samples and calculate features by sample and sensor. Then I have worked with the features of the data.\n\nI determined the number of observations per group based on memory.",
      "votes": null
    },
    {
      "id": "1117307",
      "postDate": "12/17/2020 23:01:45",
      "content": "<p>Interesting. I've got it figured out now, I think. I needed the large sample group to fit my imputer, but if I merged all 4000+ training sets into a dataframe, I ended up getting a ridiculously large one, so I chose instead to make it a 400 training set data frame (and to delete it afterwards to save memory) so that I could train the imputer. Thanks for your response!</p>",
      "rawMarkdown": "Interesting. I've got it figured out now, I think. I needed the large sample group to fit my imputer, but if I merged all 4000+ training sets into a dataframe, I ended up getting a ridiculously large one, so I chose instead to make it a 400 training set data frame (and to delete it afterwards to save memory) so that I could train the imputer. Thanks for your response!",
      "votes": null
    },
    {
      "id": "1117647",
      "postDate": "12/18/2020 09:47:37",
      "content": "<p>If you are using neural networks then you can simply reduce the batch size and problem will solved. Or you can use dimension reduction methods and can resolve this issue and also can improve the results. </p>",
      "rawMarkdown": "If you are using neural networks then you can simply reduce the batch size and problem will solved. Or you can use dimension reduction methods and can resolve this issue and also can improve the results.",
      "votes": null
    },
    {
      "id": "1117857",
      "postDate": "12/18/2020 14:22:44",
      "content": "<p>Hello Muhammad,</p>\n<p>You guessed it. I'm working on reducing the batch size to an appropriate amount. The kernel keeps restarting every time I run out of memory, but I want to go for the largest possible batch size (within about 50-100). Do you have any recommendations?</p>",
      "rawMarkdown": "Hello Muhammad,\n\nYou guessed it. I'm working on reducing the batch size to an appropriate amount. The kernel keeps restarting every time I run out of memory, but I want to go for the largest possible batch size (within about 50-100). Do you have any recommendations?",
      "votes": null
    },
    {
      "id": "1118048",
      "postDate": "12/18/2020 17:21:11",
      "content": "<p>well can you tell me what batch size while getting out of memory error? Usually we keep setting low and low batch size until issue solved.  </p>",
      "rawMarkdown": "well can you tell me what batch size while getting out of memory error? Usually we keep setting low and low batch size until issue solved.",
      "votes": null
    },
    {
      "id": "1118056",
      "postDate": "12/18/2020 17:27:20",
      "content": "<p>One more important thing, you can apply features normalization so  in this way you can also solve the issue of out of memory.  </p>",
      "rawMarkdown": "One more important thing, you can apply features normalization so  in this way you can also solve the issue of out of memory.",
      "votes": null
    },
    {
      "id": "1118057",
      "postDate": "12/18/2020 17:27:40",
      "content": "<p>Let me know the status. Thanks</p>",
      "rawMarkdown": "Let me know the status. Thanks",
      "votes": null
    },
    {
      "id": "1118555",
      "postDate": "12/19/2020 07:19:05",
      "content": "<p>its good approach as well. </p>",
      "rawMarkdown": "its good approach as well.",
      "votes": null
    },
    {
      "id": "1118636",
      "postDate": "12/19/2020 09:04:06",
      "content": "<p>Delete those variables which u r not using.</p>",
      "rawMarkdown": "Delete those variables which u r not using.",
      "votes": null
    },
    {
      "id": "1118943",
      "postDate": "12/19/2020 14:48:04",
      "content": "<p>That's what I also did.</p>",
      "rawMarkdown": "That's what I also did.",
      "votes": null
    },
    {
      "id": "1118987",
      "postDate": "12/19/2020 15:54:30",
      "content": "<p>Kyle Berdy, <br>\nDid you solved the problem or still facing issue. Let me know so i could suggest you more ?</p>",
      "rawMarkdown": "Kyle Berdy, \nDid you solved the problem or still facing issue. Let me know so i could suggest you more ?",
      "votes": null
    },
    {
      "id": "1119264",
      "postDate": "12/19/2020 21:19:26",
      "content": "<p>Yes, the problem is solved now. The only other problem I'm having is that the kernel times out after 40 minutes during training, so I'll have to decrease the amount of training.</p>",
      "rawMarkdown": "Yes, the problem is solved now. The only other problem I'm having is that the kernel times out after 40 minutes during training, so I'll have to decrease the amount of training.",
      "votes": null
    },
    {
      "id": "1119273",
      "postDate": "12/19/2020 21:23:17",
      "content": "<p>Alright great. Have a nice day !</p>",
      "rawMarkdown": "Alright great. Have a nice day !",
      "votes": null
    },
    {
      "id": "1121663",
      "postDate": "12/21/2020 19:57:21",
      "content": "<p>Thanks, you too!</p>",
      "rawMarkdown": "Thanks, you too!",
      "votes": null
    },
    {
      "id": "1138364",
      "postDate": "01/04/2021 16:33:36",
      "content": "<p>Another potential idea is to reduce batch size and use a generator.</p>",
      "rawMarkdown": "Another potential idea is to reduce batch size and use a generator.",
      "votes": null
    },
    {
      "id": "1141491",
      "postDate": "01/06/2021 18:26:06",
      "content": "<p>Hey Karthik,</p>\n<p>Thanks. Not quite certain what you mean by a generator. Could you explain a little more.</p>",
      "rawMarkdown": "Hey Karthik,\n\nThanks. Not quite certain what you mean by a generator. Could you explain a little more.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1117288,
      "author_name": "desareca",
      "author_url": "",
      "post_date": "12/17/2020 21:59:22",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/kyleberdy\" target=\"_blank\">@kyleberdy</a>,</p>\n<p>I did was take groups of samples and calculate features by sample and sensor. Then I have worked with the features of the data.</p>\n<p>I determined the number of observations per group based on memory.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1117307,
          "author_name": "kyleberdy",
          "author_url": "",
          "post_date": "12/17/2020 23:01:45",
          "content": "<p>Interesting. I've got it figured out now, I think. I needed the large sample group to fit my imputer, but if I merged all 4000+ training sets into a dataframe, I ended up getting a ridiculously large one, so I chose instead to make it a 400 training set data frame (and to delete it afterwards to save memory) so that I could train the imputer. Thanks for your response!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1118555,
          "author_name": "",
          "author_url": "",
          "post_date": "12/19/2020 07:19:05",
          "content": "<p>its good approach as well. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1117647,
      "author_name": "",
      "author_url": "",
      "post_date": "12/18/2020 09:47:37",
      "content": "<p>If you are using neural networks then you can simply reduce the batch size and problem will solved. Or you can use dimension reduction methods and can resolve this issue and also can improve the results. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1117857,
          "author_name": "kyleberdy",
          "author_url": "",
          "post_date": "12/18/2020 14:22:44",
          "content": "<p>Hello Muhammad,</p>\n<p>You guessed it. I'm working on reducing the batch size to an appropriate amount. The kernel keeps restarting every time I run out of memory, but I want to go for the largest possible batch size (within about 50-100). Do you have any recommendations?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1118048,
          "author_name": "",
          "author_url": "",
          "post_date": "12/18/2020 17:21:11",
          "content": "<p>well can you tell me what batch size while getting out of memory error? Usually we keep setting low and low batch size until issue solved.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1118056,
          "author_name": "",
          "author_url": "",
          "post_date": "12/18/2020 17:27:20",
          "content": "<p>One more important thing, you can apply features normalization so  in this way you can also solve the issue of out of memory.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1118057,
          "author_name": "",
          "author_url": "",
          "post_date": "12/18/2020 17:27:40",
          "content": "<p>Let me know the status. Thanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1118636,
      "author_name": "saurabhshahane",
      "author_url": "",
      "post_date": "12/19/2020 09:04:06",
      "content": "<p>Delete those variables which u r not using.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1118943,
          "author_name": "kyleberdy",
          "author_url": "",
          "post_date": "12/19/2020 14:48:04",
          "content": "<p>That's what I also did.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1118987,
      "author_name": "",
      "author_url": "",
      "post_date": "12/19/2020 15:54:30",
      "content": "<p>Kyle Berdy, <br>\nDid you solved the problem or still facing issue. Let me know so i could suggest you more ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1119264,
          "author_name": "kyleberdy",
          "author_url": "",
          "post_date": "12/19/2020 21:19:26",
          "content": "<p>Yes, the problem is solved now. The only other problem I'm having is that the kernel times out after 40 minutes during training, so I'll have to decrease the amount of training.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1119273,
          "author_name": "",
          "author_url": "",
          "post_date": "12/19/2020 21:23:17",
          "content": "<p>Alright great. Have a nice day !</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1121663,
          "author_name": "kyleberdy",
          "author_url": "",
          "post_date": "12/21/2020 19:57:21",
          "content": "<p>Thanks, you too!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1138364,
      "author_name": "karthikvijayraghavan",
      "author_url": "",
      "post_date": "01/04/2021 16:33:36",
      "content": "<p>Another potential idea is to reduce batch size and use a generator.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1141491,
          "author_name": "kyleberdy",
          "author_url": "",
          "post_date": "01/06/2021 18:26:06",
          "content": "<p>Hey Karthik,</p>\n<p>Thanks. Not quite certain what you mean by a generator. Could you explain a little more.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1117179": "Hello all,\n\nDue to the size of this dataset, I keep running out of RAM in the Kaggle notebooks every time I try to store the training data in a variable. Does anyone have any recommendations? \n\nThanks!",
    "1117288": "Hello @kyleberdy,\n\nI did was take groups of samples and calculate features by sample and sensor. Then I have worked with the features of the data.\n\nI determined the number of observations per group based on memory.",
    "1117307": "Interesting. I've got it figured out now, I think. I needed the large sample group to fit my imputer, but if I merged all 4000+ training sets into a dataframe, I ended up getting a ridiculously large one, so I chose instead to make it a 400 training set data frame (and to delete it afterwards to save memory) so that I could train the imputer. Thanks for your response!",
    "1117647": "If you are using neural networks then you can simply reduce the batch size and problem will solved. Or you can use dimension reduction methods and can resolve this issue and also can improve the results.",
    "1117857": "Hello Muhammad,\n\nYou guessed it. I'm working on reducing the batch size to an appropriate amount. The kernel keeps restarting every time I run out of memory, but I want to go for the largest possible batch size (within about 50-100). Do you have any recommendations?",
    "1118048": "well can you tell me what batch size while getting out of memory error? Usually we keep setting low and low batch size until issue solved.",
    "1118056": "One more important thing, you can apply features normalization so  in this way you can also solve the issue of out of memory.",
    "1118057": "Let me know the status. Thanks",
    "1118555": "its good approach as well.",
    "1118636": "Delete those variables which u r not using.",
    "1118943": "That's what I also did.",
    "1118987": "Kyle Berdy, \nDid you solved the problem or still facing issue. Let me know so i could suggest you more ?",
    "1119264": "Yes, the problem is solved now. The only other problem I'm having is that the kernel times out after 40 minutes during training, so I'll have to decrease the amount of training.",
    "1119273": "Alright great. Have a nice day !",
    "1121663": "Thanks, you too!",
    "1138364": "Another potential idea is to reduce batch size and use a generator.",
    "1141491": "Hey Karthik,\n\nThanks. Not quite certain what you mean by a generator. Could you explain a little more."
  },
  "source": "meta"
}