{
  "id": 195370,
  "title": "Memory Issues",
  "url": "/competitions/riiid-test-answer-prediction/discussion/195370",
  "author_name": "",
  "post_date": "2020-11-05T02:35:53.613518400Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I am only using 20% of the training data, then splitting into a train-test split (80/20). Trying to run a Random Forest classifier and I keep maxing out memory - how are people working around this? My data is stored in a pandas dataframe, so curious if that could be causing issues?</p>\n<p>Thanks</p>",
  "messages": [
    {
      "id": "1069860",
      "postDate": "11/05/2020 02:35:53",
      "content": "<p>I am only using 20% of the training data, then splitting into a train-test split (80/20). Trying to run a Random Forest classifier and I keep maxing out memory - how are people working around this? My data is stored in a pandas dataframe, so curious if that could be causing issues?</p>\n<p>Thanks</p>",
      "rawMarkdown": "I am only using 20% of the training data, then splitting into a train-test split (80/20). Trying to run a Random Forest classifier and I keep maxing out memory - how are people working around this? My data is stored in a pandas dataframe, so curious if that could be causing issues?\n\nThanks",
      "votes": null
    },
    {
      "id": "1070394",
      "postDate": "11/05/2020 17:55:10",
      "content": "<p>Did you try to delete all your useless variables using the <code>del</code> python function before the training of the model? </p>",
      "rawMarkdown": "Did you try to delete all your useless variables using the `del` python function before the training of the model?",
      "votes": null
    },
    {
      "id": "1070615",
      "postDate": "11/06/2020 00:34:36",
      "content": "<p>Hi Chicago,</p>\n<p>One of the challenges of this competition is working with such a large dataset - lack of memory is a common issue we're all facing. 20% of the data is still a very large amount of data   - you can create a reasonably good model using much less data . Also look into other algorithms that are less memory intensive - this why lightgbm is quite popular for this competition. </p>",
      "rawMarkdown": "Hi Chicago,\n\nOne of the challenges of this competition is working with such a large dataset - lack of memory is a common issue we're all facing. 20% of the data is still a very large amount of data   - you can create a reasonably good model using much less data . Also look into other algorithms that are less memory intensive - this why lightgbm is quite popular for this competition.",
      "votes": null
    },
    {
      "id": "1076950",
      "postDate": "11/13/2020 04:16:25",
      "content": "<p><a href=\"https://www.kaggle.com/lgreig\" target=\"_blank\">@lgreig</a> - thanks for the note, I will take a look at lightgbm. I am also down to 10% of train now, pretty slow but avoiding memory issues.</p>\n<p>Edit: been working with lightgbm, pretty nice!</p>",
      "rawMarkdown": "lgreig - thanks for the note, I will take a look at lightgbm. I am also down to 10% of train now, pretty slow but avoiding memory issues.\n\nEdit: been working with lightgbm, pretty nice!",
      "votes": null
    },
    {
      "id": "1076951",
      "postDate": "11/13/2020 04:17:44",
      "content": "<p>Thanks for the recommendation! I did do this for features I was not keeping in my model, thanks for the reminder! Any other recommendations for limiting memory usage?</p>",
      "rawMarkdown": "Thanks for the recommendation! I did do this for features I was not keeping in my model, thanks for the reminder! Any other recommendations for limiting memory usage?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1070394,
      "author_name": "matthieuplante",
      "author_url": "",
      "post_date": "11/05/2020 17:55:10",
      "content": "<p>Did you try to delete all your useless variables using the <code>del</code> python function before the training of the model? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1076951,
          "author_name": "chicagojckaggle",
          "author_url": "",
          "post_date": "11/13/2020 04:17:44",
          "content": "<p>Thanks for the recommendation! I did do this for features I was not keeping in my model, thanks for the reminder! Any other recommendations for limiting memory usage?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1070615,
      "author_name": "lgreig",
      "author_url": "",
      "post_date": "11/06/2020 00:34:36",
      "content": "<p>Hi Chicago,</p>\n<p>One of the challenges of this competition is working with such a large dataset - lack of memory is a common issue we're all facing. 20% of the data is still a very large amount of data   - you can create a reasonably good model using much less data . Also look into other algorithms that are less memory intensive - this why lightgbm is quite popular for this competition. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1076950,
      "author_name": "chicagojckaggle",
      "author_url": "",
      "post_date": "11/13/2020 04:16:25",
      "content": "<p><a href=\"https://www.kaggle.com/lgreig\" target=\"_blank\">@lgreig</a> - thanks for the note, I will take a look at lightgbm. I am also down to 10% of train now, pretty slow but avoiding memory issues.</p>\n<p>Edit: been working with lightgbm, pretty nice!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1069860": "I am only using 20% of the training data, then splitting into a train-test split (80/20). Trying to run a Random Forest classifier and I keep maxing out memory - how are people working around this? My data is stored in a pandas dataframe, so curious if that could be causing issues?\n\nThanks",
    "1070394": "Did you try to delete all your useless variables using the `del` python function before the training of the model?",
    "1070615": "Hi Chicago,\n\nOne of the challenges of this competition is working with such a large dataset - lack of memory is a common issue we're all facing. 20% of the data is still a very large amount of data   - you can create a reasonably good model using much less data . Also look into other algorithms that are less memory intensive - this why lightgbm is quite popular for this competition.",
    "1076950": "lgreig - thanks for the note, I will take a look at lightgbm. I am also down to 10% of train now, pretty slow but avoiding memory issues.\n\nEdit: been working with lightgbm, pretty nice!",
    "1076951": "Thanks for the recommendation! I did do this for features I was not keeping in my model, thanks for the reminder! Any other recommendations for limiting memory usage?"
  },
  "source": "meta"
}