{
  "id": 333342,
  "title": "Can't implement Standardization or PCA due to memory limitation",
  "url": "/competitions/amex-default-prediction/discussion/333342",
  "author_name": "",
  "post_date": "2022-06-26T03:27:44.522281300Z",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi everyone, I understand the dataset size is a big challenge and I have solved the loading data problem by transforming data type and file format.</p>\n<p>Now, the whole dataset occupied almost 5-6GB of memory but when I implement standardization or PCA work, the memory usage also reached the 16 GB limitation, and return \"Your Notebook tried to allocate more memory than available. It has been restarted\".</p>\n<p>I intended to keep more columns because removing them may lose some important information. But I still try to drop some columns by missing value percentage and correlation to solve the problem, but it didn't work.</p>\n<p>Now, my idea is to create a small train set by resampling, train models based on the small set, and repeat this process many times to improve model performance. For the test set which is larger than the train set, I plan to read part of the rows in each prediction process and repeat the above work until I get all predictions </p>\n<p>Is it a reasonable solution? Do you guys meet the same problem and how do you address or avoid the problem? </p>",
  "messages": [
    {
      "id": "1833463",
      "postDate": "06/26/2022 03:27:44",
      "content": "<p>Hi everyone, I understand the dataset size is a big challenge and I have solved the loading data problem by transforming data type and file format.</p>\n<p>Now, the whole dataset occupied almost 5-6GB of memory but when I implement standardization or PCA work, the memory usage also reached the 16 GB limitation, and return \"Your Notebook tried to allocate more memory than available. It has been restarted\".</p>\n<p>I intended to keep more columns because removing them may lose some important information. But I still try to drop some columns by missing value percentage and correlation to solve the problem, but it didn't work.</p>\n<p>Now, my idea is to create a small train set by resampling, train models based on the small set, and repeat this process many times to improve model performance. For the test set which is larger than the train set, I plan to read part of the rows in each prediction process and repeat the above work until I get all predictions </p>\n<p>Is it a reasonable solution? Do you guys meet the same problem and how do you address or avoid the problem? </p>",
      "rawMarkdown": "Hi everyone, I understand the dataset size is a big challenge and I have solved the loading data problem by transforming data type and file format.\n\nNow, the whole dataset occupied almost 5-6GB of memory but when I implement standardization or PCA work, the memory usage also reached the 16 GB limitation, and return \"Your Notebook tried to allocate more memory than available. It has been restarted\".\n\nI intended to keep more columns because removing them may lose some important information. But I still try to drop some columns by missing value percentage and correlation to solve the problem, but it didn't work.\n\nNow, my idea is to create a small train set by resampling, train models based on the small set, and repeat this process many times to improve model performance. For the test set which is larger than the train set, I plan to read part of the rows in each prediction process and repeat the above work until I get all predictions \n\nIs it a reasonable solution? Do you guys meet the same problem and how do you address or avoid the problem?",
      "votes": null
    },
    {
      "id": "1833581",
      "postDate": "06/26/2022 06:20:22",
      "content": "<p>Even I am facing the same problem. I am planning to run this on AWS cluster to resolve the problem. Does anyone have any other cheaper alternative?</p>",
      "rawMarkdown": "Even I am facing the same problem. I am planning to run this on AWS cluster to resolve the problem. Does anyone have any other cheaper alternative?",
      "votes": null
    },
    {
      "id": "1833584",
      "postDate": "06/26/2022 06:23:18",
      "content": "<p>There's a pretty decent number of things that do a temporary inflation of memory when your doing them - I don't have good practice of solutions since most of my work is local machine with big ram and big swap.   For some of the past cases it was easy to just do a few features at a time and build a new data frame from all the pieces.   So - solution 1 - do the work on local machine and put model in a kaggle data set.  BUT - Guessing you don't have a local machine :)</p>\n<p>As you noted it's not too hard to do predictions - think there are many shared notebooks that chunk the prediction.</p>\n<p>Getting PCA implies that your looking at the full feature set to do a reduction :)   </p>\n<p>Perhaps you can out-source the work !   Share your notebook for doing the process with the full file.  Than ask some kind soul with a bigger machine (like myself) to run the full data set.  I think if they (ok - I mean me) share the result (model and data) with you in a public kaggle data set (and post in this discussion the link) than I think we  would not be breaking any team/kaggle rules.  For sure a private sharing is a huge NO-NO.</p>\n<p>Hopefully someone will provide a solution that keeps the work entirely in you hands.</p>",
      "rawMarkdown": "There's a pretty decent number of things that do a temporary inflation of memory when your doing them - I don't have good practice of solutions since most of my work is local machine with big ram and big swap.   For some of the past cases it was easy to just do a few features at a time and build a new data frame from all the pieces.   So - solution 1 - do the work on local machine and put model in a kaggle data set.  BUT - Guessing you don't have a local machine :)\n\nAs you noted it's not too hard to do predictions - think there are many shared notebooks that chunk the prediction.\n\nGetting PCA implies that your looking at the full feature set to do a reduction :)   \n\nPerhaps you can out-source the work !   Share your notebook for doing the process with the full file.  Than ask some kind soul with a bigger machine (like myself) to run the full data set.  I think if they (ok - I mean me) share the result (model and data) with you in a public kaggle data set (and post in this discussion the link) than I think we  would not be breaking any team/kaggle rules.  For sure a private sharing is a huge NO-NO.\n\nHopefully someone will provide a solution that keeps the work entirely in you hands.",
      "votes": null
    },
    {
      "id": "1833586",
      "postDate": "06/26/2022 06:24:15",
      "content": "<p>Hi. I'm surprised that standardization returns a memory error. You could always do standardization yourself with a for loop like</p>\n<pre><code>all_means = {}\nall_stds = {}\nfor c in train.columns:\n    m = train[c].mean()\n    all_means[c] = m\n    s = train[c].std()\n    all_stds[c] = s\n    train[c] = train[c] - m\n    if s!=0: train[c] = train[c]/s\n</code></pre>\n<p>Then later you can use</p>\n<pre><code>for c in test.columns:\n   test[c] = test[c] - all_means[c]\n   if all_stds[c]!=0: test[c] = test[c]/all_stds[c]\n</code></pre>\n<p>If the above throws an error, then only use the first K rows like <code>m = train[c].iloc[:10_000].mean()</code> and later convert all train with <code>train[c] = train[c] - m</code></p>\n<p>Regrading PCA, you could fit the first K rows that transform everything else with that like</p>\n<pre><code>model = PCA()\nmodel.fit(train.iloc[:10_000])\ntrain = model.transform(train)\ntest = model.transform(test)\n</code></pre>\n<p>Or you can look into incremental PCA that fits PCA in batches which uses less memory.</p>",
      "rawMarkdown": "Hi. I'm surprised that standardization returns a memory error. You could always do standardization yourself with a for loop like\n\n    all_means = {}\n    all_stds = {}\n    for c in train.columns:\n        m = train[c].mean()\n        all_means[c] = m\n        s = train[c].std()\n        all_stds[c] = s\n        train[c] = train[c] - m\n        if s!=0: train[c] = train[c]/s\n\nThen later you can use\n \n    for c in test.columns:\n       test[c] = test[c] - all_means[c]\n       if all_stds[c]!=0: test[c] = test[c]/all_stds[c]\n\nIf the above throws an error, then only use the first K rows like `m = train[c].iloc[:10_000].mean()` and later convert all train with `train[c] = train[c] - m`\n\nRegrading PCA, you could fit the first K rows that transform everything else with that like\n\n    model = PCA()\n    model.fit(train.iloc[:10_000])\n    train = model.transform(train)\n    test = model.transform(test)\n\nOr you can look into incremental PCA that fits PCA in batches which uses less memory.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1833581,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "06/26/2022 06:20:22",
      "content": "<p>Even I am facing the same problem. I am planning to run this on AWS cluster to resolve the problem. Does anyone have any other cheaper alternative?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1833584,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "06/26/2022 06:23:18",
      "content": "<p>There's a pretty decent number of things that do a temporary inflation of memory when your doing them - I don't have good practice of solutions since most of my work is local machine with big ram and big swap.   For some of the past cases it was easy to just do a few features at a time and build a new data frame from all the pieces.   So - solution 1 - do the work on local machine and put model in a kaggle data set.  BUT - Guessing you don't have a local machine :)</p>\n<p>As you noted it's not too hard to do predictions - think there are many shared notebooks that chunk the prediction.</p>\n<p>Getting PCA implies that your looking at the full feature set to do a reduction :)   </p>\n<p>Perhaps you can out-source the work !   Share your notebook for doing the process with the full file.  Than ask some kind soul with a bigger machine (like myself) to run the full data set.  I think if they (ok - I mean me) share the result (model and data) with you in a public kaggle data set (and post in this discussion the link) than I think we  would not be breaking any team/kaggle rules.  For sure a private sharing is a huge NO-NO.</p>\n<p>Hopefully someone will provide a solution that keeps the work entirely in you hands.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1833586,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "06/26/2022 06:24:15",
      "content": "<p>Hi. I'm surprised that standardization returns a memory error. You could always do standardization yourself with a for loop like</p>\n<pre><code>all_means = {}\nall_stds = {}\nfor c in train.columns:\n    m = train[c].mean()\n    all_means[c] = m\n    s = train[c].std()\n    all_stds[c] = s\n    train[c] = train[c] - m\n    if s!=0: train[c] = train[c]/s\n</code></pre>\n<p>Then later you can use</p>\n<pre><code>for c in test.columns:\n   test[c] = test[c] - all_means[c]\n   if all_stds[c]!=0: test[c] = test[c]/all_stds[c]\n</code></pre>\n<p>If the above throws an error, then only use the first K rows like <code>m = train[c].iloc[:10_000].mean()</code> and later convert all train with <code>train[c] = train[c] - m</code></p>\n<p>Regrading PCA, you could fit the first K rows that transform everything else with that like</p>\n<pre><code>model = PCA()\nmodel.fit(train.iloc[:10_000])\ntrain = model.transform(train)\ntest = model.transform(test)\n</code></pre>\n<p>Or you can look into incremental PCA that fits PCA in batches which uses less memory.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1833463": "Hi everyone, I understand the dataset size is a big challenge and I have solved the loading data problem by transforming data type and file format.\n\nNow, the whole dataset occupied almost 5-6GB of memory but when I implement standardization or PCA work, the memory usage also reached the 16 GB limitation, and return \"Your Notebook tried to allocate more memory than available. It has been restarted\".\n\nI intended to keep more columns because removing them may lose some important information. But I still try to drop some columns by missing value percentage and correlation to solve the problem, but it didn't work.\n\nNow, my idea is to create a small train set by resampling, train models based on the small set, and repeat this process many times to improve model performance. For the test set which is larger than the train set, I plan to read part of the rows in each prediction process and repeat the above work until I get all predictions \n\nIs it a reasonable solution? Do you guys meet the same problem and how do you address or avoid the problem?",
    "1833581": "Even I am facing the same problem. I am planning to run this on AWS cluster to resolve the problem. Does anyone have any other cheaper alternative?",
    "1833584": "There's a pretty decent number of things that do a temporary inflation of memory when your doing them - I don't have good practice of solutions since most of my work is local machine with big ram and big swap.   For some of the past cases it was easy to just do a few features at a time and build a new data frame from all the pieces.   So - solution 1 - do the work on local machine and put model in a kaggle data set.  BUT - Guessing you don't have a local machine :)\n\nAs you noted it's not too hard to do predictions - think there are many shared notebooks that chunk the prediction.\n\nGetting PCA implies that your looking at the full feature set to do a reduction :)   \n\nPerhaps you can out-source the work !   Share your notebook for doing the process with the full file.  Than ask some kind soul with a bigger machine (like myself) to run the full data set.  I think if they (ok - I mean me) share the result (model and data) with you in a public kaggle data set (and post in this discussion the link) than I think we  would not be breaking any team/kaggle rules.  For sure a private sharing is a huge NO-NO.\n\nHopefully someone will provide a solution that keeps the work entirely in you hands.",
    "1833586": "Hi. I'm surprised that standardization returns a memory error. You could always do standardization yourself with a for loop like\n\n    all_means = {}\n    all_stds = {}\n    for c in train.columns:\n        m = train[c].mean()\n        all_means[c] = m\n        s = train[c].std()\n        all_stds[c] = s\n        train[c] = train[c] - m\n        if s!=0: train[c] = train[c]/s\n\nThen later you can use\n \n    for c in test.columns:\n       test[c] = test[c] - all_means[c]\n       if all_stds[c]!=0: test[c] = test[c]/all_stds[c]\n\nIf the above throws an error, then only use the first K rows like `m = train[c].iloc[:10_000].mean()` and later convert all train with `train[c] = train[c] - m`\n\nRegrading PCA, you could fit the first K rows that transform everything else with that like\n\n    model = PCA()\n    model.fit(train.iloc[:10_000])\n    train = model.transform(train)\n    test = model.transform(test)\n\nOr you can look into incremental PCA that fits PCA in batches which uses less memory."
  },
  "source": "meta"
}