{
  "id": 584101,
  "title": "I don't have enough RAM to realize PCA",
  "url": "/competitions/drw-crypto-market-prediction/discussion/584101",
  "author_name": "",
  "post_date": "2025-06-11T17:50:13.830479200Z",
  "votes": 5,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi, Everyone<br>\nI have a question, because when I tried to run my code, RAM isn't enough to run my code, I tried with dask, but I have the same problem, I would like know ¿how I run my code in this context?<br>\nThe context is convert data although standard scaler, and then apply PCA, or similar <br>\nWhat's your possibly solution to this problem?</p>",
  "messages": [
    {
      "id": "3222017",
      "postDate": "06/11/2025 17:50:13",
      "content": "<p>Hi, Everyone<br>\nI have a question, because when I tried to run my code, RAM isn't enough to run my code, I tried with dask, but I have the same problem, I would like know ¿how I run my code in this context?<br>\nThe context is convert data although standard scaler, and then apply PCA, or similar <br>\nWhat's your possibly solution to this problem?</p>",
      "rawMarkdown": "Hi, Everyone\nI have a question, because when I tried to run my code, RAM isn't enough to run my code, I tried with dask, but I have the same problem, I would like know ¿how I run my code in this context?\nThe context is convert data although standard scaler, and then apply PCA, or similar \nWhat's your possibly solution to this problem?",
      "votes": null
    },
    {
      "id": "3222029",
      "postDate": "06/11/2025 18:16:53",
      "content": "<p>Do you need PCA here? <a href=\"https://www.kaggle.com/danielphys\" target=\"_blank\">@danielphys</a> </p>",
      "rawMarkdown": "Do you need PCA here? @danielphys",
      "votes": null
    },
    {
      "id": "3222106",
      "postDate": "06/11/2025 20:06:09",
      "content": "<p>In my opinion, I need PCA to reduce dimensionality with all columns, or what's the best form to reduce dimensionality in this dataset? <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> </p>",
      "rawMarkdown": "In my opinion, I need PCA to reduce dimensionality with all columns, or what's the best form to reduce dimensionality in this dataset? @ravi20076",
      "votes": null
    },
    {
      "id": "3222351",
      "postDate": "06/12/2025 06:31:06",
      "content": "<p>You can use PCA here. Two important things you need to do before calling PCA() are:</p>\n<ol>\n<li>Reduce memory usage by data frames. You can use function like  reduce_mem_usage. (Available in Kaggle public notebooks).</li>\n<li>some columns contain -inifinity values. You can replace it with large negative numbers like -9999999.<br>\ntrain.replace([-np.inf], -999999999.0, inplace=True)<br>\ntest.replace([-np.inf], -999999999.0, inplace=True)</li>\n</ol>\n<p>Do share the findings of your research, please. All of us are eager to find out how PCA may improve outcomes.</p>",
      "rawMarkdown": "You can use PCA here. Two important things you need to do before calling PCA() are:\n1. Reduce memory usage by data frames. You can use function like  reduce_mem_usage. (Available in Kaggle public notebooks).\n2. some columns contain -inifinity values. You can replace it with large negative numbers like -9999999.\ntrain.replace([-np.inf], -999999999.0, inplace=True)\ntest.replace([-np.inf], -999999999.0, inplace=True)\n\nDo share the findings of your research, please. All of us are eager to find out how PCA may improve outcomes.",
      "votes": null
    },
    {
      "id": "3222355",
      "postDate": "06/12/2025 06:38:26",
      "content": "<p>Think about it - PCA has its own flaws and hardly works in most cases <a href=\"https://www.kaggle.com/danielphys\" target=\"_blank\">@danielphys</a> </p>",
      "rawMarkdown": "Think about it - PCA has its own flaws and hardly works in most cases @danielphys",
      "votes": null
    },
    {
      "id": "3222411",
      "postDate": "06/12/2025 07:26:28",
      "content": "<p>Agreed, I rarely see PCA used by top players. I would select the most important 50-100 columns to speedup testing instead of PCA.</p>",
      "rawMarkdown": "Agreed, I rarely see PCA used by top players. I would select the most important 50-100 columns to speedup testing instead of PCA.",
      "votes": null
    },
    {
      "id": "3222773",
      "postDate": "06/12/2025 13:57:25",
      "content": "<p>ok                          </p>",
      "rawMarkdown": "ok",
      "votes": null
    },
    {
      "id": "3223141",
      "postDate": "06/12/2025 22:08:40",
      "content": "<p>I remember PCA being very actively used in <a href=\"https://www.kaggle.com/competitions/lish-moa/overview\" target=\"_blank\">this competition</a>, and it kind of worked. Check top write-ups like place 1st, 2nd, 4th, and so on. Maybe you just need to use more NN with PCA features rather than GBDT.</p>",
      "rawMarkdown": "I remember PCA being very actively used in [this competition](https://www.kaggle.com/competitions/lish-moa/overview), and it kind of worked. Check top write-ups like place 1st, 2nd, 4th, and so on. Maybe you just need to use more NN with PCA features rather than GBDT.",
      "votes": null
    },
    {
      "id": "3223358",
      "postDate": "06/13/2025 07:34:27",
      "content": "<p>Use the incremental PCA, in the scikit learn <br>\nfrom sklearn.decomposition import IncrementalPCA<br>\nthis takes a few seconds and fits well </p>",
      "rawMarkdown": "Use the incremental PCA, in the scikit learn \nfrom sklearn.decomposition import IncrementalPCA\nthis takes a few seconds and fits well",
      "votes": null
    },
    {
      "id": "3226964",
      "postDate": "06/18/2025 10:07:37",
      "content": "<p>Hi! I ran into similar memory issues when applying StandardScaler followed by PCA on large datasets — especially with float-heavy parquet files. Here are a few techniques that helped me handle it efficiently:</p>\n<ol>\n<li>Use Incremental PCA<br>\nInstead of standard PCA, use a batch-based or incremental approach that processes data in chunks. It dramatically reduces memory overhead.</li>\n<li>Scale in Mini-Batches<br>\nRather than scaling the full dataset at once, fit and transform your data using mini-batches. This is especially helpful when memory is tight.</li>\n<li>Convert to float32 Early<br>\nCasting your data to float32 instead of the default float64 can save nearly 50% memory. It's a simple but often overlooked optimization.</li>\n<li>Apply Feature Selection Before PCA<br>\nUse techniques like SelectKBest, variance thresholds, or correlation filtering to reduce dimensionality before scaling. PCA works better (and faster) on a leaner dataset.</li>\n<li>Consider Saving Intermediate Outputs<br>\nIf your process fails midway, save checkpoints at key stages like after scaling or feature selection. This allows you to resume rather than restart.</li>\n<li>if youre running the code locally and its your computer hardware holding you back, consider creating a Kaggle notebook on the website as Kaggle offers you sufficient ram given you optimise your code.<br>\nThese approaches allowed me to run full pipelines even on limited RAM Kaggle environments. Hope this helps!!</li>\n</ol>",
      "rawMarkdown": "Hi! I ran into similar memory issues when applying StandardScaler followed by PCA on large datasets — especially with float-heavy parquet files. Here are a few techniques that helped me handle it efficiently:\n 1. Use Incremental PCA\nInstead of standard PCA, use a batch-based or incremental approach that processes data in chunks. It dramatically reduces memory overhead.\n 2. Scale in Mini-Batches\nRather than scaling the full dataset at once, fit and transform your data using mini-batches. This is especially helpful when memory is tight.\n 3. Convert to float32 Early\nCasting your data to float32 instead of the default float64 can save nearly 50% memory. It's a simple but often overlooked optimization.\n 4. Apply Feature Selection Before PCA\nUse techniques like SelectKBest, variance thresholds, or correlation filtering to reduce dimensionality before scaling. PCA works better (and faster) on a leaner dataset.\n 5. Consider Saving Intermediate Outputs\nIf your process fails midway, save checkpoints at key stages like after scaling or feature selection. This allows you to resume rather than restart.\n 6. if youre running the code locally and its your computer hardware holding you back, consider creating a Kaggle notebook on the website as Kaggle offers you sufficient ram given you optimise your code.\nThese approaches allowed me to run full pipelines even on limited RAM Kaggle environments. Hope this helps!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3222029,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "06/11/2025 18:16:53",
      "content": "<p>Do you need PCA here? <a href=\"https://www.kaggle.com/danielphys\" target=\"_blank\">@danielphys</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 3222106,
          "author_name": "danielphys",
          "author_url": "",
          "post_date": "06/11/2025 20:06:09",
          "content": "<p>In my opinion, I need PCA to reduce dimensionality with all columns, or what's the best form to reduce dimensionality in this dataset? <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> </p>",
          "votes": null,
          "replies": [
            {
              "id": 3222355,
              "author_name": "ravi20076",
              "author_url": "",
              "post_date": "06/12/2025 06:38:26",
              "content": "<p>Think about it - PCA has its own flaws and hardly works in most cases <a href=\"https://www.kaggle.com/danielphys\" target=\"_blank\">@danielphys</a> </p>",
              "votes": null,
              "replies": [
                {
                  "id": 3222411,
                  "author_name": "nguyennguyen599",
                  "author_url": "",
                  "post_date": "06/12/2025 07:26:28",
                  "content": "<p>Agreed, I rarely see PCA used by top players. I would select the most important 50-100 columns to speedup testing instead of PCA.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3222773,
                      "author_name": "nirupammondal",
                      "author_url": "",
                      "post_date": "06/12/2025 13:57:25",
                      "content": "<p>ok                          </p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3223141,
                          "author_name": "yekenot",
                          "author_url": "",
                          "post_date": "06/12/2025 22:08:40",
                          "content": "<p>I remember PCA being very actively used in <a href=\"https://www.kaggle.com/competitions/lish-moa/overview\" target=\"_blank\">this competition</a>, and it kind of worked. Check top write-ups like place 1st, 2nd, 4th, and so on. Maybe you just need to use more NN with PCA features rather than GBDT.</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3222351,
      "author_name": "itasps",
      "author_url": "",
      "post_date": "06/12/2025 06:31:06",
      "content": "<p>You can use PCA here. Two important things you need to do before calling PCA() are:</p>\n<ol>\n<li>Reduce memory usage by data frames. You can use function like  reduce_mem_usage. (Available in Kaggle public notebooks).</li>\n<li>some columns contain -inifinity values. You can replace it with large negative numbers like -9999999.<br>\ntrain.replace([-np.inf], -999999999.0, inplace=True)<br>\ntest.replace([-np.inf], -999999999.0, inplace=True)</li>\n</ol>\n<p>Do share the findings of your research, please. All of us are eager to find out how PCA may improve outcomes.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3223358,
      "author_name": "yoogwangpyo",
      "author_url": "",
      "post_date": "06/13/2025 07:34:27",
      "content": "<p>Use the incremental PCA, in the scikit learn <br>\nfrom sklearn.decomposition import IncrementalPCA<br>\nthis takes a few seconds and fits well </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3226964,
      "author_name": "sanketpai",
      "author_url": "",
      "post_date": "06/18/2025 10:07:37",
      "content": "<p>Hi! I ran into similar memory issues when applying StandardScaler followed by PCA on large datasets — especially with float-heavy parquet files. Here are a few techniques that helped me handle it efficiently:</p>\n<ol>\n<li>Use Incremental PCA<br>\nInstead of standard PCA, use a batch-based or incremental approach that processes data in chunks. It dramatically reduces memory overhead.</li>\n<li>Scale in Mini-Batches<br>\nRather than scaling the full dataset at once, fit and transform your data using mini-batches. This is especially helpful when memory is tight.</li>\n<li>Convert to float32 Early<br>\nCasting your data to float32 instead of the default float64 can save nearly 50% memory. It's a simple but often overlooked optimization.</li>\n<li>Apply Feature Selection Before PCA<br>\nUse techniques like SelectKBest, variance thresholds, or correlation filtering to reduce dimensionality before scaling. PCA works better (and faster) on a leaner dataset.</li>\n<li>Consider Saving Intermediate Outputs<br>\nIf your process fails midway, save checkpoints at key stages like after scaling or feature selection. This allows you to resume rather than restart.</li>\n<li>if youre running the code locally and its your computer hardware holding you back, consider creating a Kaggle notebook on the website as Kaggle offers you sufficient ram given you optimise your code.<br>\nThese approaches allowed me to run full pipelines even on limited RAM Kaggle environments. Hope this helps!!</li>\n</ol>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3222017": "Hi, Everyone\nI have a question, because when I tried to run my code, RAM isn't enough to run my code, I tried with dask, but I have the same problem, I would like know ¿how I run my code in this context?\nThe context is convert data although standard scaler, and then apply PCA, or similar \nWhat's your possibly solution to this problem?",
    "3222029": "Do you need PCA here? @danielphys",
    "3222106": "In my opinion, I need PCA to reduce dimensionality with all columns, or what's the best form to reduce dimensionality in this dataset? @ravi20076",
    "3222351": "You can use PCA here. Two important things you need to do before calling PCA() are:\n1. Reduce memory usage by data frames. You can use function like  reduce_mem_usage. (Available in Kaggle public notebooks).\n2. some columns contain -inifinity values. You can replace it with large negative numbers like -9999999.\ntrain.replace([-np.inf], -999999999.0, inplace=True)\ntest.replace([-np.inf], -999999999.0, inplace=True)\n\nDo share the findings of your research, please. All of us are eager to find out how PCA may improve outcomes.",
    "3222355": "Think about it - PCA has its own flaws and hardly works in most cases @danielphys",
    "3222411": "Agreed, I rarely see PCA used by top players. I would select the most important 50-100 columns to speedup testing instead of PCA.",
    "3222773": "ok",
    "3223141": "I remember PCA being very actively used in [this competition](https://www.kaggle.com/competitions/lish-moa/overview), and it kind of worked. Check top write-ups like place 1st, 2nd, 4th, and so on. Maybe you just need to use more NN with PCA features rather than GBDT.",
    "3223358": "Use the incremental PCA, in the scikit learn \nfrom sklearn.decomposition import IncrementalPCA\nthis takes a few seconds and fits well",
    "3226964": "Hi! I ran into similar memory issues when applying StandardScaler followed by PCA on large datasets — especially with float-heavy parquet files. Here are a few techniques that helped me handle it efficiently:\n 1. Use Incremental PCA\nInstead of standard PCA, use a batch-based or incremental approach that processes data in chunks. It dramatically reduces memory overhead.\n 2. Scale in Mini-Batches\nRather than scaling the full dataset at once, fit and transform your data using mini-batches. This is especially helpful when memory is tight.\n 3. Convert to float32 Early\nCasting your data to float32 instead of the default float64 can save nearly 50% memory. It's a simple but often overlooked optimization.\n 4. Apply Feature Selection Before PCA\nUse techniques like SelectKBest, variance thresholds, or correlation filtering to reduce dimensionality before scaling. PCA works better (and faster) on a leaner dataset.\n 5. Consider Saving Intermediate Outputs\nIf your process fails midway, save checkpoints at key stages like after scaling or feature selection. This allows you to resume rather than restart.\n 6. if youre running the code locally and its your computer hardware holding you back, consider creating a Kaggle notebook on the website as Kaggle offers you sufficient ram given you optimise your code.\nThese approaches allowed me to run full pipelines even on limited RAM Kaggle environments. Hope this helps!!"
  },
  "source": "meta"
}