{
  "id": 348377,
  "title": "Sparse matrices: PCA doesn´t work, TruncatedSVD does",
  "url": "/competitions/open-problems-multimodal/discussion/348377",
  "author_name": "Juan Smith Perera",
  "post_date": "2022-08-28T05:53:10.446000",
  "votes": 16,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I am using the idea about sparse matrices from: <a href=\"https://www.kaggle.com/code/sbunzini/reduce-memory-usage-by-95-with-sparse-matrices/comments\" target=\"_blank\">https://www.kaggle.com/code/sbunzini/reduce-memory-usage-by-95-with-sparse-matrices/comments</a></p>\n<p>I was able to run the model for CV with a restricted number of rows (10K) althught the whole training multiome data needs only 10.5GB. But to run the model we need some form of dimensionality reduction as the number of columns &gt; 228K &gt; number of rows.</p>\n<p>PCA does not work with sparse matrices. So, I found out that TruncatedSVD does.</p>\n<p>Here is the code: </p>\n<p>from sklearn.decomposition import TruncatedSVD<br>\ntruncatedSVD = TruncatedSVD(4)<br>\nmulti_train_x = truncatedSVD.fit_transform(multi_train_x)</p>\n<p>I hope it helps</p>",
  "messages": [
    {
      "id": 1916722,
      "postDate": "2022-08-28T05:53:10.447Z",
      "content": "<p>I am using the idea about sparse matrices from: <a href=\"https://www.kaggle.com/code/sbunzini/reduce-memory-usage-by-95-with-sparse-matrices/comments\" target=\"_blank\">https://www.kaggle.com/code/sbunzini/reduce-memory-usage-by-95-with-sparse-matrices/comments</a></p>\n<p>I was able to run the model for CV with a restricted number of rows (10K) althught the whole training multiome data needs only 10.5GB. But to run the model we need some form of dimensionality reduction as the number of columns &gt; 228K &gt; number of rows.</p>\n<p>PCA does not work with sparse matrices. So, I found out that TruncatedSVD does.</p>\n<p>Here is the code: </p>\n<p>from sklearn.decomposition import TruncatedSVD<br>\ntruncatedSVD = TruncatedSVD(4)<br>\nmulti_train_x = truncatedSVD.fit_transform(multi_train_x)</p>\n<p>I hope it helps</p>",
      "rawMarkdown": "I am using the idea about sparse matrices from: https://www.kaggle.com/code/sbunzini/reduce-memory-usage-by-95-with-sparse-matrices/comments\n\nI was able to run the model for CV with a restricted number of rows (10K) althught the whole training multiome data needs only 10.5GB. But to run the model we need some form of dimensionality reduction as the number of columns > 228K > number of rows.\n\nPCA does not work with sparse matrices. So, I found out that TruncatedSVD does.\n\nHere is the code: \n\nfrom sklearn.decomposition import TruncatedSVD\ntruncatedSVD = TruncatedSVD(4)\nmulti_train_x = truncatedSVD.fit_transform(multi_train_x)\n\nI hope it helps",
      "votes": 16
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1916722": "I am using the idea about sparse matrices from: https://www.kaggle.com/code/sbunzini/reduce-memory-usage-by-95-with-sparse-matrices/comments\n\nI was able to run the model for CV with a restricted number of rows (10K) althught the whole training multiome data needs only 10.5GB. But to run the model we need some form of dimensionality reduction as the number of columns > 228K > number of rows.\n\nPCA does not work with sparse matrices. So, I found out that TruncatedSVD does.\n\nHere is the code: \n\nfrom sklearn.decomposition import TruncatedSVD\ntruncatedSVD = TruncatedSVD(4)\nmulti_train_x = truncatedSVD.fit_transform(multi_train_x)\n\nI hope it helps"
  }
}