{
  "id": 350447,
  "title": "Multiome target pca",
  "url": "/competitions/open-problems-multimodal/discussion/350447",
  "author_name": "",
  "post_date": "2022-09-05T18:30:15.971228100Z",
  "votes": 21,
  "comment_count": 4,
  "views": 0,
  "content": "<p>EDIT, using truncated SVD and the inverse rotation matrix on the train targets it appears we can get a 0.68 correlation if we have perfect known targets using just n=50 components. This is a better result than I showed initially which used just a sample of the training targets with PCA. This suggests we can achieve even better scores with n &gt; 50. This is actually done in one of the better scoring notebooks, training a model to predict 128 truncated SVD multiome target components then projecting this back out to the 23,418 feature space using the truncated SVD process in reverse.</p>\n<hr>\n<p>There are 23,418 target variables in the Multiome dataset.  I think we can reduce these with PCA to get fewer targets.  Then train fewer models.  This may be helpful in the case of models that aren't good at handling multi-outputs. Then we can use the inverse of the PCA rotation matrix to blow this back out to 23,418 targets.  With known targets, I believe the maximum score you can get using this approach with the following number of principal components seems to be the following, just picked a selection of n # of components:</p>\n<p>n = 10, correlation score = 0.5726, variance explained = 8.66%<br>\nn = 340, correlation score = 0.6529, variance explained = 17.02%<br>\nn = 500, correlation score = 0.6629, variance explained = 20.02%<br>\nn = 3000, correlation score = 0.7822, variance explained = 56.62%<br>\nn = 23418, correlation score = 1.00, variance explained = 100%</p>\n<p>writing out this formula helped me:</p>\n<p>targets * rotation_matrix = target_pca_components<br>\ntarget_pca_components[,1:n]*rotation_matrix_invers[1:n,]=blown_out_targets<br>\n(to multiply matrices i used torch_mm in the torch package)</p>\n<p>Disclaimer, I've never used unsupervised learning, hopefully it wouldn't cause leakage.  Would be interested in others' thoughts on this approach.  This approach could be super standard, sorry if that's the case.</p>\n<p>Note, PCA on 23,418 is computationally expensive so I used my RTX 3090 GPU using the qrpca package in R, which uses torch with gpu acceleration on the back end (not necessary to actually use qrpca because the formula is rather simple in just torch if you look at the qrpca code).  I was able to sample 28,000 rows of multi targets, 23,418 columns which fit in my GPU memory.  It took about 11 mins to run at near 100% GPU usage.  Note the variance explained quoted above is really the variance explained on the 28,000 sample not the full train set, the correlation score however was based on the entire train target set.</p>\n<p>I got this idea from one of the solutions discussed in the workshop video from when this competition was previously hosted by 2021 NeurIPS, thank you Jiwei Liu for sharing. <a href=\"https://drive.google.com/file/d/1aQss-KyfYlzdrBQcH5joiXMlTwpG5gdf/view?usp=sharing\" target=\"_blank\">https://drive.google.com/file/d/1aQss-KyfYlzdrBQcH5joiXMlTwpG5gdf/view?usp=sharing</a>  I believe they showed a step on growing the dimensionality on a model output, but I do think details were lacking.</p>",
  "messages": [
    {
      "id": "1927564",
      "postDate": "09/05/2022 18:30:15",
      "content": "<p>EDIT, using truncated SVD and the inverse rotation matrix on the train targets it appears we can get a 0.68 correlation if we have perfect known targets using just n=50 components. This is a better result than I showed initially which used just a sample of the training targets with PCA. This suggests we can achieve even better scores with n &gt; 50. This is actually done in one of the better scoring notebooks, training a model to predict 128 truncated SVD multiome target components then projecting this back out to the 23,418 feature space using the truncated SVD process in reverse.</p>\n<hr>\n<p>There are 23,418 target variables in the Multiome dataset.  I think we can reduce these with PCA to get fewer targets.  Then train fewer models.  This may be helpful in the case of models that aren't good at handling multi-outputs. Then we can use the inverse of the PCA rotation matrix to blow this back out to 23,418 targets.  With known targets, I believe the maximum score you can get using this approach with the following number of principal components seems to be the following, just picked a selection of n # of components:</p>\n<p>n = 10, correlation score = 0.5726, variance explained = 8.66%<br>\nn = 340, correlation score = 0.6529, variance explained = 17.02%<br>\nn = 500, correlation score = 0.6629, variance explained = 20.02%<br>\nn = 3000, correlation score = 0.7822, variance explained = 56.62%<br>\nn = 23418, correlation score = 1.00, variance explained = 100%</p>\n<p>writing out this formula helped me:</p>\n<p>targets * rotation_matrix = target_pca_components<br>\ntarget_pca_components[,1:n]*rotation_matrix_invers[1:n,]=blown_out_targets<br>\n(to multiply matrices i used torch_mm in the torch package)</p>\n<p>Disclaimer, I've never used unsupervised learning, hopefully it wouldn't cause leakage.  Would be interested in others' thoughts on this approach.  This approach could be super standard, sorry if that's the case.</p>\n<p>Note, PCA on 23,418 is computationally expensive so I used my RTX 3090 GPU using the qrpca package in R, which uses torch with gpu acceleration on the back end (not necessary to actually use qrpca because the formula is rather simple in just torch if you look at the qrpca code).  I was able to sample 28,000 rows of multi targets, 23,418 columns which fit in my GPU memory.  It took about 11 mins to run at near 100% GPU usage.  Note the variance explained quoted above is really the variance explained on the 28,000 sample not the full train set, the correlation score however was based on the entire train target set.</p>\n<p>I got this idea from one of the solutions discussed in the workshop video from when this competition was previously hosted by 2021 NeurIPS, thank you Jiwei Liu for sharing. <a href=\"https://drive.google.com/file/d/1aQss-KyfYlzdrBQcH5joiXMlTwpG5gdf/view?usp=sharing\" target=\"_blank\">https://drive.google.com/file/d/1aQss-KyfYlzdrBQcH5joiXMlTwpG5gdf/view?usp=sharing</a>  I believe they showed a step on growing the dimensionality on a model output, but I do think details were lacking.</p>",
      "rawMarkdown": "EDIT, using truncated SVD and the inverse rotation matrix on the train targets it appears we can get a 0.68 correlation if we have perfect known targets using just n=50 components. This is a better result than I showed initially which used just a sample of the training targets with PCA. This suggests we can achieve even better scores with n > 50. This is actually done in one of the better scoring notebooks, training a model to predict 128 truncated SVD multiome target components then projecting this back out to the 23,418 feature space using the truncated SVD process in reverse.\n\n--------------------------------------------------------------------------------------\n\nThere are 23,418 target variables in the Multiome dataset.  I think we can reduce these with PCA to get fewer targets.  Then train fewer models.  This may be helpful in the case of models that aren't good at handling multi-outputs. Then we can use the inverse of the PCA rotation matrix to blow this back out to 23,418 targets.  With known targets, I believe the maximum score you can get using this approach with the following number of principal components seems to be the following, just picked a selection of n # of components:\n\nn = 10, correlation score = 0.5726, variance explained = 8.66%\nn = 340, correlation score = 0.6529, variance explained = 17.02%\nn = 500, correlation score = 0.6629, variance explained = 20.02%\nn = 3000, correlation score = 0.7822, variance explained = 56.62%\nn = 23418, correlation score = 1.00, variance explained = 100%\n\nwriting out this formula helped me:\n\ntargets * rotation_matrix = target_pca_components\ntarget_pca_components[,1:n]*rotation_matrix_invers[1:n,]=blown_out_targets\n(to multiply matrices i used torch_mm in the torch package)\n\nDisclaimer, I've never used unsupervised learning, hopefully it wouldn't cause leakage.  Would be interested in others' thoughts on this approach.  This approach could be super standard, sorry if that's the case.\n\nNote, PCA on 23,418 is computationally expensive so I used my RTX 3090 GPU using the qrpca package in R, which uses torch with gpu acceleration on the back end (not necessary to actually use qrpca because the formula is rather simple in just torch if you look at the qrpca code).  I was able to sample 28,000 rows of multi targets, 23,418 columns which fit in my GPU memory.  It took about 11 mins to run at near 100% GPU usage.  Note the variance explained quoted above is really the variance explained on the 28,000 sample not the full train set, the correlation score however was based on the entire train target set.\n\nI got this idea from one of the solutions discussed in the workshop video from when this competition was previously hosted by 2021 NeurIPS, thank you Jiwei Liu for sharing. https://drive.google.com/file/d/1aQss-KyfYlzdrBQcH5joiXMlTwpG5gdf/view?usp=sharing  I believe they showed a step on growing the dimensionality on a model output, but I do think details were lacking.",
      "votes": null
    },
    {
      "id": "1931031",
      "postDate": "09/08/2022 12:37:24",
      "content": "<p>I tried doing that.. Reducing the target variables using PCA, predicting and then reconstructing to the original number of dimensions.</p>\n<p>What I have found is that it's quite a bit harder to predict the value of the principal components as it is to predict the raw targets.. </p>\n<p>I think that makes some sense to me as it's more difficult to predict a value that is a linear combination of several variables, as opposed to predicting a single variable.</p>\n<p>What do you think?</p>",
      "rawMarkdown": "I tried doing that.. Reducing the target variables using PCA, predicting and then reconstructing to the original number of dimensions.\n\nWhat I have found is that it's quite a bit harder to predict the value of the principal components as it is to predict the raw targets.. \n\nI think that makes some sense to me as it's more difficult to predict a value that is a linear combination of several variables, as opposed to predicting a single variable.\n\nWhat do you think?",
      "votes": null
    },
    {
      "id": "1931789",
      "postDate": "09/09/2022 03:48:54",
      "content": "<p>I actually haven't tried yet.  I was deterred by the number of perfectly predicted targets required (500+) to produce a good score. So far I have just been using neural networks which seem well suited for so many targets.  I trust that it's hard to predict linear combinations of multiome variables, but it's also hard to predict multiome.  I intend to come back to this decision when I try my hand at gradient boosted trees.  Thanks for replying, definitely good to hear your experience in predicting pca components.</p>",
      "rawMarkdown": "I actually haven't tried yet.  I was deterred by the number of perfectly predicted targets required (500+) to produce a good score. So far I have just been using neural networks which seem well suited for so many targets.  I trust that it's hard to predict linear combinations of multiome variables, but it's also hard to predict multiome.  I intend to come back to this decision when I try my hand at gradient boosted trees.  Thanks for replying, definitely good to hear your experience in predicting pca components.",
      "votes": null
    },
    {
      "id": "1935353",
      "postDate": "09/12/2022 04:06:18",
      "content": "<p>Using truncated SVD and the inverse rotation matrix on the train targets it appears we can get a 0.68 correlation if we have perfect known targets using just n=50 components.   This is a better result than I showed initially which used just a sample of the training targets with PCA.  This suggests we can achieve even better scores with n &gt; 50.  This is actually done in one of the better scoring notebooks, training a model to predict 128 truncated SVD multiome target components then projecting this back out to the 23,418 feature space using the truncated SVD process in reverse.</p>",
      "rawMarkdown": "Using truncated SVD and the inverse rotation matrix on the train targets it appears we can get a 0.68 correlation if we have perfect known targets using just n=50 components.   This is a better result than I showed initially which used just a sample of the training targets with PCA.  This suggests we can achieve even better scores with n > 50.  This is actually done in one of the better scoring notebooks, training a model to predict 128 truncated SVD multiome target components then projecting this back out to the 23,418 feature space using the truncated SVD process in reverse.",
      "votes": null
    },
    {
      "id": "1941701",
      "postDate": "09/16/2022 08:44:11",
      "content": "<p>thanks for sharing, i will try it.</p>",
      "rawMarkdown": "thanks for sharing, i will try it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1931031,
      "author_name": "thomasuriot",
      "author_url": "",
      "post_date": "09/08/2022 12:37:24",
      "content": "<p>I tried doing that.. Reducing the target variables using PCA, predicting and then reconstructing to the original number of dimensions.</p>\n<p>What I have found is that it's quite a bit harder to predict the value of the principal components as it is to predict the raw targets.. </p>\n<p>I think that makes some sense to me as it's more difficult to predict a value that is a linear combination of several variables, as opposed to predicting a single variable.</p>\n<p>What do you think?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1931789,
          "author_name": "andyatkinson",
          "author_url": "",
          "post_date": "09/09/2022 03:48:54",
          "content": "<p>I actually haven't tried yet.  I was deterred by the number of perfectly predicted targets required (500+) to produce a good score. So far I have just been using neural networks which seem well suited for so many targets.  I trust that it's hard to predict linear combinations of multiome variables, but it's also hard to predict multiome.  I intend to come back to this decision when I try my hand at gradient boosted trees.  Thanks for replying, definitely good to hear your experience in predicting pca components.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1935353,
      "author_name": "andyatkinson",
      "author_url": "",
      "post_date": "09/12/2022 04:06:18",
      "content": "<p>Using truncated SVD and the inverse rotation matrix on the train targets it appears we can get a 0.68 correlation if we have perfect known targets using just n=50 components.   This is a better result than I showed initially which used just a sample of the training targets with PCA.  This suggests we can achieve even better scores with n &gt; 50.  This is actually done in one of the better scoring notebooks, training a model to predict 128 truncated SVD multiome target components then projecting this back out to the 23,418 feature space using the truncated SVD process in reverse.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1941701,
      "author_name": "ricopue",
      "author_url": "",
      "post_date": "09/16/2022 08:44:11",
      "content": "<p>thanks for sharing, i will try it.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1927564": "EDIT, using truncated SVD and the inverse rotation matrix on the train targets it appears we can get a 0.68 correlation if we have perfect known targets using just n=50 components. This is a better result than I showed initially which used just a sample of the training targets with PCA. This suggests we can achieve even better scores with n > 50. This is actually done in one of the better scoring notebooks, training a model to predict 128 truncated SVD multiome target components then projecting this back out to the 23,418 feature space using the truncated SVD process in reverse.\n\n--------------------------------------------------------------------------------------\n\nThere are 23,418 target variables in the Multiome dataset.  I think we can reduce these with PCA to get fewer targets.  Then train fewer models.  This may be helpful in the case of models that aren't good at handling multi-outputs. Then we can use the inverse of the PCA rotation matrix to blow this back out to 23,418 targets.  With known targets, I believe the maximum score you can get using this approach with the following number of principal components seems to be the following, just picked a selection of n # of components:\n\nn = 10, correlation score = 0.5726, variance explained = 8.66%\nn = 340, correlation score = 0.6529, variance explained = 17.02%\nn = 500, correlation score = 0.6629, variance explained = 20.02%\nn = 3000, correlation score = 0.7822, variance explained = 56.62%\nn = 23418, correlation score = 1.00, variance explained = 100%\n\nwriting out this formula helped me:\n\ntargets * rotation_matrix = target_pca_components\ntarget_pca_components[,1:n]*rotation_matrix_invers[1:n,]=blown_out_targets\n(to multiply matrices i used torch_mm in the torch package)\n\nDisclaimer, I've never used unsupervised learning, hopefully it wouldn't cause leakage.  Would be interested in others' thoughts on this approach.  This approach could be super standard, sorry if that's the case.\n\nNote, PCA on 23,418 is computationally expensive so I used my RTX 3090 GPU using the qrpca package in R, which uses torch with gpu acceleration on the back end (not necessary to actually use qrpca because the formula is rather simple in just torch if you look at the qrpca code).  I was able to sample 28,000 rows of multi targets, 23,418 columns which fit in my GPU memory.  It took about 11 mins to run at near 100% GPU usage.  Note the variance explained quoted above is really the variance explained on the 28,000 sample not the full train set, the correlation score however was based on the entire train target set.\n\nI got this idea from one of the solutions discussed in the workshop video from when this competition was previously hosted by 2021 NeurIPS, thank you Jiwei Liu for sharing. https://drive.google.com/file/d/1aQss-KyfYlzdrBQcH5joiXMlTwpG5gdf/view?usp=sharing  I believe they showed a step on growing the dimensionality on a model output, but I do think details were lacking.",
    "1931031": "I tried doing that.. Reducing the target variables using PCA, predicting and then reconstructing to the original number of dimensions.\n\nWhat I have found is that it's quite a bit harder to predict the value of the principal components as it is to predict the raw targets.. \n\nI think that makes some sense to me as it's more difficult to predict a value that is a linear combination of several variables, as opposed to predicting a single variable.\n\nWhat do you think?",
    "1931789": "I actually haven't tried yet.  I was deterred by the number of perfectly predicted targets required (500+) to produce a good score. So far I have just been using neural networks which seem well suited for so many targets.  I trust that it's hard to predict linear combinations of multiome variables, but it's also hard to predict multiome.  I intend to come back to this decision when I try my hand at gradient boosted trees.  Thanks for replying, definitely good to hear your experience in predicting pca components.",
    "1935353": "Using truncated SVD and the inverse rotation matrix on the train targets it appears we can get a 0.68 correlation if we have perfect known targets using just n=50 components.   This is a better result than I showed initially which used just a sample of the training targets with PCA.  This suggests we can achieve even better scores with n > 50.  This is actually done in one of the better scoring notebooks, training a model to predict 128 truncated SVD multiome target components then projecting this back out to the 23,418 feature space using the truncated SVD process in reverse.",
    "1941701": "thanks for sharing, i will try it."
  },
  "source": "meta"
}