{
  "id": 348233,
  "title": "Tips on Dimensionality Reduction",
  "url": "/competitions/open-problems-multimodal/discussion/348233",
  "author_name": "",
  "post_date": "2022-08-27T15:46:40.791260200Z",
  "votes": 28,
  "comment_count": 3,
  "views": 0,
  "content": "<p>The datasets in this competition have a lot of dimensions to say the least. The large number of dimensions logically makes sense since we are working with DNA, RNA, and Protein data which are often in the form of long sequences. This creates the important question: how can we decrease the number of dimensions to make the dataset easier to work with.</p>\n<h1>Handle Zeros</h1>\n<p>If you've looked at the dataset, you will notice that is is an extremely large number of zeros. There are even entire columns that only consist of zeros. </p>\n<p>Here's a tip on how to remove them.</p>\n<pre><code>all_zero_columns = (X == 0).all(axis=0)\nX = X[:,~all_zero_columns]\n</code></pre>\n<h1>Column Selection</h1>\n<p>You can also just use fewer columns. You can decide which columns to chose in a variety of ways. You can select columns randomly, but you might not get good results. You could decide using EDA. I recommend researching more advanced feature selection strategies for this type of method.</p>\n<p>Very simple method from quickstart notebook:</p>\n<pre><code>X[:,columns_to_use:]\n</code></pre>\n<h1>PCA</h1>\n<p>Principal Component Analysis (PCA) is a linear dimensionality reduction using Singular Value Decomposition of the data to project it to a lower dimensional space. PCA compresses data.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2Fb133a9ea35b1bd5f310e6166cc488527%2Fpca_pic.png?generation=1661613675406276&amp;alt=media\" alt=\"\"></p>\n<p>Example code:</p>\n<pre><code>from sklearn.decomposition import PCA\npca = PCA(n_components=n)\nX = pca.fit_transform(X)\n</code></pre>\n<h1>ICA</h1>\n<p>Independent Component Analysis (ICA) finds which vectors are independent sub-elements of your data. In other words, PCA helps to compress data and ICA helps to separate data.</p>\n<p>Example code:</p>\n<pre><code>from sklearn.decomposition import FastICA\nica = FastICA(n_components=n)\nX = ica.fit_transform(X)\n</code></pre>\n<h1>t-SNE</h1>\n<p>t-SNE is a unsupervised non-linear dimensionality reduction and data visualization technique. This could be another alternative to PCA.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2F5ad205ea489abd080be2fc863f7c570c%2Ftsne_pic%20(1).png?generation=1661615779408656&amp;alt=media\" alt=\"\"></p>\n<p>Example code:</p>\n<pre><code>from sklearn.manifold import TSNE\ntsne = TSNE(n_components=n)\nX = tsne.fit_transform(X)\n</code></pre>\n<h1>Ivis</h1>\n<p><a href=\"https://pubs.acs.org/doi/abs/10.1021/acs.jcim.0c00485\" target=\"_blank\">As you can read from this paper</a>, ivis has promising applications for biomolecular tasks. It uses a Siamese Neural Network to create embeddings and lower the number of dimensions. As a result, I predict it could have good applications in this challenge.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2Fcae2436ca9b943455f46964dc744f133%2Fivis_pic.gif?generation=1661612269532317&amp;alt=media\" alt=\"\"></p>\n<p>Example code:</p>\n<pre><code>from ivis import Ivis\nmodel = Ivis(embedding_dims=dims, k=k, batch_size=bs, epochs=ep, n_trees=n_trees)\nX = model.fit_transform(X)\n</code></pre>\n<h1>References</h1>\n<p><a href=\"https://www.kaggle.com/code/ambrosm/msci-citeseq-quickstart\" target=\"_blank\">quickstart</a> by <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a><br>\n<a href=\"https://github.com/adavoudi/msci_knn\" target=\"_blank\">knn starter</a> by <a href=\"https://www.kaggle.com/alireza9\" target=\"_blank\">@alireza9</a><br>\n<a href=\"https://scikit-learn.org/stable/modules/decomposition.html\" target=\"_blank\">sklearn decomposition</a> docs by sklearn<br>\n<a href=\"https://www.geeksforgeeks.org/difference-between-pca-vs-t-sne/\" target=\"_blank\">t-SNE info</a> on geeks for geeks site<br>\n<a href=\"https://pubs.acs.org/doi/abs/10.1021/acs.jcim.0c00485\" target=\"_blank\">ivis paper</a> by Hao Tian and Peng Tao</p>",
  "messages": [
    {
      "id": "1916067",
      "postDate": "08/27/2022 15:46:40",
      "content": "<p>The datasets in this competition have a lot of dimensions to say the least. The large number of dimensions logically makes sense since we are working with DNA, RNA, and Protein data which are often in the form of long sequences. This creates the important question: how can we decrease the number of dimensions to make the dataset easier to work with.</p>\n<h1>Handle Zeros</h1>\n<p>If you've looked at the dataset, you will notice that is is an extremely large number of zeros. There are even entire columns that only consist of zeros. </p>\n<p>Here's a tip on how to remove them.</p>\n<pre><code>all_zero_columns = (X == 0).all(axis=0)\nX = X[:,~all_zero_columns]\n</code></pre>\n<h1>Column Selection</h1>\n<p>You can also just use fewer columns. You can decide which columns to chose in a variety of ways. You can select columns randomly, but you might not get good results. You could decide using EDA. I recommend researching more advanced feature selection strategies for this type of method.</p>\n<p>Very simple method from quickstart notebook:</p>\n<pre><code>X[:,columns_to_use:]\n</code></pre>\n<h1>PCA</h1>\n<p>Principal Component Analysis (PCA) is a linear dimensionality reduction using Singular Value Decomposition of the data to project it to a lower dimensional space. PCA compresses data.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2Fb133a9ea35b1bd5f310e6166cc488527%2Fpca_pic.png?generation=1661613675406276&amp;alt=media\" alt=\"\"></p>\n<p>Example code:</p>\n<pre><code>from sklearn.decomposition import PCA\npca = PCA(n_components=n)\nX = pca.fit_transform(X)\n</code></pre>\n<h1>ICA</h1>\n<p>Independent Component Analysis (ICA) finds which vectors are independent sub-elements of your data. In other words, PCA helps to compress data and ICA helps to separate data.</p>\n<p>Example code:</p>\n<pre><code>from sklearn.decomposition import FastICA\nica = FastICA(n_components=n)\nX = ica.fit_transform(X)\n</code></pre>\n<h1>t-SNE</h1>\n<p>t-SNE is a unsupervised non-linear dimensionality reduction and data visualization technique. This could be another alternative to PCA.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2F5ad205ea489abd080be2fc863f7c570c%2Ftsne_pic%20(1).png?generation=1661615779408656&amp;alt=media\" alt=\"\"></p>\n<p>Example code:</p>\n<pre><code>from sklearn.manifold import TSNE\ntsne = TSNE(n_components=n)\nX = tsne.fit_transform(X)\n</code></pre>\n<h1>Ivis</h1>\n<p><a href=\"https://pubs.acs.org/doi/abs/10.1021/acs.jcim.0c00485\" target=\"_blank\">As you can read from this paper</a>, ivis has promising applications for biomolecular tasks. It uses a Siamese Neural Network to create embeddings and lower the number of dimensions. As a result, I predict it could have good applications in this challenge.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2Fcae2436ca9b943455f46964dc744f133%2Fivis_pic.gif?generation=1661612269532317&amp;alt=media\" alt=\"\"></p>\n<p>Example code:</p>\n<pre><code>from ivis import Ivis\nmodel = Ivis(embedding_dims=dims, k=k, batch_size=bs, epochs=ep, n_trees=n_trees)\nX = model.fit_transform(X)\n</code></pre>\n<h1>References</h1>\n<p><a href=\"https://www.kaggle.com/code/ambrosm/msci-citeseq-quickstart\" target=\"_blank\">quickstart</a> by <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a><br>\n<a href=\"https://github.com/adavoudi/msci_knn\" target=\"_blank\">knn starter</a> by <a href=\"https://www.kaggle.com/alireza9\" target=\"_blank\">@alireza9</a><br>\n<a href=\"https://scikit-learn.org/stable/modules/decomposition.html\" target=\"_blank\">sklearn decomposition</a> docs by sklearn<br>\n<a href=\"https://www.geeksforgeeks.org/difference-between-pca-vs-t-sne/\" target=\"_blank\">t-SNE info</a> on geeks for geeks site<br>\n<a href=\"https://pubs.acs.org/doi/abs/10.1021/acs.jcim.0c00485\" target=\"_blank\">ivis paper</a> by Hao Tian and Peng Tao</p>",
      "rawMarkdown": "The datasets in this competition have a lot of dimensions to say the least. The large number of dimensions logically makes sense since we are working with DNA, RNA, and Protein data which are often in the form of long sequences. This creates the important question: how can we decrease the number of dimensions to make the dataset easier to work with.\n\n# Handle Zeros\n\nIf you've looked at the dataset, you will notice that is is an extremely large number of zeros. There are even entire columns that only consist of zeros. \n\nHere's a tip on how to remove them.\n```\nall_zero_columns = (X == 0).all(axis=0)\nX = X[:,~all_zero_columns]\n```\n\n# Column Selection\n\nYou can also just use fewer columns. You can decide which columns to chose in a variety of ways. You can select columns randomly, but you might not get good results. You could decide using EDA. I recommend researching more advanced feature selection strategies for this type of method.\n\nVery simple method from quickstart notebook:\n```\nX[:,columns_to_use:]\n```\n\n# PCA\n\nPrincipal Component Analysis (PCA) is a linear dimensionality reduction using Singular Value Decomposition of the data to project it to a lower dimensional space. PCA compresses data.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2Fb133a9ea35b1bd5f310e6166cc488527%2Fpca_pic.png?generation=1661613675406276&alt=media)\n\nExample code:\n```\nfrom sklearn.decomposition import PCA\npca = PCA(n_components=n)\nX = pca.fit_transform(X)\n```\n\n# ICA\n\nIndependent Component Analysis (ICA) finds which vectors are independent sub-elements of your data. In other words, PCA helps to compress data and ICA helps to separate data.\n\nExample code:\n```\nfrom sklearn.decomposition import FastICA\nica = FastICA(n_components=n)\nX = ica.fit_transform(X)\n```\n\n# t-SNE\n\nt-SNE is a unsupervised non-linear dimensionality reduction and data visualization technique. This could be another alternative to PCA.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2F5ad205ea489abd080be2fc863f7c570c%2Ftsne_pic%20(1).png?generation=1661615779408656&alt=media)\n\nExample code:\n```\nfrom sklearn.manifold import TSNE\ntsne = TSNE(n_components=n)\nX = tsne.fit_transform(X)\n```\n\n# Ivis\n\n[As you can read from this paper](https://pubs.acs.org/doi/abs/10.1021/acs.jcim.0c00485), ivis has promising applications for biomolecular tasks. It uses a Siamese Neural Network to create embeddings and lower the number of dimensions. As a result, I predict it could have good applications in this challenge.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2Fcae2436ca9b943455f46964dc744f133%2Fivis_pic.gif?generation=1661612269532317&alt=media)\n\nExample code:\n```\nfrom ivis import Ivis\nmodel = Ivis(embedding_dims=dims, k=k, batch_size=bs, epochs=ep, n_trees=n_trees)\nX = model.fit_transform(X)\n```\n\n# References\n[quickstart](https://www.kaggle.com/code/ambrosm/msci-citeseq-quickstart) by @ambrosm\n[knn starter](https://github.com/adavoudi/msci_knn) by @alireza9\n[sklearn decomposition](https://scikit-learn.org/stable/modules/decomposition.html) docs by sklearn\n[t-SNE info](https://www.geeksforgeeks.org/difference-between-pca-vs-t-sne/) on geeks for geeks site\n[ivis paper](https://pubs.acs.org/doi/abs/10.1021/acs.jcim.0c00485) by Hao Tian and Peng Tao",
      "votes": null
    },
    {
      "id": "1916303",
      "postDate": "08/27/2022 18:29:20",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/ravishah1\" target=\"_blank\">@ravishah1</a>, for introducing ivis</p>",
      "rawMarkdown": "Thanks @ravishah1, for introducing ivis",
      "votes": null
    },
    {
      "id": "1916572",
      "postDate": "08/28/2022 02:42:35",
      "content": "<p><a href=\"https://www.kaggle.com/anityagangurde\" target=\"_blank\">@anityagangurde</a> I hope it helps! Also, I would definitely check out the second link in the references as that is where I first saw ivis used.</p>",
      "rawMarkdown": "anityagangurde I hope it helps! Also, I would definitely check out the second link in the references as that is where I first saw ivis used.",
      "votes": null
    },
    {
      "id": "1916659",
      "postDate": "08/28/2022 04:30:20",
      "content": "<p>Yes sure. Will check it out. Thanks.</p>",
      "rawMarkdown": "Yes sure. Will check it out. Thanks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1916303,
      "author_name": "anityagangurde",
      "author_url": "",
      "post_date": "08/27/2022 18:29:20",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/ravishah1\" target=\"_blank\">@ravishah1</a>, for introducing ivis</p>",
      "votes": null,
      "replies": [
        {
          "id": 1916572,
          "author_name": "ravishah1",
          "author_url": "",
          "post_date": "08/28/2022 02:42:35",
          "content": "<p><a href=\"https://www.kaggle.com/anityagangurde\" target=\"_blank\">@anityagangurde</a> I hope it helps! Also, I would definitely check out the second link in the references as that is where I first saw ivis used.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1916659,
          "author_name": "anityagangurde",
          "author_url": "",
          "post_date": "08/28/2022 04:30:20",
          "content": "<p>Yes sure. Will check it out. Thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1916067": "The datasets in this competition have a lot of dimensions to say the least. The large number of dimensions logically makes sense since we are working with DNA, RNA, and Protein data which are often in the form of long sequences. This creates the important question: how can we decrease the number of dimensions to make the dataset easier to work with.\n\n# Handle Zeros\n\nIf you've looked at the dataset, you will notice that is is an extremely large number of zeros. There are even entire columns that only consist of zeros. \n\nHere's a tip on how to remove them.\n```\nall_zero_columns = (X == 0).all(axis=0)\nX = X[:,~all_zero_columns]\n```\n\n# Column Selection\n\nYou can also just use fewer columns. You can decide which columns to chose in a variety of ways. You can select columns randomly, but you might not get good results. You could decide using EDA. I recommend researching more advanced feature selection strategies for this type of method.\n\nVery simple method from quickstart notebook:\n```\nX[:,columns_to_use:]\n```\n\n# PCA\n\nPrincipal Component Analysis (PCA) is a linear dimensionality reduction using Singular Value Decomposition of the data to project it to a lower dimensional space. PCA compresses data.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2Fb133a9ea35b1bd5f310e6166cc488527%2Fpca_pic.png?generation=1661613675406276&alt=media)\n\nExample code:\n```\nfrom sklearn.decomposition import PCA\npca = PCA(n_components=n)\nX = pca.fit_transform(X)\n```\n\n# ICA\n\nIndependent Component Analysis (ICA) finds which vectors are independent sub-elements of your data. In other words, PCA helps to compress data and ICA helps to separate data.\n\nExample code:\n```\nfrom sklearn.decomposition import FastICA\nica = FastICA(n_components=n)\nX = ica.fit_transform(X)\n```\n\n# t-SNE\n\nt-SNE is a unsupervised non-linear dimensionality reduction and data visualization technique. This could be another alternative to PCA.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2F5ad205ea489abd080be2fc863f7c570c%2Ftsne_pic%20(1).png?generation=1661615779408656&alt=media)\n\nExample code:\n```\nfrom sklearn.manifold import TSNE\ntsne = TSNE(n_components=n)\nX = tsne.fit_transform(X)\n```\n\n# Ivis\n\n[As you can read from this paper](https://pubs.acs.org/doi/abs/10.1021/acs.jcim.0c00485), ivis has promising applications for biomolecular tasks. It uses a Siamese Neural Network to create embeddings and lower the number of dimensions. As a result, I predict it could have good applications in this challenge.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6537187%2Fcae2436ca9b943455f46964dc744f133%2Fivis_pic.gif?generation=1661612269532317&alt=media)\n\nExample code:\n```\nfrom ivis import Ivis\nmodel = Ivis(embedding_dims=dims, k=k, batch_size=bs, epochs=ep, n_trees=n_trees)\nX = model.fit_transform(X)\n```\n\n# References\n[quickstart](https://www.kaggle.com/code/ambrosm/msci-citeseq-quickstart) by @ambrosm\n[knn starter](https://github.com/adavoudi/msci_knn) by @alireza9\n[sklearn decomposition](https://scikit-learn.org/stable/modules/decomposition.html) docs by sklearn\n[t-SNE info](https://www.geeksforgeeks.org/difference-between-pca-vs-t-sne/) on geeks for geeks site\n[ivis paper](https://pubs.acs.org/doi/abs/10.1021/acs.jcim.0c00485) by Hao Tian and Peng Tao",
    "1916303": "Thanks @ravishah1, for introducing ivis",
    "1916572": "anityagangurde I hope it helps! Also, I would definitely check out the second link in the references as that is where I first saw ivis used.",
    "1916659": "Yes sure. Will check it out. Thanks."
  },
  "source": "meta"
}