{
  "id": 349031,
  "title": "Sparse Matrices for MSCI: quickstart toolkit",
  "url": "/competitions/open-problems-multimodal/discussion/349031",
  "author_name": "",
  "post_date": "2022-08-31T00:54:34.762079400Z",
  "votes": 15,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>As many have pointed out, the data for this competition is rather sparse and thus can benefit from using a sparse matrix representation.</p>\n<p>To make it easier for the kagglers that are not familiars with the use of sparse matrices in ML, I have made public:</p>\n<ul>\n<li>A <a href=\"https://www.kaggle.com/datasets/fabiencrom/multimodal-single-cell-as-sparse-matrix\" target=\"_blank\">dataset</a> containing all the competition data in an efficient sparse format (i.e. <code>scipy</code>'s csr_matrix)</li>\n<li>The <a href=\"https://www.kaggle.com/code/fabiencrom/multimodal-single-cell-creating-sparse-data/\" target=\"_blank\">notebook</a> that generated this dataset</li>\n<li>A <a href=\"https://www.kaggle.com/code/fabiencrom/msci-multiome-quickstart-w-sparse-matrices\" target=\"_blank\">simple adaptation of AmbroseM'S quickstart notebook</a> to demonstrate a simple use of this sparse dataset</li>\n</ul>\n<p>When I have time, I plan to add a notebook showing how to use the sparse representation with torch and/or keras for a simple neural network.</p>\n<p>As I was writing this (I made the dataset 5 days ago, but had not taken the time to prepare it for publishing), I found out that some notebooks and discussions have already been published on this topic: <a href=\"https://www.kaggle.com/code/sbunzini/reduce-memory-usage-by-95-with-sparse-matrices/notebook#Compression-Rate\" target=\"_blank\">here</a>,  <a href=\"https://www.kaggle.com/code/jsmithperera/msci-multiome-sparse-tsvd\" target=\"_blank\">here</a>  and <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348377\" target=\"_blank\">here</a>. This will make some of the information provided slightly redundant, but hopefully it is still interesting.</p>\n<p><em>Additional informations</em>:</p>\n<p>There are many different formats for storing sparse matrices. The three most often used are COO, CSR and CSC:</p>\n<ul>\n<li>COO is the simplest to create and is often use as an intermediate form before converting to CSR or CSC. It is also the only format really supported by pytorch and CUDA.</li>\n<li>CSR is more memory efficient than COO and will also be more computation-efficient in many cases (at least on CPU, but not on GPU). Row-slicing is efficient on this format, which makes it a good choice for representing tabular data.</li>\n<li>CSC is similar to CSR, but implemented such that column-slicing is efficient.</li>\n</ul>\n<p>I am using CSR, as it is the most memory-efficient for this dataset. Note that <code>scipy</code>'s sparse matrices should be loaded with <code>scipy.sparse.load_npz</code> and not <code>numpy.load</code>.</p>\n<p>In the <a href=\"https://www.kaggle.com/datasets/fabiencrom/multimodal-single-cell-as-sparse-matrix\" target=\"_blank\">sparse CSR dataset</a>, the dataframes stored in each original <code>xxx.h5</code> files are converted into two files:</p>\n<ul>\n<li>One \"xxx_values.sparse\" file that can be loaded with <code>scipy.sparse.load_npz</code> and contains all the values of the corresponding dataframe (i.e. the result of <code>df.values</code> in a sparse CSR format)</li>\n<li>One \"xxx_idxcol.npz\" file that can be loaded with <code>numpy.load</code> and contains the values of the index and the columns of the corresponding dataframe (i.e the results of <code>df.index</code> and <code>df.columns</code> as normal arrays)</li>\n</ul>",
  "messages": [
    {
      "id": "1920165",
      "postDate": "08/31/2022 00:54:34",
      "content": "<p>Hi,</p>\n<p>As many have pointed out, the data for this competition is rather sparse and thus can benefit from using a sparse matrix representation.</p>\n<p>To make it easier for the kagglers that are not familiars with the use of sparse matrices in ML, I have made public:</p>\n<ul>\n<li>A <a href=\"https://www.kaggle.com/datasets/fabiencrom/multimodal-single-cell-as-sparse-matrix\" target=\"_blank\">dataset</a> containing all the competition data in an efficient sparse format (i.e. <code>scipy</code>'s csr_matrix)</li>\n<li>The <a href=\"https://www.kaggle.com/code/fabiencrom/multimodal-single-cell-creating-sparse-data/\" target=\"_blank\">notebook</a> that generated this dataset</li>\n<li>A <a href=\"https://www.kaggle.com/code/fabiencrom/msci-multiome-quickstart-w-sparse-matrices\" target=\"_blank\">simple adaptation of AmbroseM'S quickstart notebook</a> to demonstrate a simple use of this sparse dataset</li>\n</ul>\n<p>When I have time, I plan to add a notebook showing how to use the sparse representation with torch and/or keras for a simple neural network.</p>\n<p>As I was writing this (I made the dataset 5 days ago, but had not taken the time to prepare it for publishing), I found out that some notebooks and discussions have already been published on this topic: <a href=\"https://www.kaggle.com/code/sbunzini/reduce-memory-usage-by-95-with-sparse-matrices/notebook#Compression-Rate\" target=\"_blank\">here</a>,  <a href=\"https://www.kaggle.com/code/jsmithperera/msci-multiome-sparse-tsvd\" target=\"_blank\">here</a>  and <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348377\" target=\"_blank\">here</a>. This will make some of the information provided slightly redundant, but hopefully it is still interesting.</p>\n<p><em>Additional informations</em>:</p>\n<p>There are many different formats for storing sparse matrices. The three most often used are COO, CSR and CSC:</p>\n<ul>\n<li>COO is the simplest to create and is often use as an intermediate form before converting to CSR or CSC. It is also the only format really supported by pytorch and CUDA.</li>\n<li>CSR is more memory efficient than COO and will also be more computation-efficient in many cases (at least on CPU, but not on GPU). Row-slicing is efficient on this format, which makes it a good choice for representing tabular data.</li>\n<li>CSC is similar to CSR, but implemented such that column-slicing is efficient.</li>\n</ul>\n<p>I am using CSR, as it is the most memory-efficient for this dataset. Note that <code>scipy</code>'s sparse matrices should be loaded with <code>scipy.sparse.load_npz</code> and not <code>numpy.load</code>.</p>\n<p>In the <a href=\"https://www.kaggle.com/datasets/fabiencrom/multimodal-single-cell-as-sparse-matrix\" target=\"_blank\">sparse CSR dataset</a>, the dataframes stored in each original <code>xxx.h5</code> files are converted into two files:</p>\n<ul>\n<li>One \"xxx_values.sparse\" file that can be loaded with <code>scipy.sparse.load_npz</code> and contains all the values of the corresponding dataframe (i.e. the result of <code>df.values</code> in a sparse CSR format)</li>\n<li>One \"xxx_idxcol.npz\" file that can be loaded with <code>numpy.load</code> and contains the values of the index and the columns of the corresponding dataframe (i.e the results of <code>df.index</code> and <code>df.columns</code> as normal arrays)</li>\n</ul>",
      "rawMarkdown": "Hi,\n\nAs many have pointed out, the data for this competition is rather sparse and thus can benefit from using a sparse matrix representation.\n\nTo make it easier for the kagglers that are not familiars with the use of sparse matrices in ML, I have made public:\n- A [dataset](https://www.kaggle.com/datasets/fabiencrom/multimodal-single-cell-as-sparse-matrix) containing all the competition data in an efficient sparse format (i.e. `scipy`'s csr_matrix)\n- The [notebook](https://www.kaggle.com/code/fabiencrom/multimodal-single-cell-creating-sparse-data/) that generated this dataset\n- A [simple adaptation of AmbroseM'S quickstart notebook](https://www.kaggle.com/code/fabiencrom/msci-multiome-quickstart-w-sparse-matrices) to demonstrate a simple use of this sparse dataset\n\nWhen I have time, I plan to add a notebook showing how to use the sparse representation with torch and/or keras for a simple neural network.\n\nAs I was writing this (I made the dataset 5 days ago, but had not taken the time to prepare it for publishing), I found out that some notebooks and discussions have already been published on this topic: [here](https://www.kaggle.com/code/sbunzini/reduce-memory-usage-by-95-with-sparse-matrices/notebook#Compression-Rate),  [here](https://www.kaggle.com/code/jsmithperera/msci-multiome-sparse-tsvd)  and [here](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348377). This will make some of the information provided slightly redundant, but hopefully it is still interesting.\n\n*Additional informations*:\n\nThere are many different formats for storing sparse matrices. The three most often used are COO, CSR and CSC:\n- COO is the simplest to create and is often use as an intermediate form before converting to CSR or CSC. It is also the only format really supported by pytorch and CUDA.\n- CSR is more memory efficient than COO and will also be more computation-efficient in many cases (at least on CPU, but not on GPU). Row-slicing is efficient on this format, which makes it a good choice for representing tabular data.\n- CSC is similar to CSR, but implemented such that column-slicing is efficient.\n\nI am using CSR, as it is the most memory-efficient for this dataset. Note that `scipy`'s sparse matrices should be loaded with `scipy.sparse.load_npz` and not `numpy.load`.\n\nIn the [sparse CSR dataset](https://www.kaggle.com/datasets/fabiencrom/multimodal-single-cell-as-sparse-matrix), the dataframes stored in each original `xxx.h5` files are converted into two files:\n- One \"xxx_values.sparse\" file that can be loaded with `scipy.sparse.load_npz` and contains all the values of the corresponding dataframe (i.e. the result of `df.values` in a sparse CSR format)\n- One \"xxx_idxcol.npz\" file that can be loaded with `numpy.load` and contains the values of the index and the columns of the corresponding dataframe (i.e the results of `df.index` and `df.columns` as normal arrays)",
      "votes": null
    },
    {
      "id": "1920190",
      "postDate": "08/31/2022 01:42:59",
      "content": "<p>Actually, sbunzini had already provided a similar dataset 6 days ago <a href=\"https://www.kaggle.com/datasets/sbunzini/open-problems-msci-multiome-sparse-matrices\" target=\"_blank\">here</a>. I should have mentioned it above but I just noticed it. Sorry and kudos to him.</p>\n<p>Compared with sbuzini's dataset, the one I published was designed to be \"standalone\": it replicates all of the information in the original competition dataset, either in <code>.npz</code>, <code>.sparse</code> or <code>.parquet</code> format. (I prefered to convert everything once and for all).</p>\n<p>In practice, it means that:</p>\n<p>sbuzini's dataset only contains three files:</p>\n<ul>\n<li>test_multi_inputs_sparse.npz</li>\n<li>train_multi_targets_sparse.npz</li>\n<li>train_multi_inputs_sparse.npz</li>\n</ul>\n<p>which would correspond to the files:</p>\n<ul>\n<li>test_multi_inputs_values.sparse</li>\n<li>train_multi_targets_values.sparse</li>\n<li>train_multi_inputs_values.sparse</li>\n</ul>\n<p>in my dataset. </p>\n<p>My dataset also contain the files:</p>\n<ul>\n<li>test_multi_inputs_idxcol.npz</li>\n<li>train_multi_targets_idxcol.npz</li>\n<li>train_multi_inputs_idxcol.npz<br>\nthat gives the content of the index and column names for each original dataframe.</li>\n</ul>\n<p>In addition, my dataset also contains the corresponding files for the CITEseq dataframes (they are less sparse, but still compress reasonably well). As well as the metadat/submission info in parquet format.</p>",
      "rawMarkdown": "Actually, sbunzini had already provided a similar dataset 6 days ago [here](https://www.kaggle.com/datasets/sbunzini/open-problems-msci-multiome-sparse-matrices). I should have mentioned it above but I just noticed it. Sorry and kudos to him.\n\nCompared with sbuzini's dataset, the one I published was designed to be \"standalone\": it replicates all of the information in the original competition dataset, either in `.npz`, `.sparse` or `.parquet` format. (I prefered to convert everything once and for all).\n\nIn practice, it means that:\n\nsbuzini's dataset only contains three files:\n- test_multi_inputs_sparse.npz\n- train_multi_targets_sparse.npz\n- train_multi_inputs_sparse.npz\n\nwhich would correspond to the files:\n- test_multi_inputs_values.sparse\n- train_multi_targets_values.sparse\n- train_multi_inputs_values.sparse\n\nin my dataset. \n\nMy dataset also contain the files:\n- test_multi_inputs_idxcol.npz\n- train_multi_targets_idxcol.npz\n- train_multi_inputs_idxcol.npz\nthat gives the content of the index and column names for each original dataframe.\n\nIn addition, my dataset also contains the corresponding files for the CITEseq dataframes (they are less sparse, but still compress reasonably well). As well as the metadat/submission info in parquet format.",
      "votes": null
    },
    {
      "id": "1920802",
      "postDate": "08/31/2022 12:21:00",
      "content": "<p>Hey Fabian! Don't worry, it's always a pleasure to see different solutions to the same problem. Good work with your notebook!👍🤜</p>",
      "rawMarkdown": "Hey Fabian! Don't worry, it's always a pleasure to see different solutions to the same problem. Good work with your notebook!👍🤜",
      "votes": null
    },
    {
      "id": "1926044",
      "postDate": "09/04/2022 14:11:18",
      "content": "<p>Thank you for the kind words. Good luck to you too.</p>",
      "rawMarkdown": "Thank you for the kind words. Good luck to you too.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1920190,
      "author_name": "fabiencrom",
      "author_url": "",
      "post_date": "08/31/2022 01:42:59",
      "content": "<p>Actually, sbunzini had already provided a similar dataset 6 days ago <a href=\"https://www.kaggle.com/datasets/sbunzini/open-problems-msci-multiome-sparse-matrices\" target=\"_blank\">here</a>. I should have mentioned it above but I just noticed it. Sorry and kudos to him.</p>\n<p>Compared with sbuzini's dataset, the one I published was designed to be \"standalone\": it replicates all of the information in the original competition dataset, either in <code>.npz</code>, <code>.sparse</code> or <code>.parquet</code> format. (I prefered to convert everything once and for all).</p>\n<p>In practice, it means that:</p>\n<p>sbuzini's dataset only contains three files:</p>\n<ul>\n<li>test_multi_inputs_sparse.npz</li>\n<li>train_multi_targets_sparse.npz</li>\n<li>train_multi_inputs_sparse.npz</li>\n</ul>\n<p>which would correspond to the files:</p>\n<ul>\n<li>test_multi_inputs_values.sparse</li>\n<li>train_multi_targets_values.sparse</li>\n<li>train_multi_inputs_values.sparse</li>\n</ul>\n<p>in my dataset. </p>\n<p>My dataset also contain the files:</p>\n<ul>\n<li>test_multi_inputs_idxcol.npz</li>\n<li>train_multi_targets_idxcol.npz</li>\n<li>train_multi_inputs_idxcol.npz<br>\nthat gives the content of the index and column names for each original dataframe.</li>\n</ul>\n<p>In addition, my dataset also contains the corresponding files for the CITEseq dataframes (they are less sparse, but still compress reasonably well). As well as the metadat/submission info in parquet format.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1920802,
          "author_name": "sbunzini",
          "author_url": "",
          "post_date": "08/31/2022 12:21:00",
          "content": "<p>Hey Fabian! Don't worry, it's always a pleasure to see different solutions to the same problem. Good work with your notebook!👍🤜</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1926044,
          "author_name": "fabiencrom",
          "author_url": "",
          "post_date": "09/04/2022 14:11:18",
          "content": "<p>Thank you for the kind words. Good luck to you too.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1920165": "Hi,\n\nAs many have pointed out, the data for this competition is rather sparse and thus can benefit from using a sparse matrix representation.\n\nTo make it easier for the kagglers that are not familiars with the use of sparse matrices in ML, I have made public:\n- A [dataset](https://www.kaggle.com/datasets/fabiencrom/multimodal-single-cell-as-sparse-matrix) containing all the competition data in an efficient sparse format (i.e. `scipy`'s csr_matrix)\n- The [notebook](https://www.kaggle.com/code/fabiencrom/multimodal-single-cell-creating-sparse-data/) that generated this dataset\n- A [simple adaptation of AmbroseM'S quickstart notebook](https://www.kaggle.com/code/fabiencrom/msci-multiome-quickstart-w-sparse-matrices) to demonstrate a simple use of this sparse dataset\n\nWhen I have time, I plan to add a notebook showing how to use the sparse representation with torch and/or keras for a simple neural network.\n\nAs I was writing this (I made the dataset 5 days ago, but had not taken the time to prepare it for publishing), I found out that some notebooks and discussions have already been published on this topic: [here](https://www.kaggle.com/code/sbunzini/reduce-memory-usage-by-95-with-sparse-matrices/notebook#Compression-Rate),  [here](https://www.kaggle.com/code/jsmithperera/msci-multiome-sparse-tsvd)  and [here](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/348377). This will make some of the information provided slightly redundant, but hopefully it is still interesting.\n\n*Additional informations*:\n\nThere are many different formats for storing sparse matrices. The three most often used are COO, CSR and CSC:\n- COO is the simplest to create and is often use as an intermediate form before converting to CSR or CSC. It is also the only format really supported by pytorch and CUDA.\n- CSR is more memory efficient than COO and will also be more computation-efficient in many cases (at least on CPU, but not on GPU). Row-slicing is efficient on this format, which makes it a good choice for representing tabular data.\n- CSC is similar to CSR, but implemented such that column-slicing is efficient.\n\nI am using CSR, as it is the most memory-efficient for this dataset. Note that `scipy`'s sparse matrices should be loaded with `scipy.sparse.load_npz` and not `numpy.load`.\n\nIn the [sparse CSR dataset](https://www.kaggle.com/datasets/fabiencrom/multimodal-single-cell-as-sparse-matrix), the dataframes stored in each original `xxx.h5` files are converted into two files:\n- One \"xxx_values.sparse\" file that can be loaded with `scipy.sparse.load_npz` and contains all the values of the corresponding dataframe (i.e. the result of `df.values` in a sparse CSR format)\n- One \"xxx_idxcol.npz\" file that can be loaded with `numpy.load` and contains the values of the index and the columns of the corresponding dataframe (i.e the results of `df.index` and `df.columns` as normal arrays)",
    "1920190": "Actually, sbunzini had already provided a similar dataset 6 days ago [here](https://www.kaggle.com/datasets/sbunzini/open-problems-msci-multiome-sparse-matrices). I should have mentioned it above but I just noticed it. Sorry and kudos to him.\n\nCompared with sbuzini's dataset, the one I published was designed to be \"standalone\": it replicates all of the information in the original competition dataset, either in `.npz`, `.sparse` or `.parquet` format. (I prefered to convert everything once and for all).\n\nIn practice, it means that:\n\nsbuzini's dataset only contains three files:\n- test_multi_inputs_sparse.npz\n- train_multi_targets_sparse.npz\n- train_multi_inputs_sparse.npz\n\nwhich would correspond to the files:\n- test_multi_inputs_values.sparse\n- train_multi_targets_values.sparse\n- train_multi_inputs_values.sparse\n\nin my dataset. \n\nMy dataset also contain the files:\n- test_multi_inputs_idxcol.npz\n- train_multi_targets_idxcol.npz\n- train_multi_inputs_idxcol.npz\nthat gives the content of the index and column names for each original dataframe.\n\nIn addition, my dataset also contains the corresponding files for the CITEseq dataframes (they are less sparse, but still compress reasonably well). As well as the metadat/submission info in parquet format.",
    "1920802": "Hey Fabian! Don't worry, it's always a pleasure to see different solutions to the same problem. Good work with your notebook!👍🤜",
    "1926044": "Thank you for the kind words. Good luck to you too."
  },
  "source": "meta"
}