{
  "id": 350403,
  "title": "Memory Error while reading",
  "url": "/competitions/open-problems-multimodal/discussion/350403",
  "author_name": "",
  "post_date": "2022-09-05T15:07:01.898536500Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi, <br>\nWhen i tried reading 'block0_values' in 'train_multi_inputs' dataset, but im getting memory error. My RAM size is 8GB, any way to sort this out?</p>",
  "messages": [
    {
      "id": "1927338",
      "postDate": "09/05/2022 15:07:01",
      "content": "<p>Hi, <br>\nWhen i tried reading 'block0_values' in 'train_multi_inputs' dataset, but im getting memory error. My RAM size is 8GB, any way to sort this out?</p>",
      "rawMarkdown": "Hi, \nWhen i tried reading 'block0_values' in 'train_multi_inputs' dataset, but im getting memory error. My RAM size is 8GB, any way to sort this out?",
      "votes": null
    },
    {
      "id": "1975227",
      "postDate": "10/06/2022 16:39:48",
      "content": "<p>This dataset is too big,  I am also facing same problem even when I have my 128 GB ram</p>",
      "rawMarkdown": "This dataset is too big,  I am also facing same problem even when I have my 128 GB ram",
      "votes": null
    },
    {
      "id": "1981563",
      "postDate": "10/10/2022 21:27:02",
      "content": "<p>The only way to operate on this data is to use a <a href=\"https://www.w3schools.com/python/scipy/scipy_sparse_data.php\" target=\"_blank\">sparse dataset</a>.  <a href=\"https://www.kaggle.com/code/kirkdco/msci-multiome-sparse-datasets\" target=\"_blank\">This notebook</a> shows how I've done it.  </p>\n<p>The way I approached it is to load the data in blocks, make sparse datasets for each block, and then combine them.  I also tabulated which columns how low fraction of unique values.  I then save the sparse dataset and the columns names that have low numbers of unique values.  The Dataset class lets you use those lists to drop columns from the dataset to make them smaller.  For the complete Multiome dataset, it takes about 4.5 GB to store as a sparse dataset.</p>\n<p>A lot of people are also using <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.decomposition.TruncatedSVD.html\" target=\"_blank\">TruncatedSVD</a> for dimensionality reduction.  I have some code at the bottom that shows how to use that, too.  In that form, the data set takes around 150 MB.</p>",
      "rawMarkdown": "The only way to operate on this data is to use a [sparse dataset](https://www.w3schools.com/python/scipy/scipy_sparse_data.php).  [This notebook](https://www.kaggle.com/code/kirkdco/msci-multiome-sparse-datasets) shows how I've done it.  \n\nThe way I approached it is to load the data in blocks, make sparse datasets for each block, and then combine them.  I also tabulated which columns how low fraction of unique values.  I then save the sparse dataset and the columns names that have low numbers of unique values.  The Dataset class lets you use those lists to drop columns from the dataset to make them smaller.  For the complete Multiome dataset, it takes about 4.5 GB to store as a sparse dataset.\n\nA lot of people are also using [TruncatedSVD](https://scikit-learn.org/stable/modules/generated/sklearn.decomposition.TruncatedSVD.html) for dimensionality reduction.  I have some code at the bottom that shows how to use that, too.  In that form, the data set takes around 150 MB.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1975227,
      "author_name": "vinodsinghatjnu",
      "author_url": "",
      "post_date": "10/06/2022 16:39:48",
      "content": "<p>This dataset is too big,  I am also facing same problem even when I have my 128 GB ram</p>",
      "votes": null,
      "replies": [
        {
          "id": 1981563,
          "author_name": "kirkdco",
          "author_url": "",
          "post_date": "10/10/2022 21:27:02",
          "content": "<p>The only way to operate on this data is to use a <a href=\"https://www.w3schools.com/python/scipy/scipy_sparse_data.php\" target=\"_blank\">sparse dataset</a>.  <a href=\"https://www.kaggle.com/code/kirkdco/msci-multiome-sparse-datasets\" target=\"_blank\">This notebook</a> shows how I've done it.  </p>\n<p>The way I approached it is to load the data in blocks, make sparse datasets for each block, and then combine them.  I also tabulated which columns how low fraction of unique values.  I then save the sparse dataset and the columns names that have low numbers of unique values.  The Dataset class lets you use those lists to drop columns from the dataset to make them smaller.  For the complete Multiome dataset, it takes about 4.5 GB to store as a sparse dataset.</p>\n<p>A lot of people are also using <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.decomposition.TruncatedSVD.html\" target=\"_blank\">TruncatedSVD</a> for dimensionality reduction.  I have some code at the bottom that shows how to use that, too.  In that form, the data set takes around 150 MB.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1927338": "Hi, \nWhen i tried reading 'block0_values' in 'train_multi_inputs' dataset, but im getting memory error. My RAM size is 8GB, any way to sort this out?",
    "1975227": "This dataset is too big,  I am also facing same problem even when I have my 128 GB ram",
    "1981563": "The only way to operate on this data is to use a [sparse dataset](https://www.w3schools.com/python/scipy/scipy_sparse_data.php).  [This notebook](https://www.kaggle.com/code/kirkdco/msci-multiome-sparse-datasets) shows how I've done it.  \n\nThe way I approached it is to load the data in blocks, make sparse datasets for each block, and then combine them.  I also tabulated which columns how low fraction of unique values.  I then save the sparse dataset and the columns names that have low numbers of unique values.  The Dataset class lets you use those lists to drop columns from the dataset to make them smaller.  For the complete Multiome dataset, it takes about 4.5 GB to store as a sparse dataset.\n\nA lot of people are also using [TruncatedSVD](https://scikit-learn.org/stable/modules/generated/sklearn.decomposition.TruncatedSVD.html) for dimensionality reduction.  I have some code at the bottom that shows how to use that, too.  In that form, the data set takes around 150 MB."
  },
  "source": "meta"
}