{
  "id": 346690,
  "title": "Query specific subsets of files without loading them into RAM.",
  "url": "/competitions/open-problems-multimodal/discussion/346690",
  "author_name": "",
  "post_date": "2022-08-20T23:02:10.868352700Z",
  "votes": 12,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I have implemented a class that allows one to read specific subsets of one file (say, multiome data for one specific day and one specific donor) without the need to load all the file into RAM. I used the h5py library, as <a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a> suggested in another discussion (see: <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/344829#1905232\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/344829#1905232</a>).</p>\n<p>The class I implemented can be found in this <a href=\"https://www.kaggle.com/alekeuro/querying-specific-subsets-of-the-data-without-load\" target=\"_blank\">notebook</a>. Comments and suggestions about how to make it more efficient (or whether you find this approach useful) are most welcome. </p>",
  "messages": [
    {
      "id": "1907571",
      "postDate": "08/20/2022 23:02:10",
      "content": "<p>I have implemented a class that allows one to read specific subsets of one file (say, multiome data for one specific day and one specific donor) without the need to load all the file into RAM. I used the h5py library, as <a href=\"https://www.kaggle.com/alexandervc\" target=\"_blank\">@alexandervc</a> suggested in another discussion (see: <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/344829#1905232\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/344829#1905232</a>).</p>\n<p>The class I implemented can be found in this <a href=\"https://www.kaggle.com/alekeuro/querying-specific-subsets-of-the-data-without-load\" target=\"_blank\">notebook</a>. Comments and suggestions about how to make it more efficient (or whether you find this approach useful) are most welcome. </p>",
      "rawMarkdown": "I have implemented a class that allows one to read specific subsets of one file (say, multiome data for one specific day and one specific donor) without the need to load all the file into RAM. I used the h5py library, as @alexandervc suggested in another discussion (see: [https://www.kaggle.com/competitions/open-problems-multimodal/discussion/344829#1905232](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/344829#1905232)).\n\nThe class I implemented can be found in this [notebook](https://www.kaggle.com/alekeuro/querying-specific-subsets-of-the-data-without-load). Comments and suggestions about how to make it more efficient (or whether you find this approach useful) are most welcome.",
      "votes": null
    },
    {
      "id": "1910968",
      "postDate": "08/23/2022 20:20:39",
      "content": "<p>I just implemented a new feature in this <code>DataReader</code> class for columns selection: you can pass in </p>\n<ul>\n<li>a list of columns names (if you know them in advance)</li>\n<li>a list of ints corresponding to indices of desired columns, or </li>\n<li>you can ask for a random sample of N columns. </li>\n</ul>\n<p>I have updated the above mentioned notebook.</p>",
      "rawMarkdown": "I just implemented a new feature in this `DataReader` class for columns selection: you can pass in \n\n- a list of columns names (if you know them in advance)\n- a list of ints corresponding to indices of desired columns, or \n- you can ask for a random sample of N columns. \n\nI have updated the above mentioned notebook.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1910968,
      "author_name": "alekeuro",
      "author_url": "",
      "post_date": "08/23/2022 20:20:39",
      "content": "<p>I just implemented a new feature in this <code>DataReader</code> class for columns selection: you can pass in </p>\n<ul>\n<li>a list of columns names (if you know them in advance)</li>\n<li>a list of ints corresponding to indices of desired columns, or </li>\n<li>you can ask for a random sample of N columns. </li>\n</ul>\n<p>I have updated the above mentioned notebook.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1907571": "I have implemented a class that allows one to read specific subsets of one file (say, multiome data for one specific day and one specific donor) without the need to load all the file into RAM. I used the h5py library, as @alexandervc suggested in another discussion (see: [https://www.kaggle.com/competitions/open-problems-multimodal/discussion/344829#1905232](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/344829#1905232)).\n\nThe class I implemented can be found in this [notebook](https://www.kaggle.com/alekeuro/querying-specific-subsets-of-the-data-without-load). Comments and suggestions about how to make it more efficient (or whether you find this approach useful) are most welcome.",
    "1910968": "I just implemented a new feature in this `DataReader` class for columns selection: you can pass in \n\n- a list of columns names (if you know them in advance)\n- a list of ints corresponding to indices of desired columns, or \n- you can ask for a random sample of N columns. \n\nI have updated the above mentioned notebook."
  },
  "source": "meta"
}