{
  "id": 346045,
  "title": "How to handle huge H5 files?",
  "url": "/competitions/open-problems-multimodal/discussion/346045",
  "author_name": "",
  "post_date": "2022-08-17T17:31:02.745790400Z",
  "votes": 1,
  "comment_count": 11,
  "views": 0,
  "content": "<p>The small files that I was able to convert to parquet, reduced 20% of the size, but the file train_multi_inputs.h5, was not able to be converted, pandas requires 90GB RAM + to load in memory, and you can use chunks in pandas to read each part of the data set and convert.</p>\n<p>What's the best way to handle it?</p>\n<p>Train with \"chunks\" of data using start/stop in read_hdf?<br>\nComputers/notebooks with more ram?</p>\n<p>I'm not an expert on this type of data :(</p>",
  "messages": [
    {
      "id": "1903838",
      "postDate": "08/17/2022 17:31:02",
      "content": "<p>The small files that I was able to convert to parquet, reduced 20% of the size, but the file train_multi_inputs.h5, was not able to be converted, pandas requires 90GB RAM + to load in memory, and you can use chunks in pandas to read each part of the data set and convert.</p>\n<p>What's the best way to handle it?</p>\n<p>Train with \"chunks\" of data using start/stop in read_hdf?<br>\nComputers/notebooks with more ram?</p>\n<p>I'm not an expert on this type of data :(</p>",
      "rawMarkdown": "The small files that I was able to convert to parquet, reduced 20% of the size, but the file train_multi_inputs.h5, was not able to be converted, pandas requires 90GB RAM + to load in memory, and you can use chunks in pandas to read each part of the data set and convert.\n\nWhat's the best way to handle it?\n\nTrain with \"chunks\" of data using start/stop in read_hdf?\nComputers/notebooks with more ram?\n\nI'm not an expert on this type of data :(",
      "votes": null
    },
    {
      "id": "1903854",
      "postDate": "08/17/2022 17:40:34",
      "content": "<p>If you want to load only a fraction of the data, you can use the <code>start</code> and <code>stop</code> arguments to <a href=\"https://pandas.pydata.org/docs/reference/api/pandas.read_hdf.html\" target=\"_blank\"><code>pandas.read_hdf()</code></a>.</p>",
      "rawMarkdown": "If you want to load only a fraction of the data, you can use the `start` and `stop` arguments to [`pandas.read_hdf()`](https://pandas.pydata.org/docs/reference/api/pandas.read_hdf.html).",
      "votes": null
    },
    {
      "id": "1903986",
      "postDate": "08/17/2022 20:27:43",
      "content": "<p>Are they sparse ? Have not you tried: chunk reading + conversion to sparse format ? </p>",
      "rawMarkdown": "Are they sparse ? Have not you tried: chunk reading + conversion to sparse format ?",
      "votes": null
    },
    {
      "id": "1903997",
      "postDate": "08/17/2022 20:49:57",
      "content": "<p>With read_hdf, reading chunks didn't work because the h5 file was not written with pytables.</p>\n<p>I think if you have enough memory (~128GB) you can read inmemory and convert.</p>",
      "rawMarkdown": "With read_hdf, reading chunks didn't work because the h5 file was not written with pytables.\n\nI think if you have enough memory (~128GB) you can read inmemory and convert.",
      "votes": null
    },
    {
      "id": "1904013",
      "postDate": "08/17/2022 21:23:10",
      "content": "<p>For *.h5 files we may try to use h5py - to work with files on disk as they are in memory:<br>\nexample: <a href=\"https://www.kaggle.com/code/alexandervc/archs4-extractsave-datasets-by-gse-and-keyword\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/archs4-extractsave-datasets-by-gse-and-keyword</a><br>\n(not sure it will work for these files)</p>\n<p>if they would be h5ad then one can try to do the same with scanpy:<br>\nadata = sc.read(fn,  backed='r' )<br>\nexample: <a href=\"https://www.kaggle.com/code/alexandervc/scanpy-process-huge-files-backing-them-on-disk\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/scanpy-process-huge-files-backing-them-on-disk</a><br>\none may first try to convert to h5ad then do like that… </p>\n<p>not sure it would be helpful</p>",
      "rawMarkdown": "For *.h5 files we may try to use h5py - to work with files on disk as they are in memory:\nexample: https://www.kaggle.com/code/alexandervc/archs4-extractsave-datasets-by-gse-and-keyword\n(not sure it will work for these files)\n\nif they would be h5ad then one can try to do the same with scanpy:\nadata = sc.read(fn,  backed='r' )\nexample: https://www.kaggle.com/code/alexandervc/scanpy-process-huge-files-backing-them-on-disk\none may first try to convert to h5ad then do like that... \n\nnot sure it would be helpful",
      "votes": null
    },
    {
      "id": "1904032",
      "postDate": "08/17/2022 22:21:57",
      "content": "<p>I didn't analyse this in much depth, but I think the data is not-so-sparse. E.g., the CITEseq training feature matrix has only ~24% of entries equal to zero.</p>",
      "rawMarkdown": "I didn't analyse this in much depth, but I think the data is not-so-sparse. E.g., the CITEseq training feature matrix has only ~24% of entries equal to zero.",
      "votes": null
    },
    {
      "id": "1904404",
      "postDate": "08/18/2022 06:52:53",
      "content": "<p>Yes, but CITEseq is different data - RNA-seq. The huge files are from ATAC-seq technology - multimode. </p>",
      "rawMarkdown": "Yes, but CITEseq is different data - RNA-seq. The huge files are from ATAC-seq technology - multimode.",
      "votes": null
    },
    {
      "id": "1905227",
      "postDate": "08/18/2022 20:44:55",
      "content": "<p>h5py works for these data files:<br>\n<a href=\"https://www.kaggle.com/code/alexandervc/use-h5py-for-huge-h5-file-backing-it-on-disk\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/use-h5py-for-huge-h5-file-backing-it-on-disk</a></p>\n<p>It allows to work with files on disk \"as if they were loaded into memory\"</p>",
      "rawMarkdown": "h5py works for these data files:\nhttps://www.kaggle.com/code/alexandervc/use-h5py-for-huge-h5-file-backing-it-on-disk\n\nIt allows to work with files on disk \"as if they were loaded into memory\"",
      "votes": null
    },
    {
      "id": "1908196",
      "postDate": "08/21/2022 12:44:49",
      "content": "<p>I understand now what you were saying: ATAC-seq (multiome) input data seems to be really sparse (only around 3% of entries are non zero). </p>",
      "rawMarkdown": "I understand now what you were saying: ATAC-seq (multiome) input data seems to be really sparse (only around 3% of entries are non zero).",
      "votes": null
    },
    {
      "id": "1909802",
      "postDate": "08/23/2022 00:23:47",
      "content": "<p>if I want only a specific column without open everthing before?</p>",
      "rawMarkdown": "if I want only a specific column without open everthing before?",
      "votes": null
    },
    {
      "id": "1909887",
      "postDate": "08/23/2022 03:12:48",
      "content": "<p>(DISCLAIMER: shameless self-promotion comment here)</p>\n<p>Hi Erivan, I have implemented a class that allows you to query the files and load specific subsets (e.g., read only multiome input data for donor nb 3, for day = 2). You can find the code in this <a href=\"https://www.kaggle.com/alekeuro/querying-specific-subsets-of-the-data-without-load\" target=\"_blank\">notebook</a>.</p>\n<p>Unfortunately, this python class doesn't enable you to retrieve specific columns, but I could add that functionality if you give me some time (1-2 days). Or, you could modify the class, and implement it yourself!</p>\n<p>The class needs libraries <code>h5py</code>, <code>hdf5plugin</code> and <code>tables</code>, and uses specifically <code>h5py.File</code> class to navigate and query the files.</p>",
      "rawMarkdown": "(DISCLAIMER: shameless self-promotion comment here)\n\nHi Erivan, I have implemented a class that allows you to query the files and load specific subsets (e.g., read only multiome input data for donor nb 3, for day = 2). You can find the code in this [notebook](https://www.kaggle.com/alekeuro/querying-specific-subsets-of-the-data-without-load).\n\nUnfortunately, this python class doesn't enable you to retrieve specific columns, but I could add that functionality if you give me some time (1-2 days). Or, you could modify the class, and implement it yourself!\n\nThe class needs libraries `h5py`, `hdf5plugin` and `tables`, and uses specifically `h5py.File` class to navigate and query the files.",
      "votes": null
    },
    {
      "id": "1910965",
      "postDate": "08/23/2022 20:17:59",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/erivanoliveirajr\" target=\"_blank\">@erivanoliveirajr</a> , I just implemented this feature for columns selection: you can pass in a list of columns names (if you know them in advance), or pass a list of ints corresponding to indices of desired columns, or even ask for a random sample of N columns. I have updated the above mentioned notebook. Let me know if that's helpful.</p>",
      "rawMarkdown": "Hey @erivanoliveirajr , I just implemented this feature for columns selection: you can pass in a list of columns names (if you know them in advance), or pass a list of ints corresponding to indices of desired columns, or even ask for a random sample of N columns. I have updated the above mentioned notebook. Let me know if that's helpful.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1903854,
      "author_name": "danielburkhardt",
      "author_url": "",
      "post_date": "08/17/2022 17:40:34",
      "content": "<p>If you want to load only a fraction of the data, you can use the <code>start</code> and <code>stop</code> arguments to <a href=\"https://pandas.pydata.org/docs/reference/api/pandas.read_hdf.html\" target=\"_blank\"><code>pandas.read_hdf()</code></a>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1909802,
          "author_name": "erivanoliveirajr",
          "author_url": "",
          "post_date": "08/23/2022 00:23:47",
          "content": "<p>if I want only a specific column without open everthing before?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1909887,
          "author_name": "alekeuro",
          "author_url": "",
          "post_date": "08/23/2022 03:12:48",
          "content": "<p>(DISCLAIMER: shameless self-promotion comment here)</p>\n<p>Hi Erivan, I have implemented a class that allows you to query the files and load specific subsets (e.g., read only multiome input data for donor nb 3, for day = 2). You can find the code in this <a href=\"https://www.kaggle.com/alekeuro/querying-specific-subsets-of-the-data-without-load\" target=\"_blank\">notebook</a>.</p>\n<p>Unfortunately, this python class doesn't enable you to retrieve specific columns, but I could add that functionality if you give me some time (1-2 days). Or, you could modify the class, and implement it yourself!</p>\n<p>The class needs libraries <code>h5py</code>, <code>hdf5plugin</code> and <code>tables</code>, and uses specifically <code>h5py.File</code> class to navigate and query the files.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1910965,
          "author_name": "alekeuro",
          "author_url": "",
          "post_date": "08/23/2022 20:17:59",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/erivanoliveirajr\" target=\"_blank\">@erivanoliveirajr</a> , I just implemented this feature for columns selection: you can pass in a list of columns names (if you know them in advance), or pass a list of ints corresponding to indices of desired columns, or even ask for a random sample of N columns. I have updated the above mentioned notebook. Let me know if that's helpful.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1903986,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "08/17/2022 20:27:43",
      "content": "<p>Are they sparse ? Have not you tried: chunk reading + conversion to sparse format ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1903997,
          "author_name": "nandodmelo",
          "author_url": "",
          "post_date": "08/17/2022 20:49:57",
          "content": "<p>With read_hdf, reading chunks didn't work because the h5 file was not written with pytables.</p>\n<p>I think if you have enough memory (~128GB) you can read inmemory and convert.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1904013,
          "author_name": "alexandervc",
          "author_url": "",
          "post_date": "08/17/2022 21:23:10",
          "content": "<p>For *.h5 files we may try to use h5py - to work with files on disk as they are in memory:<br>\nexample: <a href=\"https://www.kaggle.com/code/alexandervc/archs4-extractsave-datasets-by-gse-and-keyword\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/archs4-extractsave-datasets-by-gse-and-keyword</a><br>\n(not sure it will work for these files)</p>\n<p>if they would be h5ad then one can try to do the same with scanpy:<br>\nadata = sc.read(fn,  backed='r' )<br>\nexample: <a href=\"https://www.kaggle.com/code/alexandervc/scanpy-process-huge-files-backing-them-on-disk\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/scanpy-process-huge-files-backing-them-on-disk</a><br>\none may first try to convert to h5ad then do like that… </p>\n<p>not sure it would be helpful</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1904032,
          "author_name": "alekeuro",
          "author_url": "",
          "post_date": "08/17/2022 22:21:57",
          "content": "<p>I didn't analyse this in much depth, but I think the data is not-so-sparse. E.g., the CITEseq training feature matrix has only ~24% of entries equal to zero.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1904404,
          "author_name": "alexandervc",
          "author_url": "",
          "post_date": "08/18/2022 06:52:53",
          "content": "<p>Yes, but CITEseq is different data - RNA-seq. The huge files are from ATAC-seq technology - multimode. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1905227,
          "author_name": "alexandervc",
          "author_url": "",
          "post_date": "08/18/2022 20:44:55",
          "content": "<p>h5py works for these data files:<br>\n<a href=\"https://www.kaggle.com/code/alexandervc/use-h5py-for-huge-h5-file-backing-it-on-disk\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/use-h5py-for-huge-h5-file-backing-it-on-disk</a></p>\n<p>It allows to work with files on disk \"as if they were loaded into memory\"</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1908196,
          "author_name": "alekeuro",
          "author_url": "",
          "post_date": "08/21/2022 12:44:49",
          "content": "<p>I understand now what you were saying: ATAC-seq (multiome) input data seems to be really sparse (only around 3% of entries are non zero). </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1903838": "The small files that I was able to convert to parquet, reduced 20% of the size, but the file train_multi_inputs.h5, was not able to be converted, pandas requires 90GB RAM + to load in memory, and you can use chunks in pandas to read each part of the data set and convert.\n\nWhat's the best way to handle it?\n\nTrain with \"chunks\" of data using start/stop in read_hdf?\nComputers/notebooks with more ram?\n\nI'm not an expert on this type of data :(",
    "1903854": "If you want to load only a fraction of the data, you can use the `start` and `stop` arguments to [`pandas.read_hdf()`](https://pandas.pydata.org/docs/reference/api/pandas.read_hdf.html).",
    "1903986": "Are they sparse ? Have not you tried: chunk reading + conversion to sparse format ?",
    "1903997": "With read_hdf, reading chunks didn't work because the h5 file was not written with pytables.\n\nI think if you have enough memory (~128GB) you can read inmemory and convert.",
    "1904013": "For *.h5 files we may try to use h5py - to work with files on disk as they are in memory:\nexample: https://www.kaggle.com/code/alexandervc/archs4-extractsave-datasets-by-gse-and-keyword\n(not sure it will work for these files)\n\nif they would be h5ad then one can try to do the same with scanpy:\nadata = sc.read(fn,  backed='r' )\nexample: https://www.kaggle.com/code/alexandervc/scanpy-process-huge-files-backing-them-on-disk\none may first try to convert to h5ad then do like that... \n\nnot sure it would be helpful",
    "1904032": "I didn't analyse this in much depth, but I think the data is not-so-sparse. E.g., the CITEseq training feature matrix has only ~24% of entries equal to zero.",
    "1904404": "Yes, but CITEseq is different data - RNA-seq. The huge files are from ATAC-seq technology - multimode.",
    "1905227": "h5py works for these data files:\nhttps://www.kaggle.com/code/alexandervc/use-h5py-for-huge-h5-file-backing-it-on-disk\n\nIt allows to work with files on disk \"as if they were loaded into memory\"",
    "1908196": "I understand now what you were saying: ATAC-seq (multiome) input data seems to be really sparse (only around 3% of entries are non zero).",
    "1909802": "if I want only a specific column without open everthing before?",
    "1909887": "(DISCLAIMER: shameless self-promotion comment here)\n\nHi Erivan, I have implemented a class that allows you to query the files and load specific subsets (e.g., read only multiome input data for donor nb 3, for day = 2). You can find the code in this [notebook](https://www.kaggle.com/alekeuro/querying-specific-subsets-of-the-data-without-load).\n\nUnfortunately, this python class doesn't enable you to retrieve specific columns, but I could add that functionality if you give me some time (1-2 days). Or, you could modify the class, and implement it yourself!\n\nThe class needs libraries `h5py`, `hdf5plugin` and `tables`, and uses specifically `h5py.File` class to navigate and query the files.",
    "1910965": "Hey @erivanoliveirajr , I just implemented this feature for columns selection: you can pass in a list of columns names (if you know them in advance), or pass a list of ints corresponding to indices of desired columns, or even ask for a random sample of N columns. I have updated the above mentioned notebook. Let me know if that's helpful."
  },
  "source": "meta"
}