{
  "id": 350226,
  "title": "How do I open an H5 file?",
  "url": "/competitions/open-problems-multimodal/discussion/350226",
  "author_name": "",
  "post_date": "2022-09-04T20:10:51.336707400Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hello, <br>\nI am new to this competition and I am having some trouble using the given data. So far, I can only open a few of the H5 files using pandas but all the others take up too much memory. Is there an alternative way to see the contents of an H5 file? </p>\n<p>Thank you.</p>",
  "messages": [
    {
      "id": "1926410",
      "postDate": "09/04/2022 20:10:51",
      "content": "<p>Hello, <br>\nI am new to this competition and I am having some trouble using the given data. So far, I can only open a few of the H5 files using pandas but all the others take up too much memory. Is there an alternative way to see the contents of an H5 file? </p>\n<p>Thank you.</p>",
      "rawMarkdown": "Hello, \nI am new to this competition and I am having some trouble using the given data. So far, I can only open a few of the H5 files using pandas but all the others take up too much memory. Is there an alternative way to see the contents of an H5 file? \n\nThank you.",
      "votes": null
    },
    {
      "id": "1926443",
      "postDate": "09/04/2022 20:50:00",
      "content": "<p>You don't have to read all of the information at once. You can use use the start/stops parameters. </p>\n<pre><code>targets = pd.read_hdf('../input/targets.h5', start=0, stop=6000)\n</code></pre>\n<p>You can find notebooks like </p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/digitalbro/multimodal-sc-integration-meta-resources\" target=\"_blank\">https://www.kaggle.com/code/digitalbro/multimodal-sc-integration-meta-resources</a>)<br>\nwhich provide a lot of information as well. Also, Sbunzini's response will give you a sparse version of the data, which will help when running the notebooks on Kaggle. I'd recommend looking at using Saturn Cloud </li>\n<li><a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346999\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346999</a> <br>\nfor that as it might help not to have to run into the constraints of the Kaggle VMs. Cheers!  </li>\n</ul>",
      "rawMarkdown": "You don't have to read all of the information at once. You can use use the start/stops parameters. \n\n```\ntargets = pd.read_hdf('../input/targets.h5', start=0, stop=6000)\n```\n\nYou can find notebooks like \n\n* https://www.kaggle.com/code/digitalbro/multimodal-sc-integration-meta-resources)\n\nwhich provide a lot of information as well. Also, Sbunzini's response will give you a sparse version of the data, which will help when running the notebooks on Kaggle. I'd recommend looking at using Saturn Cloud \n\n* https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346999 \n\nfor that as it might help not to have to run into the constraints of the Kaggle VMs. Cheers!",
      "votes": null
    },
    {
      "id": "1926444",
      "postDate": "09/04/2022 20:50:45",
      "content": "<p>Hey RealApex, we found that Multiome dataset can be widely compressed in CSR matrix format. Take a look at these notebooks:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/sbunzini/reduce-memory-usage-by-95-with-sparse-matrices\" target=\"_blank\">https://www.kaggle.com/code/sbunzini/reduce-memory-usage-by-95-with-sparse-matrices</a></li>\n<li><a href=\"https://www.kaggle.com/code/fabiencrom/multimodal-single-cell-creating-sparse-data\" target=\"_blank\">https://www.kaggle.com/code/fabiencrom/multimodal-single-cell-creating-sparse-data</a></li>\n</ul>",
      "rawMarkdown": "Hey RealApex, we found that Multiome dataset can be widely compressed in CSR matrix format. Take a look at these notebooks:\n\n- https://www.kaggle.com/code/sbunzini/reduce-memory-usage-by-95-with-sparse-matrices\n- https://www.kaggle.com/code/fabiencrom/multimodal-single-cell-creating-sparse-data",
      "votes": null
    },
    {
      "id": "1926563",
      "postDate": "09/05/2022 00:28:40",
      "content": "<p>Hello, I actually used the line of code you proposed above but it loads the whole dataset anyways, do you have any insight  as to why?</p>",
      "rawMarkdown": "Hello, I actually used the line of code you proposed above but it loads the whole dataset anyways, do you have any insight  as to why?",
      "votes": null
    },
    {
      "id": "1926593",
      "postDate": "09/05/2022 01:23:24",
      "content": "<p>I do not know why it does not work well for you, maybe you didn't do it properly?<br>\nHere is the code I executed and It worked well.<br>\n<code>\n!pip install tables\nimport pandas as pd\nfrom pandas import HDFStore\nx = \"../input/open-problems-multimodal/test_multi_inputs.h5\"\npd.read_hdf(x, start = 0, stop = 50)\n</code></p>",
      "rawMarkdown": "I do not know why it does not work well for you, maybe you didn't do it properly?\nHere is the code I executed and It worked well.\n`\n!pip install tables\nimport pandas as pd\nfrom pandas import HDFStore\nx = \"../input/open-problems-multimodal/test_multi_inputs.h5\"\npd.read_hdf(x, start = 0, stop = 50)\n`",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1926443,
      "author_name": "digitalbro",
      "author_url": "",
      "post_date": "09/04/2022 20:50:00",
      "content": "<p>You don't have to read all of the information at once. You can use use the start/stops parameters. </p>\n<pre><code>targets = pd.read_hdf('../input/targets.h5', start=0, stop=6000)\n</code></pre>\n<p>You can find notebooks like </p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/digitalbro/multimodal-sc-integration-meta-resources\" target=\"_blank\">https://www.kaggle.com/code/digitalbro/multimodal-sc-integration-meta-resources</a>)<br>\nwhich provide a lot of information as well. Also, Sbunzini's response will give you a sparse version of the data, which will help when running the notebooks on Kaggle. I'd recommend looking at using Saturn Cloud </li>\n<li><a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346999\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346999</a> <br>\nfor that as it might help not to have to run into the constraints of the Kaggle VMs. Cheers!  </li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1926563,
          "author_name": "marcyane7",
          "author_url": "",
          "post_date": "09/05/2022 00:28:40",
          "content": "<p>Hello, I actually used the line of code you proposed above but it loads the whole dataset anyways, do you have any insight  as to why?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1926593,
          "author_name": "realapex",
          "author_url": "",
          "post_date": "09/05/2022 01:23:24",
          "content": "<p>I do not know why it does not work well for you, maybe you didn't do it properly?<br>\nHere is the code I executed and It worked well.<br>\n<code>\n!pip install tables\nimport pandas as pd\nfrom pandas import HDFStore\nx = \"../input/open-problems-multimodal/test_multi_inputs.h5\"\npd.read_hdf(x, start = 0, stop = 50)\n</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1926444,
      "author_name": "sbunzini",
      "author_url": "",
      "post_date": "09/04/2022 20:50:45",
      "content": "<p>Hey RealApex, we found that Multiome dataset can be widely compressed in CSR matrix format. Take a look at these notebooks:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/sbunzini/reduce-memory-usage-by-95-with-sparse-matrices\" target=\"_blank\">https://www.kaggle.com/code/sbunzini/reduce-memory-usage-by-95-with-sparse-matrices</a></li>\n<li><a href=\"https://www.kaggle.com/code/fabiencrom/multimodal-single-cell-creating-sparse-data\" target=\"_blank\">https://www.kaggle.com/code/fabiencrom/multimodal-single-cell-creating-sparse-data</a></li>\n</ul>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1926410": "Hello, \nI am new to this competition and I am having some trouble using the given data. So far, I can only open a few of the H5 files using pandas but all the others take up too much memory. Is there an alternative way to see the contents of an H5 file? \n\nThank you.",
    "1926443": "You don't have to read all of the information at once. You can use use the start/stops parameters. \n\n```\ntargets = pd.read_hdf('../input/targets.h5', start=0, stop=6000)\n```\n\nYou can find notebooks like \n\n* https://www.kaggle.com/code/digitalbro/multimodal-sc-integration-meta-resources)\n\nwhich provide a lot of information as well. Also, Sbunzini's response will give you a sparse version of the data, which will help when running the notebooks on Kaggle. I'd recommend looking at using Saturn Cloud \n\n* https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346999 \n\nfor that as it might help not to have to run into the constraints of the Kaggle VMs. Cheers!",
    "1926444": "Hey RealApex, we found that Multiome dataset can be widely compressed in CSR matrix format. Take a look at these notebooks:\n\n- https://www.kaggle.com/code/sbunzini/reduce-memory-usage-by-95-with-sparse-matrices\n- https://www.kaggle.com/code/fabiencrom/multimodal-single-cell-creating-sparse-data",
    "1926563": "Hello, I actually used the line of code you proposed above but it loads the whole dataset anyways, do you have any insight  as to why?",
    "1926593": "I do not know why it does not work well for you, maybe you didn't do it properly?\nHere is the code I executed and It worked well.\n`\n!pip install tables\nimport pandas as pd\nfrom pandas import HDFStore\nx = \"../input/open-problems-multimodal/test_multi_inputs.h5\"\npd.read_hdf(x, start = 0, stop = 50)\n`"
  },
  "source": "meta"
}