{
  "id": 344829,
  "title": "To the Organizers: Why not h5ad files?",
  "url": "/competitions/open-problems-multimodal/discussion/344829",
  "author_name": "",
  "post_date": "2022-08-16T19:06:43.601498400Z",
  "votes": 14,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hi, first of all, thank you for this amazing competition, I'm a single-cell entusiast and I truly enjoy this kind of data. </p>\n<p>My question is, why using <strong>.h5</strong> instead of the standard <strong>.h5ad</strong>?</p>",
  "messages": [
    {
      "id": "1901581",
      "postDate": "08/16/2022 19:06:43",
      "content": "<p>Hi, first of all, thank you for this amazing competition, I'm a single-cell entusiast and I truly enjoy this kind of data. </p>\n<p>My question is, why using <strong>.h5</strong> instead of the standard <strong>.h5ad</strong>?</p>",
      "rawMarkdown": "Hi, first of all, thank you for this amazing competition, I'm a single-cell entusiast and I truly enjoy this kind of data. \n\nMy question is, why using **.h5** instead of the standard **.h5ad**?",
      "votes": null
    },
    {
      "id": "1901606",
      "postDate": "08/16/2022 19:23:02",
      "content": "<p>Hi Hiram!</p>\n<p>Great question, and we thought a lot about <a href=\"https://github.com/scverse/anndata\" target=\"_blank\"><code>h5ad</code></a> vs. <a href=\"https://github.com/scverse/mudata\" target=\"_blank\"><code>h5mu</code></a> vs. <a href=\"https://www.pytables.org/cookbook/inmemory_hdf5_files.html\" target=\"_blank\"><code>h5</code></a> within Open Problems and in collaboration with our data science team at Kaggle. We made this choice for two main reasons.</p>\n<ul>\n<li><strong><code>h5</code> can be directly read from pandas without requiring additional software.</strong> Using <code>h5ad</code> or <code>h5mu</code> would require competitors to use a toolkit that isn't standard in the Python data science / ML stack. Consider that scanpy / AnnData has 1.3k stars on GitHub while Pandas has 34.9k.</li>\n<li><strong>We hit issues with Muon while developing the dataset.</strong> Muon would have been a more natural choice for multimodal data, but we ended up having numerous issues with reading and writing h5mu files while working on preparing the dataset (<a href=\"https://github.com/scverse/muon/issues/57#issuecomment-1172656861\" target=\"_blank\">example</a>). </li>\n</ul>\n<p>Sadly, the single-cell stack just isn't as robust or widely used as the more foundational tools in the Numpy / Pandas / scikit-learn ecosystems.</p>\n<p>All this said, we would definitely encourage you to create a code notebook for the competition that loads the data and formats it into a different data structure! </p>\n<p>Best,<br>\nDaniel</p>",
      "rawMarkdown": "Hi Hiram!\n\nGreat question, and we thought a lot about [`h5ad`](https://github.com/scverse/anndata) vs. [`h5mu`](https://github.com/scverse/mudata) vs. [`h5`](https://www.pytables.org/cookbook/inmemory_hdf5_files.html) within Open Problems and in collaboration with our data science team at Kaggle. We made this choice for two main reasons.\n\n* **`h5` can be directly read from pandas without requiring additional software.** Using `h5ad` or `h5mu` would require competitors to use a toolkit that isn't standard in the Python data science / ML stack. Consider that scanpy / AnnData has 1.3k stars on GitHub while Pandas has 34.9k.\n* **We hit issues with Muon while developing the dataset.** Muon would have been a more natural choice for multimodal data, but we ended up having numerous issues with reading and writing h5mu files while working on preparing the dataset ([example](https://github.com/scverse/muon/issues/57#issuecomment-1172656861)). \n\nSadly, the single-cell stack just isn't as robust or widely used as the more foundational tools in the Numpy / Pandas / scikit-learn ecosystems.\n\nAll this said, we would definitely encourage you to create a code notebook for the competition that loads the data and formats it into a different data structure! \n\nBest,\nDaniel",
      "votes": null
    },
    {
      "id": "1901612",
      "postDate": "08/16/2022 19:27:13",
      "content": "<p>Hi Daniel, excelent answer, but, I'm going to invest some time trying to wrang the data for the <strong>AnnData</strong> format.</p>",
      "rawMarkdown": "Hi Daniel, excelent answer, but, I'm going to invest some time trying to wrang the data for the **AnnData** format.",
      "votes": null
    },
    {
      "id": "1901613",
      "postDate": "08/16/2022 19:27:25",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hiramcho\" target=\"_blank\">@hiramcho</a> ,</p>\n<p>We opted for the h5 format instead of h5ad to improve accessibility to the data for those who haven't encountered this sort of task before. The h5 format can be read directly with Pandas, for instance, which many people are already familiar with.</p>\n<p>I noticed from your other thread that you have some expertise in this area. You'd be welcome to create a h5ad version and share it with the community along with your notebooks, if you like. That could be great contribution to the community and a way to introduce others to single-cell data who are just getting started.</p>",
      "rawMarkdown": "Hi @hiramcho ,\n\nWe opted for the h5 format instead of h5ad to improve accessibility to the data for those who haven't encountered this sort of task before. The h5 format can be read directly with Pandas, for instance, which many people are already familiar with.\n\nI noticed from your other thread that you have some expertise in this area. You'd be welcome to create a h5ad version and share it with the community along with your notebooks, if you like. That could be great contribution to the community and a way to introduce others to single-cell data who are just getting started.",
      "votes": null
    },
    {
      "id": "1902939",
      "postDate": "08/17/2022 02:27:38",
      "content": "<p>Great discussion, <a href=\"https://www.kaggle.com/hiramcho\" target=\"_blank\">@hiramcho</a>, <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> and <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a>!<br>\nI tried to import some training data into <code>h5ad</code> and <code>h5mu</code> files, seems to work (<a href=\"https://www.kaggle.com/code/daniorerio/getting-started-scanpy-muon\" target=\"_blank\">notebook</a>). <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a>, Muon files seem to work fine so far, by the way, the issue you linked seems to concern <a href=\"https://github.com/scverse/anndata/issues/731\" target=\"_blank\">some breaking AnnData changes</a>.</p>",
      "rawMarkdown": "Great discussion, @hiramcho, @ryanholbrook and @danielburkhardt!\nI tried to import some training data into `h5ad` and `h5mu` files, seems to work ([notebook](https://www.kaggle.com/code/daniorerio/getting-started-scanpy-muon)). @danielburkhardt, Muon files seem to work fine so far, by the way, the issue you linked seems to concern [some breaking AnnData changes](https://github.com/scverse/anndata/issues/731).",
      "votes": null
    },
    {
      "id": "1903429",
      "postDate": "08/17/2022 12:28:42",
      "content": "<p>Yup! To be clear, we love muon and mudata at Open Problems, and we use these data formats every day in our work at Cellarity. The thought process here was just, \"what's going to be the most accessible format for the data science community on Kaggle.\"</p>\n<p>I love to see folks pushing more data formats to the platform.</p>",
      "rawMarkdown": "Yup! To be clear, we love muon and mudata at Open Problems, and we use these data formats every day in our work at Cellarity. The thought process here was just, \"what's going to be the most accessible format for the data science community on Kaggle.\"\n\nI love to see folks pushing more data formats to the platform.",
      "votes": null
    },
    {
      "id": "1903875",
      "postDate": "08/17/2022 18:04:39",
      "content": "<p>I think that makes sense, especially in the context of the larger community here, thanks for organizing!</p>",
      "rawMarkdown": "I think that makes sense, especially in the context of the larger community here, thanks for organizing!",
      "votes": null
    },
    {
      "id": "1905152",
      "postDate": "08/18/2022 18:59:01",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hiramcho\" target=\"_blank\">@hiramcho</a>,</p>\n<p>do you know a way to use/handle .h5ad files <em>without</em> loading them fully into memory? I looked up the anndata library, but apparently the only way to read .h5ad files implies to load all the data into memory.</p>\n<p>More generally, I am looking for a way to query these files without the need to load them. For example, it would be great to be able to load only data from a specific donor, or a combination of donor/day.</p>\n<p>Thank you in advance!</p>",
      "rawMarkdown": "Hi @hiramcho,\n\ndo you know a way to use/handle .h5ad files *without* loading them fully into memory? I looked up the anndata library, but apparently the only way to read .h5ad files implies to load all the data into memory.\n\nMore generally, I am looking for a way to query these files without the need to load them. For example, it would be great to be able to load only data from a specific donor, or a combination of donor/day.\n\nThank you in advance!",
      "votes": null
    },
    {
      "id": "1905232",
      "postDate": "08/18/2022 20:57:25",
      "content": "<p>You can do it like that:</p>\n<p>adata = sc.read(fn, backed='r' )</p>\n<p>Example: <a href=\"https://www.kaggle.com/code/alexandervc/scanpy-process-huge-files-backing-them-on-disk\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/scanpy-process-huge-files-backing-them-on-disk</a></p>\n<p>Similar can be done for h5 files, see:</p>\n<p><a href=\"https://www.kaggle.com/alexandervc/use-h5py-for-huge-h5-file-backing-it-on-disk\" target=\"_blank\">https://www.kaggle.com/alexandervc/use-h5py-for-huge-h5-file-backing-it-on-disk</a></p>",
      "rawMarkdown": "You can do it like that:\n\nadata = sc.read(fn, backed='r' )\n\nExample: https://www.kaggle.com/code/alexandervc/scanpy-process-huge-files-backing-them-on-disk\n\nSimilar can be done for h5 files, see:\n\nhttps://www.kaggle.com/alexandervc/use-h5py-for-huge-h5-file-backing-it-on-disk",
      "votes": null
    },
    {
      "id": "1906063",
      "postDate": "08/19/2022 14:36:42",
      "content": "<p>Thank you Alexander, I will give it a try</p>",
      "rawMarkdown": "Thank you Alexander, I will give it a try",
      "votes": null
    },
    {
      "id": "1928959",
      "postDate": "09/06/2022 18:07:12",
      "content": "<p><a href=\"https://www.kaggle.com/datasets/alexandervc/multimodal-singlecell-integration-related-data-01\" target=\"_blank\">https://www.kaggle.com/datasets/alexandervc/multimodal-singlecell-integration-related-data-01</a></p>\n<p>Put CITE-seq input data into single h5ad - but it is difficult to work with it with 16G memory.<br>\n(Both train and test is there). </p>",
      "rawMarkdown": "https://www.kaggle.com/datasets/alexandervc/multimodal-singlecell-integration-related-data-01\n\nPut CITE-seq input data into single h5ad - but it is difficult to work with it with 16G memory.\n(Both train and test is there).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1901606,
      "author_name": "danielburkhardt",
      "author_url": "",
      "post_date": "08/16/2022 19:23:02",
      "content": "<p>Hi Hiram!</p>\n<p>Great question, and we thought a lot about <a href=\"https://github.com/scverse/anndata\" target=\"_blank\"><code>h5ad</code></a> vs. <a href=\"https://github.com/scverse/mudata\" target=\"_blank\"><code>h5mu</code></a> vs. <a href=\"https://www.pytables.org/cookbook/inmemory_hdf5_files.html\" target=\"_blank\"><code>h5</code></a> within Open Problems and in collaboration with our data science team at Kaggle. We made this choice for two main reasons.</p>\n<ul>\n<li><strong><code>h5</code> can be directly read from pandas without requiring additional software.</strong> Using <code>h5ad</code> or <code>h5mu</code> would require competitors to use a toolkit that isn't standard in the Python data science / ML stack. Consider that scanpy / AnnData has 1.3k stars on GitHub while Pandas has 34.9k.</li>\n<li><strong>We hit issues with Muon while developing the dataset.</strong> Muon would have been a more natural choice for multimodal data, but we ended up having numerous issues with reading and writing h5mu files while working on preparing the dataset (<a href=\"https://github.com/scverse/muon/issues/57#issuecomment-1172656861\" target=\"_blank\">example</a>). </li>\n</ul>\n<p>Sadly, the single-cell stack just isn't as robust or widely used as the more foundational tools in the Numpy / Pandas / scikit-learn ecosystems.</p>\n<p>All this said, we would definitely encourage you to create a code notebook for the competition that loads the data and formats it into a different data structure! </p>\n<p>Best,<br>\nDaniel</p>",
      "votes": null,
      "replies": [
        {
          "id": 1901612,
          "author_name": "hiramcho",
          "author_url": "",
          "post_date": "08/16/2022 19:27:13",
          "content": "<p>Hi Daniel, excelent answer, but, I'm going to invest some time trying to wrang the data for the <strong>AnnData</strong> format.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1901613,
      "author_name": "ryanholbrook",
      "author_url": "",
      "post_date": "08/16/2022 19:27:25",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hiramcho\" target=\"_blank\">@hiramcho</a> ,</p>\n<p>We opted for the h5 format instead of h5ad to improve accessibility to the data for those who haven't encountered this sort of task before. The h5 format can be read directly with Pandas, for instance, which many people are already familiar with.</p>\n<p>I noticed from your other thread that you have some expertise in this area. You'd be welcome to create a h5ad version and share it with the community along with your notebooks, if you like. That could be great contribution to the community and a way to introduce others to single-cell data who are just getting started.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1902939,
      "author_name": "daniorerio",
      "author_url": "",
      "post_date": "08/17/2022 02:27:38",
      "content": "<p>Great discussion, <a href=\"https://www.kaggle.com/hiramcho\" target=\"_blank\">@hiramcho</a>, <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> and <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a>!<br>\nI tried to import some training data into <code>h5ad</code> and <code>h5mu</code> files, seems to work (<a href=\"https://www.kaggle.com/code/daniorerio/getting-started-scanpy-muon\" target=\"_blank\">notebook</a>). <a href=\"https://www.kaggle.com/danielburkhardt\" target=\"_blank\">@danielburkhardt</a>, Muon files seem to work fine so far, by the way, the issue you linked seems to concern <a href=\"https://github.com/scverse/anndata/issues/731\" target=\"_blank\">some breaking AnnData changes</a>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1903429,
          "author_name": "danielburkhardt",
          "author_url": "",
          "post_date": "08/17/2022 12:28:42",
          "content": "<p>Yup! To be clear, we love muon and mudata at Open Problems, and we use these data formats every day in our work at Cellarity. The thought process here was just, \"what's going to be the most accessible format for the data science community on Kaggle.\"</p>\n<p>I love to see folks pushing more data formats to the platform.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1903875,
          "author_name": "daniorerio",
          "author_url": "",
          "post_date": "08/17/2022 18:04:39",
          "content": "<p>I think that makes sense, especially in the context of the larger community here, thanks for organizing!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1905152,
      "author_name": "alekeuro",
      "author_url": "",
      "post_date": "08/18/2022 18:59:01",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hiramcho\" target=\"_blank\">@hiramcho</a>,</p>\n<p>do you know a way to use/handle .h5ad files <em>without</em> loading them fully into memory? I looked up the anndata library, but apparently the only way to read .h5ad files implies to load all the data into memory.</p>\n<p>More generally, I am looking for a way to query these files without the need to load them. For example, it would be great to be able to load only data from a specific donor, or a combination of donor/day.</p>\n<p>Thank you in advance!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1905232,
          "author_name": "alexandervc",
          "author_url": "",
          "post_date": "08/18/2022 20:57:25",
          "content": "<p>You can do it like that:</p>\n<p>adata = sc.read(fn, backed='r' )</p>\n<p>Example: <a href=\"https://www.kaggle.com/code/alexandervc/scanpy-process-huge-files-backing-them-on-disk\" target=\"_blank\">https://www.kaggle.com/code/alexandervc/scanpy-process-huge-files-backing-them-on-disk</a></p>\n<p>Similar can be done for h5 files, see:</p>\n<p><a href=\"https://www.kaggle.com/alexandervc/use-h5py-for-huge-h5-file-backing-it-on-disk\" target=\"_blank\">https://www.kaggle.com/alexandervc/use-h5py-for-huge-h5-file-backing-it-on-disk</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1906063,
          "author_name": "alekeuro",
          "author_url": "",
          "post_date": "08/19/2022 14:36:42",
          "content": "<p>Thank you Alexander, I will give it a try</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1928959,
      "author_name": "alexandervc",
      "author_url": "",
      "post_date": "09/06/2022 18:07:12",
      "content": "<p><a href=\"https://www.kaggle.com/datasets/alexandervc/multimodal-singlecell-integration-related-data-01\" target=\"_blank\">https://www.kaggle.com/datasets/alexandervc/multimodal-singlecell-integration-related-data-01</a></p>\n<p>Put CITE-seq input data into single h5ad - but it is difficult to work with it with 16G memory.<br>\n(Both train and test is there). </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1901581": "Hi, first of all, thank you for this amazing competition, I'm a single-cell entusiast and I truly enjoy this kind of data. \n\nMy question is, why using **.h5** instead of the standard **.h5ad**?",
    "1901606": "Hi Hiram!\n\nGreat question, and we thought a lot about [`h5ad`](https://github.com/scverse/anndata) vs. [`h5mu`](https://github.com/scverse/mudata) vs. [`h5`](https://www.pytables.org/cookbook/inmemory_hdf5_files.html) within Open Problems and in collaboration with our data science team at Kaggle. We made this choice for two main reasons.\n\n* **`h5` can be directly read from pandas without requiring additional software.** Using `h5ad` or `h5mu` would require competitors to use a toolkit that isn't standard in the Python data science / ML stack. Consider that scanpy / AnnData has 1.3k stars on GitHub while Pandas has 34.9k.\n* **We hit issues with Muon while developing the dataset.** Muon would have been a more natural choice for multimodal data, but we ended up having numerous issues with reading and writing h5mu files while working on preparing the dataset ([example](https://github.com/scverse/muon/issues/57#issuecomment-1172656861)). \n\nSadly, the single-cell stack just isn't as robust or widely used as the more foundational tools in the Numpy / Pandas / scikit-learn ecosystems.\n\nAll this said, we would definitely encourage you to create a code notebook for the competition that loads the data and formats it into a different data structure! \n\nBest,\nDaniel",
    "1901612": "Hi Daniel, excelent answer, but, I'm going to invest some time trying to wrang the data for the **AnnData** format.",
    "1901613": "Hi @hiramcho ,\n\nWe opted for the h5 format instead of h5ad to improve accessibility to the data for those who haven't encountered this sort of task before. The h5 format can be read directly with Pandas, for instance, which many people are already familiar with.\n\nI noticed from your other thread that you have some expertise in this area. You'd be welcome to create a h5ad version and share it with the community along with your notebooks, if you like. That could be great contribution to the community and a way to introduce others to single-cell data who are just getting started.",
    "1902939": "Great discussion, @hiramcho, @ryanholbrook and @danielburkhardt!\nI tried to import some training data into `h5ad` and `h5mu` files, seems to work ([notebook](https://www.kaggle.com/code/daniorerio/getting-started-scanpy-muon)). @danielburkhardt, Muon files seem to work fine so far, by the way, the issue you linked seems to concern [some breaking AnnData changes](https://github.com/scverse/anndata/issues/731).",
    "1903429": "Yup! To be clear, we love muon and mudata at Open Problems, and we use these data formats every day in our work at Cellarity. The thought process here was just, \"what's going to be the most accessible format for the data science community on Kaggle.\"\n\nI love to see folks pushing more data formats to the platform.",
    "1903875": "I think that makes sense, especially in the context of the larger community here, thanks for organizing!",
    "1905152": "Hi @hiramcho,\n\ndo you know a way to use/handle .h5ad files *without* loading them fully into memory? I looked up the anndata library, but apparently the only way to read .h5ad files implies to load all the data into memory.\n\nMore generally, I am looking for a way to query these files without the need to load them. For example, it would be great to be able to load only data from a specific donor, or a combination of donor/day.\n\nThank you in advance!",
    "1905232": "You can do it like that:\n\nadata = sc.read(fn, backed='r' )\n\nExample: https://www.kaggle.com/code/alexandervc/scanpy-process-huge-files-backing-them-on-disk\n\nSimilar can be done for h5 files, see:\n\nhttps://www.kaggle.com/alexandervc/use-h5py-for-huge-h5-file-backing-it-on-disk",
    "1906063": "Thank you Alexander, I will give it a try",
    "1928959": "https://www.kaggle.com/datasets/alexandervc/multimodal-singlecell-integration-related-data-01\n\nPut CITE-seq input data into single h5ad - but it is difficult to work with it with 16G memory.\n(Both train and test is there)."
  },
  "source": "meta"
}