{
  "id": 347865,
  "title": "multiome train divided by donors .parquet",
  "url": "/competitions/open-problems-multimodal/discussion/347865",
  "author_name": "Oleg Khudyakov",
  "post_date": "2022-08-25T17:25:28.322000",
  "votes": 9,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hello!<br>\nWe can't load original <code>train_multi_inputs.h5</code> file in memory of Kaggle notebook - ok.<br>\nBut I was not able to load it in memory on my home PC with 64GB RAM - it was unpleasant surprise.<br>\nSo i divided it into three parts.<br>\nBy bare intuition I decided to divide it by donor number.<br>\nCommon sense told me it is a reasonable way.<br>\nAlso I saved it in parquet file format.<br>\n<a href=\"https://www.kaggle.com/datasets/kaggledummie007/singlecell-integration-multiome-train-divided\" target=\"_blank\">Here is the dataset.</a><br>\nIt is more convenient for me, at least I can read separate columns.<br>\nIf someone would like to use it - You are welcome.<br>\nRegards, Oleg.</p>",
  "messages": [
    {
      "id": 1914053,
      "postDate": "2022-08-25T17:25:28.323Z",
      "content": "<p>Hello!<br>\nWe can't load original <code>train_multi_inputs.h5</code> file in memory of Kaggle notebook - ok.<br>\nBut I was not able to load it in memory on my home PC with 64GB RAM - it was unpleasant surprise.<br>\nSo i divided it into three parts.<br>\nBy bare intuition I decided to divide it by donor number.<br>\nCommon sense told me it is a reasonable way.<br>\nAlso I saved it in parquet file format.<br>\n<a href=\"https://www.kaggle.com/datasets/kaggledummie007/singlecell-integration-multiome-train-divided\" target=\"_blank\">Here is the dataset.</a><br>\nIt is more convenient for me, at least I can read separate columns.<br>\nIf someone would like to use it - You are welcome.<br>\nRegards, Oleg.</p>",
      "rawMarkdown": "Hello!\nWe can't load original `train_multi_inputs.h5` file in memory of Kaggle notebook - ok.\nBut I was not able to load it in memory on my home PC with 64GB RAM - it was unpleasant surprise.\nSo i divided it into three parts.\nBy bare intuition I decided to divide it by donor number.\nCommon sense told me it is a reasonable way.\nAlso I saved it in parquet file format.\n[Here is the dataset.](https://www.kaggle.com/datasets/kaggledummie007/singlecell-integration-multiome-train-divided)\nIt is more convenient for me, at least I can read separate columns.\nIf someone would like to use it - You are welcome.\nRegards, Oleg.",
      "votes": 8
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1914053": "Hello!\nWe can't load original `train_multi_inputs.h5` file in memory of Kaggle notebook - ok.\nBut I was not able to load it in memory on my home PC with 64GB RAM - it was unpleasant surprise.\nSo i divided it into three parts.\nBy bare intuition I decided to divide it by donor number.\nCommon sense told me it is a reasonable way.\nAlso I saved it in parquet file format.\n[Here is the dataset.](https://www.kaggle.com/datasets/kaggledummie007/singlecell-integration-multiome-train-divided)\nIt is more convenient for me, at least I can read separate columns.\nIf someone would like to use it - You are welcome.\nRegards, Oleg."
  }
}