{
  "id": 347623,
  "title": "About the memory issue with multiome dataset, help needed",
  "url": "/competitions/open-problems-multimodal/discussion/347623",
  "author_name": "",
  "post_date": "2022-08-24T22:00:52.080691100Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Considering that a small subset of the multiome training dataset will be used? <br>\nHow does one properly sample the data so as to closely match the actual dataset, and is there an optimal number of rows to be selected? Currently selecting 5000 rows works, occasionally it also faces memory issue.</p>",
  "messages": [
    {
      "id": "1912638",
      "postDate": "08/24/2022 22:00:52",
      "content": "<p>Considering that a small subset of the multiome training dataset will be used? <br>\nHow does one properly sample the data so as to closely match the actual dataset, and is there an optimal number of rows to be selected? Currently selecting 5000 rows works, occasionally it also faces memory issue.</p>",
      "rawMarkdown": "Considering that a small subset of the multiome training dataset will be used? \nHow does one properly sample the data so as to closely match the actual dataset, and is there an optimal number of rows to be selected? Currently selecting 5000 rows works, occasionally it also faces memory issue.",
      "votes": null
    },
    {
      "id": "1912650",
      "postDate": "08/24/2022 22:22:36",
      "content": "<p>The answer has been provided in this discussions<br>\n(<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346690\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346690</a>)<br>\n(<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/344829#1905232\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/344829#1905232</a>)</p>",
      "rawMarkdown": "The answer has been provided in this discussions\n(https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346690)\n(https://www.kaggle.com/competitions/open-problems-multimodal/discussion/344829#1905232)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1912650,
      "author_name": "samu2505",
      "author_url": "",
      "post_date": "08/24/2022 22:22:36",
      "content": "<p>The answer has been provided in this discussions<br>\n(<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346690\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346690</a>)<br>\n(<a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/344829#1905232\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/344829#1905232</a>)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1912638": "Considering that a small subset of the multiome training dataset will be used? \nHow does one properly sample the data so as to closely match the actual dataset, and is there an optimal number of rows to be selected? Currently selecting 5000 rows works, occasionally it also faces memory issue.",
    "1912650": "The answer has been provided in this discussions\n(https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346690)\n(https://www.kaggle.com/competitions/open-problems-multimodal/discussion/344829#1905232)"
  },
  "source": "meta"
}