{
  "id": 345958,
  "title": "How to be ready for this competition ?",
  "url": "/competitions/open-problems-multimodal/discussion/345958",
  "author_name": "",
  "post_date": "2022-08-17T10:15:35.727548200Z",
  "votes": 7,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hello</p>\n<p>I'm pretty new here.  <br>\nSeeing the volumn of data is more than 28G, I'd like to know if my GPU with 8G RAM is possible to run experiment afterwards ? <br>\nOr maybe we don't need GPU to run DL model, then 8 cores CPU with 32G RAM are possible for this competition ?<br>\nOf course, Kaggle notebook is usable. But I'd like to know about what I will need in other platforms.</p>\n<p>Thanks for helping</p>\n<p>Update: <br>\nThe competition host just annonnced:<br>\n\" free compute with 16GB GPU, 128GB RAM, and 32 CPUs:</p>\n<p>https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346999\"</p>",
  "messages": [
    {
      "id": "1903329",
      "postDate": "08/17/2022 10:15:35",
      "content": "<p>Hello</p>\n<p>I'm pretty new here.  <br>\nSeeing the volumn of data is more than 28G, I'd like to know if my GPU with 8G RAM is possible to run experiment afterwards ? <br>\nOr maybe we don't need GPU to run DL model, then 8 cores CPU with 32G RAM are possible for this competition ?<br>\nOf course, Kaggle notebook is usable. But I'd like to know about what I will need in other platforms.</p>\n<p>Thanks for helping</p>\n<p>Update: <br>\nThe competition host just annonnced:<br>\n\" free compute with 16GB GPU, 128GB RAM, and 32 CPUs:</p>\n<p>https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346999\"</p>",
      "rawMarkdown": "Hello\n\nI'm pretty new here.  \nSeeing the volumn of data is more than 28G, I'd like to know if my GPU with 8G RAM is possible to run experiment afterwards ? \nOr maybe we don't need GPU to run DL model, then 8 cores CPU with 32G RAM are possible for this competition ?\nOf course, Kaggle notebook is usable. But I'd like to know about what I will need in other platforms.\n\nThanks for helping\n\nUpdate: \nThe competition host just annonnced:\n\" free compute with 16GB GPU, 128GB RAM, and 32 CPUs:\n\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/346999\"",
      "votes": null
    },
    {
      "id": "1903493",
      "postDate": "08/17/2022 13:25:23",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/chouchouchen\" target=\"_blank\">@chouchouchen</a>!</p>\n<p>The data set volume is definitely large and it is a challenge to handle this. While this is part of the challenge as it is for researchers using such data, I can comment on your specific enquiry:</p>\n<ul>\n<li><p>There are methods that allow you to handle the data without having to load the full data in memory (e.g. <a href=\"https://docs.xarray.dev/en/v0.9.2/dask.html)\" target=\"_blank\">https://docs.xarray.dev/en/v0.9.2/dask.html)</a>. These might be of help, although I have not tried them on this specific data set.</p></li>\n<li><p>For deep learning, you usually train your models in minibatches of sizes &lt;1000. And usually, you also load only these minibatches to the GPU. So the full data set never lies on the GPU. </p></li>\n</ul>\n<p>Does that answer your question?</p>",
      "rawMarkdown": "Hi @chouchouchen!\n\nThe data set volume is definitely large and it is a challenge to handle this. While this is part of the challenge as it is for researchers using such data, I can comment on your specific enquiry:\n\n- There are methods that allow you to handle the data without having to load the full data in memory (e.g. https://docs.xarray.dev/en/v0.9.2/dask.html). These might be of help, although I have not tried them on this specific data set.\n\n- For deep learning, you usually train your models in minibatches of sizes <1000. And usually, you also load only these minibatches to the GPU. So the full data set never lies on the GPU. \n \nDoes that answer your question?",
      "votes": null
    },
    {
      "id": "1903666",
      "postDate": "08/17/2022 15:03:31",
      "content": "<p>Thanks for your answer.<br>\nI guess that I just need to have a try. Yes, there are always ways to handle huge dataset with limit resources. However, somtimes when dataset is really big, it became almost impossible to compete with others, who have way more resources. </p>\n<p>By the way, thanks for organizing the competition👍</p>",
      "rawMarkdown": "Thanks for your answer.\nI guess that I just need to have a try. Yes, there are always ways to handle huge dataset with limit resources. However, somtimes when dataset is really big, it became almost impossible to compete with others, who have way more resources. \n\nBy the way, thanks for organizing the competition👍",
      "votes": null
    },
    {
      "id": "1903723",
      "postDate": "08/17/2022 15:45:23",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/chouchouchen\" target=\"_blank\">@chouchouchen</a> , I'm also struggling with the file sizes.</p>\n<p>Besides what <a href=\"https://www.kaggle.com/peterholderrieth\" target=\"_blank\">@peterholderrieth</a> mentioned regarding training a DL model, you might consider using callbacks in your training loop. If the model gets too big, even a modest mini-batch size can yield a \"cuda out of memory error\". However, you could use a callback such as Gradient Accumulation, that allows you to use a smaller mini-batch size (say, 16) and update your weights only when a specific number of mini-batches (say, 8) have gone through the training loop. In this way, you end up with an \"effective\" mini-batch size of 16*8 = 128.</p>\n<p>Still, I find a previous bottleneck, which is handling the data efficiently <em>before</em> even thinking to train a model. I might try Peter's suggestion (xarray/dask). Another possibility I'm thinking about is to create torch Datasets by reading specific rows of data directly from the .h5 files. That's going to bee painfully slow, but at least might work..</p>",
      "rawMarkdown": "Hi @chouchouchen , I'm also struggling with the file sizes.\n\nBesides what @peterholderrieth mentioned regarding training a DL model, you might consider using callbacks in your training loop. If the model gets too big, even a modest mini-batch size can yield a \"cuda out of memory error\". However, you could use a callback such as Gradient Accumulation, that allows you to use a smaller mini-batch size (say, 16) and update your weights only when a specific number of mini-batches (say, 8) have gone through the training loop. In this way, you end up with an \"effective\" mini-batch size of 16*8 = 128.\n\nStill, I find a previous bottleneck, which is handling the data efficiently *before* even thinking to train a model. I might try Peter's suggestion (xarray/dask). Another possibility I'm thinking about is to create torch Datasets by reading specific rows of data directly from the .h5 files. That's going to bee painfully slow, but at least might work..",
      "votes": null
    },
    {
      "id": "1903733",
      "postDate": "08/17/2022 15:56:23",
      "content": "<p>Hello Alejo,<br>\nThanks for your suggestion. <br>\nYeah, handling data before model training is already a challenge. </p>",
      "rawMarkdown": "Hello Alejo,\nThanks for your suggestion. \nYeah, handling data before model training is already a challenge.",
      "votes": null
    },
    {
      "id": "1904614",
      "postDate": "08/18/2022 10:50:14",
      "content": "<p>Have you tried it(xarray/dask) yet?,it‘s worked?</p>",
      "rawMarkdown": "Have you tried it(xarray/dask) yet?,it‘s worked?",
      "votes": null
    },
    {
      "id": "1904696",
      "postDate": "08/18/2022 12:12:46",
      "content": "<p>Not yet. Will try a week later</p>",
      "rawMarkdown": "Not yet. Will try a week later",
      "votes": null
    },
    {
      "id": "1905206",
      "postDate": "08/18/2022 20:16:31",
      "content": "<p>A potential workaround could be to subset the data into smaller batches for cleaning/training, and then when you want to test your model you can divide the dataset into workable chunks and concatenate your results into a final set, Hopefully this helps you. </p>",
      "rawMarkdown": "A potential workaround could be to subset the data into smaller batches for cleaning/training, and then when you want to test your model you can divide the dataset into workable chunks and concatenate your results into a final set, Hopefully this helps you.",
      "votes": null
    },
    {
      "id": "1908204",
      "postDate": "08/21/2022 12:52:16",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/chouchouchen\" target=\"_blank\">@chouchouchen</a> ,</p>\n<p>In this <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346690\" target=\"_blank\">discussion entry</a> I talk about an approach to read subsets of *.h5 files without loading them fully into memory. There is a link to a notebook I created illustrating this approach. Essentially, I wanted to have an easy way to query these files, to get specific subsets of data (e.g., data from day 4, for donor 3, and cell_type = \"BP\"). Hopefully you might find it useful.. and sorry for the shameless self-promotion 😅</p>",
      "rawMarkdown": "Hi @chouchouchen ,\n\nIn this [discussion entry](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346690) I talk about an approach to read subsets of *.h5 files without loading them fully into memory. There is a link to a notebook I created illustrating this approach. Essentially, I wanted to have an easy way to query these files, to get specific subsets of data (e.g., data from day 4, for donor 3, and cell_type = \"BP\"). Hopefully you might find it useful.. and sorry for the shameless self-promotion 😅",
      "votes": null
    },
    {
      "id": "1908214",
      "postDate": "08/21/2022 13:05:33",
      "content": "<p>Hi thanks for your suggestions. In fact, there are already some notebooks where we can load part of data to do EDA.</p>",
      "rawMarkdown": "Hi thanks for your suggestions. In fact, there are already some notebooks where we can load part of data to do EDA.",
      "votes": null
    },
    {
      "id": "1909234",
      "postDate": "08/22/2022 12:57:19",
      "content": "<p>Hi all! To help with this, we just announced free compute with 16GB GPU, 128GB RAM, and 32 CPUs:</p>\n<p><a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346999\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346999</a></p>",
      "rawMarkdown": "Hi all! To help with this, we just announced free compute with 16GB GPU, 128GB RAM, and 32 CPUs:\n\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/346999",
      "votes": null
    },
    {
      "id": "1909257",
      "postDate": "08/22/2022 13:18:10",
      "content": "<p>Thank you ! Great news for many of us</p>",
      "rawMarkdown": "Thank you ! Great news for many of us",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1903493,
      "author_name": "peterholderrieth",
      "author_url": "",
      "post_date": "08/17/2022 13:25:23",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/chouchouchen\" target=\"_blank\">@chouchouchen</a>!</p>\n<p>The data set volume is definitely large and it is a challenge to handle this. While this is part of the challenge as it is for researchers using such data, I can comment on your specific enquiry:</p>\n<ul>\n<li><p>There are methods that allow you to handle the data without having to load the full data in memory (e.g. <a href=\"https://docs.xarray.dev/en/v0.9.2/dask.html)\" target=\"_blank\">https://docs.xarray.dev/en/v0.9.2/dask.html)</a>. These might be of help, although I have not tried them on this specific data set.</p></li>\n<li><p>For deep learning, you usually train your models in minibatches of sizes &lt;1000. And usually, you also load only these minibatches to the GPU. So the full data set never lies on the GPU. </p></li>\n</ul>\n<p>Does that answer your question?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1903666,
          "author_name": "chouchouchen",
          "author_url": "",
          "post_date": "08/17/2022 15:03:31",
          "content": "<p>Thanks for your answer.<br>\nI guess that I just need to have a try. Yes, there are always ways to handle huge dataset with limit resources. However, somtimes when dataset is really big, it became almost impossible to compete with others, who have way more resources. </p>\n<p>By the way, thanks for organizing the competition👍</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1903723,
          "author_name": "alekeuro",
          "author_url": "",
          "post_date": "08/17/2022 15:45:23",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/chouchouchen\" target=\"_blank\">@chouchouchen</a> , I'm also struggling with the file sizes.</p>\n<p>Besides what <a href=\"https://www.kaggle.com/peterholderrieth\" target=\"_blank\">@peterholderrieth</a> mentioned regarding training a DL model, you might consider using callbacks in your training loop. If the model gets too big, even a modest mini-batch size can yield a \"cuda out of memory error\". However, you could use a callback such as Gradient Accumulation, that allows you to use a smaller mini-batch size (say, 16) and update your weights only when a specific number of mini-batches (say, 8) have gone through the training loop. In this way, you end up with an \"effective\" mini-batch size of 16*8 = 128.</p>\n<p>Still, I find a previous bottleneck, which is handling the data efficiently <em>before</em> even thinking to train a model. I might try Peter's suggestion (xarray/dask). Another possibility I'm thinking about is to create torch Datasets by reading specific rows of data directly from the .h5 files. That's going to bee painfully slow, but at least might work..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1903733,
          "author_name": "chouchouchen",
          "author_url": "",
          "post_date": "08/17/2022 15:56:23",
          "content": "<p>Hello Alejo,<br>\nThanks for your suggestion. <br>\nYeah, handling data before model training is already a challenge. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1904614,
          "author_name": "kunmingxie",
          "author_url": "",
          "post_date": "08/18/2022 10:50:14",
          "content": "<p>Have you tried it(xarray/dask) yet?,it‘s worked?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1904696,
          "author_name": "chouchouchen",
          "author_url": "",
          "post_date": "08/18/2022 12:12:46",
          "content": "<p>Not yet. Will try a week later</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1908204,
          "author_name": "alekeuro",
          "author_url": "",
          "post_date": "08/21/2022 12:52:16",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/chouchouchen\" target=\"_blank\">@chouchouchen</a> ,</p>\n<p>In this <a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346690\" target=\"_blank\">discussion entry</a> I talk about an approach to read subsets of *.h5 files without loading them fully into memory. There is a link to a notebook I created illustrating this approach. Essentially, I wanted to have an easy way to query these files, to get specific subsets of data (e.g., data from day 4, for donor 3, and cell_type = \"BP\"). Hopefully you might find it useful.. and sorry for the shameless self-promotion 😅</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1908214,
          "author_name": "chouchouchen",
          "author_url": "",
          "post_date": "08/21/2022 13:05:33",
          "content": "<p>Hi thanks for your suggestions. In fact, there are already some notebooks where we can load part of data to do EDA.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1905206,
      "author_name": "cameronriley",
      "author_url": "",
      "post_date": "08/18/2022 20:16:31",
      "content": "<p>A potential workaround could be to subset the data into smaller batches for cleaning/training, and then when you want to test your model you can divide the dataset into workable chunks and concatenate your results into a final set, Hopefully this helps you. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1909234,
      "author_name": "danielburkhardt",
      "author_url": "",
      "post_date": "08/22/2022 12:57:19",
      "content": "<p>Hi all! To help with this, we just announced free compute with 16GB GPU, 128GB RAM, and 32 CPUs:</p>\n<p><a href=\"https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346999\" target=\"_blank\">https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346999</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1909257,
          "author_name": "chouchouchen",
          "author_url": "",
          "post_date": "08/22/2022 13:18:10",
          "content": "<p>Thank you ! Great news for many of us</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1903329": "Hello\n\nI'm pretty new here.  \nSeeing the volumn of data is more than 28G, I'd like to know if my GPU with 8G RAM is possible to run experiment afterwards ? \nOr maybe we don't need GPU to run DL model, then 8 cores CPU with 32G RAM are possible for this competition ?\nOf course, Kaggle notebook is usable. But I'd like to know about what I will need in other platforms.\n\nThanks for helping\n\nUpdate: \nThe competition host just annonnced:\n\" free compute with 16GB GPU, 128GB RAM, and 32 CPUs:\n\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/346999\"",
    "1903493": "Hi @chouchouchen!\n\nThe data set volume is definitely large and it is a challenge to handle this. While this is part of the challenge as it is for researchers using such data, I can comment on your specific enquiry:\n\n- There are methods that allow you to handle the data without having to load the full data in memory (e.g. https://docs.xarray.dev/en/v0.9.2/dask.html). These might be of help, although I have not tried them on this specific data set.\n\n- For deep learning, you usually train your models in minibatches of sizes <1000. And usually, you also load only these minibatches to the GPU. So the full data set never lies on the GPU. \n \nDoes that answer your question?",
    "1903666": "Thanks for your answer.\nI guess that I just need to have a try. Yes, there are always ways to handle huge dataset with limit resources. However, somtimes when dataset is really big, it became almost impossible to compete with others, who have way more resources. \n\nBy the way, thanks for organizing the competition👍",
    "1903723": "Hi @chouchouchen , I'm also struggling with the file sizes.\n\nBesides what @peterholderrieth mentioned regarding training a DL model, you might consider using callbacks in your training loop. If the model gets too big, even a modest mini-batch size can yield a \"cuda out of memory error\". However, you could use a callback such as Gradient Accumulation, that allows you to use a smaller mini-batch size (say, 16) and update your weights only when a specific number of mini-batches (say, 8) have gone through the training loop. In this way, you end up with an \"effective\" mini-batch size of 16*8 = 128.\n\nStill, I find a previous bottleneck, which is handling the data efficiently *before* even thinking to train a model. I might try Peter's suggestion (xarray/dask). Another possibility I'm thinking about is to create torch Datasets by reading specific rows of data directly from the .h5 files. That's going to bee painfully slow, but at least might work..",
    "1903733": "Hello Alejo,\nThanks for your suggestion. \nYeah, handling data before model training is already a challenge.",
    "1904614": "Have you tried it(xarray/dask) yet?,it‘s worked?",
    "1904696": "Not yet. Will try a week later",
    "1905206": "A potential workaround could be to subset the data into smaller batches for cleaning/training, and then when you want to test your model you can divide the dataset into workable chunks and concatenate your results into a final set, Hopefully this helps you.",
    "1908204": "Hi @chouchouchen ,\n\nIn this [discussion entry](https://www.kaggle.com/competitions/open-problems-multimodal/discussion/346690) I talk about an approach to read subsets of *.h5 files without loading them fully into memory. There is a link to a notebook I created illustrating this approach. Essentially, I wanted to have an easy way to query these files, to get specific subsets of data (e.g., data from day 4, for donor 3, and cell_type = \"BP\"). Hopefully you might find it useful.. and sorry for the shameless self-promotion 😅",
    "1908214": "Hi thanks for your suggestions. In fact, there are already some notebooks where we can load part of data to do EDA.",
    "1909234": "Hi all! To help with this, we just announced free compute with 16GB GPU, 128GB RAM, and 32 CPUs:\n\nhttps://www.kaggle.com/competitions/open-problems-multimodal/discussion/346999",
    "1909257": "Thank you ! Great news for many of us"
  },
  "source": "meta"
}