{
  "id": 224073,
  "title": "I really can't figure out how to work with this 150GB+ data",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/224073",
  "author_name": "Joseph Assaker",
  "post_date": "2021-03-06T20:20:39.764000",
  "votes": 4,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Given that the data is HUGE (150GB+), and storage in an interactive Kaggle session is relatively very small, doing any kind of preprocessing would result in either memory or storage issues. Also, it is frankly impossible for me to download it locally and work with it. That's when I usually refer back to Google colab (so that I can download big datasets and save my intermediate results on drive). However, even colab has a limit of ~70GB of storage. Could there any other way to get the data in smaller chunks?</p>\n<p>The issue of data size is quite literally crippling me. I managed to run some proof of concepts on a very small subset of data here on Kaggle, however moving to the next step of using the whole data just seems impossible for me.</p>",
  "messages": [
    {
      "id": 1228835,
      "postDate": "2021-03-06T20:20:39.763Z",
      "content": "<p>Given that the data is HUGE (150GB+), and storage in an interactive Kaggle session is relatively very small, doing any kind of preprocessing would result in either memory or storage issues. Also, it is frankly impossible for me to download it locally and work with it. That's when I usually refer back to Google colab (so that I can download big datasets and save my intermediate results on drive). However, even colab has a limit of ~70GB of storage. Could there any other way to get the data in smaller chunks?</p>\n<p>The issue of data size is quite literally crippling me. I managed to run some proof of concepts on a very small subset of data here on Kaggle, however moving to the next step of using the whole data just seems impossible for me.</p>",
      "rawMarkdown": "Given that the data is HUGE (150GB+), and storage in an interactive Kaggle session is relatively very small, doing any kind of preprocessing would result in either memory or storage issues. Also, it is frankly impossible for me to download it locally and work with it. That's when I usually refer back to Google colab (so that I can download big datasets and save my intermediate results on drive). However, even colab has a limit of ~70GB of storage. Could there any other way to get the data in smaller chunks?\n\nThe issue of data size is quite literally crippling me. I managed to run some proof of concepts on a very small subset of data here on Kaggle, however moving to the next step of using the whole data just seems impossible for me.",
      "votes": 3
    },
    {
      "id": 1258989,
      "postDate": "2021-04-01T04:30:46.347Z",
      "content": "<p>This kaggle dataset may help. You can use kaggle kernel to down sample it or 8bit compress it for smaller size.<br>\n<a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/229839\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/229839</a></p>",
      "rawMarkdown": "This kaggle dataset may help. You can use kaggle kernel to down sample it or 8bit compress it for smaller size.\nhttps://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/229839"
    },
    {
      "id": 1230628,
      "postDate": "2021-03-08T09:42:21.747Z",
      "content": "<p>Hi, if it may help:</p>\n<p>I made tf-records for single labeled cropped cell </p>\n<p><a href=\"https://www.kaggle.com/lucamtb/hpa-sell-segments-tfrecords\" target=\"_blank\">https://www.kaggle.com/lucamtb/hpa-sell-segments-tfrecords</a></p>\n<p>I got the cropped cells from here:</p>\n<p><a href=\"https://www.kaggle.com/thedrcat/hpa-cell-tiles-sample-balanced-dataset\" target=\"_blank\">https://www.kaggle.com/thedrcat/hpa-cell-tiles-sample-balanced-dataset</a></p>\n<p>Working with TPU I have no memory issue</p>",
      "rawMarkdown": "Hi, if it may help:\n\nI made tf-records for single labeled cropped cell \n\nhttps://www.kaggle.com/lucamtb/hpa-sell-segments-tfrecords\n\nI got the cropped cells from here:\n\nhttps://www.kaggle.com/thedrcat/hpa-cell-tiles-sample-balanced-dataset\n\nWorking with TPU I have no memory issue"
    },
    {
      "id": 1229026,
      "postDate": "2021-03-07T01:57:14.837Z",
      "content": "<p>Do you still have memory issues if you do data preprocessing on the fly? In my case, I haven't had any issues as long as I am reading the images batch by batch and performing operations on them before sending them to the model for training. </p>",
      "rawMarkdown": "Do you still have memory issues if you do data preprocessing on the fly? In my case, I haven't had any issues as long as I am reading the images batch by batch and performing operations on them before sending them to the model for training. ",
      "replies": [
        {
          "id": 1229279,
          "postDate": "2021-03-07T08:19:17.267Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/adeebabbas\" target=\"_blank\">@adeebabbas</a> for you kind help, however, this solution suffers from lots of limitations. Currently, my preprocessing phase partially consists in cropping individual cells from an image and saving them as <code>n</code> training samples. Also, if this on-the-fly idea were to be implementable, it will just render the training <strong>that much slower</strong> (i.e., instead of running preprocessing once, and training a multitude of times on the preprocessed data). </p>",
          "rawMarkdown": "Thank you @adeebabbas for you kind help, however, this solution suffers from lots of limitations. Currently, my preprocessing phase partially consists in cropping individual cells from an image and saving them as `n` training samples. Also, if this on-the-fly idea were to be implementable, it will just render the training **that much slower** (i.e., instead of running preprocessing once, and training a multitude of times on the preprocessed data). ",
          "replies": [
            {
              "id": 1229389,
              "postDate": "2021-03-07T10:01:15.777Z",
              "content": "<p>There are dataset available with the masks and also with the cropped cells, just find them on kaggle</p>",
              "rawMarkdown": "There are dataset available with the masks and also with the cropped cells, just find them on kaggle"
            }
          ]
        },
        {
          "id": 1229533,
          "postDate": "2021-03-07T12:20:18.830Z",
          "content": "<p><a href=\"https://www.kaggle.com/alexanderriedel\" target=\"_blank\">@alexanderriedel</a> As mentioned in my comment above, my data preprocessing consists <strong>partially</strong> in cropping individual cells. Many more steps are performed in my preprocessing.</p>",
          "rawMarkdown": "@alexanderriedel As mentioned in my comment above, my data preprocessing consists **partially** in cropping individual cells. Many more steps are performed in my preprocessing.",
          "replies": [
            {
              "id": 1229546,
              "postDate": "2021-03-07T12:24:30.027Z",
              "content": "<p>yes i understand, but do you really need the fullsize images for all of this?</p>",
              "rawMarkdown": "yes i understand, but do you really need the fullsize images for all of this?"
            }
          ]
        },
        {
          "id": 1230638,
          "postDate": "2021-03-08T10:01:49.883Z",
          "content": "<p><a href=\"https://www.kaggle.com/alexanderriedel\" target=\"_blank\">@alexanderriedel</a> Unfortunately yes. I also need access to each of the 4 channels (which most datasets you've mentioned excluding the Yellow channel).</p>",
          "rawMarkdown": "@alexanderriedel Unfortunately yes. I also need access to each of the 4 channels (which most datasets you've mentioned excluding the Yellow channel).",
          "replies": [
            {
              "id": 1230780,
              "postDate": "2021-03-08T13:00:09.667Z",
              "content": "<p>some stuff you could try: get colab pro (140GB storage) and google drive with 200GB storage. download the dataset to your google drive, split it in parts and import parts of the dataset to your colab pro. be beware of quota limits between colab and google drive for transferring many files, thats why they advise you to zip everything before transferring between colab and drive</p>",
              "rawMarkdown": "some stuff you could try: get colab pro (140GB storage) and google drive with 200GB storage. download the dataset to your google drive, split it in parts and import parts of the dataset to your colab pro. be beware of quota limits between colab and google drive for transferring many files, thats why they advise you to zip everything before transferring between colab and drive",
              "votes": 1
            }
          ]
        },
        {
          "id": 1230931,
          "postDate": "2021-03-08T14:49:03.333Z",
          "content": "<p>If you use the Dataloader from Pytorch, training won't be affected at all. as Pytorch prepares future batches while training is being done. Apart from that, you can use the CPU for preprocessing while one of your epochs is running. If you can share the preprocessing code, I might be able to help better. </p>",
          "rawMarkdown": "If you use the Dataloader from Pytorch, training won't be affected at all. as Pytorch prepares future batches while training is being done. Apart from that, you can use the CPU for preprocessing while one of your epochs is running. If you can share the preprocessing code, I might be able to help better. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1232652,
      "postDate": "2021-03-09T22:39:51.407Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1258989,
      "author_name": "seefun",
      "author_url": "",
      "post_date": "2021-04-01T04:30:46.347000",
      "content": "<p>This kaggle dataset may help. You can use kaggle kernel to down sample it or 8bit compress it for smaller size.<br>\n<a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/229839\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/229839</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1230628,
      "author_name": "LucaMTB",
      "author_url": "",
      "post_date": "2021-03-08T09:42:21.747000",
      "content": "<p>Hi, if it may help:</p>\n<p>I made tf-records for single labeled cropped cell </p>\n<p><a href=\"https://www.kaggle.com/lucamtb/hpa-sell-segments-tfrecords\" target=\"_blank\">https://www.kaggle.com/lucamtb/hpa-sell-segments-tfrecords</a></p>\n<p>I got the cropped cells from here:</p>\n<p><a href=\"https://www.kaggle.com/thedrcat/hpa-cell-tiles-sample-balanced-dataset\" target=\"_blank\">https://www.kaggle.com/thedrcat/hpa-cell-tiles-sample-balanced-dataset</a></p>\n<p>Working with TPU I have no memory issue</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1229026,
      "author_name": "adeeb10abbas",
      "author_url": "",
      "post_date": "2021-03-07T01:57:14.837000",
      "content": "<p>Do you still have memory issues if you do data preprocessing on the fly? In my case, I haven't had any issues as long as I am reading the images batch by batch and performing operations on them before sending them to the model for training. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1229279,
          "author_name": "Joseph Assaker",
          "author_url": "",
          "post_date": "2021-03-07T08:19:17.267000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/adeebabbas\" target=\"_blank\">@adeebabbas</a> for you kind help, however, this solution suffers from lots of limitations. Currently, my preprocessing phase partially consists in cropping individual cells from an image and saving them as <code>n</code> training samples. Also, if this on-the-fly idea were to be implementable, it will just render the training <strong>that much slower</strong> (i.e., instead of running preprocessing once, and training a multitude of times on the preprocessed data). </p>",
          "votes": 0,
          "replies": [
            {
              "id": 1229389,
              "author_name": "Alexander Riedel",
              "author_url": "",
              "post_date": "2021-03-07T10:01:15.777000",
              "content": "<p>There are dataset available with the masks and also with the cropped cells, just find them on kaggle</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 1229533,
          "author_name": "Joseph Assaker",
          "author_url": "",
          "post_date": "2021-03-07T12:20:18.830000",
          "content": "<p><a href=\"https://www.kaggle.com/alexanderriedel\" target=\"_blank\">@alexanderriedel</a> As mentioned in my comment above, my data preprocessing consists <strong>partially</strong> in cropping individual cells. Many more steps are performed in my preprocessing.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 1229546,
              "author_name": "Alexander Riedel",
              "author_url": "",
              "post_date": "2021-03-07T12:24:30.027000",
              "content": "<p>yes i understand, but do you really need the fullsize images for all of this?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 1230638,
          "author_name": "Joseph Assaker",
          "author_url": "",
          "post_date": "2021-03-08T10:01:49.883000",
          "content": "<p><a href=\"https://www.kaggle.com/alexanderriedel\" target=\"_blank\">@alexanderriedel</a> Unfortunately yes. I also need access to each of the 4 channels (which most datasets you've mentioned excluding the Yellow channel).</p>",
          "votes": 0,
          "replies": [
            {
              "id": 1230780,
              "author_name": "Alexander Riedel",
              "author_url": "",
              "post_date": "2021-03-08T13:00:09.667000",
              "content": "<p>some stuff you could try: get colab pro (140GB storage) and google drive with 200GB storage. download the dataset to your google drive, split it in parts and import parts of the dataset to your colab pro. be beware of quota limits between colab and google drive for transferring many files, thats why they advise you to zip everything before transferring between colab and drive</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 1230931,
          "author_name": "adeeb10abbas",
          "author_url": "",
          "post_date": "2021-03-08T14:49:03.333000",
          "content": "<p>If you use the Dataloader from Pytorch, training won't be affected at all. as Pytorch prepares future batches while training is being done. Apart from that, you can use the CPU for preprocessing while one of your epochs is running. If you can share the preprocessing code, I might be able to help better. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1232652,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-09T22:39:51.407000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1228835": "Given that the data is HUGE (150GB+), and storage in an interactive Kaggle session is relatively very small, doing any kind of preprocessing would result in either memory or storage issues. Also, it is frankly impossible for me to download it locally and work with it. That's when I usually refer back to Google colab (so that I can download big datasets and save my intermediate results on drive). However, even colab has a limit of ~70GB of storage. Could there any other way to get the data in smaller chunks?\n\nThe issue of data size is quite literally crippling me. I managed to run some proof of concepts on a very small subset of data here on Kaggle, however moving to the next step of using the whole data just seems impossible for me.",
    "1258989": "This kaggle dataset may help. You can use kaggle kernel to down sample it or 8bit compress it for smaller size.\nhttps://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/229839",
    "1230628": "Hi, if it may help:\n\nI made tf-records for single labeled cropped cell \n\nhttps://www.kaggle.com/lucamtb/hpa-sell-segments-tfrecords\n\nI got the cropped cells from here:\n\nhttps://www.kaggle.com/thedrcat/hpa-cell-tiles-sample-balanced-dataset\n\nWorking with TPU I have no memory issue",
    "1229026": "Do you still have memory issues if you do data preprocessing on the fly? In my case, I haven't had any issues as long as I am reading the images batch by batch and performing operations on them before sending them to the model for training. ",
    "1232652": ""
  }
}