{
  "id": 156794,
  "title": "Saving preprocessed data - Kernel Output limit size",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/156794",
  "author_name": "",
  "post_date": "2020-06-07T21:53:48.933358700Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I have found another forum post from a few years back that said to preprocess in one notebook and save the processed images as output files and then train on a seperate notebook and use those preprocessed files as input. However, it's not clear if the 5gb working limit applies.</p>\n\n<p>The jpeg train images are 25gb originally and will be less when resized but even if the average rescaling is 50% then it'll still leave us with 6.25gb which is over the limit.</p>\n\n<p>I don't have the capacity to run the model on my own setup, hence wanting to run it on Kaggle but doing this preprocessing everytime I want to test a small adjustment to the model is clearly not efficient or feasible with the 30 hour CPU limit.</p>\n\n<p>Thanks.</p>",
  "messages": [
    {
      "id": "877713",
      "postDate": "06/07/2020 21:53:48",
      "content": "<p>I have found another forum post from a few years back that said to preprocess in one notebook and save the processed images as output files and then train on a seperate notebook and use those preprocessed files as input. However, it's not clear if the 5gb working limit applies.</p>\n\n<p>The jpeg train images are 25gb originally and will be less when resized but even if the average rescaling is 50% then it'll still leave us with 6.25gb which is over the limit.</p>\n\n<p>I don't have the capacity to run the model on my own setup, hence wanting to run it on Kaggle but doing this preprocessing everytime I want to test a small adjustment to the model is clearly not efficient or feasible with the 30 hour CPU limit.</p>\n\n<p>Thanks.</p>",
      "rawMarkdown": "I have found another forum post from a few years back that said to preprocess in one notebook and save the processed images as output files and then train on a seperate notebook and use those preprocessed files as input. However, it's not clear if the 5gb working limit applies.\n\nThe jpeg train images are 25gb originally and will be less when resized but even if the average rescaling is 50% then it'll still leave us with 6.25gb which is over the limit.\n\nI don't have the capacity to run the model on my own setup, hence wanting to run it on Kaggle but doing this preprocessing everytime I want to test a small adjustment to the model is clearly not efficient or feasible with the 30 hour CPU limit.\n\nThanks.",
      "votes": null
    },
    {
      "id": "877794",
      "postDate": "06/08/2020 01:37:27",
      "content": "<p>If you center square crop all the train and test images, and resize them to 704x704 and use default JPEG 95% compression, it all fits in under 5GB. I posted a 768x768 TFRecords dataset <a href=\"https://www.kaggle.com/cdeotte/melanoma-768x768\">here</a> which is only 5.3GB. And 512x512 TFRecords dataset <a href=\"https://www.kaggle.com/cdeotte/melanoma-512x512\">here</a> is only 2.6GB</p>",
      "rawMarkdown": "If you center square crop all the train and test images, and resize them to 704x704 and use default JPEG 95% compression, it all fits in under 5GB. I posted a 768x768 TFRecords dataset [here][1] which is only 5.3GB. And 512x512 TFRecords dataset [here][2] is only 2.6GB\n\n[1]: https://www.kaggle.com/cdeotte/melanoma-768x768\n[2]: https://www.kaggle.com/cdeotte/melanoma-512x512",
      "votes": null
    },
    {
      "id": "877989",
      "postDate": "06/08/2020 07:14:42",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> i was doing some preprocessing and im saving it in zip file then i commit the notebook and i got the output with a file size of 100Mb.But still i didnt got a \"New Dataset\" icon on the output</p>",
      "rawMarkdown": "cdeotte i was doing some preprocessing and im saving it in zip file then i commit the notebook and i got the output with a file size of 100Mb.But still i didnt got a \"New Dataset\" icon on the output",
      "votes": null
    },
    {
      "id": "878578",
      "postDate": "06/08/2020 16:41:58",
      "content": "<p>Kaggle notebooks keep changing appearance. Have your scrolled down and looked for button in bottom left just above the Execution Info section? (In plot below you see a New Version button which would be New Dataset)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F612b743b782a99bb09fd53c06b9aa69e%2FScreen%20Shot%202020-06-08%20at%209.40.20%20AM.png?generation=1591634441749767&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Kaggle notebooks keep changing appearance. Have your scrolled down and looked for button in bottom left just above the Execution Info section? (In plot below you see a New Version button which would be New Dataset)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F612b743b782a99bb09fd53c06b9aa69e%2FScreen%20Shot%202020-06-08%20at%209.40.20%20AM.png?generation=1591634441749767&amp;alt=media)",
      "votes": null
    },
    {
      "id": "878594",
      "postDate": "06/08/2020 16:57:59",
      "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a>.\nJust one more thing. I couldnt get a gcs path of the kernel output file, after i have created a dataset out of that kernel.Its taking too long to get that.</p>",
      "rawMarkdown": "Thanks @cdeotte.\nJust one more thing. I couldnt get a gcs path of the kernel output file, after i have created a dataset out of that kernel.Its taking too long to get that.",
      "votes": null
    },
    {
      "id": "878604",
      "postDate": "06/08/2020 17:03:03",
      "content": "<p>I've seen this before. There is a weird bug at Kaggle that comes and goes. We should tell Kaggle.</p>\n\n<p>For a few days last week you couldn't use</p>\n\n<pre><code>GCS_PATH = KaggleDatasets().get_gcs_path('siim-isic-melanoma-classification')\n</code></pre>\n\n<p>And had to use </p>\n\n<pre><code>GCS_PATH = 'gs://kds-0b2c68d2b2fa4692fcffc1029c606b32dd6a88de8d6da08fcd30d0c4'\n</code></pre>\n\n<p>instead. However that weird link is temporary because it no longer works. But now the first one with <code>siim-isic-melanoma-classification</code> works again. Sometime is wrong with getting <code>gcs_path</code>. And the problem seems to come and go. I posted a discussion <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155963\">here</a></p>",
      "rawMarkdown": "I've seen this before. There is a weird bug at Kaggle that comes and goes. We should tell Kaggle.\n\nFor a few days last week you couldn't use\n\n    GCS_PATH = KaggleDatasets().get_gcs_path('siim-isic-melanoma-classification')\n\nAnd had to use \n\n    GCS_PATH = 'gs://kds-0b2c68d2b2fa4692fcffc1029c606b32dd6a88de8d6da08fcd30d0c4'\n\ninstead. However that weird link is temporary because it no longer works. But now the first one with `siim-isic-melanoma-classification` works again. Sometime is wrong with getting `gcs_path`. And the problem seems to come and go. I posted a discussion [here][1]\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155963",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 877794,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "06/08/2020 01:37:27",
      "content": "<p>If you center square crop all the train and test images, and resize them to 704x704 and use default JPEG 95% compression, it all fits in under 5GB. I posted a 768x768 TFRecords dataset <a href=\"https://www.kaggle.com/cdeotte/melanoma-768x768\">here</a> which is only 5.3GB. And 512x512 TFRecords dataset <a href=\"https://www.kaggle.com/cdeotte/melanoma-512x512\">here</a> is only 2.6GB</p>",
      "votes": null,
      "replies": [
        {
          "id": 877989,
          "author_name": "msharuk589",
          "author_url": "",
          "post_date": "06/08/2020 07:14:42",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> i was doing some preprocessing and im saving it in zip file then i commit the notebook and i got the output with a file size of 100Mb.But still i didnt got a \"New Dataset\" icon on the output</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 878578,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "06/08/2020 16:41:58",
          "content": "<p>Kaggle notebooks keep changing appearance. Have your scrolled down and looked for button in bottom left just above the Execution Info section? (In plot below you see a New Version button which would be New Dataset)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F612b743b782a99bb09fd53c06b9aa69e%2FScreen%20Shot%202020-06-08%20at%209.40.20%20AM.png?generation=1591634441749767&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 878594,
          "author_name": "msharuk589",
          "author_url": "",
          "post_date": "06/08/2020 16:57:59",
          "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a>.\nJust one more thing. I couldnt get a gcs path of the kernel output file, after i have created a dataset out of that kernel.Its taking too long to get that.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 878604,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "06/08/2020 17:03:03",
          "content": "<p>I've seen this before. There is a weird bug at Kaggle that comes and goes. We should tell Kaggle.</p>\n\n<p>For a few days last week you couldn't use</p>\n\n<pre><code>GCS_PATH = KaggleDatasets().get_gcs_path('siim-isic-melanoma-classification')\n</code></pre>\n\n<p>And had to use </p>\n\n<pre><code>GCS_PATH = 'gs://kds-0b2c68d2b2fa4692fcffc1029c606b32dd6a88de8d6da08fcd30d0c4'\n</code></pre>\n\n<p>instead. However that weird link is temporary because it no longer works. But now the first one with <code>siim-isic-melanoma-classification</code> works again. Sometime is wrong with getting <code>gcs_path</code>. And the problem seems to come and go. I posted a discussion <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155963\">here</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "877713": "I have found another forum post from a few years back that said to preprocess in one notebook and save the processed images as output files and then train on a seperate notebook and use those preprocessed files as input. However, it's not clear if the 5gb working limit applies.\n\nThe jpeg train images are 25gb originally and will be less when resized but even if the average rescaling is 50% then it'll still leave us with 6.25gb which is over the limit.\n\nI don't have the capacity to run the model on my own setup, hence wanting to run it on Kaggle but doing this preprocessing everytime I want to test a small adjustment to the model is clearly not efficient or feasible with the 30 hour CPU limit.\n\nThanks.",
    "877794": "If you center square crop all the train and test images, and resize them to 704x704 and use default JPEG 95% compression, it all fits in under 5GB. I posted a 768x768 TFRecords dataset [here][1] which is only 5.3GB. And 512x512 TFRecords dataset [here][2] is only 2.6GB\n\n[1]: https://www.kaggle.com/cdeotte/melanoma-768x768\n[2]: https://www.kaggle.com/cdeotte/melanoma-512x512",
    "877989": "cdeotte i was doing some preprocessing and im saving it in zip file then i commit the notebook and i got the output with a file size of 100Mb.But still i didnt got a \"New Dataset\" icon on the output",
    "878578": "Kaggle notebooks keep changing appearance. Have your scrolled down and looked for button in bottom left just above the Execution Info section? (In plot below you see a New Version button which would be New Dataset)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F612b743b782a99bb09fd53c06b9aa69e%2FScreen%20Shot%202020-06-08%20at%209.40.20%20AM.png?generation=1591634441749767&amp;alt=media)",
    "878594": "Thanks @cdeotte.\nJust one more thing. I couldnt get a gcs path of the kernel output file, after i have created a dataset out of that kernel.Its taking too long to get that.",
    "878604": "I've seen this before. There is a weird bug at Kaggle that comes and goes. We should tell Kaggle.\n\nFor a few days last week you couldn't use\n\n    GCS_PATH = KaggleDatasets().get_gcs_path('siim-isic-melanoma-classification')\n\nAnd had to use \n\n    GCS_PATH = 'gs://kds-0b2c68d2b2fa4692fcffc1029c606b32dd6a88de8d6da08fcd30d0c4'\n\ninstead. However that weird link is temporary because it no longer works. But now the first one with `siim-isic-melanoma-classification` works again. Sometime is wrong with getting `gcs_path`. And the problem seems to come and go. I posted a discussion [here][1]\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155963"
  },
  "source": "meta"
}