{
  "id": 154281,
  "title": "Data loading is slow ?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/154281",
  "author_name": "Elbanan",
  "post_date": "2020-05-27T23:12:46.458000",
  "votes": 9,
  "comment_count": 54,
  "views": 0,
  "content": "<p>Is anyone else experiencing slow loading of data in notebooks or is it just me ? In the screenshot below, I am using only 20% of the supplied training data. This is the time needed for a forward pass through ResNet50. Processor utilization is super high.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2196129%2F66b2a243ea15e01830a2377d30c12369%2FScreenshot%202020-05-27%2022.11.26.png?generation=1590631909985239&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 864266,
      "postDate": "2020-05-27T23:12:46.460Z",
      "content": "<p>Is anyone else experiencing slow loading of data in notebooks or is it just me ? In the screenshot below, I am using only 20% of the supplied training data. This is the time needed for a forward pass through ResNet50. Processor utilization is super high.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2196129%2F66b2a243ea15e01830a2377d30c12369%2FScreenshot%202020-05-27%2022.11.26.png?generation=1590631909985239&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Is anyone else experiencing slow loading of data in notebooks or is it just me ? In the screenshot below, I am using only 20% of the supplied training data. This is the time needed for a forward pass through ResNet50. Processor utilization is super high.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2196129%2F66b2a243ea15e01830a2377d30c12369%2FScreenshot%202020-05-27%2022.11.26.png?generation=1590631909985239&amp;alt=media)\n",
      "votes": 9
    },
    {
      "id": 864419,
      "postDate": "2020-05-28T01:55:05.620Z",
      "content": "<p>Same ☹️ </p>\n\n<p><strong>UPDATE:</strong>\nThis is the ticket - <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154347\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154347</a>. 👍</p>",
      "rawMarkdown": "Same ☹️ \n\n**UPDATE:**\nThis is the ticket - https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154347. 👍",
      "votes": 1
    },
    {
      "id": 864368,
      "postDate": "2020-05-28T00:58:10.727Z",
      "content": "<p>Experiencing the same here. Super slow download.</p>",
      "rawMarkdown": "Experiencing the same here. Super slow download.",
      "votes": 1
    },
    {
      "id": 957749,
      "postDate": "2020-08-04T14:33:35.783Z",
      "content": "<p>Google Colab seems to have upgraded to TF 2.3 this night as well, which is causing issues. Perhaps both are related?</p>",
      "rawMarkdown": "Google Colab seems to have upgraded to TF 2.3 this night as well, which is causing issues. Perhaps both are related?",
      "replies": [
        {
          "id": 957998,
          "postDate": "2020-08-04T17:22:27.520Z",
          "content": "<p>By any chance have you found a way to make colab works ? </p>",
          "rawMarkdown": "By any chance have you found a way to make colab works ? "
        },
        {
          "id": 958831,
          "postDate": "2020-08-05T06:52:08.497Z",
          "content": "<p>I haven't but <a href=\"/gdonchyts\">@gdonchyts</a> has and posted it here: <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172357\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172357</a></p>",
          "rawMarkdown": "I haven't but @gdonchyts has and posted it here: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172357"
        }
      ]
    },
    {
      "id": 957299,
      "postDate": "2020-08-04T07:52:43.990Z",
      "content": "<p>Hi <a href=\"/herbison\">@herbison</a> Data loading is taking a lot of time in my notebook today and wasting my TPU quota. I am using <a href=\"/cdeotte\">@cdeotte</a>  datasets which is posted here: <a href=\"https://www.kaggle.com/cdeotte/isic2019-384x384\">https://www.kaggle.com/cdeotte/isic2019-384x384</a></p>",
      "rawMarkdown": "Hi @herbison Data loading is taking a lot of time in my notebook today and wasting my TPU quota. I am using @cdeotte  datasets which is posted here: https://www.kaggle.com/cdeotte/isic2019-384x384",
      "replies": [
        {
          "id": 957648,
          "postDate": "2020-08-04T13:17:04.467Z",
          "content": "<p><a href=\"/abdurrehman245\">@abdurrehman245</a> what part is slow? My comments below detail how it takes some investigation to figure out what 'exactly' is slow (file reading from disk, image transformation, actual learning code?).</p>",
          "rawMarkdown": "@abdurrehman245 what part is slow? My comments below detail how it takes some investigation to figure out what 'exactly' is slow (file reading from disk, image transformation, actual learning code?)."
        },
        {
          "id": 957703,
          "postDate": "2020-08-04T13:57:16.763Z",
          "content": "<p><a href=\"/herbison\">@herbison</a> thanks for the response.</p>\n\n<p>I am using the below code to get gcs paths and which is taking more than <code>2mins on TPU</code> now but taking less than<code>2 secs on GPU</code>:</p>\n\n<p><code>\n GCS_PATH[i] = KaggleDatasets().get_gcs_path('melanoma-384x384')\n    GCS_PATH2[i] = KaggleDatasets().get_gcs_path('isic2019-384x384')\n    GCS_PATH3[i] = KaggleDatasets().get_gcs_path('malignant-v2-384x384')\n</code>\nDatasets path:\n<a href=\"https://www.kaggle.com/cdeotte/melanoma-384x384\">https://www.kaggle.com/cdeotte/melanoma-384x384</a>\n<a href=\"https://www.kaggle.com/cdeotte/isic2019-384x384\">https://www.kaggle.com/cdeotte/isic2019-384x384</a>\n<a href=\"https://www.kaggle.com/cdeotte/malignant-v2-384x384\">https://www.kaggle.com/cdeotte/malignant-v2-384x384</a></p>\n\n<p>Also, when I start reading tfrecords using data pipeline, that part stuck. I waited for <code>10mins</code>for reading tfrecords but there was no response so I switch off my TPU to avoid the wastage of quota but yesterday the same code was working fine and it was taking <code>less than 1min</code> to read all the tfrecords from all mentioned datasets.</p>",
          "rawMarkdown": "@herbison thanks for the response.\n\nI am using the below code to get gcs paths and which is taking more than `2mins on TPU` now but taking less than` 2 secs on GPU`:\n\n   ```\n GCS_PATH[i] = KaggleDatasets().get_gcs_path('melanoma-384x384')\n    GCS_PATH2[i] = KaggleDatasets().get_gcs_path('isic2019-384x384')\n    GCS_PATH3[i] = KaggleDatasets().get_gcs_path('malignant-v2-384x384')\n```\nDatasets path:\nhttps://www.kaggle.com/cdeotte/melanoma-384x384\nhttps://www.kaggle.com/cdeotte/isic2019-384x384\nhttps://www.kaggle.com/cdeotte/malignant-v2-384x384\n\nAlso, when I start reading tfrecords using data pipeline, that part stuck. I waited for `10mins `for reading tfrecords but there was no response so I switch off my TPU to avoid the wastage of quota but yesterday the same code was working fine and it was taking `less than 1min` to read all the tfrecords from all mentioned datasets."
        },
        {
          "id": 957717,
          "postDate": "2020-08-04T14:05:50.623Z",
          "content": "<p><a href=\"https://www.kaggle.com/ifigotin\" target=\"_blank\">@ifigotin</a> Any ideas why these would be slow?</p>\n<p>My first guess for get<em>gcs</em>path would just be that you got 'lucky/unlucky' about expiration/recaching to GCS. Running those from a CPU session before trying on TPU should always make sure they are cached on GCS ahead of time.</p>",
          "rawMarkdown": "@ifigotin Any ideas why these would be slow?\n\nMy first guess for get_gcs_path would just be that you got 'lucky/unlucky' about expiration/recaching to GCS. Running those from a CPU session before trying on TPU should always make sure they are cached on GCS ahead of time."
        },
        {
          "id": 957746,
          "postDate": "2020-08-04T14:28:58.383Z",
          "content": "<p><a href=\"/herbison\">@herbison</a> I just ran the <code>get_gcs_path</code> on CPU which took <code>6.28 secs</code> and after that I ran the same on TPU which is still taking <code>2 min 20 secs</code> and that's a huge difference.</p>",
          "rawMarkdown": "@herbison I just ran the `get_gcs_path` on CPU which took `6.28 secs` and after that I ran the same on TPU which is still taking `2 min 20 secs` and that's a huge difference."
        },
        {
          "id": 957895,
          "postDate": "2020-08-04T16:09:41.680Z",
          "content": "<p><a href=\"/abdurrehman245\">@abdurrehman245</a> when you start a new session (say you switched from CPU to TPU), you may also end up in a different GCP region. That new region may not necessarily have the data cached in a GCS bucket yet. </p>\n\n<p>If however, you see this issue consistently on the same dataset during the same day, please let us know.</p>",
          "rawMarkdown": "@abdurrehman245 when you start a new session (say you switched from CPU to TPU), you may also end up in a different GCP region. That new region may not necessarily have the data cached in a GCS bucket yet. \n\nIf however, you see this issue consistently on the same dataset during the same day, please let us know."
        },
        {
          "id": 957964,
          "postDate": "2020-08-04T16:50:48.007Z",
          "content": "<p>I am working on TPU since a month but this problem never happened on any session restart.</p>\n\n<p><a href=\"/ifigotin\">@ifigotin</a> what solution do you suggest as right now I am stuck and can't train my model as data loading before model training stucks and wastes a lot of TPU quota.</p>",
          "rawMarkdown": "I am working on TPU since a month but this problem never happened on any session restart.\n\n@ifigotin what solution do you suggest as right now I am stuck and can't train my model as data loading before model training stucks and wastes a lot of TPU quota."
        },
        {
          "id": 957986,
          "postDate": "2020-08-04T17:09:51.323Z",
          "content": "<p><a href=\"/abdurrehman245\">@abdurrehman245</a> could you share a notebook with us so that we can try to reproduce? Do you experience slowness all the time right now?</p>\n\n<p>Btw, if the issue is with the caching of the dataset in a new region, only some of the requests (usually the first ones that trigger caching) will see the slowness initially; so you may have not personally seen slowness for months, if your requests were not the first ones to trigger it. </p>",
          "rawMarkdown": "@abdurrehman245 could you share a notebook with us so that we can try to reproduce? Do you experience slowness all the time right now?\n\nBtw, if the issue is with the caching of the dataset in a new region, only some of the requests (usually the first ones that trigger caching) will see the slowness initially; so you may have not personally seen slowness for months, if your requests were not the first ones to trigger it. "
        },
        {
          "id": 958002,
          "postDate": "2020-08-04T17:24:36.300Z",
          "content": "<p><a href=\"/ifigotin\">@ifigotin</a> I met these problems on google colab too. Since this night data loading is slow on TPU mode. \nSeveral other people also experienced this problem today.</p>",
          "rawMarkdown": "@ifigotin I met these problems on google colab too. Since this night data loading is slow on TPU mode. \nSeveral other people also experienced this problem today."
        },
        {
          "id": 958013,
          "postDate": "2020-08-04T17:34:06.753Z",
          "content": "<p><a href=\"/speedwagon\">@speedwagon</a> thanks for letting us know. I did check this particular dataset, and I see that it is fully cached at the moment in all the regions. However, a simple test of <code>KaggleDatasets().get_gcs_path()</code> shows me that this call is slower on a TPU enabled session (i.e. milliseconds vs about 8 seconds in my test). We are investigating. </p>",
          "rawMarkdown": "@speedwagon thanks for letting us know. I did check this particular dataset, and I see that it is fully cached at the moment in all the regions. However, a simple test of `KaggleDatasets().get_gcs_path()` shows me that this call is slower on a TPU enabled session (i.e. milliseconds vs about 8 seconds in my test). We are investigating. ",
          "votes": 1
        },
        {
          "id": 958812,
          "postDate": "2020-08-05T06:41:59.537Z",
          "content": "<p><a href=\"/ifigotin\">@ifigotin</a> please let us know when you have any update on your side. \nThanks.</p>",
          "rawMarkdown": "@ifigotin please let us know when you have any update on your side. \nThanks."
        },
        {
          "id": 958905,
          "postDate": "2020-08-05T07:57:21.137Z",
          "content": "<p><a href=\"/ifigotin\">@ifigotin</a> <a href=\"/speedwagon\">@speedwagon</a> <a href=\"/abdurrehman245\">@abdurrehman245</a> <a href=\"/herbison\">@herbison</a> \nIt seems that something is wrong with the TF dependencies. I read that there was TF2.3 installed yesterday.\nAnyway, the following helped me to get runtime back to normal in interactive session, so you may want to try it too:\n<code>!pip install tensorflow~=2.2.0 tensorflow_gcs_config~=2.2.0</code></p>",
          "rawMarkdown": "@ifigotin @speedwagon @abdurrehman245 @herbison \nIt seems that something is wrong with the TF dependencies. I read that there was TF2.3 installed yesterday.\nAnyway, the following helped me to get runtime back to normal in interactive session, so you may want to try it too:\n`!pip install tensorflow~=2.2.0 tensorflow_gcs_config~=2.2.0`",
          "votes": 2
        },
        {
          "id": 959372,
          "postDate": "2020-08-05T14:39:00.190Z",
          "content": "<p><a href=\"/abdurrehman245\">@abdurrehman245</a> we found a network configuration issue last night while debugging this and corrected it. Thanks for reporting this, it should be looking better now.</p>",
          "rawMarkdown": "@abdurrehman245 we found a network configuration issue last night while debugging this and corrected it. Thanks for reporting this, it should be looking better now.",
          "votes": 1
        },
        {
          "id": 959373,
          "postDate": "2020-08-05T14:39:32.433Z",
          "content": "<p><a href=\"/abdurrehman245\">@abdurrehman245</a>, <a href=\"/stanislavblinov\">@stanislavblinov</a> we identified the issue yesterday, and the fix should already be in place. The issue was not related to TF, but to our networking stack. What users saw on Colab may be a different issue. </p>\n\n<p><code>KaggleDatsets().get_gcs_path()</code> should hopefully be fast now.</p>",
          "rawMarkdown": "@abdurrehman245, @stanislavblinov we identified the issue yesterday, and the fix should already be in place. The issue was not related to TF, but to our networking stack. What users saw on Colab may be a different issue. \n\n`KaggleDatsets().get_gcs_path()` should hopefully be fast now.",
          "votes": 1
        },
        {
          "id": 959377,
          "postDate": "2020-08-05T14:40:53.227Z",
          "content": "<p><a href=\"/stanislavblinov\">@stanislavblinov</a> see my answer above. Re TF 2.3, I don't believe we preinstall it yet on Kaggle images (as of today).</p>",
          "rawMarkdown": "@stanislavblinov see my answer above. Re TF 2.3, I don't believe we preinstall it yet on Kaggle images (as of today).",
          "votes": 1
        },
        {
          "id": 959418,
          "postDate": "2020-08-05T15:06:43.023Z",
          "content": "<p>Thanks for investigating it!</p>",
          "rawMarkdown": "Thanks for investigating it!"
        },
        {
          "id": 959570,
          "postDate": "2020-08-05T17:44:12.373Z",
          "content": "<p><a href=\"/herbison\">@herbison</a> <a href=\"/ifigotin\">@ifigotin</a> thanks for nice and quick response. Yup the issue has been resolved.</p>",
          "rawMarkdown": "@herbison @ifigotin thanks for nice and quick response. Yup the issue has been resolved.",
          "votes": 1
        }
      ]
    },
    {
      "id": 868428,
      "postDate": "2020-05-31T08:19:38.697Z",
      "content": "<p>I thought I was the only one until I saw this, it really is irritating, I want my model to train for 30 epochs, but only in the 1 epoch it has taken over an hour to get past 33%. :/</p>",
      "rawMarkdown": "I thought I was the only one until I saw this, it really is irritating, I want my model to train for 30 epochs, but only in the 1 epoch it has taken over an hour to get past 33%. :/"
    },
    {
      "id": 865543,
      "postDate": "2020-05-28T17:32:46.987Z",
      "content": "<p>This is too much slow.</p>",
      "rawMarkdown": "This is too much slow."
    },
    {
      "id": 864788,
      "postDate": "2020-05-28T07:28:25.647Z",
      "content": "<p>good</p>",
      "rawMarkdown": "good"
    },
    {
      "id": 864486,
      "postDate": "2020-05-28T03:07:44.880Z",
      "content": "<p>Edit nevermind</p>",
      "rawMarkdown": "Edit nevermind\n",
      "replies": [
        {
          "id": 864489,
          "postDate": "2020-05-28T03:11:35.393Z",
          "content": "<p><a href=\"/arroqc\">@arroqc</a> Can you throw this in a notebook and share it so that it loads an example image showing this problem? I can take a deeper look then.</p>",
          "rawMarkdown": "@arroqc Can you throw this in a notebook and share it so that it loads an example image showing this problem? I can take a deeper look then."
        },
        {
          "id": 864493,
          "postDate": "2020-05-28T03:14:53.977Z",
          "content": "<p>I tried using CPU only. I got the same timing. Running single pass via my pipeline for feature extraction for the whole dataset takes almost 5 hours using either CPU or GPU. I still have a feeling it has something to do with data transfer. I will try to dig around and see what might be causing that.</p>",
          "rawMarkdown": "I tried using CPU only. I got the same timing. Running single pass via my pipeline for feature extraction for the whole dataset takes almost 5 hours using either CPU or GPU. I still have a feeling it has something to do with data transfer. I will try to dig around and see what might be causing that."
        },
        {
          "id": 864498,
          "postDate": "2020-05-28T03:20:10.987Z",
          "content": "<p>Yeah just reading using cat 1 file took close to 1s I'll take a look.</p>\n\n<p><code>fn = \"/kaggle/input/siim-isic-melanoma-classification/jpeg/train/ISIC_0077735.jpg\"</code></p>\n\n<p><code>%%timeit</code>\n<code>!cat $fn &amp;gt; /dev/null</code></p>\n\n<p><code>741 ms ± 5.33 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)</code></p>\n\n<p>Edit: the file isn't super small though, I'll see what kind of speed we should expect.\n<code>!du -h $fn</code></p>\n\n<p><code>1.3M   /kaggle/input/siim-isic-melanoma-classification/jpeg/train/ISIC_0077735.jpg</code></p>",
          "rawMarkdown": "Yeah just reading using cat 1 file took close to 1s I'll take a look.\n\n`fn = \"/kaggle/input/siim-isic-melanoma-classification/jpeg/train/ISIC_0077735.jpg\"`\n\n`%%timeit`\n`!cat $fn &gt; /dev/null`\n\n`741 ms ± 5.33 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)`\n\nEdit: the file isn't super small though, I'll see what kind of speed we should expect.\n`!du -h $fn`\n\n`1.3M\t/kaggle/input/siim-isic-melanoma-classification/jpeg/train/ISIC_0077735.jpg`",
          "votes": 1
        },
        {
          "id": 864516,
          "postDate": "2020-05-28T03:38:19.917Z",
          "content": "<p>Actually it's the loading that takes time. Maybe because of how large the images are... This would mean having to preprocess and save first I guess.</p>\n\n<p>You can easily try with this code (and splitting the %%timeit in individual cells):</p>\n\n<p>```\nimport cv2\nimport PIL.Image as Image\nfrom torchvision import transforms</p>\n\n<p>image_fn = '/kaggle/input/siim-isic-melanoma-classification/jpeg/train/ISIC_2637011.jpg'\nimage = Image.open(image_fn)\nimage.resize((224, 224)).save('savetest.jpg')</p>\n\n<p>%%timeit\nimage = Image.open(image_fn)\ndown = transforms.ToTensor()(image)</p>\n\n<p>%%timeit\nimage = Image.open(image_fn)\ndown = image.resize((224, 224))\ndown = transforms.ToTensor()(down)</p>\n\n<p>%%timeit\nimage = Image.open('savetest.jpg')\ndown = transforms.ToTensor()(down)</p>\n\n<p>%%timeit\nimage = cv2.imread(image_fn)\ndown = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\ndown = cv2.resize(down, (224, 224))\ndown = transforms.ToTensor()(down)\n```</p>\n\n<p>Note that cv2 pipeline is actually 3 times faster here... but still too slow at 224ms. Saving first make it drop to acceptable level. Still surprise by how slow loading a single image is though...</p>",
          "rawMarkdown": "Actually it's the loading that takes time. Maybe because of how large the images are... This would mean having to preprocess and save first I guess.\n\nYou can easily try with this code (and splitting the %%timeit in individual cells):\n\n```\nimport cv2\nimport PIL.Image as Image\nfrom torchvision import transforms\n\nimage_fn = '/kaggle/input/siim-isic-melanoma-classification/jpeg/train/ISIC_2637011.jpg'\nimage = Image.open(image_fn)\nimage.resize((224, 224)).save('savetest.jpg')\n\n%%timeit\nimage = Image.open(image_fn)\ndown = transforms.ToTensor()(image)\n\n%%timeit\nimage = Image.open(image_fn)\ndown = image.resize((224, 224))\ndown = transforms.ToTensor()(down)\n\n%%timeit\nimage = Image.open('savetest.jpg')\ndown = transforms.ToTensor()(down)\n\n%%timeit\nimage = cv2.imread(image_fn)\ndown = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\ndown = cv2.resize(down, (224, 224))\ndown = transforms.ToTensor()(down)\n```\n\nNote that cv2 pipeline is actually 3 times faster here... but still too slow at 224ms. Saving first make it drop to acceptable level. Still surprise by how slow loading a single image is though...",
          "votes": 2
        },
        {
          "id": 864590,
          "postDate": "2020-05-28T04:51:21.403Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 864596,
          "postDate": "2020-05-28T04:55:14.323Z",
          "content": "<p>I tried two more tests:\n* I tried a different dataset, loading a 20Mb image\n* I tried fetching just the first byte (using head -c 1 instead of cat)</p>\n\n<p>Both of those also took 700ms which means it's not the amount of bytes transferred but the latency to first byte. I'll chat with my team about this tomorrow to figure out what's going on. I gotta get some sleep though it's late in my TZ... 😴 </p>",
          "rawMarkdown": "I tried two more tests:\n* I tried a different dataset, loading a 20Mb image\n* I tried fetching just the first byte (using head -c 1 instead of cat)\n\nBoth of those also took 700ms which means it's not the amount of bytes transferred but the latency to first byte. I'll chat with my team about this tomorrow to figure out what's going on. I gotta get some sleep though it's late in my TZ... 😴 ",
          "votes": 3
        },
        {
          "id": 864987,
          "postDate": "2020-05-28T10:16:24.490Z",
          "content": "<p>For info, I tried the same command as you did on my own server : \n<code>\nfn = \"jpeg/train/ISIC_0077735.jpg\"\n%%timeit\n!cat $fn &gt; /dev/null\n117 ms ± 1.26 ms per loop (mean ± std. dev. of 7 runs, 10 loops each)\n</code></p>\n\n<p>Seems much quicker.</p>",
          "rawMarkdown": "For info, I tried the same command as you did on my own server : \n```\nfn = \"jpeg/train/ISIC_0077735.jpg\"\n%%timeit\n!cat $fn &gt; /dev/null\n117 ms ± 1.26 ms per loop (mean ± std. dev. of 7 runs, 10 loops each)\n```\n\nSeems much quicker."
        },
        {
          "id": 865490,
          "postDate": "2020-05-28T16:51:51.530Z",
          "content": "<p>Yeah i've preprocessed everything to be 224x224 png files but it is still slow... 10s per batch instead of 32s. Something is definetly fishy. Obviously no problem on my local machine but I want to publish a kernel :P</p>",
          "rawMarkdown": "Yeah i've preprocessed everything to be 224x224 png files but it is still slow... 10s per batch instead of 32s. Something is definetly fishy. Obviously no problem on my local machine but I want to publish a kernel :P"
        },
        {
          "id": 865506,
          "postDate": "2020-05-28T17:08:00.830Z",
          "content": "<p>It does look like the Image libraries are the problem, after doing a bunch more tests, it looks like the reason my %%time was slow for cat was because booting the shell is really slow, the actual copy part is fast.</p>\n\n<p>Edit: the libraries are slow even when working on the data in memory.</p>\n\n<p><a href=\"/arroqc\">@arroqc</a> <a href=\"/elbanan\">@elbanan</a> </p>\n\n<p>Example: <a href=\"https://www.kaggle.com/herbison/fast-load-and-resize?scriptVersionId=34992706\">https://www.kaggle.com/herbison/fast-load-and-resize?scriptVersionId=34992706</a></p>",
          "rawMarkdown": "It does look like the Image libraries are the problem, after doing a bunch more tests, it looks like the reason my %%time was slow for cat was because booting the shell is really slow, the actual copy part is fast.\n\nEdit: the libraries are slow even when working on the data in memory.\n\n@arroqc @elbanan \n\nExample: https://www.kaggle.com/herbison/fast-load-and-resize?scriptVersionId=34992706",
          "votes": 1
        },
        {
          "id": 865510,
          "postDate": "2020-05-28T17:09:48.733Z",
          "content": "<p><a href=\"/herbison\">@herbison</a> I have a 404 on that link</p>",
          "rawMarkdown": "@herbison I have a 404 on that link"
        },
        {
          "id": 865520,
          "postDate": "2020-05-28T17:14:25.570Z",
          "content": "<p><a href=\"/arroqc\">@arroqc</a> Updated it and made public, it turns out I messed up time %%time magic, but it does show that it's Image that's slow, not reading the data.</p>",
          "rawMarkdown": "@arroqc Updated it and made public, it turns out I messed up time %%time magic, but it does show that it's Image that's slow, not reading the data.",
          "votes": 1
        },
        {
          "id": 865536,
          "postDate": "2020-05-28T17:28:42.993Z",
          "content": "<p>Thanks i'll give it a try</p>",
          "rawMarkdown": "Thanks i'll give it a try"
        },
        {
          "id": 865577,
          "postDate": "2020-05-28T17:54:42.843Z",
          "content": "<p><a href=\"/herbison\">@herbison</a> I've created a dataset of small 224x224 images in png. It obviously load faster and is acceptable for me to continue. Edit: nevermind</p>",
          "rawMarkdown": "@herbison I've created a dataset of small 224x224 images in png. It obviously load faster and is acceptable for me to continue. Edit: nevermind"
        },
        {
          "id": 865585,
          "postDate": "2020-05-28T18:08:34.047Z",
          "content": "<p><a href=\"/arroqc\">@arroqc</a> Just to confirm, are you saying the 224x224 dataset is still slow (10s per batch)?</p>",
          "rawMarkdown": "@arroqc Just to confirm, are you saying the 224x224 dataset is still slow (10s per batch)?"
        },
        {
          "id": 865590,
          "postDate": "2020-05-28T18:12:39.507Z",
          "content": "<p>No sorry. I've fixed that problem. For some reason it was still running on CPU rather than GPU. </p>\n\n<p>Thanks team.</p>",
          "rawMarkdown": "No sorry. I've fixed that problem. For some reason it was still running on CPU rather than GPU. \n\nThanks team.",
          "votes": 1
        }
      ]
    },
    {
      "id": 864431,
      "postDate": "2020-05-28T02:04:52.760Z",
      "content": "<p>It is driving me crazy. Creating batches takes forever. I think the bottleneck is data transfer from data repo/storage to the virtual machines. </p>\n\n<p>Any feedback from @Kaggle team , <a href=\"/juliaelliott\">@juliaelliott</a>  would be appreciated . Thanks a lot !</p>",
      "rawMarkdown": "It is driving me crazy. Creating batches takes forever. I think the bottleneck is data transfer from data repo/storage to the virtual machines. \n\nAny feedback from @Kaggle team , @juliaelliott  would be appreciated . Thanks a lot !",
      "replies": [
        {
          "id": 864440,
          "postDate": "2020-05-28T02:09:44.537Z",
          "content": "<p>I have same problem.</p>\n\n<p>What is really weird is that Loading images is fast, Resizing an image is fast but Loading + Resizing is 10 times slower than the individual operation Oo</p>",
          "rawMarkdown": "I have same problem.\n\nWhat is really weird is that Loading images is fast, Resizing an image is fast but Loading + Resizing is 10 times slower than the individual operation Oo"
        },
        {
          "id": 864475,
          "postDate": "2020-05-28T02:51:55.983Z",
          "content": "<p><a href=\"/elbanan\">@elbanan</a> I'm not seeing anything immediately in the storage that would indicate a problem (read ops/latency looks normal).</p>\n\n<p><a href=\"/arroqc\">@arroqc</a> comments also lead me to believe it's CPU, since loading the images is fine, but doing a processing operation with a load is slower. </p>\n\n<p>If your CPU is pinned at 200% then it's more likely you are CPU-bottlenecked than disk read bottlenecked. Not sure exactly what you're doing but looks like your GPU isn't in use either (though maybe it's too early in your pipeline). When GPU is enabled, you only get 2 CPU cores, but when it's disabled you get 4 CPU cores. So if you're CPU-constrained and don't need GPU you're better off on a CPU machine.</p>",
          "rawMarkdown": "@elbanan I'm not seeing anything immediately in the storage that would indicate a problem (read ops/latency looks normal).\n\n@arroqc comments also lead me to believe it's CPU, since loading the images is fine, but doing a processing operation with a load is slower. \n\nIf your CPU is pinned at 200% then it's more likely you are CPU-bottlenecked than disk read bottlenecked. Not sure exactly what you're doing but looks like your GPU isn't in use either (though maybe it's too early in your pipeline). When GPU is enabled, you only get 2 CPU cores, but when it's disabled you get 4 CPU cores. So if you're CPU-constrained and don't need GPU you're better off on a CPU machine.",
          "votes": 1
        },
        {
          "id": 864480,
          "postDate": "2020-05-28T02:55:35.027Z",
          "content": "<p><a href=\"/herbison\">@herbison</a> I used gpu for 30s until I noticed it took literally 30s to load a batch of 32 then stopped and tried to investigate. The images are large for sure but loading + resizing shouldn't take a full second.</p>",
          "rawMarkdown": "@herbison I used gpu for 30s until I noticed it took literally 30s to load a batch of 32 then stopped and tried to investigate. The images are large for sure but loading + resizing shouldn't take a full second."
        },
        {
          "id": 864483,
          "postDate": "2020-05-28T03:01:57.313Z",
          "content": "<p><a href=\"/herbison\">@herbison</a> Thanks  for getting back to me. I will try going the CPU route now although this is unusual for me. I have used this pipeline in the past with multiple datasets and never had similar problem unless there's delay in data transfer. </p>\n\n<p>The pipeline actually loads the data in batches, resizes and does a single pass via CNN for feature extraction.</p>",
          "rawMarkdown": "@herbison Thanks  for getting back to me. I will try going the CPU route now although this is unusual for me. I have used this pipeline in the past with multiple datasets and never had similar problem unless there's delay in data transfer. \n\nThe pipeline actually loads the data in batches, resizes and does a single pass via CNN for feature extraction."
        }
      ]
    },
    {
      "id": 864397,
      "postDate": "2020-05-28T01:30:58.357Z",
      "content": "<p>Downloading using google colab is faster. Then, the data can be downloaded from colab.</p>",
      "rawMarkdown": "Downloading using google colab is faster. Then, the data can be downloaded from colab.",
      "replies": [
        {
          "id": 864481,
          "postDate": "2020-05-28T02:57:45.947Z",
          "content": "<p><a href=\"/shayekh\">@shayekh</a> <a href=\"/vikasvpatil\">@vikasvpatil</a> Can you clarify what you mean by download? It's rare that datasets are actually downloaded inside Kaggle notebooks, usually they only need to be 'added' to your notebook via the \"Add Data\" feature?</p>",
          "rawMarkdown": "@shayekh @vikasvpatil Can you clarify what you mean by download? It's rare that datasets are actually downloaded inside Kaggle notebooks, usually they only need to be 'added' to your notebook via the \"Add Data\" feature?"
        },
        {
          "id": 864545,
          "postDate": "2020-05-28T04:07:53.370Z",
          "content": "<p><a href=\"/herbison\">@herbison</a> I meant for training in local machine, not applicable for the question here. The details were added in the question later.</p>",
          "rawMarkdown": "@herbison I meant for training in local machine, not applicable for the question here. The details were added in the question later.",
          "votes": 1
        },
        {
          "id": 865826,
          "postDate": "2020-05-28T23:12:54.987Z",
          "content": "<p><a href=\"/herbison\">@herbison</a> Thanks for the followup. I was planning to train it on a local machine. So was downloading the the files from Kaggle which took some time. Seems like if we use the Kaggle TPU, its much faster. Is it the norm with CV problems to use Kaggle's TPU or GPU? </p>",
          "rawMarkdown": "@herbison Thanks for the followup. I was planning to train it on a local machine. So was downloading the the files from Kaggle which took some time. Seems like if we use the Kaggle TPU, its much faster. Is it the norm with CV problems to use Kaggle's TPU or GPU? "
        },
        {
          "id": 865857,
          "postDate": "2020-05-29T00:05:21.213Z",
          "content": "<p><a href=\"/vikasvpatil\">@vikasvpatil</a>  TPUs are significantly faster than GPUs for certain tasks like this one. This competition dataset includes tfrecord files for easy use with the TPU, and those tfrecord files have already been properly sized for the TPU (unlike the raw images which are fairly large which is probably contributing to why it's slow to download the dataset / process the raw image files).</p>",
          "rawMarkdown": "@vikasvpatil  TPUs are significantly faster than GPUs for certain tasks like this one. This competition dataset includes tfrecord files for easy use with the TPU, and those tfrecord files have already been properly sized for the TPU (unlike the raw images which are fairly large which is probably contributing to why it's slow to download the dataset / process the raw image files).",
          "votes": 2
        },
        {
          "id": 865869,
          "postDate": "2020-05-29T00:15:46.907Z",
          "content": "<p><a href=\"/herbison\">@herbison</a> - Nice response.  This is the ticket.  Here is another discussion that points to the the same idea - <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154347\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154347</a>. 👍 </p>",
          "rawMarkdown": "@herbison - Nice response.  This is the ticket.  Here is another discussion that points to the the same idea - https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154347. 👍 ",
          "votes": 2
        },
        {
          "id": 865890,
          "postDate": "2020-05-29T00:46:37.627Z",
          "content": "<p><a href=\"/herbison\">@herbison</a> This is helpful. Appreciate your response. 👍 </p>",
          "rawMarkdown": "@herbison This is helpful. Appreciate your response. 👍 "
        }
      ]
    },
    {
      "id": 866114,
      "postDate": "2020-05-29T05:50:07.240Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 864419,
      "author_name": "Matt Yates",
      "author_url": "",
      "post_date": "2020-05-28T01:55:05.620000",
      "content": "<p>Same ☹️ </p>\n\n<p><strong>UPDATE:</strong>\nThis is the ticket - <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154347\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154347</a>. 👍</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 864368,
      "author_name": "Vikas V Patil",
      "author_url": "",
      "post_date": "2020-05-28T00:58:10.727000",
      "content": "<p>Experiencing the same here. Super slow download.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 957749,
      "author_name": "Gilles Vandewiele",
      "author_url": "",
      "post_date": "2020-08-04T14:33:35.783000",
      "content": "<p>Google Colab seems to have upgraded to TF 2.3 this night as well, which is causing issues. Perhaps both are related?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 957998,
          "author_name": "jayjay",
          "author_url": "",
          "post_date": "2020-08-04T17:22:27.520000",
          "content": "<p>By any chance have you found a way to make colab works ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 958831,
          "author_name": "Gilles Vandewiele",
          "author_url": "",
          "post_date": "2020-08-05T06:52:08.497000",
          "content": "<p>I haven't but <a href=\"/gdonchyts\">@gdonchyts</a> has and posted it here: <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172357\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172357</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 957299,
      "author_name": "Abdur Rehman",
      "author_url": "",
      "post_date": "2020-08-04T07:52:43.990000",
      "content": "<p>Hi <a href=\"/herbison\">@herbison</a> Data loading is taking a lot of time in my notebook today and wasting my TPU quota. I am using <a href=\"/cdeotte\">@cdeotte</a>  datasets which is posted here: <a href=\"https://www.kaggle.com/cdeotte/isic2019-384x384\">https://www.kaggle.com/cdeotte/isic2019-384x384</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 957648,
          "author_name": "Dustin",
          "author_url": "",
          "post_date": "2020-08-04T13:17:04.467000",
          "content": "<p><a href=\"/abdurrehman245\">@abdurrehman245</a> what part is slow? My comments below detail how it takes some investigation to figure out what 'exactly' is slow (file reading from disk, image transformation, actual learning code?).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 957703,
          "author_name": "Abdur Rehman",
          "author_url": "",
          "post_date": "2020-08-04T13:57:16.763000",
          "content": "<p><a href=\"/herbison\">@herbison</a> thanks for the response.</p>\n\n<p>I am using the below code to get gcs paths and which is taking more than <code>2mins on TPU</code> now but taking less than<code>2 secs on GPU</code>:</p>\n\n<p><code>\n GCS_PATH[i] = KaggleDatasets().get_gcs_path('melanoma-384x384')\n    GCS_PATH2[i] = KaggleDatasets().get_gcs_path('isic2019-384x384')\n    GCS_PATH3[i] = KaggleDatasets().get_gcs_path('malignant-v2-384x384')\n</code>\nDatasets path:\n<a href=\"https://www.kaggle.com/cdeotte/melanoma-384x384\">https://www.kaggle.com/cdeotte/melanoma-384x384</a>\n<a href=\"https://www.kaggle.com/cdeotte/isic2019-384x384\">https://www.kaggle.com/cdeotte/isic2019-384x384</a>\n<a href=\"https://www.kaggle.com/cdeotte/malignant-v2-384x384\">https://www.kaggle.com/cdeotte/malignant-v2-384x384</a></p>\n\n<p>Also, when I start reading tfrecords using data pipeline, that part stuck. I waited for <code>10mins</code>for reading tfrecords but there was no response so I switch off my TPU to avoid the wastage of quota but yesterday the same code was working fine and it was taking <code>less than 1min</code> to read all the tfrecords from all mentioned datasets.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 957717,
          "author_name": "Dustin",
          "author_url": "",
          "post_date": "2020-08-04T14:05:50.623000",
          "content": "<p><a href=\"https://www.kaggle.com/ifigotin\" target=\"_blank\">@ifigotin</a> Any ideas why these would be slow?</p>\n<p>My first guess for get<em>gcs</em>path would just be that you got 'lucky/unlucky' about expiration/recaching to GCS. Running those from a CPU session before trying on TPU should always make sure they are cached on GCS ahead of time.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 957746,
          "author_name": "Abdur Rehman",
          "author_url": "",
          "post_date": "2020-08-04T14:28:58.383000",
          "content": "<p><a href=\"/herbison\">@herbison</a> I just ran the <code>get_gcs_path</code> on CPU which took <code>6.28 secs</code> and after that I ran the same on TPU which is still taking <code>2 min 20 secs</code> and that's a huge difference.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 957895,
          "author_name": "Ilya Figotin",
          "author_url": "",
          "post_date": "2020-08-04T16:09:41.680000",
          "content": "<p><a href=\"/abdurrehman245\">@abdurrehman245</a> when you start a new session (say you switched from CPU to TPU), you may also end up in a different GCP region. That new region may not necessarily have the data cached in a GCS bucket yet. </p>\n\n<p>If however, you see this issue consistently on the same dataset during the same day, please let us know.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 957964,
          "author_name": "Abdur Rehman",
          "author_url": "",
          "post_date": "2020-08-04T16:50:48.007000",
          "content": "<p>I am working on TPU since a month but this problem never happened on any session restart.</p>\n\n<p><a href=\"/ifigotin\">@ifigotin</a> what solution do you suggest as right now I am stuck and can't train my model as data loading before model training stucks and wastes a lot of TPU quota.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 957986,
          "author_name": "Ilya Figotin",
          "author_url": "",
          "post_date": "2020-08-04T17:09:51.323000",
          "content": "<p><a href=\"/abdurrehman245\">@abdurrehman245</a> could you share a notebook with us so that we can try to reproduce? Do you experience slowness all the time right now?</p>\n\n<p>Btw, if the issue is with the caching of the dataset in a new region, only some of the requests (usually the first ones that trigger caching) will see the slowness initially; so you may have not personally seen slowness for months, if your requests were not the first ones to trigger it. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 958002,
          "author_name": "Vladislav Bakhteev",
          "author_url": "",
          "post_date": "2020-08-04T17:24:36.300000",
          "content": "<p><a href=\"/ifigotin\">@ifigotin</a> I met these problems on google colab too. Since this night data loading is slow on TPU mode. \nSeveral other people also experienced this problem today.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 958013,
          "author_name": "Ilya Figotin",
          "author_url": "",
          "post_date": "2020-08-04T17:34:06.753000",
          "content": "<p><a href=\"/speedwagon\">@speedwagon</a> thanks for letting us know. I did check this particular dataset, and I see that it is fully cached at the moment in all the regions. However, a simple test of <code>KaggleDatasets().get_gcs_path()</code> shows me that this call is slower on a TPU enabled session (i.e. milliseconds vs about 8 seconds in my test). We are investigating. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 958812,
          "author_name": "Abdur Rehman",
          "author_url": "",
          "post_date": "2020-08-05T06:41:59.537000",
          "content": "<p><a href=\"/ifigotin\">@ifigotin</a> please let us know when you have any update on your side. \nThanks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 958905,
          "author_name": "Stanislav Blinov",
          "author_url": "",
          "post_date": "2020-08-05T07:57:21.137000",
          "content": "<p><a href=\"/ifigotin\">@ifigotin</a> <a href=\"/speedwagon\">@speedwagon</a> <a href=\"/abdurrehman245\">@abdurrehman245</a> <a href=\"/herbison\">@herbison</a> \nIt seems that something is wrong with the TF dependencies. I read that there was TF2.3 installed yesterday.\nAnyway, the following helped me to get runtime back to normal in interactive session, so you may want to try it too:\n<code>!pip install tensorflow~=2.2.0 tensorflow_gcs_config~=2.2.0</code></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 959372,
          "author_name": "Dustin",
          "author_url": "",
          "post_date": "2020-08-05T14:39:00.190000",
          "content": "<p><a href=\"/abdurrehman245\">@abdurrehman245</a> we found a network configuration issue last night while debugging this and corrected it. Thanks for reporting this, it should be looking better now.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 959373,
          "author_name": "Ilya Figotin",
          "author_url": "",
          "post_date": "2020-08-05T14:39:32.433000",
          "content": "<p><a href=\"/abdurrehman245\">@abdurrehman245</a>, <a href=\"/stanislavblinov\">@stanislavblinov</a> we identified the issue yesterday, and the fix should already be in place. The issue was not related to TF, but to our networking stack. What users saw on Colab may be a different issue. </p>\n\n<p><code>KaggleDatsets().get_gcs_path()</code> should hopefully be fast now.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 959377,
          "author_name": "Ilya Figotin",
          "author_url": "",
          "post_date": "2020-08-05T14:40:53.227000",
          "content": "<p><a href=\"/stanislavblinov\">@stanislavblinov</a> see my answer above. Re TF 2.3, I don't believe we preinstall it yet on Kaggle images (as of today).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 959418,
          "author_name": "Stanislav Blinov",
          "author_url": "",
          "post_date": "2020-08-05T15:06:43.023000",
          "content": "<p>Thanks for investigating it!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 959570,
          "author_name": "Abdur Rehman",
          "author_url": "",
          "post_date": "2020-08-05T17:44:12.373000",
          "content": "<p><a href=\"/herbison\">@herbison</a> <a href=\"/ifigotin\">@ifigotin</a> thanks for nice and quick response. Yup the issue has been resolved.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 868428,
      "author_name": "Gajendra Saraswat",
      "author_url": "",
      "post_date": "2020-05-31T08:19:38.697000",
      "content": "<p>I thought I was the only one until I saw this, it really is irritating, I want my model to train for 30 epochs, but only in the 1 epoch it has taken over an hour to get past 33%. :/</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 865543,
      "author_name": "Margi",
      "author_url": "",
      "post_date": "2020-05-28T17:32:46.987000",
      "content": "<p>This is too much slow.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 864788,
      "author_name": "Vadym Samilenko",
      "author_url": "",
      "post_date": "2020-05-28T07:28:25.647000",
      "content": "<p>good</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 864486,
      "author_name": "Arnaud Roussel",
      "author_url": "",
      "post_date": "2020-05-28T03:07:44.880000",
      "content": "<p>Edit nevermind</p>",
      "votes": 0,
      "replies": [
        {
          "id": 864489,
          "author_name": "Dustin",
          "author_url": "",
          "post_date": "2020-05-28T03:11:35.393000",
          "content": "<p><a href=\"/arroqc\">@arroqc</a> Can you throw this in a notebook and share it so that it loads an example image showing this problem? I can take a deeper look then.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 864493,
          "author_name": "Elbanan",
          "author_url": "",
          "post_date": "2020-05-28T03:14:53.977000",
          "content": "<p>I tried using CPU only. I got the same timing. Running single pass via my pipeline for feature extraction for the whole dataset takes almost 5 hours using either CPU or GPU. I still have a feeling it has something to do with data transfer. I will try to dig around and see what might be causing that.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 864498,
          "author_name": "Dustin",
          "author_url": "",
          "post_date": "2020-05-28T03:20:10.987000",
          "content": "<p>Yeah just reading using cat 1 file took close to 1s I'll take a look.</p>\n\n<p><code>fn = \"/kaggle/input/siim-isic-melanoma-classification/jpeg/train/ISIC_0077735.jpg\"</code></p>\n\n<p><code>%%timeit</code>\n<code>!cat $fn &amp;gt; /dev/null</code></p>\n\n<p><code>741 ms ± 5.33 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)</code></p>\n\n<p>Edit: the file isn't super small though, I'll see what kind of speed we should expect.\n<code>!du -h $fn</code></p>\n\n<p><code>1.3M   /kaggle/input/siim-isic-melanoma-classification/jpeg/train/ISIC_0077735.jpg</code></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 864516,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-05-28T03:38:19.917000",
          "content": "<p>Actually it's the loading that takes time. Maybe because of how large the images are... This would mean having to preprocess and save first I guess.</p>\n\n<p>You can easily try with this code (and splitting the %%timeit in individual cells):</p>\n\n<p>```\nimport cv2\nimport PIL.Image as Image\nfrom torchvision import transforms</p>\n\n<p>image_fn = '/kaggle/input/siim-isic-melanoma-classification/jpeg/train/ISIC_2637011.jpg'\nimage = Image.open(image_fn)\nimage.resize((224, 224)).save('savetest.jpg')</p>\n\n<p>%%timeit\nimage = Image.open(image_fn)\ndown = transforms.ToTensor()(image)</p>\n\n<p>%%timeit\nimage = Image.open(image_fn)\ndown = image.resize((224, 224))\ndown = transforms.ToTensor()(down)</p>\n\n<p>%%timeit\nimage = Image.open('savetest.jpg')\ndown = transforms.ToTensor()(down)</p>\n\n<p>%%timeit\nimage = cv2.imread(image_fn)\ndown = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\ndown = cv2.resize(down, (224, 224))\ndown = transforms.ToTensor()(down)\n```</p>\n\n<p>Note that cv2 pipeline is actually 3 times faster here... but still too slow at 224ms. Saving first make it drop to acceptable level. Still surprise by how slow loading a single image is though...</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 864590,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-05-28T04:51:21.403000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 864596,
          "author_name": "Dustin",
          "author_url": "",
          "post_date": "2020-05-28T04:55:14.323000",
          "content": "<p>I tried two more tests:\n* I tried a different dataset, loading a 20Mb image\n* I tried fetching just the first byte (using head -c 1 instead of cat)</p>\n\n<p>Both of those also took 700ms which means it's not the amount of bytes transferred but the latency to first byte. I'll chat with my team about this tomorrow to figure out what's going on. I gotta get some sleep though it's late in my TZ... 😴 </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 864987,
          "author_name": "olivier",
          "author_url": "",
          "post_date": "2020-05-28T10:16:24.490000",
          "content": "<p>For info, I tried the same command as you did on my own server : \n<code>\nfn = \"jpeg/train/ISIC_0077735.jpg\"\n%%timeit\n!cat $fn &gt; /dev/null\n117 ms ± 1.26 ms per loop (mean ± std. dev. of 7 runs, 10 loops each)\n</code></p>\n\n<p>Seems much quicker.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 865490,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-05-28T16:51:51.530000",
          "content": "<p>Yeah i've preprocessed everything to be 224x224 png files but it is still slow... 10s per batch instead of 32s. Something is definetly fishy. Obviously no problem on my local machine but I want to publish a kernel :P</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 865506,
          "author_name": "Dustin",
          "author_url": "",
          "post_date": "2020-05-28T17:08:00.830000",
          "content": "<p>It does look like the Image libraries are the problem, after doing a bunch more tests, it looks like the reason my %%time was slow for cat was because booting the shell is really slow, the actual copy part is fast.</p>\n\n<p>Edit: the libraries are slow even when working on the data in memory.</p>\n\n<p><a href=\"/arroqc\">@arroqc</a> <a href=\"/elbanan\">@elbanan</a> </p>\n\n<p>Example: <a href=\"https://www.kaggle.com/herbison/fast-load-and-resize?scriptVersionId=34992706\">https://www.kaggle.com/herbison/fast-load-and-resize?scriptVersionId=34992706</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 865510,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-05-28T17:09:48.733000",
          "content": "<p><a href=\"/herbison\">@herbison</a> I have a 404 on that link</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 865520,
          "author_name": "Dustin",
          "author_url": "",
          "post_date": "2020-05-28T17:14:25.570000",
          "content": "<p><a href=\"/arroqc\">@arroqc</a> Updated it and made public, it turns out I messed up time %%time magic, but it does show that it's Image that's slow, not reading the data.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 865536,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-05-28T17:28:42.993000",
          "content": "<p>Thanks i'll give it a try</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 865577,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-05-28T17:54:42.843000",
          "content": "<p><a href=\"/herbison\">@herbison</a> I've created a dataset of small 224x224 images in png. It obviously load faster and is acceptable for me to continue. Edit: nevermind</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 865585,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2020-05-28T18:08:34.047000",
          "content": "<p><a href=\"/arroqc\">@arroqc</a> Just to confirm, are you saying the 224x224 dataset is still slow (10s per batch)?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 865590,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-05-28T18:12:39.507000",
          "content": "<p>No sorry. I've fixed that problem. For some reason it was still running on CPU rather than GPU. </p>\n\n<p>Thanks team.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 864431,
      "author_name": "Elbanan",
      "author_url": "",
      "post_date": "2020-05-28T02:04:52.760000",
      "content": "<p>It is driving me crazy. Creating batches takes forever. I think the bottleneck is data transfer from data repo/storage to the virtual machines. </p>\n\n<p>Any feedback from @Kaggle team , <a href=\"/juliaelliott\">@juliaelliott</a>  would be appreciated . Thanks a lot !</p>",
      "votes": 0,
      "replies": [
        {
          "id": 864440,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-05-28T02:09:44.537000",
          "content": "<p>I have same problem.</p>\n\n<p>What is really weird is that Loading images is fast, Resizing an image is fast but Loading + Resizing is 10 times slower than the individual operation Oo</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 864475,
          "author_name": "Dustin",
          "author_url": "",
          "post_date": "2020-05-28T02:51:55.983000",
          "content": "<p><a href=\"/elbanan\">@elbanan</a> I'm not seeing anything immediately in the storage that would indicate a problem (read ops/latency looks normal).</p>\n\n<p><a href=\"/arroqc\">@arroqc</a> comments also lead me to believe it's CPU, since loading the images is fine, but doing a processing operation with a load is slower. </p>\n\n<p>If your CPU is pinned at 200% then it's more likely you are CPU-bottlenecked than disk read bottlenecked. Not sure exactly what you're doing but looks like your GPU isn't in use either (though maybe it's too early in your pipeline). When GPU is enabled, you only get 2 CPU cores, but when it's disabled you get 4 CPU cores. So if you're CPU-constrained and don't need GPU you're better off on a CPU machine.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 864480,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-05-28T02:55:35.027000",
          "content": "<p><a href=\"/herbison\">@herbison</a> I used gpu for 30s until I noticed it took literally 30s to load a batch of 32 then stopped and tried to investigate. The images are large for sure but loading + resizing shouldn't take a full second.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 864483,
          "author_name": "Elbanan",
          "author_url": "",
          "post_date": "2020-05-28T03:01:57.313000",
          "content": "<p><a href=\"/herbison\">@herbison</a> Thanks  for getting back to me. I will try going the CPU route now although this is unusual for me. I have used this pipeline in the past with multiple datasets and never had similar problem unless there's delay in data transfer. </p>\n\n<p>The pipeline actually loads the data in batches, resizes and does a single pass via CNN for feature extraction.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 864397,
      "author_name": "Shayekh Islam",
      "author_url": "",
      "post_date": "2020-05-28T01:30:58.357000",
      "content": "<p>Downloading using google colab is faster. Then, the data can be downloaded from colab.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 864481,
          "author_name": "Dustin",
          "author_url": "",
          "post_date": "2020-05-28T02:57:45.947000",
          "content": "<p><a href=\"/shayekh\">@shayekh</a> <a href=\"/vikasvpatil\">@vikasvpatil</a> Can you clarify what you mean by download? It's rare that datasets are actually downloaded inside Kaggle notebooks, usually they only need to be 'added' to your notebook via the \"Add Data\" feature?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 864545,
          "author_name": "Shayekh Islam",
          "author_url": "",
          "post_date": "2020-05-28T04:07:53.370000",
          "content": "<p><a href=\"/herbison\">@herbison</a> I meant for training in local machine, not applicable for the question here. The details were added in the question later.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 865826,
          "author_name": "Vikas V Patil",
          "author_url": "",
          "post_date": "2020-05-28T23:12:54.987000",
          "content": "<p><a href=\"/herbison\">@herbison</a> Thanks for the followup. I was planning to train it on a local machine. So was downloading the the files from Kaggle which took some time. Seems like if we use the Kaggle TPU, its much faster. Is it the norm with CV problems to use Kaggle's TPU or GPU? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 865857,
          "author_name": "Dustin",
          "author_url": "",
          "post_date": "2020-05-29T00:05:21.213000",
          "content": "<p><a href=\"/vikasvpatil\">@vikasvpatil</a>  TPUs are significantly faster than GPUs for certain tasks like this one. This competition dataset includes tfrecord files for easy use with the TPU, and those tfrecord files have already been properly sized for the TPU (unlike the raw images which are fairly large which is probably contributing to why it's slow to download the dataset / process the raw image files).</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 865869,
          "author_name": "Matt Yates",
          "author_url": "",
          "post_date": "2020-05-29T00:15:46.907000",
          "content": "<p><a href=\"/herbison\">@herbison</a> - Nice response.  This is the ticket.  Here is another discussion that points to the the same idea - <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154347\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154347</a>. 👍 </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 865890,
          "author_name": "Vikas V Patil",
          "author_url": "",
          "post_date": "2020-05-29T00:46:37.627000",
          "content": "<p><a href=\"/herbison\">@herbison</a> This is helpful. Appreciate your response. 👍 </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 866114,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-29T05:50:07.240000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "864266": "Is anyone else experiencing slow loading of data in notebooks or is it just me ? In the screenshot below, I am using only 20% of the supplied training data. This is the time needed for a forward pass through ResNet50. Processor utilization is super high.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2196129%2F66b2a243ea15e01830a2377d30c12369%2FScreenshot%202020-05-27%2022.11.26.png?generation=1590631909985239&amp;alt=media)\n",
    "864419": "Same ☹️ \n\n**UPDATE:**\nThis is the ticket - https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154347. 👍",
    "864368": "Experiencing the same here. Super slow download.",
    "957749": "Google Colab seems to have upgraded to TF 2.3 this night as well, which is causing issues. Perhaps both are related?",
    "957299": "Hi @herbison Data loading is taking a lot of time in my notebook today and wasting my TPU quota. I am using @cdeotte  datasets which is posted here: https://www.kaggle.com/cdeotte/isic2019-384x384",
    "868428": "I thought I was the only one until I saw this, it really is irritating, I want my model to train for 30 epochs, but only in the 1 epoch it has taken over an hour to get past 33%. :/",
    "865543": "This is too much slow.",
    "864788": "good",
    "864486": "Edit nevermind\n",
    "864431": "It is driving me crazy. Creating batches takes forever. I think the bottleneck is data transfer from data repo/storage to the virtual machines. \n\nAny feedback from @Kaggle team , @juliaelliott  would be appreciated . Thanks a lot !",
    "864397": "Downloading using google colab is faster. Then, the data can be downloaded from colab.",
    "866114": ""
  }
}