{
  "id": 116261,
  "title": "Trick to Speed Up Data Loaders",
  "url": "/competitions/understanding_cloud_organization/discussion/116261",
  "author_name": "Chris Deotte",
  "post_date": "2019-11-08T01:10:23.016000",
  "votes": 46,
  "comment_count": 29,
  "views": 0,
  "content": "<p>As the competition deadline approaches everyone wants their experiments to run faster so we can test more ideas.</p>\n\n<p>I would like to point out that many public data loaders can actually run up to 4 times faster! (And hence epochs can train up to 4 times faster!!)</p>\n\n<p>As Ryches pointed out previously <a href=\"https://www.kaggle.com/c/understanding_cloud_organization/discussion/108079\">here</a>, we should resize everything (images and masks) ahead of time and then our data loader won't have to do it before each batch. Ryches published a 4x reduced size dataset <a href=\"https://www.kaggle.com/ryches/understanding-clouds-resized\">here</a>. And Phung published a notebook which can pre-process any size reduction <a href=\"https://www.kaggle.com/phunghieu/dataset-preparation-resize-images\">here</a></p>\n\n<p>Alternatively if you must resize arrays in your data loader try using NumPy indexing like <code>img_quarter = img[::4,::4,:]</code> and <code>mask_quarter = mask[::4,::4,:]</code>. Currently my experiments run at 3 minute epochs!!</p>\n\n<p>UPDATE: I posted a Kaggle dataset <a href=\"https://www.kaggle.com/c/understanding_cloud_organization/discussion/108079\">here</a> with 6 different rescaled sizes, <code>384x576</code>, <code>320x480</code>, <code>256x384</code>, <code>192x288</code>, <code>128x192</code>, and <code>64x96</code>. Run your experiments with these smaller sizes to save time.</p>",
  "messages": [
    {
      "id": 668093,
      "postDate": "2019-11-08T01:10:23.017Z",
      "content": "<p>As the competition deadline approaches everyone wants their experiments to run faster so we can test more ideas.</p>\n\n<p>I would like to point out that many public data loaders can actually run up to 4 times faster! (And hence epochs can train up to 4 times faster!!)</p>\n\n<p>As Ryches pointed out previously <a href=\"https://www.kaggle.com/c/understanding_cloud_organization/discussion/108079\">here</a>, we should resize everything (images and masks) ahead of time and then our data loader won't have to do it before each batch. Ryches published a 4x reduced size dataset <a href=\"https://www.kaggle.com/ryches/understanding-clouds-resized\">here</a>. And Phung published a notebook which can pre-process any size reduction <a href=\"https://www.kaggle.com/phunghieu/dataset-preparation-resize-images\">here</a></p>\n\n<p>Alternatively if you must resize arrays in your data loader try using NumPy indexing like <code>img_quarter = img[::4,::4,:]</code> and <code>mask_quarter = mask[::4,::4,:]</code>. Currently my experiments run at 3 minute epochs!!</p>\n\n<p>UPDATE: I posted a Kaggle dataset <a href=\"https://www.kaggle.com/c/understanding_cloud_organization/discussion/108079\">here</a> with 6 different rescaled sizes, <code>384x576</code>, <code>320x480</code>, <code>256x384</code>, <code>192x288</code>, <code>128x192</code>, and <code>64x96</code>. Run your experiments with these smaller sizes to save time.</p>",
      "rawMarkdown": "As the competition deadline approaches everyone wants their experiments to run faster so we can test more ideas.\n\nI would like to point out that many public data loaders can actually run up to 4 times faster! (And hence epochs can train up to 4 times faster!!)\n\nAs Ryches pointed out previously [here][1], we should resize everything (images and masks) ahead of time and then our data loader won't have to do it before each batch. Ryches published a 4x reduced size dataset [here][2]. And Phung published a notebook which can pre-process any size reduction [here][3]\n\nAlternatively if you must resize arrays in your data loader try using NumPy indexing like `img_quarter = img[::4,::4,:]` and `mask_quarter = mask[::4,::4,:]`. Currently my experiments run at 3 minute epochs!!\n\nUPDATE: I posted a Kaggle dataset [here][1] with 6 different rescaled sizes, `384x576`, `320x480`, `256x384`, `192x288`, `128x192`, and `64x96`. Run your experiments with these smaller sizes to save time.\n\n[1]: https://www.kaggle.com/c/understanding_cloud_organization/discussion/108079\n[2]: https://www.kaggle.com/ryches/understanding-clouds-resized\n[3]: https://www.kaggle.com/phunghieu/dataset-preparation-resize-images\n[4]: https://www.kaggle.com/cdeotte/cloud-images-resized\n\n",
      "votes": 46
    },
    {
      "id": 668100,
      "postDate": "2019-11-08T01:32:32.513Z",
      "content": "<p>Resizing previously has saved me a lot of time, but one extra information, I've also posted <a href=\"https://www.kaggle.com/dimitreoliveira/cloud-segmentation-with-utility-scripts-and-keras\">a kernel with resizing code</a> (Pre-process data section), mine use parallel preprocessing, maybe it can reduce the time to resize images.</p>\n\n<p>And the epoch time depends heavily on the model (U-net, FPN) and backbone, what configuration you use for the 3 minutes epochs?</p>",
      "rawMarkdown": "Resizing previously has saved me a lot of time, but one extra information, I've also posted [a kernel with resizing code](https://www.kaggle.com/dimitreoliveira/cloud-segmentation-with-utility-scripts-and-keras) (Pre-process data section), mine use parallel preprocessing, maybe it can reduce the time to resize images.\n\nAnd the epoch time depends heavily on the model (U-net, FPN) and backbone, what configuration you use for the 3 minutes epochs?",
      "votes": 3,
      "replies": [
        {
          "id": 668108,
          "postDate": "2019-11-08T01:44:53.760Z",
          "content": "<p>I didn't notice that when I read your notebook before. Nice job. I see that your notebook calls your utility script <a href=\"https://www.kaggle.com/dimitreoliveira/cloud-images-segmentation-utillity-script\">here</a>.</p>\n\n<p>I'm achieving 3 minutes epochs with custom made U-nets, pretrained ImageNet backbones, and 5x reduction on the fly, i.e. <code>img = img[::5,::5,:]</code> and <code>mask = mask[::5,::5,:]</code>. Also my data loader does augmentation rotation, scale, shift, and flips. I use albumentations for rotation and NumPy for others.</p>",
          "rawMarkdown": "I didn't notice that when I read your notebook before. Nice job. I see that your notebook calls your utility script [here][1].\n\nI'm achieving 3 minutes epochs with custom made U-nets, pretrained ImageNet backbones, and 5x reduction on the fly, i.e. `img = img[::5,::5,:]` and `mask = mask[::5,::5,:]`. Also my data loader does augmentation rotation, scale, shift, and flips. I use albumentations for rotation and NumPy for others.\n\n[1]: https://www.kaggle.com/dimitreoliveira/cloud-images-segmentation-utillity-script",
          "votes": 1
        },
        {
          "id": 668393,
          "postDate": "2019-11-08T11:27:42.003Z",
          "content": "<p>threChris,I wonder if this would lose some information to do <code>img=img[::5,::5,:]</code> cause the <code>resize</code> function has many different ways to reduce pixels</p>",
          "rawMarkdown": "threChris,I wonder if this would lose some information to do `img=img[::5,::5,:]` cause the `resize` function has many different ways to reduce pixels",
          "votes": 3
        },
        {
          "id": 668408,
          "postDate": "2019-11-08T11:52:00.683Z",
          "content": "<p>Just now I stopped to think about this way of resizing a numpy array, this is really clever, but as <a href=\"/sj626591833\">@sj626591833</a> said, there are many ways to resize images, I think this \"step index\" is closer to a linear resizing, For those that don't got how it works, what it does is to \"avoid\" some rows and columns, is you use <code>img=img[::5,::5,:]</code> it will get the first row and column and jump over the next 5.</p>",
          "rawMarkdown": "Just now I stopped to think about this way of resizing a numpy array, this is really clever, but as @sj626591833 said, there are many ways to resize images, I think this \"step index\" is closer to a linear resizing, For those that don't got how it works, what it does is to \"avoid\" some rows and columns, is you use `img=img[::5,::5,:]` it will get the first row and column and jump over the next 5.",
          "votes": 1
        },
        {
          "id": 668677,
          "postDate": "2019-11-08T17:53:15.047Z",
          "content": "<p><a href=\"/sj626591833\">@sj626591833</a> <a href=\"/dimitreoliveira\">@dimitreoliveira</a>  Great point. I realize that using <code>img[::5,::5,:]</code> is different than using <code>cv2.INTER_AREA</code>. The former replaces a <code>5x5</code> pixel crop with a single pixel value, whereas the later replaces a <code>5x5</code> pixel crop with the average of 25 pixel values. I tested both and my CV is actually a little better using <code>img[::5,::5,:]</code>.</p>\n\n<p>(But note it may not be better for everyone. Everyone should determine which downsampling method is best for them by comparing CV scores. But it is true that <code>img[::5,::5,:]</code> is super fast!).</p>",
          "rawMarkdown": "@sj626591833 @dimitreoliveira  Great point. I realize that using `img[::5,::5,:]` is different than using `cv2.INTER_AREA`. The former replaces a `5x5` pixel crop with a single pixel value, whereas the later replaces a `5x5` pixel crop with the average of 25 pixel values. I tested both and my CV is actually a little better using `img[::5,::5,:]`.\n\n(But note it may not be better for everyone. Everyone should determine which downsampling method is best for them by comparing CV scores. But it is true that `img[::5,::5,:]` is super fast!).",
          "votes": 3
        },
        {
          "id": 668719,
          "postDate": "2019-11-08T19:11:34.420Z",
          "content": "<p>Kinda off-topic, but if anyone could give me a hand, I just need 3 more votes on the <a href=\"https://www.kaggle.com/dimitreoliveira/cloud-segmentation-with-utility-scripts-and-keras/comments#668712\">kernel that I just linked</a> with all the pre-process, resizing and stuff to get a silver and become a kernels master 😄 , thanks in advance.</p>",
          "rawMarkdown": "Kinda off-topic, but if anyone could give me a hand, I just need 3 more votes on the [kernel that I just linked](https://www.kaggle.com/dimitreoliveira/cloud-segmentation-with-utility-scripts-and-keras/comments#668712) with all the pre-process, resizing and stuff to get a silver and become a kernels master 😄 , thanks in advance.",
          "votes": 2
        },
        {
          "id": 668730,
          "postDate": "2019-11-08T19:23:53.433Z",
          "content": "<p>That kernel is awesome. I'm surprised that it doesn't have more votes already. I upvoted it. </p>\n\n<p>I encourage everyone to upvote it too <a href=\"https://www.kaggle.com/dimitreoliveira/cloud-segmentation-with-utility-scripts-and-keras/\">here</a>. It is a perfect template for this competition which includes everything neccessary for building and training a segmentation model. </p>",
          "rawMarkdown": "That kernel is awesome. I'm surprised that it doesn't have more votes already. I upvoted it. \n  \nI encourage everyone to upvote it too [here][1]. It is a perfect template for this competition which includes everything neccessary for building and training a segmentation model. \n\n[1]: https://www.kaggle.com/dimitreoliveira/cloud-segmentation-with-utility-scripts-and-keras/",
          "votes": 1
        },
        {
          "id": 668738,
          "postDate": "2019-11-08T19:35:04.013Z",
          "content": "<p>Thanks for the help Chris</p>",
          "rawMarkdown": "Thanks for the help Chris",
          "votes": 1
        },
        {
          "id": 668875,
          "postDate": "2019-11-09T01:51:04.327Z",
          "content": "<p><a href=\"/dimitreoliveira\">@dimitreoliveira</a>  Congratulations on becoming Kaggle Kernels Master !! Thanks for sharing so many great notebooks.</p>",
          "rawMarkdown": "@dimitreoliveira  Congratulations on becoming Kaggle Kernels Master !! Thanks for sharing so many great notebooks.",
          "votes": 2
        },
        {
          "id": 669066,
          "postDate": "2019-11-09T12:25:51.897Z",
          "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a> , this has been a great journey.</p>",
          "rawMarkdown": "Thanks @cdeotte , this has been a great journey.",
          "votes": 1
        }
      ]
    },
    {
      "id": 684193,
      "postDate": "2019-11-29T11:35:08.507Z",
      "content": "<p>nice</p>",
      "rawMarkdown": "nice",
      "votes": 1
    },
    {
      "id": 683980,
      "postDate": "2019-11-29T05:13:06.510Z",
      "content": "<p>nice</p>",
      "rawMarkdown": "nice",
      "votes": 1
    },
    {
      "id": 683171,
      "postDate": "2019-11-28T07:00:30.720Z",
      "content": "<p>Cool!</p>",
      "rawMarkdown": "Cool!",
      "votes": 1
    },
    {
      "id": 682604,
      "postDate": "2019-11-27T16:57:53.210Z",
      "content": "<p>valuable insight</p>",
      "rawMarkdown": "valuable insight",
      "votes": 1
    },
    {
      "id": 682381,
      "postDate": "2019-11-27T11:05:54.007Z",
      "content": "<p>nice</p>",
      "rawMarkdown": "nice",
      "votes": 1
    },
    {
      "id": 682325,
      "postDate": "2019-11-27T09:07:02.157Z",
      "content": "<p>Great work..!</p>",
      "rawMarkdown": "Great work..!",
      "votes": 1
    },
    {
      "id": 679950,
      "postDate": "2019-11-23T18:00:19.237Z",
      "content": "<p>Helpful. Struggling with the data loading for image classification. Great Work.</p>",
      "rawMarkdown": "Helpful. Struggling with the data loading for image classification. Great Work.",
      "votes": 1
    },
    {
      "id": 678967,
      "postDate": "2019-11-22T05:33:34.813Z",
      "content": "<p>nice</p>",
      "rawMarkdown": "nice",
      "votes": 1
    },
    {
      "id": 675540,
      "postDate": "2019-11-18T07:48:19.613Z",
      "content": "<p>great work!!!</p>",
      "rawMarkdown": "great work!!!",
      "votes": 1
    },
    {
      "id": 672285,
      "postDate": "2019-11-13T18:04:22.797Z",
      "content": "<p>nice work!!!</p>",
      "rawMarkdown": "nice work!!!",
      "votes": 1
    },
    {
      "id": 670129,
      "postDate": "2019-11-11T03:47:10.170Z",
      "content": "<p>Pretty cool !!</p>",
      "rawMarkdown": "Pretty cool !!",
      "votes": 1
    },
    {
      "id": 669439,
      "postDate": "2019-11-10T04:36:47.963Z",
      "content": "<p>Don't forget to resize mask too😎  Some part of my code was hard coded with (1400, 2100)....</p>",
      "rawMarkdown": "Don't forget to resize mask too😎  Some part of my code was hard coded with (1400, 2100)....",
      "votes": 2,
      "replies": [
        {
          "id": 672288,
          "postDate": "2019-11-13T18:13:30.957Z",
          "content": "<p>Yes, great warning. We need to update our calls to <code>rle2mask</code> function to use the new size.</p>",
          "rawMarkdown": "Yes, great warning. We need to update our calls to `rle2mask` function to use the new size.",
          "votes": 1
        }
      ]
    },
    {
      "id": 675535,
      "postDate": "2019-11-18T07:37:32.207Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 673609,
      "postDate": "2019-11-15T08:03:36.123Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 686336,
      "postDate": "2019-12-03T03:26:13.793Z",
      "content": "<p>Thanks! ^_^</p>",
      "rawMarkdown": "Thanks! ^_^",
      "votes": 1
    },
    {
      "id": 682488,
      "postDate": "2019-11-27T13:59:56.457Z",
      "content": "<p>That's actually really helpful! Thanks</p>",
      "rawMarkdown": "That's actually really helpful! Thanks",
      "votes": 1
    },
    {
      "id": 682188,
      "postDate": "2019-11-27T04:02:41.577Z",
      "content": "<p>It helps, thank you.</p>",
      "rawMarkdown": "It helps, thank you.",
      "votes": 1
    },
    {
      "id": 669956,
      "postDate": "2019-11-10T18:49:37.360Z",
      "content": "<p>Thank you for sharing.</p>",
      "rawMarkdown": "Thank you for sharing.",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 668100,
      "author_name": "DimitreOliveira",
      "author_url": "",
      "post_date": "2019-11-08T01:32:32.513000",
      "content": "<p>Resizing previously has saved me a lot of time, but one extra information, I've also posted <a href=\"https://www.kaggle.com/dimitreoliveira/cloud-segmentation-with-utility-scripts-and-keras\">a kernel with resizing code</a> (Pre-process data section), mine use parallel preprocessing, maybe it can reduce the time to resize images.</p>\n\n<p>And the epoch time depends heavily on the model (U-net, FPN) and backbone, what configuration you use for the 3 minutes epochs?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 668108,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2019-11-08T01:44:53.760000",
          "content": "<p>I didn't notice that when I read your notebook before. Nice job. I see that your notebook calls your utility script <a href=\"https://www.kaggle.com/dimitreoliveira/cloud-images-segmentation-utillity-script\">here</a>.</p>\n\n<p>I'm achieving 3 minutes epochs with custom made U-nets, pretrained ImageNet backbones, and 5x reduction on the fly, i.e. <code>img = img[::5,::5,:]</code> and <code>mask = mask[::5,::5,:]</code>. Also my data loader does augmentation rotation, scale, shift, and flips. I use albumentations for rotation and NumPy for others.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 668393,
          "author_name": "Overfit Queen",
          "author_url": "",
          "post_date": "2019-11-08T11:27:42.003000",
          "content": "<p>threChris,I wonder if this would lose some information to do <code>img=img[::5,::5,:]</code> cause the <code>resize</code> function has many different ways to reduce pixels</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 668408,
          "author_name": "DimitreOliveira",
          "author_url": "",
          "post_date": "2019-11-08T11:52:00.683000",
          "content": "<p>Just now I stopped to think about this way of resizing a numpy array, this is really clever, but as <a href=\"/sj626591833\">@sj626591833</a> said, there are many ways to resize images, I think this \"step index\" is closer to a linear resizing, For those that don't got how it works, what it does is to \"avoid\" some rows and columns, is you use <code>img=img[::5,::5,:]</code> it will get the first row and column and jump over the next 5.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 668677,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2019-11-08T17:53:15.047000",
          "content": "<p><a href=\"/sj626591833\">@sj626591833</a> <a href=\"/dimitreoliveira\">@dimitreoliveira</a>  Great point. I realize that using <code>img[::5,::5,:]</code> is different than using <code>cv2.INTER_AREA</code>. The former replaces a <code>5x5</code> pixel crop with a single pixel value, whereas the later replaces a <code>5x5</code> pixel crop with the average of 25 pixel values. I tested both and my CV is actually a little better using <code>img[::5,::5,:]</code>.</p>\n\n<p>(But note it may not be better for everyone. Everyone should determine which downsampling method is best for them by comparing CV scores. But it is true that <code>img[::5,::5,:]</code> is super fast!).</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 668719,
          "author_name": "DimitreOliveira",
          "author_url": "",
          "post_date": "2019-11-08T19:11:34.420000",
          "content": "<p>Kinda off-topic, but if anyone could give me a hand, I just need 3 more votes on the <a href=\"https://www.kaggle.com/dimitreoliveira/cloud-segmentation-with-utility-scripts-and-keras/comments#668712\">kernel that I just linked</a> with all the pre-process, resizing and stuff to get a silver and become a kernels master 😄 , thanks in advance.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 668730,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2019-11-08T19:23:53.433000",
          "content": "<p>That kernel is awesome. I'm surprised that it doesn't have more votes already. I upvoted it. </p>\n\n<p>I encourage everyone to upvote it too <a href=\"https://www.kaggle.com/dimitreoliveira/cloud-segmentation-with-utility-scripts-and-keras/\">here</a>. It is a perfect template for this competition which includes everything neccessary for building and training a segmentation model. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 668738,
          "author_name": "DimitreOliveira",
          "author_url": "",
          "post_date": "2019-11-08T19:35:04.013000",
          "content": "<p>Thanks for the help Chris</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 668875,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2019-11-09T01:51:04.327000",
          "content": "<p><a href=\"/dimitreoliveira\">@dimitreoliveira</a>  Congratulations on becoming Kaggle Kernels Master !! Thanks for sharing so many great notebooks.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 669066,
          "author_name": "DimitreOliveira",
          "author_url": "",
          "post_date": "2019-11-09T12:25:51.897000",
          "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a> , this has been a great journey.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 684193,
      "author_name": "Rushikesh Pawar",
      "author_url": "",
      "post_date": "2019-11-29T11:35:08.507000",
      "content": "<p>nice</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 683980,
      "author_name": "Abhishek Pathak",
      "author_url": "",
      "post_date": "2019-11-29T05:13:06.510000",
      "content": "<p>nice</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 683171,
      "author_name": "Artem Petrushenko",
      "author_url": "",
      "post_date": "2019-11-28T07:00:30.720000",
      "content": "<p>Cool!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 682604,
      "author_name": "dara singh",
      "author_url": "",
      "post_date": "2019-11-27T16:57:53.210000",
      "content": "<p>valuable insight</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 682381,
      "author_name": "asheensyam",
      "author_url": "",
      "post_date": "2019-11-27T11:05:54.007000",
      "content": "<p>nice</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 682325,
      "author_name": "Dev Jadhav",
      "author_url": "",
      "post_date": "2019-11-27T09:07:02.157000",
      "content": "<p>Great work..!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 679950,
      "author_name": "Aadarsh Choudhary",
      "author_url": "",
      "post_date": "2019-11-23T18:00:19.237000",
      "content": "<p>Helpful. Struggling with the data loading for image classification. Great Work.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 678967,
      "author_name": "JMPLVA",
      "author_url": "",
      "post_date": "2019-11-22T05:33:34.813000",
      "content": "<p>nice</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 675540,
      "author_name": "Mallikarjun v Sajjan",
      "author_url": "",
      "post_date": "2019-11-18T07:48:19.613000",
      "content": "<p>great work!!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 672285,
      "author_name": "Dheer",
      "author_url": "",
      "post_date": "2019-11-13T18:04:22.797000",
      "content": "<p>nice work!!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 670129,
      "author_name": "smart_pointer",
      "author_url": "",
      "post_date": "2019-11-11T03:47:10.170000",
      "content": "<p>Pretty cool !!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 669439,
      "author_name": "Camaro",
      "author_url": "",
      "post_date": "2019-11-10T04:36:47.963000",
      "content": "<p>Don't forget to resize mask too😎  Some part of my code was hard coded with (1400, 2100)....</p>",
      "votes": 2,
      "replies": [
        {
          "id": 672288,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2019-11-13T18:13:30.957000",
          "content": "<p>Yes, great warning. We need to update our calls to <code>rle2mask</code> function to use the new size.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 675535,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-18T07:37:32.207000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 673609,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-15T08:03:36.123000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 686336,
      "author_name": "Azwad Abid",
      "author_url": "",
      "post_date": "2019-12-03T03:26:13.793000",
      "content": "<p>Thanks! ^_^</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 682488,
      "author_name": "DanTe",
      "author_url": "",
      "post_date": "2019-11-27T13:59:56.457000",
      "content": "<p>That's actually really helpful! Thanks</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 682188,
      "author_name": "Yuming Chen",
      "author_url": "",
      "post_date": "2019-11-27T04:02:41.577000",
      "content": "<p>It helps, thank you.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 669956,
      "author_name": "Prachi",
      "author_url": "",
      "post_date": "2019-11-10T18:49:37.360000",
      "content": "<p>Thank you for sharing.</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "668093": "As the competition deadline approaches everyone wants their experiments to run faster so we can test more ideas.\n\nI would like to point out that many public data loaders can actually run up to 4 times faster! (And hence epochs can train up to 4 times faster!!)\n\nAs Ryches pointed out previously [here][1], we should resize everything (images and masks) ahead of time and then our data loader won't have to do it before each batch. Ryches published a 4x reduced size dataset [here][2]. And Phung published a notebook which can pre-process any size reduction [here][3]\n\nAlternatively if you must resize arrays in your data loader try using NumPy indexing like `img_quarter = img[::4,::4,:]` and `mask_quarter = mask[::4,::4,:]`. Currently my experiments run at 3 minute epochs!!\n\nUPDATE: I posted a Kaggle dataset [here][1] with 6 different rescaled sizes, `384x576`, `320x480`, `256x384`, `192x288`, `128x192`, and `64x96`. Run your experiments with these smaller sizes to save time.\n\n[1]: https://www.kaggle.com/c/understanding_cloud_organization/discussion/108079\n[2]: https://www.kaggle.com/ryches/understanding-clouds-resized\n[3]: https://www.kaggle.com/phunghieu/dataset-preparation-resize-images\n[4]: https://www.kaggle.com/cdeotte/cloud-images-resized\n\n",
    "668100": "Resizing previously has saved me a lot of time, but one extra information, I've also posted [a kernel with resizing code](https://www.kaggle.com/dimitreoliveira/cloud-segmentation-with-utility-scripts-and-keras) (Pre-process data section), mine use parallel preprocessing, maybe it can reduce the time to resize images.\n\nAnd the epoch time depends heavily on the model (U-net, FPN) and backbone, what configuration you use for the 3 minutes epochs?",
    "684193": "nice",
    "683980": "nice",
    "683171": "Cool!",
    "682604": "valuable insight",
    "682381": "nice",
    "682325": "Great work..!",
    "679950": "Helpful. Struggling with the data loading for image classification. Great Work.",
    "678967": "nice",
    "675540": "great work!!!",
    "672285": "nice work!!!",
    "670129": "Pretty cool !!",
    "669439": "Don't forget to resize mask too😎  Some part of my code was hard coded with (1400, 2100)....",
    "675535": "",
    "673609": "",
    "686336": "Thanks! ^_^",
    "682488": "That's actually really helpful! Thanks",
    "682188": "It helps, thank you.",
    "669956": "Thank you for sharing."
  }
}