{
  "id": 108079,
  "title": "Resized dataset",
  "url": "/competitions/understanding_cloud_organization/discussion/108079",
  "author_name": "ryches",
  "post_date": "2019-09-08T23:57:20.357000",
  "votes": 26,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I see that most people are doing the resize on the fly over and over again for each epoch. This resize process is quite computationally expensive. Here is a resized dataset scaled down to 350, 525 already. This will significantly reduce the time of each epoch training. </p>\n\n<p><a href=\"https://www.kaggle.com/ryches/understanding-clouds-resized\">https://www.kaggle.com/ryches/understanding-clouds-resized</a></p>",
  "messages": [
    {
      "id": 621777,
      "postDate": "2019-09-08T23:57:20.357Z",
      "content": "<p>I see that most people are doing the resize on the fly over and over again for each epoch. This resize process is quite computationally expensive. Here is a resized dataset scaled down to 350, 525 already. This will significantly reduce the time of each epoch training. </p>\n\n<p><a href=\"https://www.kaggle.com/ryches/understanding-clouds-resized\">https://www.kaggle.com/ryches/understanding-clouds-resized</a></p>",
      "rawMarkdown": "I see that most people are doing the resize on the fly over and over again for each epoch. This resize process is quite computationally expensive. Here is a resized dataset scaled down to 350, 525 already. This will significantly reduce the time of each epoch training. \n\nhttps://www.kaggle.com/ryches/understanding-clouds-resized",
      "votes": 26
    },
    {
      "id": 622755,
      "postDate": "2019-09-10T03:40:28.920Z",
      "content": "<p>Utilizing this dataset I was able to turbo charge andrew's pytorch kernel and get .638. I also threw nvidia apex in there so I could do the largest models available in the segmentation models package with a decently large batch size. \n<a href=\"https://www.kaggle.com/ryches/turbo-charging-andrew-s-pytorch?scriptVersionId=20370396\">https://www.kaggle.com/ryches/turbo-charging-andrew-s-pytorch?scriptVersionId=20370396</a></p>",
      "rawMarkdown": "Utilizing this dataset I was able to turbo charge andrew's pytorch kernel and get .638. I also threw nvidia apex in there so I could do the largest models available in the segmentation models package with a decently large batch size. \nhttps://www.kaggle.com/ryches/turbo-charging-andrew-s-pytorch?scriptVersionId=20370396",
      "votes": 3
    },
    {
      "id": 644528,
      "postDate": "2019-10-09T00:07:37.240Z",
      "content": "<p>How did you generate and save the mask?\nI find it is different from mask = rledecode(maskrle) then mask = resize_it350(mask) . \nThere are many small points in masks of your data set.</p>",
      "rawMarkdown": "How did you generate and save the mask?\nI find it is different from mask = rledecode(maskrle) then mask = resize_it350(mask) . \nThere are many small points in masks of your data set.",
      "replies": [
        {
          "id": 644566,
          "postDate": "2019-10-09T01:47:37.873Z",
          "content": "<p>It should be just those two steps. I'd have to go back and review the exact code though. I saw the same and thought it was just artifacts of resizing or something like that</p>",
          "rawMarkdown": "It should be just those two steps. I'd have to go back and review the exact code though. I saw the same and thought it was just artifacts of resizing or something like that"
        },
        {
          "id": 645361,
          "postDate": "2019-10-10T03:22:33.863Z",
          "content": "<p>it is possibly caused by data type change during saving.</p>",
          "rawMarkdown": "it is possibly caused by data type change during saving."
        }
      ]
    },
    {
      "id": 638711,
      "postDate": "2019-10-02T10:26:18.620Z",
      "content": "<p>Shouldn't height and width of input images be divisible by 32 for segmentation models? I get this error using Keras: ValueError: A <code>Concatenate</code> layer requires inputs with matching shapes except for the concat axis.</p>",
      "rawMarkdown": "Shouldn't height and width of input images be divisible by 32 for segmentation models? I get this error using Keras: ValueError: A `Concatenate` layer requires inputs with matching shapes except for the concat axis.",
      "replies": [
        {
          "id": 638955,
          "postDate": "2019-10-02T15:57:11.430Z",
          "content": "<p>Here is what I do:</p>\n\n<ul>\n<li>pick a height that is divisible by 64</li>\n<li>derive width: int(h * 1.5)</li>\n</ul>\n\n<p>Aspect ratio is identical to the original aspect ratio.</p>\n\n<p>I need the height divisible by 64 to make the width divisible by 32.</p>",
          "rawMarkdown": "Here is what I do:\n\n- pick a height that is divisible by 64\n- derive width: int(h * 1.5)\n\nAspect ratio is identical to the original aspect ratio.\n\nI need the height divisible by 64 to make the width divisible by 32."
        },
        {
          "id": 639177,
          "postDate": "2019-10-02T20:56:08.467Z",
          "content": "<p>It does if you use certain architectures. The goal of this dataset is just to get it down to the target size. Since our predictions have to be 350x525 they'll need to be resized to that at some point. And better to be 350x525 once instead of downscaling from much higher resolution down to that everytime for every epoch</p>",
          "rawMarkdown": "It does if you use certain architectures. The goal of this dataset is just to get it down to the target size. Since our predictions have to be 350x525 they'll need to be resized to that at some point. And better to be 350x525 once instead of downscaling from much higher resolution down to that everytime for every epoch"
        },
        {
          "id": 640072,
          "postDate": "2019-10-03T19:47:19.300Z",
          "content": "<p>Of course I do downscaling prior to ML and store the images, like everyone does. And I rescale the result to 350x525 after training, so this is a negligible overhead. During training, I use images of various sizes read directly from disk with no additional resizing.</p>",
          "rawMarkdown": "Of course I do downscaling prior to ML and store the images, like everyone does. And I rescale the result to 350x525 after training, so this is a negligible overhead. During training, I use images of various sizes read directly from disk with no additional resizing."
        }
      ]
    },
    {
      "id": 628345,
      "postDate": "2019-09-17T07:34:11.977Z",
      "content": "<p>Thankyou!!</p>",
      "rawMarkdown": "Thankyou!!"
    },
    {
      "id": 628121,
      "postDate": "2019-09-16T20:26:56.167Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 621983,
      "postDate": "2019-09-09T06:19:21.423Z",
      "content": "<p>Thanks a ton!</p>",
      "rawMarkdown": "Thanks a ton!"
    }
  ],
  "comments": [
    {
      "id": 622755,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2019-09-10T03:40:28.920000",
      "content": "<p>Utilizing this dataset I was able to turbo charge andrew's pytorch kernel and get .638. I also threw nvidia apex in there so I could do the largest models available in the segmentation models package with a decently large batch size. \n<a href=\"https://www.kaggle.com/ryches/turbo-charging-andrew-s-pytorch?scriptVersionId=20370396\">https://www.kaggle.com/ryches/turbo-charging-andrew-s-pytorch?scriptVersionId=20370396</a></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 644528,
      "author_name": "abnerzhang",
      "author_url": "",
      "post_date": "2019-10-09T00:07:37.240000",
      "content": "<p>How did you generate and save the mask?\nI find it is different from mask = rledecode(maskrle) then mask = resize_it350(mask) . \nThere are many small points in masks of your data set.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 644566,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2019-10-09T01:47:37.873000",
          "content": "<p>It should be just those two steps. I'd have to go back and review the exact code though. I saw the same and thought it was just artifacts of resizing or something like that</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 645361,
          "author_name": "abnerzhang",
          "author_url": "",
          "post_date": "2019-10-10T03:22:33.863000",
          "content": "<p>it is possibly caused by data type change during saving.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 638711,
      "author_name": "Socratis Gkelios",
      "author_url": "",
      "post_date": "2019-10-02T10:26:18.620000",
      "content": "<p>Shouldn't height and width of input images be divisible by 32 for segmentation models? I get this error using Keras: ValueError: A <code>Concatenate</code> layer requires inputs with matching shapes except for the concat axis.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 638955,
          "author_name": "pete",
          "author_url": "",
          "post_date": "2019-10-02T15:57:11.430000",
          "content": "<p>Here is what I do:</p>\n\n<ul>\n<li>pick a height that is divisible by 64</li>\n<li>derive width: int(h * 1.5)</li>\n</ul>\n\n<p>Aspect ratio is identical to the original aspect ratio.</p>\n\n<p>I need the height divisible by 64 to make the width divisible by 32.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 639177,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2019-10-02T20:56:08.467000",
          "content": "<p>It does if you use certain architectures. The goal of this dataset is just to get it down to the target size. Since our predictions have to be 350x525 they'll need to be resized to that at some point. And better to be 350x525 once instead of downscaling from much higher resolution down to that everytime for every epoch</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 640072,
          "author_name": "pete",
          "author_url": "",
          "post_date": "2019-10-03T19:47:19.300000",
          "content": "<p>Of course I do downscaling prior to ML and store the images, like everyone does. And I rescale the result to 350x525 after training, so this is a negligible overhead. During training, I use images of various sizes read directly from disk with no additional resizing.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 628345,
      "author_name": "Anthony Wynne",
      "author_url": "",
      "post_date": "2019-09-17T07:34:11.977000",
      "content": "<p>Thankyou!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 628121,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-16T20:26:56.167000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 621983,
      "author_name": "Mighty Rains",
      "author_url": "",
      "post_date": "2019-09-09T06:19:21.423000",
      "content": "<p>Thanks a ton!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "621777": "I see that most people are doing the resize on the fly over and over again for each epoch. This resize process is quite computationally expensive. Here is a resized dataset scaled down to 350, 525 already. This will significantly reduce the time of each epoch training. \n\nhttps://www.kaggle.com/ryches/understanding-clouds-resized",
    "622755": "Utilizing this dataset I was able to turbo charge andrew's pytorch kernel and get .638. I also threw nvidia apex in there so I could do the largest models available in the segmentation models package with a decently large batch size. \nhttps://www.kaggle.com/ryches/turbo-charging-andrew-s-pytorch?scriptVersionId=20370396",
    "644528": "How did you generate and save the mask?\nI find it is different from mask = rledecode(maskrle) then mask = resize_it350(mask) . \nThere are many small points in masks of your data set.",
    "638711": "Shouldn't height and width of input images be divisible by 32 for segmentation models? I get this error using Keras: ValueError: A `Concatenate` layer requires inputs with matching shapes except for the concat axis.",
    "628345": "Thankyou!!",
    "628121": "",
    "621983": "Thanks a ton!"
  }
}