{
  "id": 115115,
  "title": "Train with crops, Predict with full images",
  "url": "/competitions/understanding_cloud_organization/discussion/115115",
  "author_name": "",
  "post_date": "2019-10-31T12:27:03.885711200Z",
  "votes": 28,
  "comment_count": 19,
  "views": 0,
  "content": "<p>I would like to share a cool trick that I learned in Steel Comp. Since segmentation models use all convolutions you can change the input size whenever you want even <strong>during</strong> or <strong>after</strong> training. (More explanation <a href=\"https://www.kaggle.com/c/severstal-steel-defect-detection/discussion/114321\">here</a>).</p>\n\n<p>If you use Qubvel's GitHub segmentation models, Keras <a href=\"https://github.com/qubvel/segmentation_models\">here</a>, PyTorch <a href=\"https://github.com/qubvel/segmentation_models.pytorch\">here</a>. Then define your input size as <code>(None, None, 3)</code> like this</p>\n\n<pre><code>from segmentation_models import Unet\nmodel = Unet('resnet34', input_shape=(None, None, 3), classes=4, \n                    activation='sigmoid')\n</code></pre>\n\n<p>This allows us to be more creative with our training. For example, you can train a few epochs with <code>384x384</code> crops, then train a few epochs with <code>512x512</code> crops, then train with <code>512x1024</code> crops. And then after your model has been trained, you can use it to predict full images of size <code>1400x2100</code>. Pretty cool, huh?</p>\n\n<pre><code>model.fit(X_train, y_train)\n</code></pre>\n\n<p>Above <code>X_train</code> has dimensions <code>(batch_size, height, width, 3)</code>. Each time you call <code>fit()</code> or <code>predict()</code>,  the sizes <code>height, width</code> can be any values divisible by 32. Since GPU memory is limited, using random crops allows us to increase <code>batch_size</code> while decreasing <code>height, width</code>. This allows our network to see more images (different annotators) in one batch. </p>\n\n<p>Additionally instead of random crops, you can selectively choose crops if you wish to teach your network something specific. This was helpful in Steel Comp for networks to learn rare defects but may not be as helpful here. I posted a starter kernel <a href=\"https://www.kaggle.com/cdeotte/train-with-crops-cv-0-60\">here</a>. Enjoy!!</p>\n\n############################\n\n<p><strong>Important Note</strong>: that when I say \"crops\", I mean you cut random rectangles from the original image therefore your network sees a \"localized region\". This is different than \"resizing\" which rescales images and your network still sees the \"everything\". Note that the crops have the same resolution as the full image. </p>\n\n<p>For example, if the full image is <code>400x600</code> pixels and 4 inch by 6 inch. When you cut <code>100x100</code> pixel crops of size 1 inch by inch, both the full image and crops have 100 pixels per inch. If your network learns that <code>Flowers</code> fit in boxes of 50 pixels (0.5 inch) from the crops, this is still true for the full images.</p>",
  "messages": [
    {
      "id": "662320",
      "postDate": "10/31/2019 12:27:03",
      "content": "<p>I would like to share a cool trick that I learned in Steel Comp. Since segmentation models use all convolutions you can change the input size whenever you want even <strong>during</strong> or <strong>after</strong> training. (More explanation <a href=\"https://www.kaggle.com/c/severstal-steel-defect-detection/discussion/114321\">here</a>).</p>\n\n<p>If you use Qubvel's GitHub segmentation models, Keras <a href=\"https://github.com/qubvel/segmentation_models\">here</a>, PyTorch <a href=\"https://github.com/qubvel/segmentation_models.pytorch\">here</a>. Then define your input size as <code>(None, None, 3)</code> like this</p>\n\n<pre><code>from segmentation_models import Unet\nmodel = Unet('resnet34', input_shape=(None, None, 3), classes=4, \n                    activation='sigmoid')\n</code></pre>\n\n<p>This allows us to be more creative with our training. For example, you can train a few epochs with <code>384x384</code> crops, then train a few epochs with <code>512x512</code> crops, then train with <code>512x1024</code> crops. And then after your model has been trained, you can use it to predict full images of size <code>1400x2100</code>. Pretty cool, huh?</p>\n\n<pre><code>model.fit(X_train, y_train)\n</code></pre>\n\n<p>Above <code>X_train</code> has dimensions <code>(batch_size, height, width, 3)</code>. Each time you call <code>fit()</code> or <code>predict()</code>,  the sizes <code>height, width</code> can be any values divisible by 32. Since GPU memory is limited, using random crops allows us to increase <code>batch_size</code> while decreasing <code>height, width</code>. This allows our network to see more images (different annotators) in one batch. </p>\n\n<p>Additionally instead of random crops, you can selectively choose crops if you wish to teach your network something specific. This was helpful in Steel Comp for networks to learn rare defects but may not be as helpful here. I posted a starter kernel <a href=\"https://www.kaggle.com/cdeotte/train-with-crops-cv-0-60\">here</a>. Enjoy!!</p>\n\n############################\n\n<p><strong>Important Note</strong>: that when I say \"crops\", I mean you cut random rectangles from the original image therefore your network sees a \"localized region\". This is different than \"resizing\" which rescales images and your network still sees the \"everything\". Note that the crops have the same resolution as the full image. </p>\n\n<p>For example, if the full image is <code>400x600</code> pixels and 4 inch by 6 inch. When you cut <code>100x100</code> pixel crops of size 1 inch by inch, both the full image and crops have 100 pixels per inch. If your network learns that <code>Flowers</code> fit in boxes of 50 pixels (0.5 inch) from the crops, this is still true for the full images.</p>",
      "rawMarkdown": "I would like to share a cool trick that I learned in Steel Comp. Since segmentation models use all convolutions you can change the input size whenever you want even **during** or **after** training. (More explanation [here][2]).\n \nIf you use Qubvel's GitHub segmentation models, Keras [here][3], PyTorch [here][4]. Then define your input size as `(None, None, 3)` like this\n\n    from segmentation_models import Unet\n    model = Unet('resnet34', input_shape=(None, None, 3), classes=4, \n                        activation='sigmoid')\n\nThis allows us to be more creative with our training. For example, you can train a few epochs with `384x384` crops, then train a few epochs with `512x512` crops, then train with `512x1024` crops. And then after your model has been trained, you can use it to predict full images of size `1400x2100`. Pretty cool, huh?\n\n    model.fit(X_train, y_train)\n\nAbove `X_train` has dimensions `(batch_size, height, width, 3)`. Each time you call `fit()` or `predict()`,  the sizes `height, width` can be any values divisible by 32. Since GPU memory is limited, using random crops allows us to increase `batch_size` while decreasing `height, width`. This allows our network to see more images (different annotators) in one batch. \n\nAdditionally instead of random crops, you can selectively choose crops if you wish to teach your network something specific. This was helpful in Steel Comp for networks to learn rare defects but may not be as helpful here. I posted a starter kernel [here][1]. Enjoy!!\n   \n################################## \n**Important Note**: that when I say \"crops\", I mean you cut random rectangles from the original image therefore your network sees a \"localized region\". This is different than \"resizing\" which rescales images and your network still sees the \"everything\". Note that the crops have the same resolution as the full image. \n\nFor example, if the full image is `400x600` pixels and 4 inch by 6 inch. When you cut `100x100` pixel crops of size 1 inch by inch, both the full image and crops have 100 pixels per inch. If your network learns that `Flowers` fit in boxes of 50 pixels (0.5 inch) from the crops, this is still true for the full images.\n\n[1]: https://www.kaggle.com/cdeotte/train-with-crops-cv-0-60\n[2]: https://www.kaggle.com/c/severstal-steel-defect-detection/discussion/114321\n[3]: https://github.com/qubvel/segmentation_models\n[4]: https://github.com/qubvel/segmentation_models.pytorch",
      "votes": null
    },
    {
      "id": "662392",
      "postDate": "10/31/2019 14:03:02",
      "content": "<p>cool but not with such bad labels</p>",
      "rawMarkdown": "cool but not with such bad labels",
      "votes": null
    },
    {
      "id": "662453",
      "postDate": "10/31/2019 15:21:21",
      "content": "<p>I think the bad labels would only be a problem if the crops are too small, this way the model might not be able to learn the overall cloud formation structure.</p>",
      "rawMarkdown": "I think the bad labels would only be a problem if the crops are too small, this way the model might not be able to learn the overall cloud formation structure.",
      "votes": null
    },
    {
      "id": "662464",
      "postDate": "10/31/2019 15:28:43",
      "content": "<p>With my tests I can say the same... too noisy labels and some masks are pretty much the whole image. May be better thinking is this competition as object detection.</p>",
      "rawMarkdown": "With my tests I can say the same... too noisy labels and some masks are pretty much the whole image. May be better thinking is this competition as object detection.",
      "votes": null
    },
    {
      "id": "662484",
      "postDate": "10/31/2019 15:55:46",
      "content": "<p>&gt; I think the bad labels would only be a problem if the crops are too small, this way the model might not be able to learn the overall cloud formation structure.</p>\n\n<p>Yes. For example I can use random crops of size <code>1300x2000</code> from the original <code>1400x2100</code> then it becomes a form of data augmentation (random shifting). No reason why bad labels should prevent me from doing that.</p>\n\n<p>But even training a few epochs with small crops (like 256x256) have value. For example they learn what density of clouds equate to what type. For example if the area is mostly blue then it is more likely Sugar. If the area is mostly white then it is more likely Flower. </p>",
      "rawMarkdown": "&gt; I think the bad labels would only be a problem if the crops are too small, this way the model might not be able to learn the overall cloud formation structure.\n\nYes. For example I can use random crops of size `1300x2000` from the original `1400x2100` then it becomes a form of data augmentation (random shifting). No reason why bad labels should prevent me from doing that.\n\nBut even training a few epochs with small crops (like 256x256) have value. For example they learn what density of clouds equate to what type. For example if the area is mostly blue then it is more likely Sugar. If the area is mostly white then it is more likely Flower.",
      "votes": null
    },
    {
      "id": "662521",
      "postDate": "10/31/2019 16:45:52",
      "content": "<p>I use Kera's \"zoom_range\" with the same intent as you said about \"random shifting\", maybe a lazy way to train with cropped images is to just apply zoom_range to images with 100%, instead of just creating a new dataset with cropped images.</p>",
      "rawMarkdown": "I use Kera's \"zoom_range\" with the same intent as you said about \"random shifting\", maybe a lazy way to train with cropped images is to just apply zoom_range to images with 100%, instead of just creating a new dataset with cropped images.",
      "votes": null
    },
    {
      "id": "662527",
      "postDate": "10/31/2019 16:49:43",
      "content": "<p>I don't create a new dataset. I have a Keras data generator (PyTorch data loader) that makes random crops on the fly.</p>\n\n<p>The main trade off is that I can use a large batch size with a smaller crop size. Versus a small batch size with a larger crop size. This trade off is interesting. I'm wondering whether it is somewhat mathematically equivalent.</p>",
      "rawMarkdown": "I don't create a new dataset. I have a Keras data generator (PyTorch data loader) that makes random crops on the fly.\n\nThe main trade off is that I can use a large batch size with a smaller crop size. Versus a small batch size with a larger crop size. This trade off is interesting. I'm wondering whether it is somewhat mathematically equivalent.",
      "votes": null
    },
    {
      "id": "662538",
      "postDate": "10/31/2019 17:03:17",
      "content": "<p>I'm thinking now that maybe creating a dataset of crops maybe is better, because with the generator you still have to load the images with large size and then resize when feeding the model, having the images already resized may let us use even bigger batches.</p>",
      "rawMarkdown": "I'm thinking now that maybe creating a dataset of crops maybe is better, because with the generator you still have to load the images with large size and then resize when feeding the model, having the images already resized may let us use even bigger batches.",
      "votes": null
    },
    {
      "id": "662542",
      "postDate": "10/31/2019 17:07:19",
      "content": "<p>I think we need to use small batches because the output of the data generator needs to fit in GPU memory. I don't think the input matters. The data generator itself loads images from disk and processes them one by one in CPU. (The GPU processes things batch by batch).</p>",
      "rawMarkdown": "I think we need to use small batches because the output of the data generator needs to fit in GPU memory. I don't think the input matters. The data generator itself loads images from disk and processes them one by one in CPU. (The GPU processes things batch by batch).",
      "votes": null
    },
    {
      "id": "662557",
      "postDate": "10/31/2019 17:28:25",
      "content": "<p>This makes sense, thanks.</p>",
      "rawMarkdown": "This makes sense, thanks.",
      "votes": null
    },
    {
      "id": "662715",
      "postDate": "10/31/2019 22:38:41",
      "content": "<p>I've tried this after seeing that it worked for Steel Comp. Unfortunately my results seem to be worse. I tried a 512x512 crop from the original image size and saw a decrease in CV/LB of about 0.03. The crops were probably too small relative to the labels (some labels take up the entire image).</p>\n\n<p>I'm now trying resizing to 704x1056 and then taking 704x704 crops which means I'll be feeding in 2/3 of the image.</p>\n\n<p><code>Python\ndef get_training_augmentation():\n    train_transform = [\n        albu.HorizontalFlip(p=0.5),\n        albu.VerticalFlip(p=0.5),\n        albu.ShiftScaleRotate(scale_limit=(0.1, 0.1), rotate_limit=45, p=0.5, border_mode=0),\n        albu.GridDistortion(p=0.5),\n        albu.RandomContrast(limit=0.3, p=0.5),\n        albu.Resize(704, 1056),\n        albu.RandomCrop(704, 704)\n    ]\n</code></p>",
      "rawMarkdown": "I've tried this after seeing that it worked for Steel Comp. Unfortunately my results seem to be worse. I tried a 512x512 crop from the original image size and saw a decrease in CV/LB of about 0.03. The crops were probably too small relative to the labels (some labels take up the entire image).\n\nI'm now trying resizing to 704x1056 and then taking 704x704 crops which means I'll be feeding in 2/3 of the image.\n\n```Python\ndef get_training_augmentation():\n    train_transform = [\n        albu.HorizontalFlip(p=0.5),\n        albu.VerticalFlip(p=0.5),\n        albu.ShiftScaleRotate(scale_limit=(0.1, 0.1), rotate_limit=45, p=0.5, border_mode=0),\n        albu.GridDistortion(p=0.5),\n        albu.RandomContrast(limit=0.3, p=0.5),\n        albu.Resize(704, 1056),\n        albu.RandomCrop(704, 704)\n    ]\n```",
      "votes": null
    },
    {
      "id": "662723",
      "postDate": "10/31/2019 22:56:43",
      "content": "<p>Great idea. That's what I'm trying at this very moment too. When doing crops, there are two rectangles sizes to choose. First you choose which size you will resize the original images to. Next you choose which size crop you will use.</p>\n\n<p>In my starter kernel linked to this discussion, I resize original images to <code>700x1050</code> and use crops of <code>352x512</code>. Therefore each crop was <code>1/4</code> of the original image. And the original images where downsampled by a factor of 2. Currently I'm trying other combinations offline.</p>",
      "rawMarkdown": "Great idea. That's what I'm trying at this very moment too. When doing crops, there are two rectangles sizes to choose. First you choose which size you will resize the original images to. Next you choose which size crop you will use.\n\nIn my starter kernel linked to this discussion, I resize original images to `700x1050` and use crops of `352x512`. Therefore each crop was `1/4` of the original image. And the original images where downsampled by a factor of 2. Currently I'm trying other combinations offline.",
      "votes": null
    },
    {
      "id": "662909",
      "postDate": "11/01/2019 07:18:39",
      "content": "<p>Or you can give a try AdamAccumulate to combine losses for multiple passes  . In pytorch Gradient Accumulation is little simpler . I had some improvement with batch accumulation till 64 (dataloader batchsize is 4)</p>",
      "rawMarkdown": "Or you can give a try AdamAccumulate to combine losses for multiple passes  . In pytorch Gradient Accumulation is little simpler . I had some improvement with batch accumulation till 64 (dataloader batchsize is 4)",
      "votes": null
    },
    {
      "id": "663062",
      "postDate": "11/01/2019 12:07:57",
      "content": "<p><a href=\"/phoenix9032\">@phoenix9032</a> agreed, in my experiments, gradient accumulation works nice, the loss goes down a lot more smoothly.</p>",
      "rawMarkdown": "phoenix9032 agreed, in my experiments, gradient accumulation works nice, the loss goes down a lot more smoothly.",
      "votes": null
    },
    {
      "id": "663084",
      "postDate": "11/01/2019 12:37:27",
      "content": "<p>My understanding of this competition so far is , \na) I have joined little too late \nb) This competition so far is testing the basics for me . No fancy stuff .\n  e.g. 1. Can I do a good classification of similar looking images \n         2. How good am I in hyperparameter tuning ? Learning Rate , Batch Size\n         3. Can I understand how basic image processing works .  etc.\n         4. Do I know how to do a proper CV-LB split .   etc .</p>",
      "rawMarkdown": "My understanding of this competition so far is , \na) I have joined little too late \nb) This competition so far is testing the basics for me . No fancy stuff .\n  e.g. 1. Can I do a good classification of similar looking images \n         2. How good am I in hyperparameter tuning ? Learning Rate , Batch Size\n         3. Can I understand how basic image processing works .  etc.\n         4. Do I know how to do a proper CV-LB split .   etc .",
      "votes": null
    },
    {
      "id": "664935",
      "postDate": "11/04/2019 12:29:26",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a>  do you make changes to your inference logic, like cropping 4 or 3 parts from the image and then predicting separately on the crops and give the output? \nMy current solution is not using random crop, just using a resize. But I am doing inference on <code>525, 350</code> as per andrew's kernel. \nDid you see improvement in inference on whole image vs inference on smaller image?</p>",
      "rawMarkdown": "cdeotte  do you make changes to your inference logic, like cropping 4 or 3 parts from the image and then predicting separately on the crops and give the output? \nMy current solution is not using random crop, just using a resize. But I am doing inference on `525, 350` as per andrew's kernel. \nDid you see improvement in inference on whole image vs inference on smaller image?",
      "votes": null
    },
    {
      "id": "665028",
      "postDate": "11/04/2019 14:53:25",
      "content": "<p>I'm not sure I understand your first question. Let's say that you resize all the images to <code>1024x704</code>. You can train with random crops of <code>512x352</code> and <code>batch_size=8</code>. Then inference is a single prediction with <code>batch_size=1</code> and <code>1024x704</code>. You do not need to make multiple predictions and stitch them together. Even though the network was trained with <code>512x352</code>, you can predict <code>1024x704</code> all at once.</p>\n\n<p>This let's you use larger original images but still have large batch sizes (using crops). Your network can utilize the additional information in the larger original image. However, so far I have not increased my CV nor LB using more than an original image of <code>512x352</code> same as you. </p>\n\n<p>(A second advantage is that you can use specific crops instead of random crops to control what your network sees/learns. I have not gained an advantage from this yet either).</p>",
      "rawMarkdown": "I'm not sure I understand your first question. Let's say that you resize all the images to `1024x704`. You can train with random crops of `512x352` and `batch_size=8`. Then inference is a single prediction with `batch_size=1` and `1024x704`. You do not need to make multiple predictions and stitch them together. Even though the network was trained with `512x352`, you can predict `1024x704` all at once.\n\nThis let's you use larger original images but still have large batch sizes (using crops). Your network can utilize the additional information in the larger original image. However, so far I have not increased my CV nor LB using more than an original image of `512x352` same as you. \n\n(A second advantage is that you can use specific crops instead of random crops to control what your network sees/learns. I have not gained an advantage from this yet either).",
      "votes": null
    },
    {
      "id": "670689",
      "postDate": "11/11/2019 18:30:21",
      "content": "<p>Hi <a href=\"/cdeotte\">@cdeotte</a> , when we crop all these training images, will that not affect the encoded pixles? how would that vary? what if only a part of the mask is present in the crop? will that be considered as a mask? I feel like I am missing some details here, could you help me? thanks.</p>",
      "rawMarkdown": "Hi @cdeotte , when we crop all these training images, will that not affect the encoded pixles? how would that vary? what if only a part of the mask is present in the crop? will that be considered as a mask? I feel like I am missing some details here, could you help me? thanks.",
      "votes": null
    },
    {
      "id": "670701",
      "postDate": "11/11/2019 18:49:48",
      "content": "<p>You crop the mask to match the image.</p>\n\n<p>Here is an example. (1) First resize all your train images, test images, and <strong>train masks</strong> to some size, say <code>320x480</code>. (2) Next, choose a crop size, say <code>256x256</code>. (3) Now for each image, randomly extract a <code>256x256</code> crop. Also extract a <code>256x256</code> crop from the mask in the <strong>same</strong> <code>(x,y)</code> <strong>location</strong> as image crop. (4) Feed the crop and matching mask into your network. (5) When you are done training, feed the full test <code>320x480</code> into network and you will get <code>320x480</code> mask out. (6) Lastly convert those <code>320x480</code> mask to <code>350x525</code> so that you can submit them to Kaggle.</p>",
      "rawMarkdown": "You crop the mask to match the image.\n\nHere is an example. (1) First resize all your train images, test images, and **train masks** to some size, say `320x480`. (2) Next, choose a crop size, say `256x256`. (3) Now for each image, randomly extract a `256x256` crop. Also extract a `256x256` crop from the mask in the **same** `(x,y)` **location** as image crop. (4) Feed the crop and matching mask into your network. (5) When you are done training, feed the full test `320x480` into network and you will get `320x480` mask out. (6) Lastly convert those `320x480` mask to `350x525` so that you can submit them to Kaggle.",
      "votes": null
    },
    {
      "id": "670902",
      "postDate": "11/12/2019 02:45:40",
      "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a> , will check this, and thanks for the kernel, going through it now. </p>",
      "rawMarkdown": "Thanks @cdeotte , will check this, and thanks for the kernel, going through it now.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 662392,
      "author_name": "maksimovka",
      "author_url": "",
      "post_date": "10/31/2019 14:03:02",
      "content": "<p>cool but not with such bad labels</p>",
      "votes": null,
      "replies": [
        {
          "id": 662453,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "10/31/2019 15:21:21",
          "content": "<p>I think the bad labels would only be a problem if the crops are too small, this way the model might not be able to learn the overall cloud formation structure.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 662464,
          "author_name": "igormunizims",
          "author_url": "",
          "post_date": "10/31/2019 15:28:43",
          "content": "<p>With my tests I can say the same... too noisy labels and some masks are pretty much the whole image. May be better thinking is this competition as object detection.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 662484,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "10/31/2019 15:55:46",
          "content": "<p>&gt; I think the bad labels would only be a problem if the crops are too small, this way the model might not be able to learn the overall cloud formation structure.</p>\n\n<p>Yes. For example I can use random crops of size <code>1300x2000</code> from the original <code>1400x2100</code> then it becomes a form of data augmentation (random shifting). No reason why bad labels should prevent me from doing that.</p>\n\n<p>But even training a few epochs with small crops (like 256x256) have value. For example they learn what density of clouds equate to what type. For example if the area is mostly blue then it is more likely Sugar. If the area is mostly white then it is more likely Flower. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 662521,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "10/31/2019 16:45:52",
          "content": "<p>I use Kera's \"zoom_range\" with the same intent as you said about \"random shifting\", maybe a lazy way to train with cropped images is to just apply zoom_range to images with 100%, instead of just creating a new dataset with cropped images.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 662527,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "10/31/2019 16:49:43",
          "content": "<p>I don't create a new dataset. I have a Keras data generator (PyTorch data loader) that makes random crops on the fly.</p>\n\n<p>The main trade off is that I can use a large batch size with a smaller crop size. Versus a small batch size with a larger crop size. This trade off is interesting. I'm wondering whether it is somewhat mathematically equivalent.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 662538,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "10/31/2019 17:03:17",
          "content": "<p>I'm thinking now that maybe creating a dataset of crops maybe is better, because with the generator you still have to load the images with large size and then resize when feeding the model, having the images already resized may let us use even bigger batches.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 662542,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "10/31/2019 17:07:19",
          "content": "<p>I think we need to use small batches because the output of the data generator needs to fit in GPU memory. I don't think the input matters. The data generator itself loads images from disk and processes them one by one in CPU. (The GPU processes things batch by batch).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 662557,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "10/31/2019 17:28:25",
          "content": "<p>This makes sense, thanks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 662909,
          "author_name": "phoenix9032",
          "author_url": "",
          "post_date": "11/01/2019 07:18:39",
          "content": "<p>Or you can give a try AdamAccumulate to combine losses for multiple passes  . In pytorch Gradient Accumulation is little simpler . I had some improvement with batch accumulation till 64 (dataloader batchsize is 4)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 663062,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "11/01/2019 12:07:57",
          "content": "<p><a href=\"/phoenix9032\">@phoenix9032</a> agreed, in my experiments, gradient accumulation works nice, the loss goes down a lot more smoothly.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 663084,
          "author_name": "phoenix9032",
          "author_url": "",
          "post_date": "11/01/2019 12:37:27",
          "content": "<p>My understanding of this competition so far is , \na) I have joined little too late \nb) This competition so far is testing the basics for me . No fancy stuff .\n  e.g. 1. Can I do a good classification of similar looking images \n         2. How good am I in hyperparameter tuning ? Learning Rate , Batch Size\n         3. Can I understand how basic image processing works .  etc.\n         4. Do I know how to do a proper CV-LB split .   etc .</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 662715,
      "author_name": "joshvarty",
      "author_url": "",
      "post_date": "10/31/2019 22:38:41",
      "content": "<p>I've tried this after seeing that it worked for Steel Comp. Unfortunately my results seem to be worse. I tried a 512x512 crop from the original image size and saw a decrease in CV/LB of about 0.03. The crops were probably too small relative to the labels (some labels take up the entire image).</p>\n\n<p>I'm now trying resizing to 704x1056 and then taking 704x704 crops which means I'll be feeding in 2/3 of the image.</p>\n\n<p><code>Python\ndef get_training_augmentation():\n    train_transform = [\n        albu.HorizontalFlip(p=0.5),\n        albu.VerticalFlip(p=0.5),\n        albu.ShiftScaleRotate(scale_limit=(0.1, 0.1), rotate_limit=45, p=0.5, border_mode=0),\n        albu.GridDistortion(p=0.5),\n        albu.RandomContrast(limit=0.3, p=0.5),\n        albu.Resize(704, 1056),\n        albu.RandomCrop(704, 704)\n    ]\n</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 662723,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "10/31/2019 22:56:43",
          "content": "<p>Great idea. That's what I'm trying at this very moment too. When doing crops, there are two rectangles sizes to choose. First you choose which size you will resize the original images to. Next you choose which size crop you will use.</p>\n\n<p>In my starter kernel linked to this discussion, I resize original images to <code>700x1050</code> and use crops of <code>352x512</code>. Therefore each crop was <code>1/4</code> of the original image. And the original images where downsampled by a factor of 2. Currently I'm trying other combinations offline.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 664935,
          "author_name": "axel81",
          "author_url": "",
          "post_date": "11/04/2019 12:29:26",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a>  do you make changes to your inference logic, like cropping 4 or 3 parts from the image and then predicting separately on the crops and give the output? \nMy current solution is not using random crop, just using a resize. But I am doing inference on <code>525, 350</code> as per andrew's kernel. \nDid you see improvement in inference on whole image vs inference on smaller image?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 665028,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "11/04/2019 14:53:25",
          "content": "<p>I'm not sure I understand your first question. Let's say that you resize all the images to <code>1024x704</code>. You can train with random crops of <code>512x352</code> and <code>batch_size=8</code>. Then inference is a single prediction with <code>batch_size=1</code> and <code>1024x704</code>. You do not need to make multiple predictions and stitch them together. Even though the network was trained with <code>512x352</code>, you can predict <code>1024x704</code> all at once.</p>\n\n<p>This let's you use larger original images but still have large batch sizes (using crops). Your network can utilize the additional information in the larger original image. However, so far I have not increased my CV nor LB using more than an original image of <code>512x352</code> same as you. </p>\n\n<p>(A second advantage is that you can use specific crops instead of random crops to control what your network sees/learns. I have not gained an advantage from this yet either).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 670689,
          "author_name": "harshahat02",
          "author_url": "",
          "post_date": "11/11/2019 18:30:21",
          "content": "<p>Hi <a href=\"/cdeotte\">@cdeotte</a> , when we crop all these training images, will that not affect the encoded pixles? how would that vary? what if only a part of the mask is present in the crop? will that be considered as a mask? I feel like I am missing some details here, could you help me? thanks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 670701,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "11/11/2019 18:49:48",
          "content": "<p>You crop the mask to match the image.</p>\n\n<p>Here is an example. (1) First resize all your train images, test images, and <strong>train masks</strong> to some size, say <code>320x480</code>. (2) Next, choose a crop size, say <code>256x256</code>. (3) Now for each image, randomly extract a <code>256x256</code> crop. Also extract a <code>256x256</code> crop from the mask in the <strong>same</strong> <code>(x,y)</code> <strong>location</strong> as image crop. (4) Feed the crop and matching mask into your network. (5) When you are done training, feed the full test <code>320x480</code> into network and you will get <code>320x480</code> mask out. (6) Lastly convert those <code>320x480</code> mask to <code>350x525</code> so that you can submit them to Kaggle.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 670902,
          "author_name": "harshahat02",
          "author_url": "",
          "post_date": "11/12/2019 02:45:40",
          "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a> , will check this, and thanks for the kernel, going through it now. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "662320": "I would like to share a cool trick that I learned in Steel Comp. Since segmentation models use all convolutions you can change the input size whenever you want even **during** or **after** training. (More explanation [here][2]).\n \nIf you use Qubvel's GitHub segmentation models, Keras [here][3], PyTorch [here][4]. Then define your input size as `(None, None, 3)` like this\n\n    from segmentation_models import Unet\n    model = Unet('resnet34', input_shape=(None, None, 3), classes=4, \n                        activation='sigmoid')\n\nThis allows us to be more creative with our training. For example, you can train a few epochs with `384x384` crops, then train a few epochs with `512x512` crops, then train with `512x1024` crops. And then after your model has been trained, you can use it to predict full images of size `1400x2100`. Pretty cool, huh?\n\n    model.fit(X_train, y_train)\n\nAbove `X_train` has dimensions `(batch_size, height, width, 3)`. Each time you call `fit()` or `predict()`,  the sizes `height, width` can be any values divisible by 32. Since GPU memory is limited, using random crops allows us to increase `batch_size` while decreasing `height, width`. This allows our network to see more images (different annotators) in one batch. \n\nAdditionally instead of random crops, you can selectively choose crops if you wish to teach your network something specific. This was helpful in Steel Comp for networks to learn rare defects but may not be as helpful here. I posted a starter kernel [here][1]. Enjoy!!\n   \n################################## \n**Important Note**: that when I say \"crops\", I mean you cut random rectangles from the original image therefore your network sees a \"localized region\". This is different than \"resizing\" which rescales images and your network still sees the \"everything\". Note that the crops have the same resolution as the full image. \n\nFor example, if the full image is `400x600` pixels and 4 inch by 6 inch. When you cut `100x100` pixel crops of size 1 inch by inch, both the full image and crops have 100 pixels per inch. If your network learns that `Flowers` fit in boxes of 50 pixels (0.5 inch) from the crops, this is still true for the full images.\n\n[1]: https://www.kaggle.com/cdeotte/train-with-crops-cv-0-60\n[2]: https://www.kaggle.com/c/severstal-steel-defect-detection/discussion/114321\n[3]: https://github.com/qubvel/segmentation_models\n[4]: https://github.com/qubvel/segmentation_models.pytorch",
    "662392": "cool but not with such bad labels",
    "662453": "I think the bad labels would only be a problem if the crops are too small, this way the model might not be able to learn the overall cloud formation structure.",
    "662464": "With my tests I can say the same... too noisy labels and some masks are pretty much the whole image. May be better thinking is this competition as object detection.",
    "662484": "&gt; I think the bad labels would only be a problem if the crops are too small, this way the model might not be able to learn the overall cloud formation structure.\n\nYes. For example I can use random crops of size `1300x2000` from the original `1400x2100` then it becomes a form of data augmentation (random shifting). No reason why bad labels should prevent me from doing that.\n\nBut even training a few epochs with small crops (like 256x256) have value. For example they learn what density of clouds equate to what type. For example if the area is mostly blue then it is more likely Sugar. If the area is mostly white then it is more likely Flower.",
    "662521": "I use Kera's \"zoom_range\" with the same intent as you said about \"random shifting\", maybe a lazy way to train with cropped images is to just apply zoom_range to images with 100%, instead of just creating a new dataset with cropped images.",
    "662527": "I don't create a new dataset. I have a Keras data generator (PyTorch data loader) that makes random crops on the fly.\n\nThe main trade off is that I can use a large batch size with a smaller crop size. Versus a small batch size with a larger crop size. This trade off is interesting. I'm wondering whether it is somewhat mathematically equivalent.",
    "662538": "I'm thinking now that maybe creating a dataset of crops maybe is better, because with the generator you still have to load the images with large size and then resize when feeding the model, having the images already resized may let us use even bigger batches.",
    "662542": "I think we need to use small batches because the output of the data generator needs to fit in GPU memory. I don't think the input matters. The data generator itself loads images from disk and processes them one by one in CPU. (The GPU processes things batch by batch).",
    "662557": "This makes sense, thanks.",
    "662715": "I've tried this after seeing that it worked for Steel Comp. Unfortunately my results seem to be worse. I tried a 512x512 crop from the original image size and saw a decrease in CV/LB of about 0.03. The crops were probably too small relative to the labels (some labels take up the entire image).\n\nI'm now trying resizing to 704x1056 and then taking 704x704 crops which means I'll be feeding in 2/3 of the image.\n\n```Python\ndef get_training_augmentation():\n    train_transform = [\n        albu.HorizontalFlip(p=0.5),\n        albu.VerticalFlip(p=0.5),\n        albu.ShiftScaleRotate(scale_limit=(0.1, 0.1), rotate_limit=45, p=0.5, border_mode=0),\n        albu.GridDistortion(p=0.5),\n        albu.RandomContrast(limit=0.3, p=0.5),\n        albu.Resize(704, 1056),\n        albu.RandomCrop(704, 704)\n    ]\n```",
    "662723": "Great idea. That's what I'm trying at this very moment too. When doing crops, there are two rectangles sizes to choose. First you choose which size you will resize the original images to. Next you choose which size crop you will use.\n\nIn my starter kernel linked to this discussion, I resize original images to `700x1050` and use crops of `352x512`. Therefore each crop was `1/4` of the original image. And the original images where downsampled by a factor of 2. Currently I'm trying other combinations offline.",
    "662909": "Or you can give a try AdamAccumulate to combine losses for multiple passes  . In pytorch Gradient Accumulation is little simpler . I had some improvement with batch accumulation till 64 (dataloader batchsize is 4)",
    "663062": "phoenix9032 agreed, in my experiments, gradient accumulation works nice, the loss goes down a lot more smoothly.",
    "663084": "My understanding of this competition so far is , \na) I have joined little too late \nb) This competition so far is testing the basics for me . No fancy stuff .\n  e.g. 1. Can I do a good classification of similar looking images \n         2. How good am I in hyperparameter tuning ? Learning Rate , Batch Size\n         3. Can I understand how basic image processing works .  etc.\n         4. Do I know how to do a proper CV-LB split .   etc .",
    "664935": "cdeotte  do you make changes to your inference logic, like cropping 4 or 3 parts from the image and then predicting separately on the crops and give the output? \nMy current solution is not using random crop, just using a resize. But I am doing inference on `525, 350` as per andrew's kernel. \nDid you see improvement in inference on whole image vs inference on smaller image?",
    "665028": "I'm not sure I understand your first question. Let's say that you resize all the images to `1024x704`. You can train with random crops of `512x352` and `batch_size=8`. Then inference is a single prediction with `batch_size=1` and `1024x704`. You do not need to make multiple predictions and stitch them together. Even though the network was trained with `512x352`, you can predict `1024x704` all at once.\n\nThis let's you use larger original images but still have large batch sizes (using crops). Your network can utilize the additional information in the larger original image. However, so far I have not increased my CV nor LB using more than an original image of `512x352` same as you. \n\n(A second advantage is that you can use specific crops instead of random crops to control what your network sees/learns. I have not gained an advantage from this yet either).",
    "670689": "Hi @cdeotte , when we crop all these training images, will that not affect the encoded pixles? how would that vary? what if only a part of the mask is present in the crop? will that be considered as a mask? I feel like I am missing some details here, could you help me? thanks.",
    "670701": "You crop the mask to match the image.\n\nHere is an example. (1) First resize all your train images, test images, and **train masks** to some size, say `320x480`. (2) Next, choose a crop size, say `256x256`. (3) Now for each image, randomly extract a `256x256` crop. Also extract a `256x256` crop from the mask in the **same** `(x,y)` **location** as image crop. (4) Feed the crop and matching mask into your network. (5) When you are done training, feed the full test `320x480` into network and you will get `320x480` mask out. (6) Lastly convert those `320x480` mask to `350x525` so that you can submit them to Kaggle.",
    "670902": "Thanks @cdeotte , will check this, and thanks for the kernel, going through it now."
  },
  "source": "meta"
}