{
  "id": 103733,
  "title": "Properly use Data Augmentation",
  "url": "/competitions/recursion-cellular-image-classification/discussion/103733",
  "author_name": "",
  "post_date": "2019-08-11T10:56:14.465778100Z",
  "votes": 4,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Up until now I did not use any form of data augmentation. Today I tried adding horizontal flipping. I'm using 6-channel images, 512x512, ResNet18 (starting from a pre-trained model), Adam with 0.0001 LR, Cosine Annealing: all of this in PyTorch. I also standardize each image by channel (with a MinMax Scaler).</p>\n\n<p>Without any augmentation, the mean validation accuracy (0.1% dataset) is around 0.009. If I had horizontal flipping, it drops to 0.001 (so basically random) and it does not increase (the loss is actually increasing a little). Has anybody else seen this?</p>\n\n<p>For completeness, here's my code to do horizontal flipping with 6-channel images:</p>\n\n<p>```\nclass HorizontalFlipChannels(object):</p>\n\n<pre><code>def __init__(self, p=0.5):\n    self.p = p\n\ndef transform(self, tensor):\n    transformed_channels = []\n\n    if random.random() &amp;lt; self.p:\n\n        for idx, channel in enumerate(tensor):\n            channel = transforms.ToPILImage()(channel)\n            channel = TF.hflip(channel)\n            channel = transforms.ToTensor()(channel)\n\n            transformed_channels.append(channel)\n\n        tensor = torch.cat(transformed_channels)\n\n    return tensor\n</code></pre>\n\n<p>```</p>\n\n<p>and for scaling:</p>\n\n<p>```\nclass MinMaxScaler(object):\n    \"\"\"\n    Transforms each channel to the range [0, 1].\n    \"\"\"</p>\n\n<pre><code>def __call__(self, tensor):\n\n    for ch in tensor:\n        scale = 1.0 / (ch.max() - ch.min())\n        ch.mul_(scale).sub_(ch.min().mul_(scale))\n\n    return tensor\n</code></pre>\n\n<p>```</p>",
  "messages": [
    {
      "id": "596819",
      "postDate": "08/11/2019 10:56:14",
      "content": "<p>Up until now I did not use any form of data augmentation. Today I tried adding horizontal flipping. I'm using 6-channel images, 512x512, ResNet18 (starting from a pre-trained model), Adam with 0.0001 LR, Cosine Annealing: all of this in PyTorch. I also standardize each image by channel (with a MinMax Scaler).</p>\n\n<p>Without any augmentation, the mean validation accuracy (0.1% dataset) is around 0.009. If I had horizontal flipping, it drops to 0.001 (so basically random) and it does not increase (the loss is actually increasing a little). Has anybody else seen this?</p>\n\n<p>For completeness, here's my code to do horizontal flipping with 6-channel images:</p>\n\n<p>```\nclass HorizontalFlipChannels(object):</p>\n\n<pre><code>def __init__(self, p=0.5):\n    self.p = p\n\ndef transform(self, tensor):\n    transformed_channels = []\n\n    if random.random() &amp;lt; self.p:\n\n        for idx, channel in enumerate(tensor):\n            channel = transforms.ToPILImage()(channel)\n            channel = TF.hflip(channel)\n            channel = transforms.ToTensor()(channel)\n\n            transformed_channels.append(channel)\n\n        tensor = torch.cat(transformed_channels)\n\n    return tensor\n</code></pre>\n\n<p>```</p>\n\n<p>and for scaling:</p>\n\n<p>```\nclass MinMaxScaler(object):\n    \"\"\"\n    Transforms each channel to the range [0, 1].\n    \"\"\"</p>\n\n<pre><code>def __call__(self, tensor):\n\n    for ch in tensor:\n        scale = 1.0 / (ch.max() - ch.min())\n        ch.mul_(scale).sub_(ch.min().mul_(scale))\n\n    return tensor\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "Up until now I did not use any form of data augmentation. Today I tried adding horizontal flipping. I'm using 6-channel images, 512x512, ResNet18 (starting from a pre-trained model), Adam with 0.0001 LR, Cosine Annealing: all of this in PyTorch. I also standardize each image by channel (with a MinMax Scaler).\n\nWithout any augmentation, the mean validation accuracy (0.1% dataset) is around 0.009. If I had horizontal flipping, it drops to 0.001 (so basically random) and it does not increase (the loss is actually increasing a little). Has anybody else seen this?\n\nFor completeness, here's my code to do horizontal flipping with 6-channel images:\n\n```\nclass HorizontalFlipChannels(object):\n    \n    def __init__(self, p=0.5):\n        self.p = p\n    \n    def transform(self, tensor):\n        transformed_channels = []\n        \n        if random.random() &lt; self.p:\n            \n            for idx, channel in enumerate(tensor):\n                channel = transforms.ToPILImage()(channel)\n                channel = TF.hflip(channel)\n                channel = transforms.ToTensor()(channel)\n                \n                transformed_channels.append(channel)\n            \n            tensor = torch.cat(transformed_channels)\n\n        return tensor\n```\n\nand for scaling:\n\n```\nclass MinMaxScaler(object):\n    \"\"\"\n    Transforms each channel to the range [0, 1].\n    \"\"\"\n    \n    def __call__(self, tensor):\n        \n        for ch in tensor:\n            scale = 1.0 / (ch.max() - ch.min())\n            ch.mul_(scale).sub_(ch.min().mul_(scale))\n        \n        return tensor\n```",
      "votes": null
    },
    {
      "id": "596915",
      "postDate": "08/11/2019 14:00:50",
      "content": "<p>It's nice that you're trying to write everything yourself (it's a good way to learn). However, it's very easy to make small mistakes so I'd recommend you using an augmentation library like albumentations, imgaug or even just the torch.vision package. There's no much point on reinventing the wheel, and you can put your time on experimenting other things.\nAnother tip, would be to have a good look on all top scoring public kernels and discussions here and see what people are doing. Generally they will tell you things that didn't work and what did! That can save you a lot of time!</p>\n\n<p>Having said that, if you get everything right into the network, it should learn quickly to get at least LB:0.1 \nand 0.2 and 0.3 shouldn't be too hard, even without any augmentation at all.</p>",
      "rawMarkdown": "It's nice that you're trying to write everything yourself (it's a good way to learn). However, it's very easy to make small mistakes so I'd recommend you using an augmentation library like albumentations, imgaug or even just the torch.vision package. There's no much point on reinventing the wheel, and you can put your time on experimenting other things.\nAnother tip, would be to have a good look on all top scoring public kernels and discussions here and see what people are doing. Generally they will tell you things that didn't work and what did! That can save you a lot of time!\n\nHaving said that, if you get everything right into the network, it should learn quickly to get at least LB:0.1 \nand 0.2 and 0.3 shouldn't be too hard, even without any augmentation at all.",
      "votes": null
    },
    {
      "id": "596925",
      "postDate": "08/11/2019 14:09:51",
      "content": "<p>Thanks <a href=\"/hmendonca\">@hmendonca</a> for the reply.\nThe reason I'm writing everything myself is that if I use what <code>torchvision</code> provides but with 6-channel images, it does not work (the output is not a 6-channel image anymore). So that's why I'm using per-channel transformations.</p>\n\n<p>I actually had many looks at the public kernels! And I really do not know what I'm doing wrong.</p>",
      "rawMarkdown": "Thanks @hmendonca for the reply.\nThe reason I'm writing everything myself is that if I use what `torchvision` provides but with 6-channel images, it does not work (the output is not a 6-channel image anymore). So that's why I'm using per-channel transformations.\n\nI actually had many looks at the public kernels! And I really do not know what I'm doing wrong.",
      "votes": null
    },
    {
      "id": "596952",
      "postDate": "08/11/2019 15:13:56",
      "content": "<p>Lorenzo, I'm using 6-channel inputs and I augment them with albumentation no problem :)</p>",
      "rawMarkdown": "Lorenzo, I'm using 6-channel inputs and I augment them with albumentation no problem :)",
      "votes": null
    },
    {
      "id": "596954",
      "postDate": "08/11/2019 15:17:22",
      "content": "<p>I just today found out about these libraries. They seem very handy! Thanks.</p>",
      "rawMarkdown": "I just today found out about these libraries. They seem very handy! Thanks.",
      "votes": null
    },
    {
      "id": "602119",
      "postDate": "08/18/2019 15:54:14",
      "content": "<p>Definitely albumentation is great, <a href=\"https://github.com/albu/albumentations\">https://github.com/albu/albumentations</a> ... you could also consider working with RGB first (3 channels) and use PyTorch transforms for a faster iteration loop. </p>\n\n<p>Training on six channels + two sites with 512px images is quite time consuming.</p>",
      "rawMarkdown": "Definitely albumentation is great, https://github.com/albu/albumentations ... you could also consider working with RGB first (3 channels) and use PyTorch transforms for a faster iteration loop. \n\nTraining on six channels + two sites with 512px images is quite time consuming.",
      "votes": null
    },
    {
      "id": "602242",
      "postDate": "08/18/2019 20:10:42",
      "content": "<p>Also, you can apply the bag of tricks as in <a href=\"https://arxiv.org/abs/1812.01187\">https://arxiv.org/abs/1812.01187</a> . I implemented it as-is wo/ PCA noise and got improvements on my CV and LB scores. If you use 512px images consider cropping 448px crops, if you use 256px images consider cropping 224px crops.</p>",
      "rawMarkdown": "Also, you can apply the bag of tricks as in https://arxiv.org/abs/1812.01187 . I implemented it as-is wo/ PCA noise and got improvements on my CV and LB scores. If you use 512px images consider cropping 448px crops, if you use 256px images consider cropping 224px crops.",
      "votes": null
    },
    {
      "id": "602380",
      "postDate": "08/19/2019 02:53:24",
      "content": "<p><a href=\"/giuliasavorgnan\">@giuliasavorgnan</a> <a href=\"/michelml\">@michelml</a> Thanks both for the answers. I'll resize the images since I need to experiment faster. I tried using Albumentations, but the training phase is now very slow. Have you noticed something similar when using their transformations? Thanks.</p>",
      "rawMarkdown": "giuliasavorgnan @michelml Thanks both for the answers. I'll resize the images since I need to experiment faster. I tried using Albumentations, but the training phase is now very slow. Have you noticed something similar when using their transformations? Thanks.",
      "votes": null
    },
    {
      "id": "604768",
      "postDate": "08/21/2019 19:00:18",
      "content": "<p>I finally don't use albumentation, I use variations of pytorch's transforms to be able to work with 6-channels properly. Sorry to have mislead you on that one.</p>",
      "rawMarkdown": "I finally don't use albumentation, I use variations of pytorch's transforms to be able to work with 6-channels properly. Sorry to have mislead you on that one.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 596915,
      "author_name": "hmendonca",
      "author_url": "",
      "post_date": "08/11/2019 14:00:50",
      "content": "<p>It's nice that you're trying to write everything yourself (it's a good way to learn). However, it's very easy to make small mistakes so I'd recommend you using an augmentation library like albumentations, imgaug or even just the torch.vision package. There's no much point on reinventing the wheel, and you can put your time on experimenting other things.\nAnother tip, would be to have a good look on all top scoring public kernels and discussions here and see what people are doing. Generally they will tell you things that didn't work and what did! That can save you a lot of time!</p>\n\n<p>Having said that, if you get everything right into the network, it should learn quickly to get at least LB:0.1 \nand 0.2 and 0.3 shouldn't be too hard, even without any augmentation at all.</p>",
      "votes": null,
      "replies": [
        {
          "id": 596925,
          "author_name": "lorenzofabbri92",
          "author_url": "",
          "post_date": "08/11/2019 14:09:51",
          "content": "<p>Thanks <a href=\"/hmendonca\">@hmendonca</a> for the reply.\nThe reason I'm writing everything myself is that if I use what <code>torchvision</code> provides but with 6-channel images, it does not work (the output is not a 6-channel image anymore). So that's why I'm using per-channel transformations.</p>\n\n<p>I actually had many looks at the public kernels! And I really do not know what I'm doing wrong.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 596952,
          "author_name": "giuliasavorgnan",
          "author_url": "",
          "post_date": "08/11/2019 15:13:56",
          "content": "<p>Lorenzo, I'm using 6-channel inputs and I augment them with albumentation no problem :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 596954,
          "author_name": "lorenzofabbri92",
          "author_url": "",
          "post_date": "08/11/2019 15:17:22",
          "content": "<p>I just today found out about these libraries. They seem very handy! Thanks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 602119,
          "author_name": "michelml",
          "author_url": "",
          "post_date": "08/18/2019 15:54:14",
          "content": "<p>Definitely albumentation is great, <a href=\"https://github.com/albu/albumentations\">https://github.com/albu/albumentations</a> ... you could also consider working with RGB first (3 channels) and use PyTorch transforms for a faster iteration loop. </p>\n\n<p>Training on six channels + two sites with 512px images is quite time consuming.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 602242,
          "author_name": "michelml",
          "author_url": "",
          "post_date": "08/18/2019 20:10:42",
          "content": "<p>Also, you can apply the bag of tricks as in <a href=\"https://arxiv.org/abs/1812.01187\">https://arxiv.org/abs/1812.01187</a> . I implemented it as-is wo/ PCA noise and got improvements on my CV and LB scores. If you use 512px images consider cropping 448px crops, if you use 256px images consider cropping 224px crops.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 602380,
          "author_name": "lorenzofabbri92",
          "author_url": "",
          "post_date": "08/19/2019 02:53:24",
          "content": "<p><a href=\"/giuliasavorgnan\">@giuliasavorgnan</a> <a href=\"/michelml\">@michelml</a> Thanks both for the answers. I'll resize the images since I need to experiment faster. I tried using Albumentations, but the training phase is now very slow. Have you noticed something similar when using their transformations? Thanks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 604768,
          "author_name": "michelml",
          "author_url": "",
          "post_date": "08/21/2019 19:00:18",
          "content": "<p>I finally don't use albumentation, I use variations of pytorch's transforms to be able to work with 6-channels properly. Sorry to have mislead you on that one.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "596819": "Up until now I did not use any form of data augmentation. Today I tried adding horizontal flipping. I'm using 6-channel images, 512x512, ResNet18 (starting from a pre-trained model), Adam with 0.0001 LR, Cosine Annealing: all of this in PyTorch. I also standardize each image by channel (with a MinMax Scaler).\n\nWithout any augmentation, the mean validation accuracy (0.1% dataset) is around 0.009. If I had horizontal flipping, it drops to 0.001 (so basically random) and it does not increase (the loss is actually increasing a little). Has anybody else seen this?\n\nFor completeness, here's my code to do horizontal flipping with 6-channel images:\n\n```\nclass HorizontalFlipChannels(object):\n    \n    def __init__(self, p=0.5):\n        self.p = p\n    \n    def transform(self, tensor):\n        transformed_channels = []\n        \n        if random.random() &lt; self.p:\n            \n            for idx, channel in enumerate(tensor):\n                channel = transforms.ToPILImage()(channel)\n                channel = TF.hflip(channel)\n                channel = transforms.ToTensor()(channel)\n                \n                transformed_channels.append(channel)\n            \n            tensor = torch.cat(transformed_channels)\n\n        return tensor\n```\n\nand for scaling:\n\n```\nclass MinMaxScaler(object):\n    \"\"\"\n    Transforms each channel to the range [0, 1].\n    \"\"\"\n    \n    def __call__(self, tensor):\n        \n        for ch in tensor:\n            scale = 1.0 / (ch.max() - ch.min())\n            ch.mul_(scale).sub_(ch.min().mul_(scale))\n        \n        return tensor\n```",
    "596915": "It's nice that you're trying to write everything yourself (it's a good way to learn). However, it's very easy to make small mistakes so I'd recommend you using an augmentation library like albumentations, imgaug or even just the torch.vision package. There's no much point on reinventing the wheel, and you can put your time on experimenting other things.\nAnother tip, would be to have a good look on all top scoring public kernels and discussions here and see what people are doing. Generally they will tell you things that didn't work and what did! That can save you a lot of time!\n\nHaving said that, if you get everything right into the network, it should learn quickly to get at least LB:0.1 \nand 0.2 and 0.3 shouldn't be too hard, even without any augmentation at all.",
    "596925": "Thanks @hmendonca for the reply.\nThe reason I'm writing everything myself is that if I use what `torchvision` provides but with 6-channel images, it does not work (the output is not a 6-channel image anymore). So that's why I'm using per-channel transformations.\n\nI actually had many looks at the public kernels! And I really do not know what I'm doing wrong.",
    "596952": "Lorenzo, I'm using 6-channel inputs and I augment them with albumentation no problem :)",
    "596954": "I just today found out about these libraries. They seem very handy! Thanks.",
    "602119": "Definitely albumentation is great, https://github.com/albu/albumentations ... you could also consider working with RGB first (3 channels) and use PyTorch transforms for a faster iteration loop. \n\nTraining on six channels + two sites with 512px images is quite time consuming.",
    "602242": "Also, you can apply the bag of tricks as in https://arxiv.org/abs/1812.01187 . I implemented it as-is wo/ PCA noise and got improvements on my CV and LB scores. If you use 512px images consider cropping 448px crops, if you use 256px images consider cropping 224px crops.",
    "602380": "giuliasavorgnan @michelml Thanks both for the answers. I'll resize the images since I need to experiment faster. I tried using Albumentations, but the training phase is now very slow. Have you noticed something similar when using their transformations? Thanks.",
    "604768": "I finally don't use albumentation, I use variations of pytorch's transforms to be able to work with 6-channels properly. Sorry to have mislead you on that one."
  },
  "source": "meta"
}