{
  "id": 160495,
  "title": "Different Data Augmentation Techniques",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/160495",
  "author_name": "Khizar Hussain",
  "post_date": "2020-06-21T12:55:57.971000",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi, I am a newbie to CNNs and Kaggle so I might be asking pretty obvious questions...\nThere are 584 malignant images in train dataset, whereas 32,524 for benign images.\nThis means if I augment my data I have to make 56 copies for each malignant image by augmentation.\nWhat techniques are you guys using for data augmentation ? I used ImageDataGenerator from keras.preprocessing.Image but had terrible results for malignant images.\nAs you can see I completely overfitted my model by augmenting each malignant image 58 times :P\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2534874%2F379e86f8fda1b7fe7c55099401a86f3c%2F14F11CE4-695E-4158-9644-EDB42B189CE1.png.jpg?generation=1592744036142825&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2534874%2Fa792614f23807c66f385507e1fe6eec0%2FCF612092-3D7F-4AC7-B67C-BA117FBB4E31.png.jpg?generation=1592744037379738&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 895552,
      "postDate": "2020-06-21T12:55:57.970Z",
      "content": "<p>Hi, I am a newbie to CNNs and Kaggle so I might be asking pretty obvious questions...\nThere are 584 malignant images in train dataset, whereas 32,524 for benign images.\nThis means if I augment my data I have to make 56 copies for each malignant image by augmentation.\nWhat techniques are you guys using for data augmentation ? I used ImageDataGenerator from keras.preprocessing.Image but had terrible results for malignant images.\nAs you can see I completely overfitted my model by augmenting each malignant image 58 times :P\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2534874%2F379e86f8fda1b7fe7c55099401a86f3c%2F14F11CE4-695E-4158-9644-EDB42B189CE1.png.jpg?generation=1592744036142825&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2534874%2Fa792614f23807c66f385507e1fe6eec0%2FCF612092-3D7F-4AC7-B67C-BA117FBB4E31.png.jpg?generation=1592744037379738&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hi, I am a newbie to CNNs and Kaggle so I might be asking pretty obvious questions...\nThere are 584 malignant images in train dataset, whereas 32,524 for benign images.\nThis means if I augment my data I have to make 56 copies for each malignant image by augmentation.\nWhat techniques are you guys using for data augmentation ? I used ImageDataGenerator from keras.preprocessing.Image but had terrible results for malignant images.\nAs you can see I completely overfitted my model by augmenting each malignant image 58 times :P\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2534874%2F379e86f8fda1b7fe7c55099401a86f3c%2F14F11CE4-695E-4158-9644-EDB42B189CE1.png.jpg?generation=1592744036142825&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2534874%2Fa792614f23807c66f385507e1fe6eec0%2FCF612092-3D7F-4AC7-B67C-BA117FBB4E31.png.jpg?generation=1592744037379738&amp;alt=media)\n",
      "votes": 2
    },
    {
      "id": 895583,
      "postDate": "2020-06-21T13:24:06.650Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 895604,
          "postDate": "2020-06-21T13:42:46.600Z",
          "content": "<p>Haha okay, thanks for responding\nActually I got really bad accuracy for malignant test images using this :/</p>",
          "rawMarkdown": "Haha okay, thanks for responding\nActually I got really bad accuracy for malignant test images using this :/"
        },
        {
          "id": 895721,
          "postDate": "2020-06-21T15:03:49.560Z",
          "content": "<p>Always augment both classes, or the model will only learn to recognize the augmentations (which corresponds to only the positive class in your training dataset).</p>",
          "rawMarkdown": "Always augment both classes, or the model will only learn to recognize the augmentations (which corresponds to only the positive class in your training dataset).",
          "votes": 3
        },
        {
          "id": 895733,
          "postDate": "2020-06-21T15:11:28.377Z",
          "content": "<p>If I do that won't it make my model to learn the augmented features as well like the pixels stretching in this case due to shifting...Then I will need to augment the test images as well which is a bad idea.\nOr should I leave the horizontal and vertical shift and only focus on zoom, flip and intensity values to generate the augmentation ?</p>",
          "rawMarkdown": "If I do that won't it make my model to learn the augmented features as well like the pixels stretching in this case due to shifting...Then I will need to augment the test images as well which is a bad idea.\nOr should I leave the horizontal and vertical shift and only focus on zoom, flip and intensity values to generate the augmentation ?"
        },
        {
          "id": 895754,
          "postDate": "2020-06-21T15:28:45.007Z",
          "content": "<p>Your CNN will learn to ignore those stretched pixels as it will not contain any information whether it is a positive/negative image. If you only have those stretched pixels on your positive images, then it contains a lot of information ;). So you can definitely rotate IMO.</p>\n\n<p>Augmenting your test images is not a bad idea btw, it is often done as a way to lower the variance of your predictions. You can generate predictions for different augmented versions of your test images and take the mean. This is called <strong>test-time-augmentation</strong>.</p>",
          "rawMarkdown": "Your CNN will learn to ignore those stretched pixels as it will not contain any information whether it is a positive/negative image. If you only have those stretched pixels on your positive images, then it contains a lot of information ;). So you can definitely rotate IMO.\n\nAugmenting your test images is not a bad idea btw, it is often done as a way to lower the variance of your predictions. You can generate predictions for different augmented versions of your test images and take the mean. This is called **test-time-augmentation**.",
          "votes": 1
        },
        {
          "id": 895824,
          "postDate": "2020-06-21T16:22:40.960Z",
          "content": "<p>Woah I didnt know about test-time-augmentation.\nThank you so much man :D I will try again by augmenting the entire dataset including the benign images. </p>",
          "rawMarkdown": "Woah I didnt know about test-time-augmentation.\nThank you so much man :D I will try again by augmenting the entire dataset including the benign images. "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 895583,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-21T13:24:06.650000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 895604,
          "author_name": "Khizar Hussain",
          "author_url": "",
          "post_date": "2020-06-21T13:42:46.600000",
          "content": "<p>Haha okay, thanks for responding\nActually I got really bad accuracy for malignant test images using this :/</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 895721,
          "author_name": "Gilles Vandewiele",
          "author_url": "",
          "post_date": "2020-06-21T15:03:49.560000",
          "content": "<p>Always augment both classes, or the model will only learn to recognize the augmentations (which corresponds to only the positive class in your training dataset).</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 895733,
          "author_name": "Khizar Hussain",
          "author_url": "",
          "post_date": "2020-06-21T15:11:28.377000",
          "content": "<p>If I do that won't it make my model to learn the augmented features as well like the pixels stretching in this case due to shifting...Then I will need to augment the test images as well which is a bad idea.\nOr should I leave the horizontal and vertical shift and only focus on zoom, flip and intensity values to generate the augmentation ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 895754,
          "author_name": "Gilles Vandewiele",
          "author_url": "",
          "post_date": "2020-06-21T15:28:45.007000",
          "content": "<p>Your CNN will learn to ignore those stretched pixels as it will not contain any information whether it is a positive/negative image. If you only have those stretched pixels on your positive images, then it contains a lot of information ;). So you can definitely rotate IMO.</p>\n\n<p>Augmenting your test images is not a bad idea btw, it is often done as a way to lower the variance of your predictions. You can generate predictions for different augmented versions of your test images and take the mean. This is called <strong>test-time-augmentation</strong>.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 895824,
          "author_name": "Khizar Hussain",
          "author_url": "",
          "post_date": "2020-06-21T16:22:40.960000",
          "content": "<p>Woah I didnt know about test-time-augmentation.\nThank you so much man :D I will try again by augmenting the entire dataset including the benign images. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "895552": "Hi, I am a newbie to CNNs and Kaggle so I might be asking pretty obvious questions...\nThere are 584 malignant images in train dataset, whereas 32,524 for benign images.\nThis means if I augment my data I have to make 56 copies for each malignant image by augmentation.\nWhat techniques are you guys using for data augmentation ? I used ImageDataGenerator from keras.preprocessing.Image but had terrible results for malignant images.\nAs you can see I completely overfitted my model by augmenting each malignant image 58 times :P\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2534874%2F379e86f8fda1b7fe7c55099401a86f3c%2F14F11CE4-695E-4158-9644-EDB42B189CE1.png.jpg?generation=1592744036142825&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2534874%2Fa792614f23807c66f385507e1fe6eec0%2FCF612092-3D7F-4AC7-B67C-BA117FBB4E31.png.jpg?generation=1592744037379738&amp;alt=media)\n",
    "895583": ""
  }
}