{
  "id": 154791,
  "title": "How should I deal with this highly imbalanced data? Any Suggestions🙏",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/154791",
  "author_name": "somuSan",
  "post_date": "2020-05-29T20:24:42.224000",
  "votes": 13,
  "comment_count": 14,
  "views": 0,
  "content": "<p>I was trying to perform EDA on this competition dataset and I have found that there are only 584 images are available for malignant, compare to that the images available with respect to benign class are so many in number (32542). it is nearly more than 50 times of the malignant class.\n <a href=\"https://imgur.com/s4eJaH3\"><img src=\"https://i.imgur.com/s4eJaH3.png\"></a></p>\n\n<p>can anybody help me with this, how to handle this level of imbalancement in the data? I would like to have your valuable advice on this.</p>",
  "messages": [
    {
      "id": 866965,
      "postDate": "2020-05-29T20:24:42.223Z",
      "content": "<p>I was trying to perform EDA on this competition dataset and I have found that there are only 584 images are available for malignant, compare to that the images available with respect to benign class are so many in number (32542). it is nearly more than 50 times of the malignant class.\n <a href=\"https://imgur.com/s4eJaH3\"><img src=\"https://i.imgur.com/s4eJaH3.png\"></a></p>\n\n<p>can anybody help me with this, how to handle this level of imbalancement in the data? I would like to have your valuable advice on this.</p>",
      "rawMarkdown": "I was trying to perform EDA on this competition dataset and I have found that there are only 584 images are available for malignant, compare to that the images available with respect to benign class are so many in number (32542). it is nearly more than 50 times of the malignant class.\n <a href=\"https://imgur.com/s4eJaH3\"><img src=\"https://i.imgur.com/s4eJaH3.png\"></a>\n\n\ncan anybody help me with this, how to handle this level of imbalancement in the data? I would like to have your valuable advice on this.",
      "votes": 13
    },
    {
      "id": 871405,
      "postDate": "2020-06-02T10:34:16.317Z",
      "content": "<p>Adding on to what everyone has written, some important ideas are : \n1. Use Focal Loss and not BCE. Focal Loss is used for image segmentation when there are lot of background pixels and fewer foreground pixels. Using them would be a better strategy\n2. Augmentation should be done very carefully as not all of our usual strategies would work especially color ones.\n3. Use SMOTE which basically oversamples from the smaller class and tries to balance it.\n4. Weighted Sampler in pytorch is also a good idea</p>",
      "rawMarkdown": "Adding on to what everyone has written, some important ideas are : \n1. Use Focal Loss and not BCE. Focal Loss is used for image segmentation when there are lot of background pixels and fewer foreground pixels. Using them would be a better strategy\n2. Augmentation should be done very carefully as not all of our usual strategies would work especially color ones.\n3. Use SMOTE which basically oversamples from the smaller class and tries to balance it.\n4. Weighted Sampler in pytorch is also a good idea",
      "votes": 3,
      "replies": [
        {
          "id": 871667,
          "postDate": "2020-06-02T15:00:47.607Z",
          "content": "<p><a href=\"/darthgera\">@darthgera</a> can you provide me any link for the 1st point, that you have mentioned.</p>",
          "rawMarkdown": "@darthgera can you provide me any link for the 1st point, that you have mentioned."
        },
        {
          "id": 871724,
          "postDate": "2020-06-02T15:41:14.250Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 871891,
          "postDate": "2020-06-02T17:58:43.513Z",
          "content": "<p>Here are the arxiv link and pytorch code for Focal Loss </p>\n\n<p><a href=\"https://arxiv.org/abs/1708.02002\">Arxiv Link</a></p>\n\n<p><a href=\"https://www.kaggle.com/c/tgs-salt-identification-challenge/discussion/65938\">Check this for pytorch implementation</a></p>",
          "rawMarkdown": "Here are the arxiv link and pytorch code for Focal Loss \n\n[Arxiv Link](https://arxiv.org/abs/1708.02002)\n\n[Check this for pytorch implementation](https://www.kaggle.com/c/tgs-salt-identification-challenge/discussion/65938)\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 867993,
      "postDate": "2020-05-30T20:03:23.920Z",
      "content": "<p>Hi, <a href=\"/soumya9977\">@soumya9977</a> </p>\n\n<p>Here are a few things we can consider to address the balance issue with heavy augmentation:\n- weighted loss \n- external data\n- synthetic sample of minority class </p>",
      "rawMarkdown": "Hi, @soumya9977 \n\nHere are a few things we can consider to address the balance issue with heavy augmentation:\n- weighted loss \n- external data\n- synthetic sample of minority class ",
      "votes": 3
    },
    {
      "id": 867560,
      "postDate": "2020-05-30T12:09:50.680Z",
      "content": "<p>1) Generate New Samples using GANs : <a href=\"https://medium.com/health-data-science/using-generative-adversarial-networks-gans-for-data-augmentation-in-colorectal-images-565deda07a22\">GANs for Data Augmentation</a>\n2) Use External Data to Upsample the minority class (check External Data Thread)\n3) Update Loss functions : Focal Loss/Weighted Loss \n4) Autoencoders (use majority class during training and then use reconstruction error to find the minority class)</p>",
      "rawMarkdown": "1) Generate New Samples using GANs : [GANs for Data Augmentation](https://medium.com/health-data-science/using-generative-adversarial-networks-gans-for-data-augmentation-in-colorectal-images-565deda07a22)\n2) Use External Data to Upsample the minority class (check External Data Thread)\n3) Update Loss functions : Focal Loss/Weighted Loss \n4) Autoencoders (use majority class during training and then use reconstruction error to find the minority class)",
      "votes": 3
    },
    {
      "id": 866982,
      "postDate": "2020-05-29T21:08:27.610Z",
      "content": "<p>Ideas:\n* Use a rebalanced loss where the weight of a class is #Otherclass / N\n* You can try focal loss\n* Oversample the minority class (and apply heavy augmentation).</p>",
      "rawMarkdown": "Ideas:\n* Use a rebalanced loss where the weight of a class is #Otherclass / N\n* You can try focal loss\n* Oversample the minority class (and apply heavy augmentation).",
      "votes": 4,
      "replies": [
        {
          "id": 871302,
          "postDate": "2020-06-02T08:59:37.157Z",
          "content": "<p>Could you say which loss works better, focal loss or weighted class loss?\nThanks in advance.</p>",
          "rawMarkdown": "Could you say which loss works better, focal loss or weighted class loss?\nThanks in advance.",
          "votes": 1
        }
      ]
    },
    {
      "id": 869738,
      "postDate": "2020-06-01T08:48:53.780Z",
      "content": "<p>Unbalanced data might seem to be introducing a bias in your model but it cannot be causation for poor accuracy results.\nArtificially introducing balance in your dataset will result in your model not learning comprehensively about the dataset and hence leading to poor predictions.\nA hypothesis could be tested after creating balance and checking out the normalized confusion matrix for your model. I faced the same problem while dealing with my <a href=\"https://www.kaggle.com/navinmundhra/beginners-mistakes-01-multiclass-lgbm-tuned\">notebook here</a>.\nAfter a bit of research, <a href=\"https://matloff.wordpress.com/2015/09/29/unbalanced-data-is-a-problem-no-balanced-data-is-worse/\">this article</a> explained to me why unbalanced data might not be a reason for my model giving poor results.\nI hope you find it helpful. Also, do check out my first post in my new series named Beginners' Mistakes over <a href=\"https://www.kaggle.com/navinmundhra/beginners-mistakes-01-multiclass-lgbm-tuned\">here</a>. Thank you :)</p>",
      "rawMarkdown": "Unbalanced data might seem to be introducing a bias in your model but it cannot be causation for poor accuracy results.\nArtificially introducing balance in your dataset will result in your model not learning comprehensively about the dataset and hence leading to poor predictions.\nA hypothesis could be tested after creating balance and checking out the normalized confusion matrix for your model. I faced the same problem while dealing with my [notebook here](https://www.kaggle.com/navinmundhra/beginners-mistakes-01-multiclass-lgbm-tuned).\nAfter a bit of research, [this article](https://matloff.wordpress.com/2015/09/29/unbalanced-data-is-a-problem-no-balanced-data-is-worse/) explained to me why unbalanced data might not be a reason for my model giving poor results.\nI hope you find it helpful. Also, do check out my first post in my new series named Beginners' Mistakes over [here](https://www.kaggle.com/navinmundhra/beginners-mistakes-01-multiclass-lgbm-tuned). Thank you :)",
      "votes": 1
    },
    {
      "id": 867889,
      "postDate": "2020-05-30T17:54:05.530Z",
      "content": "<p>You can use a rebalanced weight loss for each class as <a href=\"/arroqc\">@arroqc</a> commented, or apply Upsample / Over-sampling methods like:</p>\n\n<ul>\n<li>Random minority over-sampling with replacement</li>\n<li>Synthetic Minority Over-sampling Technique</li>\n<li>Borderline SMOTE-1</li>\n<li>Borderline SMOTE-2</li>\n<li>Adaptive synthetic sampling approach for imbalanced learning</li>\n<li>Generation of synthetic data by Randomly Over Sampling Examples</li>\n</ul>\n\n<p>For more information see: <a href=\"https://github.com/tidymodels/themis\">https://github.com/tidymodels/themis</a> (material in R but there must be the same methods for Python)</p>",
      "rawMarkdown": "You can use a rebalanced weight loss for each class as @arroqc commented, or apply Upsample / Over-sampling methods like:\n\n- Random minority over-sampling with replacement\n- Synthetic Minority Over-sampling Technique\n- Borderline SMOTE-1\n- Borderline SMOTE-2\n- Adaptive synthetic sampling approach for imbalanced learning\n- Generation of synthetic data by Randomly Over Sampling Examples\n\nFor more information see: [https://github.com/tidymodels/themis](https://github.com/tidymodels/themis) (material in R but there must be the same methods for Python)"
    },
    {
      "id": 867851,
      "postDate": "2020-05-30T17:19:05.450Z",
      "content": "<p>You could rotate malignant at 6 degrees steps to get 35 040 samples?</p>",
      "rawMarkdown": "You could rotate malignant at 6 degrees steps to get 35 040 samples?"
    },
    {
      "id": 867225,
      "postDate": "2020-05-30T05:31:15.560Z",
      "content": "<p>Use SMOTE. But its not a very good practice to balance the data to 50-50 percent. Try improving it to max 85-15 or 80-20...</p>",
      "rawMarkdown": "Use SMOTE. But its not a very good practice to balance the data to 50-50 percent. Try improving it to max 85-15 or 80-20..."
    },
    {
      "id": 867012,
      "postDate": "2020-05-29T22:23:37.467Z",
      "content": "<p>You can randomly duplicate examples from minority class and adding them to the training dataset, which is random oversampling. Here is the <a href=\"https://imbalanced-learn.readthedocs.io/en/stable/generated/imblearn.over_sampling.RandomOverSampler.html\">python package</a> to do it.</p>",
      "rawMarkdown": "You can randomly duplicate examples from minority class and adding them to the training dataset, which is random oversampling. Here is the [python package](https://imbalanced-learn.readthedocs.io/en/stable/generated/imblearn.over_sampling.RandomOverSampler.html) to do it."
    },
    {
      "id": 868006,
      "postDate": "2020-05-30T20:13:25.343Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 871405,
      "author_name": "darthgera123",
      "author_url": "",
      "post_date": "2020-06-02T10:34:16.317000",
      "content": "<p>Adding on to what everyone has written, some important ideas are : \n1. Use Focal Loss and not BCE. Focal Loss is used for image segmentation when there are lot of background pixels and fewer foreground pixels. Using them would be a better strategy\n2. Augmentation should be done very carefully as not all of our usual strategies would work especially color ones.\n3. Use SMOTE which basically oversamples from the smaller class and tries to balance it.\n4. Weighted Sampler in pytorch is also a good idea</p>",
      "votes": 3,
      "replies": [
        {
          "id": 871667,
          "author_name": "somuSan",
          "author_url": "",
          "post_date": "2020-06-02T15:00:47.607000",
          "content": "<p><a href=\"/darthgera\">@darthgera</a> can you provide me any link for the 1st point, that you have mentioned.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 871724,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-02T15:41:14.250000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 871891,
          "author_name": "Utsav Nandi",
          "author_url": "",
          "post_date": "2020-06-02T17:58:43.513000",
          "content": "<p>Here are the arxiv link and pytorch code for Focal Loss </p>\n\n<p><a href=\"https://arxiv.org/abs/1708.02002\">Arxiv Link</a></p>\n\n<p><a href=\"https://www.kaggle.com/c/tgs-salt-identification-challenge/discussion/65938\">Check this for pytorch implementation</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 867993,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-05-30T20:03:23.920000",
      "content": "<p>Hi, <a href=\"/soumya9977\">@soumya9977</a> </p>\n\n<p>Here are a few things we can consider to address the balance issue with heavy augmentation:\n- weighted loss \n- external data\n- synthetic sample of minority class </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 867560,
      "author_name": "Jimit Shah",
      "author_url": "",
      "post_date": "2020-05-30T12:09:50.680000",
      "content": "<p>1) Generate New Samples using GANs : <a href=\"https://medium.com/health-data-science/using-generative-adversarial-networks-gans-for-data-augmentation-in-colorectal-images-565deda07a22\">GANs for Data Augmentation</a>\n2) Use External Data to Upsample the minority class (check External Data Thread)\n3) Update Loss functions : Focal Loss/Weighted Loss \n4) Autoencoders (use majority class during training and then use reconstruction error to find the minority class)</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 866982,
      "author_name": "Arnaud Roussel",
      "author_url": "",
      "post_date": "2020-05-29T21:08:27.610000",
      "content": "<p>Ideas:\n* Use a rebalanced loss where the weight of a class is #Otherclass / N\n* You can try focal loss\n* Oversample the minority class (and apply heavy augmentation).</p>",
      "votes": 4,
      "replies": [
        {
          "id": 871302,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-02T08:59:37.157000",
          "content": "<p>Could you say which loss works better, focal loss or weighted class loss?\nThanks in advance.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 869738,
      "author_name": "Navin",
      "author_url": "",
      "post_date": "2020-06-01T08:48:53.780000",
      "content": "<p>Unbalanced data might seem to be introducing a bias in your model but it cannot be causation for poor accuracy results.\nArtificially introducing balance in your dataset will result in your model not learning comprehensively about the dataset and hence leading to poor predictions.\nA hypothesis could be tested after creating balance and checking out the normalized confusion matrix for your model. I faced the same problem while dealing with my <a href=\"https://www.kaggle.com/navinmundhra/beginners-mistakes-01-multiclass-lgbm-tuned\">notebook here</a>.\nAfter a bit of research, <a href=\"https://matloff.wordpress.com/2015/09/29/unbalanced-data-is-a-problem-no-balanced-data-is-worse/\">this article</a> explained to me why unbalanced data might not be a reason for my model giving poor results.\nI hope you find it helpful. Also, do check out my first post in my new series named Beginners' Mistakes over <a href=\"https://www.kaggle.com/navinmundhra/beginners-mistakes-01-multiclass-lgbm-tuned\">here</a>. Thank you :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 867889,
      "author_name": "Fellipe Gomes",
      "author_url": "",
      "post_date": "2020-05-30T17:54:05.530000",
      "content": "<p>You can use a rebalanced weight loss for each class as <a href=\"/arroqc\">@arroqc</a> commented, or apply Upsample / Over-sampling methods like:</p>\n\n<ul>\n<li>Random minority over-sampling with replacement</li>\n<li>Synthetic Minority Over-sampling Technique</li>\n<li>Borderline SMOTE-1</li>\n<li>Borderline SMOTE-2</li>\n<li>Adaptive synthetic sampling approach for imbalanced learning</li>\n<li>Generation of synthetic data by Randomly Over Sampling Examples</li>\n</ul>\n\n<p>For more information see: <a href=\"https://github.com/tidymodels/themis\">https://github.com/tidymodels/themis</a> (material in R but there must be the same methods for Python)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 867851,
      "author_name": "Tord Malmgren",
      "author_url": "",
      "post_date": "2020-05-30T17:19:05.450000",
      "content": "<p>You could rotate malignant at 6 degrees steps to get 35 040 samples?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 867225,
      "author_name": "Pratik Asarkar",
      "author_url": "",
      "post_date": "2020-05-30T05:31:15.560000",
      "content": "<p>Use SMOTE. But its not a very good practice to balance the data to 50-50 percent. Try improving it to max 85-15 or 80-20...</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 867012,
      "author_name": "yohan.chung",
      "author_url": "",
      "post_date": "2020-05-29T22:23:37.467000",
      "content": "<p>You can randomly duplicate examples from minority class and adding them to the training dataset, which is random oversampling. Here is the <a href=\"https://imbalanced-learn.readthedocs.io/en/stable/generated/imblearn.over_sampling.RandomOverSampler.html\">python package</a> to do it.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 868006,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-30T20:13:25.343000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "866965": "I was trying to perform EDA on this competition dataset and I have found that there are only 584 images are available for malignant, compare to that the images available with respect to benign class are so many in number (32542). it is nearly more than 50 times of the malignant class.\n <a href=\"https://imgur.com/s4eJaH3\"><img src=\"https://i.imgur.com/s4eJaH3.png\"></a>\n\n\ncan anybody help me with this, how to handle this level of imbalancement in the data? I would like to have your valuable advice on this.",
    "871405": "Adding on to what everyone has written, some important ideas are : \n1. Use Focal Loss and not BCE. Focal Loss is used for image segmentation when there are lot of background pixels and fewer foreground pixels. Using them would be a better strategy\n2. Augmentation should be done very carefully as not all of our usual strategies would work especially color ones.\n3. Use SMOTE which basically oversamples from the smaller class and tries to balance it.\n4. Weighted Sampler in pytorch is also a good idea",
    "867993": "Hi, @soumya9977 \n\nHere are a few things we can consider to address the balance issue with heavy augmentation:\n- weighted loss \n- external data\n- synthetic sample of minority class ",
    "867560": "1) Generate New Samples using GANs : [GANs for Data Augmentation](https://medium.com/health-data-science/using-generative-adversarial-networks-gans-for-data-augmentation-in-colorectal-images-565deda07a22)\n2) Use External Data to Upsample the minority class (check External Data Thread)\n3) Update Loss functions : Focal Loss/Weighted Loss \n4) Autoencoders (use majority class during training and then use reconstruction error to find the minority class)",
    "866982": "Ideas:\n* Use a rebalanced loss where the weight of a class is #Otherclass / N\n* You can try focal loss\n* Oversample the minority class (and apply heavy augmentation).",
    "869738": "Unbalanced data might seem to be introducing a bias in your model but it cannot be causation for poor accuracy results.\nArtificially introducing balance in your dataset will result in your model not learning comprehensively about the dataset and hence leading to poor predictions.\nA hypothesis could be tested after creating balance and checking out the normalized confusion matrix for your model. I faced the same problem while dealing with my [notebook here](https://www.kaggle.com/navinmundhra/beginners-mistakes-01-multiclass-lgbm-tuned).\nAfter a bit of research, [this article](https://matloff.wordpress.com/2015/09/29/unbalanced-data-is-a-problem-no-balanced-data-is-worse/) explained to me why unbalanced data might not be a reason for my model giving poor results.\nI hope you find it helpful. Also, do check out my first post in my new series named Beginners' Mistakes over [here](https://www.kaggle.com/navinmundhra/beginners-mistakes-01-multiclass-lgbm-tuned). Thank you :)",
    "867889": "You can use a rebalanced weight loss for each class as @arroqc commented, or apply Upsample / Over-sampling methods like:\n\n- Random minority over-sampling with replacement\n- Synthetic Minority Over-sampling Technique\n- Borderline SMOTE-1\n- Borderline SMOTE-2\n- Adaptive synthetic sampling approach for imbalanced learning\n- Generation of synthetic data by Randomly Over Sampling Examples\n\nFor more information see: [https://github.com/tidymodels/themis](https://github.com/tidymodels/themis) (material in R but there must be the same methods for Python)",
    "867851": "You could rotate malignant at 6 degrees steps to get 35 040 samples?",
    "867225": "Use SMOTE. But its not a very good practice to balance the data to 50-50 percent. Try improving it to max 85-15 or 80-20...",
    "867012": "You can randomly duplicate examples from minority class and adding them to the training dataset, which is random oversampling. Here is the [python package](https://imbalanced-learn.readthedocs.io/en/stable/generated/imblearn.over_sampling.RandomOverSampler.html) to do it.",
    "868006": ""
  }
}