{
  "id": 163422,
  "title": "How to deal with class imbalance?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/163422",
  "author_name": "",
  "post_date": "2020-07-02T00:55:27.196606700Z",
  "votes": 4,
  "comment_count": 14,
  "views": 0,
  "content": "<p>So i was wondering which techniques you guys used for managing the 55:1 imbalance in the training set. I'm currently ensembling different models with not approach for this issue, but I was considering over sampling the possitive melanomas by doing random transformations and saving them in a new dataset, for later retraining, but augmenting each picture 55 times seems like not such a good idea. I'm looking for suggestions in this matter. Thanks in advance!</p>",
  "messages": [
    {
      "id": "911710",
      "postDate": "07/02/2020 00:55:27",
      "content": "<p>So i was wondering which techniques you guys used for managing the 55:1 imbalance in the training set. I'm currently ensembling different models with not approach for this issue, but I was considering over sampling the possitive melanomas by doing random transformations and saving them in a new dataset, for later retraining, but augmenting each picture 55 times seems like not such a good idea. I'm looking for suggestions in this matter. Thanks in advance!</p>",
      "rawMarkdown": "So i was wondering which techniques you guys used for managing the 55:1 imbalance in the training set. I'm currently ensembling different models with not approach for this issue, but I was considering over sampling the possitive melanomas by doing random transformations and saving them in a new dataset, for later retraining, but augmenting each picture 55 times seems like not such a good idea. I'm looking for suggestions in this matter. Thanks in advance!",
      "votes": null
    },
    {
      "id": "911771",
      "postDate": "07/02/2020 02:34:56",
      "content": "<p>I think you can use <strong>class weight</strong> during training. \n<code>python\nfrom sklearn.utils import class_weight\n</code>\n<a href=\"https://stackoverflow.com/questions/30972029/how-does-the-class-weight-parameter-in-scikit-learn-work\">https://stackoverflow.com/questions/30972029/how-does-the-class-weight-parameter-in-scikit-learn-work</a></p>",
      "rawMarkdown": "I think you can use **class weight** during training. \n```python\nfrom sklearn.utils import class_weight\n```\nhttps://stackoverflow.com/questions/30972029/how-does-the-class-weight-parameter-in-scikit-learn-work",
      "votes": null
    },
    {
      "id": "911778",
      "postDate": "07/02/2020 02:50:07",
      "content": "<p>Thanks for your response! To what i understand, class weights are mainly used in non images classifications. Have you experienced any upgrade on your score by using class weights? I really want some empirical knowledge so I can jump over to the best approaches and not waste time training all over again :)</p>",
      "rawMarkdown": "Thanks for your response! To what i understand, class weights are mainly used in non images classifications. Have you experienced any upgrade on your score by using class weights? I really want some empirical knowledge so I can jump over to the best approaches and not waste time training all over again :)",
      "votes": null
    },
    {
      "id": "911782",
      "postDate": "07/02/2020 02:51:23",
      "content": "<p>Also, I'm wondering how to implement this weights in a pytorch pipeline. If you know or have an example I'll greatly appreciate it!</p>",
      "rawMarkdown": "Also, I'm wondering how to implement this weights in a pytorch pipeline. If you know or have an example I'll greatly appreciate it!",
      "votes": null
    },
    {
      "id": "911801",
      "postDate": "07/02/2020 03:06:23",
      "content": "<p>IMO, whether or not to oversample depends on the imbalance in the test set. If the test set has the same imbalance then perhaps oversampling will not help, but if the test set has more postive cases, then oversampling will help.</p>",
      "rawMarkdown": "IMO, whether or not to oversample depends on the imbalance in the test set. If the test set has the same imbalance then perhaps oversampling will not help, but if the test set has more postive cases, then oversampling will help.",
      "votes": null
    },
    {
      "id": "911840",
      "postDate": "07/02/2020 04:03:34",
      "content": "<p>Yea i thought the same thing, but then what would you recommend for the imbalance, leaving it the way it is, or maybe punishing FalseNegatives in some way for making the trade/off fair in the model. And if you choose the second option, what kind of strategies would you recommend?</p>",
      "rawMarkdown": "Yea i thought the same thing, but then what would you recommend for the imbalance, leaving it the way it is, or maybe punishing FalseNegatives in some way for making the trade/off fair in the model. And if you choose the second option, what kind of strategies would you recommend?",
      "votes": null
    },
    {
      "id": "911877",
      "postDate": "07/02/2020 04:41:03",
      "content": "<p>According to Sirish <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161497\">here</a>, he believes that the proportion of positive to negative in the test set is 2x the ratio in the train set. So maybe we should try upsampling 2x. (Or setting sample weight = 2.0 )</p>",
      "rawMarkdown": "According to Sirish [here][1], he believes that the proportion of positive to negative in the test set is 2x the ratio in the train set. So maybe we should try upsampling 2x. (Or setting sample weight = 2.0 )\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161497",
      "votes": null
    },
    {
      "id": "911884",
      "postDate": "07/02/2020 04:50:54",
      "content": "<p>Thanks for the update! Also i wanted to ask, if I upload my response with a sigmoid activation, is it consider a bad practice if I round my response to certain thresholds? Lets say all values below .1 are consider 0s and all values above .85 are consider 1s. I have a feeling this will help out on the classification, guess I could waste one upload to try it out, but I dont wan't to develop bad habits.</p>",
      "rawMarkdown": "Thanks for the update! Also i wanted to ask, if I upload my response with a sigmoid activation, is it consider a bad practice if I round my response to certain thresholds? Lets say all values below .1 are consider 0s and all values above .85 are consider 1s. I have a feeling this will help out on the classification, guess I could waste one upload to try it out, but I dont wan't to develop bad habits.",
      "votes": null
    },
    {
      "id": "911902",
      "postDate": "07/02/2020 05:14:44",
      "content": "<p><a href=\"/guillermocampollo\">@guillermocampollo</a> \nI got your point and I am too beginner to help at this, sorry about that.\nI am using TF 2.x and the implementation was straightforward. </p>",
      "rawMarkdown": "guillermocampollo \nI got your point and I am too beginner to help at this, sorry about that.\nI am using TF 2.x and the implementation was straightforward.",
      "votes": null
    },
    {
      "id": "911913",
      "postDate": "07/02/2020 05:24:54",
      "content": "<p>oh BTW, I did not try one without class weight too. I'll try that shortly!</p>",
      "rawMarkdown": "oh BTW, I did not try one without class weight too. I'll try that shortly!",
      "votes": null
    },
    {
      "id": "912235",
      "postDate": "07/02/2020 10:41:47",
      "content": "<p>You can try using Synthetic minority over-sampling technique(SMOTE) and oversample just the minority class. Doing that should be good !</p>",
      "rawMarkdown": "You can try using Synthetic minority over-sampling technique(SMOTE) and oversample just the minority class. Doing that should be good !",
      "votes": null
    },
    {
      "id": "912353",
      "postDate": "07/02/2020 12:25:10",
      "content": "<p>You could try to use the focal loss as well, which is nothing but a generalization of the normal cross entropy but you can weight more the mis/hard classification in a dynamic way instead of giving a static weight to the minority class.\nHere is the paper which describes it better: <a href=\"https://arxiv.org/abs/1708.02002\">https://arxiv.org/abs/1708.02002</a></p>",
      "rawMarkdown": "You could try to use the focal loss as well, which is nothing but a generalization of the normal cross entropy but you can weight more the mis/hard classification in a dynamic way instead of giving a static weight to the minority class.\nHere is the paper which describes it better: https://arxiv.org/abs/1708.02002",
      "votes": null
    },
    {
      "id": "912544",
      "postDate": "07/02/2020 14:59:20",
      "content": "<p>In general that will lower your AUC score but in some situations it may help. The best way to test that idea is do it on your OOF. (i.e. train on 80% of train data and predict the other 20%. Then calculate AUC locally. Then use <code>new_pred = np.clip(pred,0.1,0.85)</code> and calculate AUC on that and see which is higher. Or do this test on your full 5 fold oof).</p>",
      "rawMarkdown": "In general that will lower your AUC score but in some situations it may help. The best way to test that idea is do it on your OOF. (i.e. train on 80% of train data and predict the other 20%. Then calculate AUC locally. Then use `new_pred = np.clip(pred,0.1,0.85)` and calculate AUC on that and see which is higher. Or do this test on your full 5 fold oof).",
      "votes": null
    },
    {
      "id": "912634",
      "postDate": "07/02/2020 16:06:21",
      "content": "<p>My suggestion not related to class imbalance, but specifically some comments from above:\n&gt; is it consider a bad practice if I round my response to certain thresholds?\n&gt; new_pred = np.clip(pred,0.1,0.85)</p>\n\n<p>is that there is <strong>no need to clip</strong> when using \"AUC\" metric.</p>\n\n<p>The goal is not to Classify and predict labels (0s and 1s), but really to <strong>separate</strong> both classes i.e. ensure that highest scores are assigned to the 220-330 malignant images, and the rest of the Benign images (10982-malignants) have some random score lesser than the score assigned to malignant images.</p>\n\n<p>Note: AUC metric definition does not care about the exact score given to any specific image, just that when you pick a random M and a random B, the M should have a higher score compared to B.</p>\n\n<p>This is how I kept improving my score - by identifying ways to <strong>separate</strong> both classes M and B. See example techniques from earlier posts <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156282\">Supervised Contrastive Learning</a></p>",
      "rawMarkdown": "My suggestion not related to class imbalance, but specifically some comments from above:\n&gt; is it consider a bad practice if I round my response to certain thresholds?\n&gt; new_pred = np.clip(pred,0.1,0.85)\n\nis that there is **no need to clip** when using \"AUC\" metric.\n\nThe goal is not to Classify and predict labels (0s and 1s), but really to **separate** both classes i.e. ensure that highest scores are assigned to the 220-330 malignant images, and the rest of the Benign images (10982-malignants) have some random score lesser than the score assigned to malignant images.\n\nNote: AUC metric definition does not care about the exact score given to any specific image, just that when you pick a random M and a random B, the M should have a higher score compared to B.\n\nThis is how I kept improving my score - by identifying ways to **separate** both classes M and B. See example techniques from earlier posts [Supervised Contrastive Learning](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156282)",
      "votes": null
    },
    {
      "id": "912742",
      "postDate": "07/02/2020 17:31:33",
      "content": "<p>Ok this helps a lot, so the importance of your scores is proportional to your scores? meaning that it is relative to your own mean and std from your scores and not to the classic 0s and 1s classification?</p>",
      "rawMarkdown": "Ok this helps a lot, so the importance of your scores is proportional to your scores? meaning that it is relative to your own mean and std from your scores and not to the classic 0s and 1s classification?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 911771,
      "author_name": "bayartsogtya",
      "author_url": "",
      "post_date": "07/02/2020 02:34:56",
      "content": "<p>I think you can use <strong>class weight</strong> during training. \n<code>python\nfrom sklearn.utils import class_weight\n</code>\n<a href=\"https://stackoverflow.com/questions/30972029/how-does-the-class-weight-parameter-in-scikit-learn-work\">https://stackoverflow.com/questions/30972029/how-does-the-class-weight-parameter-in-scikit-learn-work</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 911778,
          "author_name": "guillermocampollo",
          "author_url": "",
          "post_date": "07/02/2020 02:50:07",
          "content": "<p>Thanks for your response! To what i understand, class weights are mainly used in non images classifications. Have you experienced any upgrade on your score by using class weights? I really want some empirical knowledge so I can jump over to the best approaches and not waste time training all over again :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 911782,
          "author_name": "guillermocampollo",
          "author_url": "",
          "post_date": "07/02/2020 02:51:23",
          "content": "<p>Also, I'm wondering how to implement this weights in a pytorch pipeline. If you know or have an example I'll greatly appreciate it!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 911902,
          "author_name": "bayartsogtya",
          "author_url": "",
          "post_date": "07/02/2020 05:14:44",
          "content": "<p><a href=\"/guillermocampollo\">@guillermocampollo</a> \nI got your point and I am too beginner to help at this, sorry about that.\nI am using TF 2.x and the implementation was straightforward. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 911913,
          "author_name": "bayartsogtya",
          "author_url": "",
          "post_date": "07/02/2020 05:24:54",
          "content": "<p>oh BTW, I did not try one without class weight too. I'll try that shortly!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 911801,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "07/02/2020 03:06:23",
      "content": "<p>IMO, whether or not to oversample depends on the imbalance in the test set. If the test set has the same imbalance then perhaps oversampling will not help, but if the test set has more postive cases, then oversampling will help.</p>",
      "votes": null,
      "replies": [
        {
          "id": 911840,
          "author_name": "guillermocampollo",
          "author_url": "",
          "post_date": "07/02/2020 04:03:34",
          "content": "<p>Yea i thought the same thing, but then what would you recommend for the imbalance, leaving it the way it is, or maybe punishing FalseNegatives in some way for making the trade/off fair in the model. And if you choose the second option, what kind of strategies would you recommend?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 911877,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/02/2020 04:41:03",
          "content": "<p>According to Sirish <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161497\">here</a>, he believes that the proportion of positive to negative in the test set is 2x the ratio in the train set. So maybe we should try upsampling 2x. (Or setting sample weight = 2.0 )</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 911884,
          "author_name": "guillermocampollo",
          "author_url": "",
          "post_date": "07/02/2020 04:50:54",
          "content": "<p>Thanks for the update! Also i wanted to ask, if I upload my response with a sigmoid activation, is it consider a bad practice if I round my response to certain thresholds? Lets say all values below .1 are consider 0s and all values above .85 are consider 1s. I have a feeling this will help out on the classification, guess I could waste one upload to try it out, but I dont wan't to develop bad habits.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 912544,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/02/2020 14:59:20",
          "content": "<p>In general that will lower your AUC score but in some situations it may help. The best way to test that idea is do it on your OOF. (i.e. train on 80% of train data and predict the other 20%. Then calculate AUC locally. Then use <code>new_pred = np.clip(pred,0.1,0.85)</code> and calculate AUC on that and see which is higher. Or do this test on your full 5 fold oof).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 912235,
      "author_name": "manikhindwan",
      "author_url": "",
      "post_date": "07/02/2020 10:41:47",
      "content": "<p>You can try using Synthetic minority over-sampling technique(SMOTE) and oversample just the minority class. Doing that should be good !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 912353,
      "author_name": "pranavkasela",
      "author_url": "",
      "post_date": "07/02/2020 12:25:10",
      "content": "<p>You could try to use the focal loss as well, which is nothing but a generalization of the normal cross entropy but you can weight more the mis/hard classification in a dynamic way instead of giving a static weight to the minority class.\nHere is the paper which describes it better: <a href=\"https://arxiv.org/abs/1708.02002\">https://arxiv.org/abs/1708.02002</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 912634,
      "author_name": "sirishks",
      "author_url": "",
      "post_date": "07/02/2020 16:06:21",
      "content": "<p>My suggestion not related to class imbalance, but specifically some comments from above:\n&gt; is it consider a bad practice if I round my response to certain thresholds?\n&gt; new_pred = np.clip(pred,0.1,0.85)</p>\n\n<p>is that there is <strong>no need to clip</strong> when using \"AUC\" metric.</p>\n\n<p>The goal is not to Classify and predict labels (0s and 1s), but really to <strong>separate</strong> both classes i.e. ensure that highest scores are assigned to the 220-330 malignant images, and the rest of the Benign images (10982-malignants) have some random score lesser than the score assigned to malignant images.</p>\n\n<p>Note: AUC metric definition does not care about the exact score given to any specific image, just that when you pick a random M and a random B, the M should have a higher score compared to B.</p>\n\n<p>This is how I kept improving my score - by identifying ways to <strong>separate</strong> both classes M and B. See example techniques from earlier posts <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156282\">Supervised Contrastive Learning</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 912742,
          "author_name": "guillermocampollo",
          "author_url": "",
          "post_date": "07/02/2020 17:31:33",
          "content": "<p>Ok this helps a lot, so the importance of your scores is proportional to your scores? meaning that it is relative to your own mean and std from your scores and not to the classic 0s and 1s classification?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "911710": "So i was wondering which techniques you guys used for managing the 55:1 imbalance in the training set. I'm currently ensembling different models with not approach for this issue, but I was considering over sampling the possitive melanomas by doing random transformations and saving them in a new dataset, for later retraining, but augmenting each picture 55 times seems like not such a good idea. I'm looking for suggestions in this matter. Thanks in advance!",
    "911771": "I think you can use **class weight** during training. \n```python\nfrom sklearn.utils import class_weight\n```\nhttps://stackoverflow.com/questions/30972029/how-does-the-class-weight-parameter-in-scikit-learn-work",
    "911778": "Thanks for your response! To what i understand, class weights are mainly used in non images classifications. Have you experienced any upgrade on your score by using class weights? I really want some empirical knowledge so I can jump over to the best approaches and not waste time training all over again :)",
    "911782": "Also, I'm wondering how to implement this weights in a pytorch pipeline. If you know or have an example I'll greatly appreciate it!",
    "911801": "IMO, whether or not to oversample depends on the imbalance in the test set. If the test set has the same imbalance then perhaps oversampling will not help, but if the test set has more postive cases, then oversampling will help.",
    "911840": "Yea i thought the same thing, but then what would you recommend for the imbalance, leaving it the way it is, or maybe punishing FalseNegatives in some way for making the trade/off fair in the model. And if you choose the second option, what kind of strategies would you recommend?",
    "911877": "According to Sirish [here][1], he believes that the proportion of positive to negative in the test set is 2x the ratio in the train set. So maybe we should try upsampling 2x. (Or setting sample weight = 2.0 )\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161497",
    "911884": "Thanks for the update! Also i wanted to ask, if I upload my response with a sigmoid activation, is it consider a bad practice if I round my response to certain thresholds? Lets say all values below .1 are consider 0s and all values above .85 are consider 1s. I have a feeling this will help out on the classification, guess I could waste one upload to try it out, but I dont wan't to develop bad habits.",
    "911902": "guillermocampollo \nI got your point and I am too beginner to help at this, sorry about that.\nI am using TF 2.x and the implementation was straightforward.",
    "911913": "oh BTW, I did not try one without class weight too. I'll try that shortly!",
    "912235": "You can try using Synthetic minority over-sampling technique(SMOTE) and oversample just the minority class. Doing that should be good !",
    "912353": "You could try to use the focal loss as well, which is nothing but a generalization of the normal cross entropy but you can weight more the mis/hard classification in a dynamic way instead of giving a static weight to the minority class.\nHere is the paper which describes it better: https://arxiv.org/abs/1708.02002",
    "912544": "In general that will lower your AUC score but in some situations it may help. The best way to test that idea is do it on your OOF. (i.e. train on 80% of train data and predict the other 20%. Then calculate AUC locally. Then use `new_pred = np.clip(pred,0.1,0.85)` and calculate AUC on that and see which is higher. Or do this test on your full 5 fold oof).",
    "912634": "My suggestion not related to class imbalance, but specifically some comments from above:\n&gt; is it consider a bad practice if I round my response to certain thresholds?\n&gt; new_pred = np.clip(pred,0.1,0.85)\n\nis that there is **no need to clip** when using \"AUC\" metric.\n\nThe goal is not to Classify and predict labels (0s and 1s), but really to **separate** both classes i.e. ensure that highest scores are assigned to the 220-330 malignant images, and the rest of the Benign images (10982-malignants) have some random score lesser than the score assigned to malignant images.\n\nNote: AUC metric definition does not care about the exact score given to any specific image, just that when you pick a random M and a random B, the M should have a higher score compared to B.\n\nThis is how I kept improving my score - by identifying ways to **separate** both classes M and B. See example techniques from earlier posts [Supervised Contrastive Learning](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/156282)",
    "912742": "Ok this helps a lot, so the importance of your scores is proportional to your scores? meaning that it is relative to your own mean and std from your scores and not to the classic 0s and 1s classification?"
  },
  "source": "meta"
}