{
  "id": 166167,
  "title": "No focal loss? Then how do you handle class imbalance?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/166167",
  "author_name": "",
  "post_date": "2020-07-12T03:18:27.208384300Z",
  "votes": 3,
  "comment_count": 18,
  "views": 0,
  "content": "<p>I'm really puzzled by this subject.\nThe dataset is heavily unbalanced. And my initial tests showed that the improvements using focal_loss were significant. I did study the theory behind focal_loss and I was (am) convinced that it is a good solution (yes, you need to find the right hyper-params...). <br>\nI also tried with weights in fit() call, but results were poor or nothing.\nThen, I have seen some posts claiming the contrary, in other words,  that it is better with BCE. \nI have retried this night switching from FL to BCE and my results are (single model, no ensemble, no metadata)\n- FL: 0.935\n- BCE: 0.91</p>\n\n<p>what do you think?</p>",
  "messages": [
    {
      "id": "925367",
      "postDate": "07/12/2020 03:18:27",
      "content": "<p>I'm really puzzled by this subject.\nThe dataset is heavily unbalanced. And my initial tests showed that the improvements using focal_loss were significant. I did study the theory behind focal_loss and I was (am) convinced that it is a good solution (yes, you need to find the right hyper-params...). <br>\nI also tried with weights in fit() call, but results were poor or nothing.\nThen, I have seen some posts claiming the contrary, in other words,  that it is better with BCE. \nI have retried this night switching from FL to BCE and my results are (single model, no ensemble, no metadata)\n- FL: 0.935\n- BCE: 0.91</p>\n\n<p>what do you think?</p>",
      "rawMarkdown": "I'm really puzzled by this subject.\nThe dataset is heavily unbalanced. And my initial tests showed that the improvements using focal_loss were significant. I did study the theory behind focal_loss and I was (am) convinced that it is a good solution (yes, you need to find the right hyper-params...).  \nI also tried with weights in fit() call, but results were poor or nothing.\nThen, I have seen some posts claiming the contrary, in other words,  that it is better with BCE. \nI have retried this night switching from FL to BCE and my results are (single model, no ensemble, no metadata)\n- FL: 0.935\n- BCE: 0.91\n\nwhat do you think?",
      "votes": null
    },
    {
      "id": "925721",
      "postDate": "07/12/2020 09:02:01",
      "content": "<p>Weighted Sampler maybe? Oversampling? Undersampling? </p>\n\n<p>I use Focal Loss, but these would be my best guess. I don't think BCE accomodates for class imbalance on it's own so I believe there would be something extra that they must be doing. </p>",
      "rawMarkdown": "Weighted Sampler maybe? Oversampling? Undersampling? \n\nI use Focal Loss, but these would be my best guess. I don't think BCE accomodates for class imbalance on it's own so I believe there would be something extra that they must be doing.",
      "votes": null
    },
    {
      "id": "925773",
      "postDate": "07/12/2020 10:00:35",
      "content": "<p>Yes, in general. But in this case, I don't think undersampling the majority class is feasible, and it is hard to oversample. Class weight? I thought it could be an option, but have tried several times with no benefits. I suspect that simply giving weight based on class proportion doesn't work, you need to be more \"creative\".</p>",
      "rawMarkdown": "Yes, in general. But in this case, I don't think undersampling the majority class is feasible, and it is hard to oversample. Class weight? I thought it could be an option, but have tried several times with no benefits. I suspect that simply giving weight based on class proportion doesn't work, you need to be more \"creative\".",
      "votes": null
    },
    {
      "id": "925794",
      "postDate": "07/12/2020 10:24:38",
      "content": "<p>Check out this note book for weighted class sampler <a href=\"https://www.kaggle.com/gopidurgaprasad/pytorch-weightedclasssampler\">https://www.kaggle.com/gopidurgaprasad/pytorch-weightedclasssampler</a></p>",
      "rawMarkdown": "Check out this note book for weighted class sampler https://www.kaggle.com/gopidurgaprasad/pytorch-weightedclasssampler",
      "votes": null
    },
    {
      "id": "925886",
      "postDate": "07/12/2020 11:13:00",
      "content": "<p>BCE with label smoothing works better in this competition - ls=0.05 seems to be the way to go</p>",
      "rawMarkdown": "BCE with label smoothing works better in this competition - ls=0.05 seems to be the way to go",
      "votes": null
    },
    {
      "id": "925914",
      "postDate": "07/12/2020 11:36:09",
      "content": "<p><a href=\"/romanweilguny\">@romanweilguny</a> Hey could you share a snippet showing how to implement this??</p>",
      "rawMarkdown": "romanweilguny Hey could you share a snippet showing how to implement this??",
      "votes": null
    },
    {
      "id": "925983",
      "postDate": "07/12/2020 12:12:45",
      "content": "<p>for tensorflow:\nopt = tf.keras.optimizers.Adam(learning_rate=0.001)\nloss = tf.keras.losses.BinaryCrossentropy(label_smoothing=0.05) \nmodel.compile(optimizer=opt,loss=loss,metrics=['AUC'])</p>\n\n<p>for PyTorch there should be something similar ... search the pytorch  notebooks</p>",
      "rawMarkdown": "for tensorflow:\nopt = tf.keras.optimizers.Adam(learning_rate=0.001)\nloss = tf.keras.losses.BinaryCrossentropy(label_smoothing=0.05) \nmodel.compile(optimizer=opt,loss=loss,metrics=['AUC'])\n\nfor PyTorch there should be something similar ... search the pytorch  notebooks",
      "votes": null
    },
    {
      "id": "926007",
      "postDate": "07/12/2020 12:34:54",
      "content": "<p>Can we use focal loss with label smoothing?? <a href=\"/romanweilguny\">@romanweilguny</a> </p>",
      "rawMarkdown": "Can we use focal loss with label smoothing?? @romanweilguny",
      "votes": null
    },
    {
      "id": "926271",
      "postDate": "07/12/2020 15:58:05",
      "content": "<p>I think we are both doing something wrong while switching from one loss to another (hyperparamters tuning) because we got the exact opposite results. I got 0.93 with BCE and 0.91 with focal loss. \nI believe that focal loss should give a slightly better score because it takes care of imbalanced data but the difference shouldn't be this big. \nWhat kind of hyperparamters tweaking do you do when switching from focal loss to BCE? </p>",
      "rawMarkdown": "I think we are both doing something wrong while switching from one loss to another (hyperparamters tuning) because we got the exact opposite results. I got 0.93 with BCE and 0.91 with focal loss. \nI believe that focal loss should give a slightly better score because it takes care of imbalanced data but the difference shouldn't be this big. \nWhat kind of hyperparamters tweaking do you do when switching from focal loss to BCE?",
      "votes": null
    },
    {
      "id": "926297",
      "postDate": "07/12/2020 16:30:00",
      "content": "<p>it's also highly possible that you are both overfitting</p>",
      "rawMarkdown": "it's also highly possible that you are both overfitting",
      "votes": null
    },
    {
      "id": "926311",
      "postDate": "07/12/2020 16:34:45",
      "content": "<p>I think I overfitted in most of my experiments either</p>",
      "rawMarkdown": "I think I overfitted in most of my experiments either",
      "votes": null
    },
    {
      "id": "926315",
      "postDate": "07/12/2020 16:36:04",
      "content": "<p>I don't know</p>",
      "rawMarkdown": "I don't know",
      "votes": null
    },
    {
      "id": "926587",
      "postDate": "07/12/2020 19:45:34",
      "content": "<p>I use BCE with Label Smoothing. That seems better for my EFFNET model</p>",
      "rawMarkdown": "I use BCE with Label Smoothing. That seems better for my EFFNET model",
      "votes": null
    },
    {
      "id": "926635",
      "postDate": "07/12/2020 20:13:48",
      "content": "<p>My params:\n- BCE: label smoothing = 0.05\n- Focal Loss: alfa = 0.75, gamma = 2</p>\n\n<p>I did some test changing alfa.</p>\n\n<p>The only fixed point for me is that if I switch to BCE I get worse CV. Really worse. I need to think about it.</p>",
      "rawMarkdown": "My params:\n- BCE: label smoothing = 0.05\n- Focal Loss: alfa = 0.75, gamma = 2\n\nI did some test changing alfa.\n\nThe only fixed point for me is that if I switch to BCE I get worse CV. Really worse. I need to think about it.",
      "votes": null
    },
    {
      "id": "926813",
      "postDate": "07/13/2020 02:53:38",
      "content": "<p><a href=\"/msharuk589\">@msharuk589</a>  <a href=\"/romanweilguny\">@romanweilguny</a> \nThis notebook might be helpful for you! \n<a href=\"https://www.kaggle.com/jimitshah777/bilinear-efficientnet-focal-loss-label-smoothing\">https://www.kaggle.com/jimitshah777/bilinear-efficientnet-focal-loss-label-smoothing</a></p>",
      "rawMarkdown": "msharuk589  @romanweilguny \nThis notebook might be helpful for you! \nhttps://www.kaggle.com/jimitshah777/bilinear-efficientnet-focal-loss-label-smoothing",
      "votes": null
    },
    {
      "id": "926818",
      "postDate": "07/13/2020 02:55:46",
      "content": "<p>I also did the same experiments and got 0.941 (BCE + ls=0.05) and 0.929 (focal loss alpha = 0.75, gamma = 2). I'm tring different hyper parameters.</p>",
      "rawMarkdown": "I also did the same experiments and got 0.941 (BCE + ls=0.05) and 0.929 (focal loss alpha = 0.75, gamma = 2). I'm tring different hyper parameters.",
      "votes": null
    },
    {
      "id": "927042",
      "postDate": "07/13/2020 07:01:03",
      "content": "<p>Maybe the answer is in the article pointed by <a href=\"/syumei\">@syumei</a> (<a href=\"https://www.kaggle.com/jimitshah777/bilinear-efficientnet-focal-loss-label-smoothing\">https://www.kaggle.com/jimitshah777/bilinear-efficientnet-focal-loss-label-smoothing</a>) In the article the author finds that it gets similar score with focal loss and with BCE + label smoothing. It is also interesting that with label smoothing you can make your model less overconfident. It should help when the private leaderboard comes.</p>",
      "rawMarkdown": "Maybe the answer is in the article pointed by @syumei (https://www.kaggle.com/jimitshah777/bilinear-efficientnet-focal-loss-label-smoothing) In the article the author finds that it gets similar score with focal loss and with BCE + label smoothing. It is also interesting that with label smoothing you can make your model less overconfident. It should help when the private leaderboard comes.",
      "votes": null
    },
    {
      "id": "927585",
      "postDate": "07/13/2020 13:50:49",
      "content": "<p><a href=\"/luigisaetta\">@luigisaetta</a> \nI found some bugs in my previous experiment. After I fix them, I got a better score with focal loss!</p>",
      "rawMarkdown": "luigisaetta \nI found some bugs in my previous experiment. After I fix them, I got a better score with focal loss!",
      "votes": null
    },
    {
      "id": "927667",
      "postDate": "07/13/2020 14:25:39",
      "content": "<p>I just tried a dual loss model, 2 heads. One for Labelsmoothing and the other for focal loss. It increased my local CV by 0.005 which is not negligible! </p>\n\n<p>Also, I think dual loss models tend to overfit less! Needs more testing tho! :)</p>",
      "rawMarkdown": "I just tried a dual loss model, 2 heads. One for Labelsmoothing and the other for focal loss. It increased my local CV by 0.005 which is not negligible! \n\nAlso, I think dual loss models tend to overfit less! Needs more testing tho! :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 925721,
      "author_name": "aroraaman",
      "author_url": "",
      "post_date": "07/12/2020 09:02:01",
      "content": "<p>Weighted Sampler maybe? Oversampling? Undersampling? </p>\n\n<p>I use Focal Loss, but these would be my best guess. I don't think BCE accomodates for class imbalance on it's own so I believe there would be something extra that they must be doing. </p>",
      "votes": null,
      "replies": [
        {
          "id": 925773,
          "author_name": "luigisaetta",
          "author_url": "",
          "post_date": "07/12/2020 10:00:35",
          "content": "<p>Yes, in general. But in this case, I don't think undersampling the majority class is feasible, and it is hard to oversample. Class weight? I thought it could be an option, but have tried several times with no benefits. I suspect that simply giving weight based on class proportion doesn't work, you need to be more \"creative\".</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 925794,
      "author_name": "gopidurgaprasad",
      "author_url": "",
      "post_date": "07/12/2020 10:24:38",
      "content": "<p>Check out this note book for weighted class sampler <a href=\"https://www.kaggle.com/gopidurgaprasad/pytorch-weightedclasssampler\">https://www.kaggle.com/gopidurgaprasad/pytorch-weightedclasssampler</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 925886,
      "author_name": "romanweilguny",
      "author_url": "",
      "post_date": "07/12/2020 11:13:00",
      "content": "<p>BCE with label smoothing works better in this competition - ls=0.05 seems to be the way to go</p>",
      "votes": null,
      "replies": [
        {
          "id": 925914,
          "author_name": "msharuk589",
          "author_url": "",
          "post_date": "07/12/2020 11:36:09",
          "content": "<p><a href=\"/romanweilguny\">@romanweilguny</a> Hey could you share a snippet showing how to implement this??</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 925983,
          "author_name": "romanweilguny",
          "author_url": "",
          "post_date": "07/12/2020 12:12:45",
          "content": "<p>for tensorflow:\nopt = tf.keras.optimizers.Adam(learning_rate=0.001)\nloss = tf.keras.losses.BinaryCrossentropy(label_smoothing=0.05) \nmodel.compile(optimizer=opt,loss=loss,metrics=['AUC'])</p>\n\n<p>for PyTorch there should be something similar ... search the pytorch  notebooks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 926007,
          "author_name": "msharuk589",
          "author_url": "",
          "post_date": "07/12/2020 12:34:54",
          "content": "<p>Can we use focal loss with label smoothing?? <a href=\"/romanweilguny\">@romanweilguny</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 926315,
          "author_name": "romanweilguny",
          "author_url": "",
          "post_date": "07/12/2020 16:36:04",
          "content": "<p>I don't know</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 926813,
          "author_name": "syumei",
          "author_url": "",
          "post_date": "07/13/2020 02:53:38",
          "content": "<p><a href=\"/msharuk589\">@msharuk589</a>  <a href=\"/romanweilguny\">@romanweilguny</a> \nThis notebook might be helpful for you! \n<a href=\"https://www.kaggle.com/jimitshah777/bilinear-efficientnet-focal-loss-label-smoothing\">https://www.kaggle.com/jimitshah777/bilinear-efficientnet-focal-loss-label-smoothing</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 926271,
      "author_name": "amiiiney",
      "author_url": "",
      "post_date": "07/12/2020 15:58:05",
      "content": "<p>I think we are both doing something wrong while switching from one loss to another (hyperparamters tuning) because we got the exact opposite results. I got 0.93 with BCE and 0.91 with focal loss. \nI believe that focal loss should give a slightly better score because it takes care of imbalanced data but the difference shouldn't be this big. \nWhat kind of hyperparamters tweaking do you do when switching from focal loss to BCE? </p>",
      "votes": null,
      "replies": [
        {
          "id": 926297,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "07/12/2020 16:30:00",
          "content": "<p>it's also highly possible that you are both overfitting</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 926311,
          "author_name": "romanweilguny",
          "author_url": "",
          "post_date": "07/12/2020 16:34:45",
          "content": "<p>I think I overfitted in most of my experiments either</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 926635,
          "author_name": "luigisaetta",
          "author_url": "",
          "post_date": "07/12/2020 20:13:48",
          "content": "<p>My params:\n- BCE: label smoothing = 0.05\n- Focal Loss: alfa = 0.75, gamma = 2</p>\n\n<p>I did some test changing alfa.</p>\n\n<p>The only fixed point for me is that if I switch to BCE I get worse CV. Really worse. I need to think about it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 926818,
          "author_name": "syumei",
          "author_url": "",
          "post_date": "07/13/2020 02:55:46",
          "content": "<p>I also did the same experiments and got 0.941 (BCE + ls=0.05) and 0.929 (focal loss alpha = 0.75, gamma = 2). I'm tring different hyper parameters.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 927042,
          "author_name": "luigisaetta",
          "author_url": "",
          "post_date": "07/13/2020 07:01:03",
          "content": "<p>Maybe the answer is in the article pointed by <a href=\"/syumei\">@syumei</a> (<a href=\"https://www.kaggle.com/jimitshah777/bilinear-efficientnet-focal-loss-label-smoothing\">https://www.kaggle.com/jimitshah777/bilinear-efficientnet-focal-loss-label-smoothing</a>) In the article the author finds that it gets similar score with focal loss and with BCE + label smoothing. It is also interesting that with label smoothing you can make your model less overconfident. It should help when the private leaderboard comes.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 927585,
          "author_name": "syumei",
          "author_url": "",
          "post_date": "07/13/2020 13:50:49",
          "content": "<p><a href=\"/luigisaetta\">@luigisaetta</a> \nI found some bugs in my previous experiment. After I fix them, I got a better score with focal loss!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 927667,
          "author_name": "yannmajewski",
          "author_url": "",
          "post_date": "07/13/2020 14:25:39",
          "content": "<p>I just tried a dual loss model, 2 heads. One for Labelsmoothing and the other for focal loss. It increased my local CV by 0.005 which is not negligible! </p>\n\n<p>Also, I think dual loss models tend to overfit less! Needs more testing tho! :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 926587,
      "author_name": "doanquanvietnamca",
      "author_url": "",
      "post_date": "07/12/2020 19:45:34",
      "content": "<p>I use BCE with Label Smoothing. That seems better for my EFFNET model</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "925367": "I'm really puzzled by this subject.\nThe dataset is heavily unbalanced. And my initial tests showed that the improvements using focal_loss were significant. I did study the theory behind focal_loss and I was (am) convinced that it is a good solution (yes, you need to find the right hyper-params...).  \nI also tried with weights in fit() call, but results were poor or nothing.\nThen, I have seen some posts claiming the contrary, in other words,  that it is better with BCE. \nI have retried this night switching from FL to BCE and my results are (single model, no ensemble, no metadata)\n- FL: 0.935\n- BCE: 0.91\n\nwhat do you think?",
    "925721": "Weighted Sampler maybe? Oversampling? Undersampling? \n\nI use Focal Loss, but these would be my best guess. I don't think BCE accomodates for class imbalance on it's own so I believe there would be something extra that they must be doing.",
    "925773": "Yes, in general. But in this case, I don't think undersampling the majority class is feasible, and it is hard to oversample. Class weight? I thought it could be an option, but have tried several times with no benefits. I suspect that simply giving weight based on class proportion doesn't work, you need to be more \"creative\".",
    "925794": "Check out this note book for weighted class sampler https://www.kaggle.com/gopidurgaprasad/pytorch-weightedclasssampler",
    "925886": "BCE with label smoothing works better in this competition - ls=0.05 seems to be the way to go",
    "925914": "romanweilguny Hey could you share a snippet showing how to implement this??",
    "925983": "for tensorflow:\nopt = tf.keras.optimizers.Adam(learning_rate=0.001)\nloss = tf.keras.losses.BinaryCrossentropy(label_smoothing=0.05) \nmodel.compile(optimizer=opt,loss=loss,metrics=['AUC'])\n\nfor PyTorch there should be something similar ... search the pytorch  notebooks",
    "926007": "Can we use focal loss with label smoothing?? @romanweilguny",
    "926271": "I think we are both doing something wrong while switching from one loss to another (hyperparamters tuning) because we got the exact opposite results. I got 0.93 with BCE and 0.91 with focal loss. \nI believe that focal loss should give a slightly better score because it takes care of imbalanced data but the difference shouldn't be this big. \nWhat kind of hyperparamters tweaking do you do when switching from focal loss to BCE?",
    "926297": "it's also highly possible that you are both overfitting",
    "926311": "I think I overfitted in most of my experiments either",
    "926315": "I don't know",
    "926587": "I use BCE with Label Smoothing. That seems better for my EFFNET model",
    "926635": "My params:\n- BCE: label smoothing = 0.05\n- Focal Loss: alfa = 0.75, gamma = 2\n\nI did some test changing alfa.\n\nThe only fixed point for me is that if I switch to BCE I get worse CV. Really worse. I need to think about it.",
    "926813": "msharuk589  @romanweilguny \nThis notebook might be helpful for you! \nhttps://www.kaggle.com/jimitshah777/bilinear-efficientnet-focal-loss-label-smoothing",
    "926818": "I also did the same experiments and got 0.941 (BCE + ls=0.05) and 0.929 (focal loss alpha = 0.75, gamma = 2). I'm tring different hyper parameters.",
    "927042": "Maybe the answer is in the article pointed by @syumei (https://www.kaggle.com/jimitshah777/bilinear-efficientnet-focal-loss-label-smoothing) In the article the author finds that it gets similar score with focal loss and with BCE + label smoothing. It is also interesting that with label smoothing you can make your model less overconfident. It should help when the private leaderboard comes.",
    "927585": "luigisaetta \nI found some bugs in my previous experiment. After I fix them, I got a better score with focal loss!",
    "927667": "I just tried a dual loss model, 2 heads. One for Labelsmoothing and the other for focal loss. It increased my local CV by 0.005 which is not negligible! \n\nAlso, I think dual loss models tend to overfit less! Needs more testing tho! :)"
  },
  "source": "meta"
}