{
  "id": 73199,
  "title": "Interesting overfitting problem",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/73199",
  "author_name": "",
  "post_date": "2018-11-30T18:33:29.545773200Z",
  "votes": 2,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Hi all,\n - I am having an interesting problem (I think) :\n - I am using stratified k-fold validation as discussed somewhere in the forums\n - I am monitoring the focal loss on train and validation\n - I am monitoring f1 score at 0.5 threshold on validation</p>\n\n<p>My model is overfitting strongly when it comes to loss but is improving for a very long time when it comes to f1. Please look at the attached chart:</p>\n\n<p><img src=\"https://i.postimg.cc/k4LTvNTD/overfit.png\" alt=\"overfit\"></p>\n\n<p>I figured that it would be a good idea to monitor the distribution of predictions on validation. This is the distribution after 20 epochs when it is not yet overfitting:</p>\n\n<p><img src=\"https://i.postimg.cc/XJjHv7p7/epoch-20-class-0.png\" alt=\"epoch 20\"></p>\n\n<p>And this is how it looks like after 60 epochs when it already is:</p>\n\n<p><img src=\"https://i.postimg.cc/kD3WBffF/epoch-60-class0.png\" alt=\"epoch 60\"></p>\n\n<p>If you wanna check the evolution of prediction distributions for all the labels and all epochs please go here.\nLooking at those charts, I am guessing, that it learns (overfits) to the distribution of labels. \nWhat do you think? \nHow would you combat that?</p>\n\n<p>PS: I couldn't find anywhere how to plot the images that are simply attached to the post. Do you know how to do that?</p>",
  "messages": [
    {
      "id": "430629",
      "postDate": "11/30/2018 18:33:29",
      "content": "<p>Hi all,\n - I am having an interesting problem (I think) :\n - I am using stratified k-fold validation as discussed somewhere in the forums\n - I am monitoring the focal loss on train and validation\n - I am monitoring f1 score at 0.5 threshold on validation</p>\n\n<p>My model is overfitting strongly when it comes to loss but is improving for a very long time when it comes to f1. Please look at the attached chart:</p>\n\n<p><img src=\"https://i.postimg.cc/k4LTvNTD/overfit.png\" alt=\"overfit\"></p>\n\n<p>I figured that it would be a good idea to monitor the distribution of predictions on validation. This is the distribution after 20 epochs when it is not yet overfitting:</p>\n\n<p><img src=\"https://i.postimg.cc/XJjHv7p7/epoch-20-class-0.png\" alt=\"epoch 20\"></p>\n\n<p>And this is how it looks like after 60 epochs when it already is:</p>\n\n<p><img src=\"https://i.postimg.cc/kD3WBffF/epoch-60-class0.png\" alt=\"epoch 60\"></p>\n\n<p>If you wanna check the evolution of prediction distributions for all the labels and all epochs please go here.\nLooking at those charts, I am guessing, that it learns (overfits) to the distribution of labels. \nWhat do you think? \nHow would you combat that?</p>\n\n<p>PS: I couldn't find anywhere how to plot the images that are simply attached to the post. Do you know how to do that?</p>",
      "rawMarkdown": "Hi all,\n - I am having an interesting problem (I think) :\n - I am using stratified k-fold validation as discussed somewhere in the forums\n - I am monitoring the focal loss on train and validation\n - I am monitoring f1 score at 0.5 threshold on validation\n\nMy model is overfitting strongly when it comes to loss but is improving for a very long time when it comes to f1. Please look at the attached chart:\n\n![overfit](https://i.postimg.cc/k4LTvNTD/overfit.png)\n\nI figured that it would be a good idea to monitor the distribution of predictions on validation. This is the distribution after 20 epochs when it is not yet overfitting:\n\n![epoch 20](https://i.postimg.cc/XJjHv7p7/epoch-20-class-0.png)\n\nAnd this is how it looks like after 60 epochs when it already is:\n\n![epoch 60](https://i.postimg.cc/kD3WBffF/epoch-60-class0.png)\n\nIf you wanna check the evolution of prediction distributions for all the labels and all epochs please go here.\nLooking at those charts, I am guessing, that it learns (overfits) to the distribution of labels. \nWhat do you think? \nHow would you combat that?\n\nPS: I couldn't find anywhere how to plot the images that are simply attached to the post. Do you know how to do that?",
      "votes": null
    },
    {
      "id": "430634",
      "postDate": "11/30/2018 18:46:32",
      "content": "<blockquote>\n  <p>PS: I couldn't find anywhere how to plot the images that are simply attached to the post. Do you know how to do that?</p>\n</blockquote>\n\n<p><a href=\"/jakubczakon\">@jakubczakon</a> This is not directly answering your question, but it may be useful. What I do is upload images to one of free storage sites - try <a href=\"https://postimages.org/\"><strong>this</strong></a>. They give you a link to your file which you paste after clicking on \"Image\" icon when composing messages (a shortcut for it is CTRL+G). Below is your overfit image.</p>\n\n<p><img src=\"https://i.postimg.cc/cHQDW7mf/overfit.png\" alt=\"enter image description here\"></p>\n\n<p>Beware that this hasn't been working for me reproducibly in the last few days. Not sure if it is a browser issue or something else.</p>",
      "rawMarkdown": "&gt; PS: I couldn't find anywhere how to plot the images that are simply attached to the post. Do you know how to do that?\n\n@jakubczakon This is not directly answering your question, but it may be useful. What I do is upload images to one of free storage sites - try [__this__](https://postimages.org/). They give you a link to your file which you paste after clicking on \"Image\" icon when composing messages (a shortcut for it is CTRL+G). Below is your overfit image.\n\n![enter image description here][1]\n\nBeware that this hasn't been working for me reproducibly in the last few days. Not sure if it is a browser issue or something else.\n\n  [1]: https://i.postimg.cc/cHQDW7mf/overfit.png",
      "votes": null
    },
    {
      "id": "430645",
      "postDate": "11/30/2018 19:21:13",
      "content": "<p>Thanks for pasting the image!</p>\n\n<p>I used to upload stuff to gist and then use the link but I figured there has to be a way to do it by pointing to the attachment.</p>",
      "rawMarkdown": "Thanks for pasting the image!\n\nI used to upload stuff to gist and then use the link but I figured there has to be a way to do it by pointing to the attachment.",
      "votes": null
    },
    {
      "id": "430743",
      "postDate": "11/30/2018 22:50:16",
      "content": "<p>Yes it looks like you are overfitting, your FL on val is increasing which doesn't look good... </p>\n\n<p>For the F1 did you compute the number of predictions &gt; 0.5 after 20 epochs, and after 60 epochs ? My supposition is that those number would be similar in both case, and that would mean that after 20 epochs your model already classified your val set properly. After the 20 first epochs, it's just trying to predict number closer to 0 and 1 to reduce the focal loss but that doesn't improve your F1.</p>\n\n<p>To avoid overfitting you can try using :\n   - data augmentation, \n   - increase dropout \n   - use a more basic architecture for your model\n   - increase image resolution </p>",
      "rawMarkdown": "Yes it looks like you are overfitting, your FL on val is increasing which doesn't look good... \n\nFor the F1 did you compute the number of predictions &gt; 0.5 after 20 epochs, and after 60 epochs ? My supposition is that those number would be similar in both case, and that would mean that after 20 epochs your model already classified your val set properly. After the 20 first epochs, it's just trying to predict number closer to 0 and 1 to reduce the focal loss but that doesn't improve your F1.\n\nTo avoid overfitting you can try using :\n   - data augmentation, \n   - increase dropout \n   - use a more basic architecture for your model\n   - increase image resolution",
      "votes": null
    },
    {
      "id": "430771",
      "postDate": "11/30/2018 23:33:21",
      "content": "<p>I saw similar behavior before I added data augmentation and increased the dropout layers.</p>",
      "rawMarkdown": "I saw similar behavior before I added data augmentation and increased the dropout layers.",
      "votes": null
    },
    {
      "id": "430920",
      "postDate": "12/01/2018 08:08:31",
      "content": "<p>Thanks for suggestions Antoine,</p>\n\n<p>If what you said:</p>\n\n<blockquote>\n  <p>it's just trying to predict number closer to 0 and 1 to reduce the focal loss but that doesn't improve your F1.</p>\n</blockquote>\n\n<p>was happenning I would understand, it's just overfitting and let's combat that. What I find intriguing is that </p>\n\n<p><strong>F1 on validation keeps improving for the next 100 epochs</strong>.</p>",
      "rawMarkdown": "Thanks for suggestions Antoine,\n\nIf what you said:\n\n&gt; it's just trying to predict number closer to 0 and 1 to reduce the focal loss but that doesn't improve your F1.\n\nwas happenning I would understand, it's just overfitting and let's combat that. What I find intriguing is that \n\n **F1 on validation keeps improving for the next 100 epochs**.",
      "votes": null
    },
    {
      "id": "430922",
      "postDate": "12/01/2018 08:11:39",
      "content": "<p>Thanks Brian.</p>\n\n<p>For now, I am training on 224 x 224 without augmentation (I do have some dropout).\nIf it helps, which I suppose it will I would still not understand get why f1 keeps improving though.</p>",
      "rawMarkdown": "Thanks Brian.\n\nFor now, I am training on 224 x 224 without augmentation (I do have some dropout).\nIf it helps, which I suppose it will I would still not understand get why f1 keeps improving though.",
      "votes": null
    },
    {
      "id": "430980",
      "postDate": "12/01/2018 11:14:46",
      "content": "<p>I have the simplest augmentations, e.g. rotations and flips and still my model overfits although the validation F1 increases. </p>\n\n<p>In particular, the overfitted model yielded 0.406 on LB and 0.66 locally (I used F1 score for selection). In contrast, when I use the actual focal loss, I get 0.446 on LB and 0.64 locally.</p>\n\n<p>As you can see from the numbers above, it might be better to use the actual loss to select your snapshots on your validation set.</p>",
      "rawMarkdown": "I have the simplest augmentations, e.g. rotations and flips and still my model overfits although the validation F1 increases. \n\nIn particular, the overfitted model yielded 0.406 on LB and 0.66 locally (I used F1 score for selection). In contrast, when I use the actual focal loss, I get 0.446 on LB and 0.64 locally.\n\nAs you can see from the numbers above, it might be better to use the actual loss to select your snapshots on your validation set.",
      "votes": null
    },
    {
      "id": "430994",
      "postDate": "12/01/2018 11:31:06",
      "content": "<p><a href=\"/jakubczakon\">@jakubczakon</a> - what is the difference between the blue and red distributions - is one train and one validation?</p>",
      "rawMarkdown": "jakubczakon - what is the difference between the blue and red distributions - is one train and one validation?",
      "votes": null
    },
    {
      "id": "431020",
      "postDate": "12/01/2018 12:58:56",
      "content": "<p>red is true label 0 blue is true label 1.</p>",
      "rawMarkdown": "red is true label 0 blue is true label 1.",
      "votes": null
    },
    {
      "id": "431058",
      "postDate": "12/01/2018 14:53:11",
      "content": "<p>I agree it's strange that your F1 doesn't decrease as well... but it increase very very slowly, and F1 is quite unstable especially for rare class.</p>",
      "rawMarkdown": "I agree it's strange that your F1 doesn't decrease as well... but it increase very very slowly, and F1 is quite unstable especially for rare class.",
      "votes": null
    },
    {
      "id": "431166",
      "postDate": "12/01/2018 17:39:52",
      "content": "<p>Ah, I see now. </p>\n\n<p>I've actually been thinking about this recently. I think it's definitely a problem that you have a distribution that isn't properly bimodal - this is primarily why we have such sensitivity to the threshold...the model is unable to clearly distinguish between classes and so there is no obvious cut-off (or it's unstable). I suffer from this issue too. You are effectively trading precision and recall against each other in the middle of the chart. </p>\n\n<p>I thought about using a loss function (e.g. lovasz) which encourages a clear separation of the distributions but given the class imbalance I doubt it will be overly helpful. </p>",
      "rawMarkdown": "Ah, I see now. \n\nI've actually been thinking about this recently. I think it's definitely a problem that you have a distribution that isn't properly bimodal - this is primarily why we have such sensitivity to the threshold...the model is unable to clearly distinguish between classes and so there is no obvious cut-off (or it's unstable). I suffer from this issue too. You are effectively trading precision and recall against each other in the middle of the chart. \n\nI thought about using a loss function (e.g. lovasz) which encourages a clear separation of the distributions but given the class imbalance I doubt it will be overly helpful.",
      "votes": null
    },
    {
      "id": "431862",
      "postDate": "12/03/2018 03:08:47",
      "content": "<p>emm, actually, i found that sometimes larger validation loss can also lead to better LB. this is quite wired...</p>",
      "rawMarkdown": "emm, actually, i found that sometimes larger validation loss can also lead to better LB. this is quite wired...",
      "votes": null
    },
    {
      "id": "433424",
      "postDate": "12/05/2018 04:03:47",
      "content": "<p>Still, I don't see how he gets a monotonically increasing F1 if he is not using that metric to select snapshots (which I don't think he is).</p>",
      "rawMarkdown": "Still, I don't see how he gets a monotonically increasing F1 if he is not using that metric to select snapshots (which I don't think he is).",
      "votes": null
    },
    {
      "id": "433598",
      "postDate": "12/05/2018 08:03:25",
      "content": "<p>I guess the reason relates to the use of BCE as a loss function (if that is what is being used). <a href=\"/jakubczakon\">@jakubczakon</a> said 'monitoring' focal loss, implying it's not actually being used?</p>\n\n<p>Over-fitting BCE will mean you over-fit the popular classes but these are only 1/28 of the f1 and by over-fitting them you are probably capturing some structure that's helping predict the rarer classes better, thus boosting f1.</p>",
      "rawMarkdown": "I guess the reason relates to the use of BCE as a loss function (if that is what is being used). @jakubczakon said 'monitoring' focal loss, implying it's not actually being used?\n\nOver-fitting BCE will mean you over-fit the popular classes but these are only 1/28 of the f1 and by over-fitting them you are probably capturing some structure that's helping predict the rarer classes better, thus boosting f1.",
      "votes": null
    },
    {
      "id": "433623",
      "postDate": "12/05/2018 08:44:56",
      "content": "<p>Sorry for miscommunicating.</p>\n\n<p>I am using focal as my loss.\nI am monitoring both focal and f1 on the validation set.</p>",
      "rawMarkdown": "Sorry for miscommunicating.\n\nI am using focal as my loss.\nI am monitoring both focal and f1 on the validation set.",
      "votes": null
    },
    {
      "id": "433627",
      "postDate": "12/05/2018 08:54:43",
      "content": "<p>Hmmmmmmm, well I still think it's possibly related to this misalignment between the loss function being optimized and the f1 metric. </p>",
      "rawMarkdown": "Hmmmmmmm, well I still think it's possibly related to this misalignment between the loss function being optimized and the f1 metric.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 430634,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "11/30/2018 18:46:32",
      "content": "<blockquote>\n  <p>PS: I couldn't find anywhere how to plot the images that are simply attached to the post. Do you know how to do that?</p>\n</blockquote>\n\n<p><a href=\"/jakubczakon\">@jakubczakon</a> This is not directly answering your question, but it may be useful. What I do is upload images to one of free storage sites - try <a href=\"https://postimages.org/\"><strong>this</strong></a>. They give you a link to your file which you paste after clicking on \"Image\" icon when composing messages (a shortcut for it is CTRL+G). Below is your overfit image.</p>\n\n<p><img src=\"https://i.postimg.cc/cHQDW7mf/overfit.png\" alt=\"enter image description here\"></p>\n\n<p>Beware that this hasn't been working for me reproducibly in the last few days. Not sure if it is a browser issue or something else.</p>",
      "votes": null,
      "replies": [
        {
          "id": 430645,
          "author_name": "jakubczakon",
          "author_url": "",
          "post_date": "11/30/2018 19:21:13",
          "content": "<p>Thanks for pasting the image!</p>\n\n<p>I used to upload stuff to gist and then use the link but I figured there has to be a way to do it by pointing to the attachment.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 430743,
      "author_name": "areveillon",
      "author_url": "",
      "post_date": "11/30/2018 22:50:16",
      "content": "<p>Yes it looks like you are overfitting, your FL on val is increasing which doesn't look good... </p>\n\n<p>For the F1 did you compute the number of predictions &gt; 0.5 after 20 epochs, and after 60 epochs ? My supposition is that those number would be similar in both case, and that would mean that after 20 epochs your model already classified your val set properly. After the 20 first epochs, it's just trying to predict number closer to 0 and 1 to reduce the focal loss but that doesn't improve your F1.</p>\n\n<p>To avoid overfitting you can try using :\n   - data augmentation, \n   - increase dropout \n   - use a more basic architecture for your model\n   - increase image resolution </p>",
      "votes": null,
      "replies": [
        {
          "id": 430920,
          "author_name": "jakubczakon",
          "author_url": "",
          "post_date": "12/01/2018 08:08:31",
          "content": "<p>Thanks for suggestions Antoine,</p>\n\n<p>If what you said:</p>\n\n<blockquote>\n  <p>it's just trying to predict number closer to 0 and 1 to reduce the focal loss but that doesn't improve your F1.</p>\n</blockquote>\n\n<p>was happenning I would understand, it's just overfitting and let's combat that. What I find intriguing is that </p>\n\n<p><strong>F1 on validation keeps improving for the next 100 epochs</strong>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 431058,
          "author_name": "areveillon",
          "author_url": "",
          "post_date": "12/01/2018 14:53:11",
          "content": "<p>I agree it's strange that your F1 doesn't decrease as well... but it increase very very slowly, and F1 is quite unstable especially for rare class.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 430771,
      "author_name": "ldm314",
      "author_url": "",
      "post_date": "11/30/2018 23:33:21",
      "content": "<p>I saw similar behavior before I added data augmentation and increased the dropout layers.</p>",
      "votes": null,
      "replies": [
        {
          "id": 430922,
          "author_name": "jakubczakon",
          "author_url": "",
          "post_date": "12/01/2018 08:11:39",
          "content": "<p>Thanks Brian.</p>\n\n<p>For now, I am training on 224 x 224 without augmentation (I do have some dropout).\nIf it helps, which I suppose it will I would still not understand get why f1 keeps improving though.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 430980,
          "author_name": "alekseit",
          "author_url": "",
          "post_date": "12/01/2018 11:14:46",
          "content": "<p>I have the simplest augmentations, e.g. rotations and flips and still my model overfits although the validation F1 increases. </p>\n\n<p>In particular, the overfitted model yielded 0.406 on LB and 0.66 locally (I used F1 score for selection). In contrast, when I use the actual focal loss, I get 0.446 on LB and 0.64 locally.</p>\n\n<p>As you can see from the numbers above, it might be better to use the actual loss to select your snapshots on your validation set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 431862,
          "author_name": "zjucor",
          "author_url": "",
          "post_date": "12/03/2018 03:08:47",
          "content": "<p>emm, actually, i found that sometimes larger validation loss can also lead to better LB. this is quite wired...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 430994,
      "author_name": "maw501",
      "author_url": "",
      "post_date": "12/01/2018 11:31:06",
      "content": "<p><a href=\"/jakubczakon\">@jakubczakon</a> - what is the difference between the blue and red distributions - is one train and one validation?</p>",
      "votes": null,
      "replies": [
        {
          "id": 431020,
          "author_name": "jakubczakon",
          "author_url": "",
          "post_date": "12/01/2018 12:58:56",
          "content": "<p>red is true label 0 blue is true label 1.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 431166,
          "author_name": "maw501",
          "author_url": "",
          "post_date": "12/01/2018 17:39:52",
          "content": "<p>Ah, I see now. </p>\n\n<p>I've actually been thinking about this recently. I think it's definitely a problem that you have a distribution that isn't properly bimodal - this is primarily why we have such sensitivity to the threshold...the model is unable to clearly distinguish between classes and so there is no obvious cut-off (or it's unstable). I suffer from this issue too. You are effectively trading precision and recall against each other in the middle of the chart. </p>\n\n<p>I thought about using a loss function (e.g. lovasz) which encourages a clear separation of the distributions but given the class imbalance I doubt it will be overly helpful. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 433424,
          "author_name": "petewills",
          "author_url": "",
          "post_date": "12/05/2018 04:03:47",
          "content": "<p>Still, I don't see how he gets a monotonically increasing F1 if he is not using that metric to select snapshots (which I don't think he is).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 433598,
          "author_name": "maw501",
          "author_url": "",
          "post_date": "12/05/2018 08:03:25",
          "content": "<p>I guess the reason relates to the use of BCE as a loss function (if that is what is being used). <a href=\"/jakubczakon\">@jakubczakon</a> said 'monitoring' focal loss, implying it's not actually being used?</p>\n\n<p>Over-fitting BCE will mean you over-fit the popular classes but these are only 1/28 of the f1 and by over-fitting them you are probably capturing some structure that's helping predict the rarer classes better, thus boosting f1.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 433623,
          "author_name": "jakubczakon",
          "author_url": "",
          "post_date": "12/05/2018 08:44:56",
          "content": "<p>Sorry for miscommunicating.</p>\n\n<p>I am using focal as my loss.\nI am monitoring both focal and f1 on the validation set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 433627,
          "author_name": "maw501",
          "author_url": "",
          "post_date": "12/05/2018 08:54:43",
          "content": "<p>Hmmmmmmm, well I still think it's possibly related to this misalignment between the loss function being optimized and the f1 metric. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "430629": "Hi all,\n - I am having an interesting problem (I think) :\n - I am using stratified k-fold validation as discussed somewhere in the forums\n - I am monitoring the focal loss on train and validation\n - I am monitoring f1 score at 0.5 threshold on validation\n\nMy model is overfitting strongly when it comes to loss but is improving for a very long time when it comes to f1. Please look at the attached chart:\n\n![overfit](https://i.postimg.cc/k4LTvNTD/overfit.png)\n\nI figured that it would be a good idea to monitor the distribution of predictions on validation. This is the distribution after 20 epochs when it is not yet overfitting:\n\n![epoch 20](https://i.postimg.cc/XJjHv7p7/epoch-20-class-0.png)\n\nAnd this is how it looks like after 60 epochs when it already is:\n\n![epoch 60](https://i.postimg.cc/kD3WBffF/epoch-60-class0.png)\n\nIf you wanna check the evolution of prediction distributions for all the labels and all epochs please go here.\nLooking at those charts, I am guessing, that it learns (overfits) to the distribution of labels. \nWhat do you think? \nHow would you combat that?\n\nPS: I couldn't find anywhere how to plot the images that are simply attached to the post. Do you know how to do that?",
    "430634": "&gt; PS: I couldn't find anywhere how to plot the images that are simply attached to the post. Do you know how to do that?\n\n@jakubczakon This is not directly answering your question, but it may be useful. What I do is upload images to one of free storage sites - try [__this__](https://postimages.org/). They give you a link to your file which you paste after clicking on \"Image\" icon when composing messages (a shortcut for it is CTRL+G). Below is your overfit image.\n\n![enter image description here][1]\n\nBeware that this hasn't been working for me reproducibly in the last few days. Not sure if it is a browser issue or something else.\n\n  [1]: https://i.postimg.cc/cHQDW7mf/overfit.png",
    "430645": "Thanks for pasting the image!\n\nI used to upload stuff to gist and then use the link but I figured there has to be a way to do it by pointing to the attachment.",
    "430743": "Yes it looks like you are overfitting, your FL on val is increasing which doesn't look good... \n\nFor the F1 did you compute the number of predictions &gt; 0.5 after 20 epochs, and after 60 epochs ? My supposition is that those number would be similar in both case, and that would mean that after 20 epochs your model already classified your val set properly. After the 20 first epochs, it's just trying to predict number closer to 0 and 1 to reduce the focal loss but that doesn't improve your F1.\n\nTo avoid overfitting you can try using :\n   - data augmentation, \n   - increase dropout \n   - use a more basic architecture for your model\n   - increase image resolution",
    "430771": "I saw similar behavior before I added data augmentation and increased the dropout layers.",
    "430920": "Thanks for suggestions Antoine,\n\nIf what you said:\n\n&gt; it's just trying to predict number closer to 0 and 1 to reduce the focal loss but that doesn't improve your F1.\n\nwas happenning I would understand, it's just overfitting and let's combat that. What I find intriguing is that \n\n **F1 on validation keeps improving for the next 100 epochs**.",
    "430922": "Thanks Brian.\n\nFor now, I am training on 224 x 224 without augmentation (I do have some dropout).\nIf it helps, which I suppose it will I would still not understand get why f1 keeps improving though.",
    "430980": "I have the simplest augmentations, e.g. rotations and flips and still my model overfits although the validation F1 increases. \n\nIn particular, the overfitted model yielded 0.406 on LB and 0.66 locally (I used F1 score for selection). In contrast, when I use the actual focal loss, I get 0.446 on LB and 0.64 locally.\n\nAs you can see from the numbers above, it might be better to use the actual loss to select your snapshots on your validation set.",
    "430994": "jakubczakon - what is the difference between the blue and red distributions - is one train and one validation?",
    "431020": "red is true label 0 blue is true label 1.",
    "431058": "I agree it's strange that your F1 doesn't decrease as well... but it increase very very slowly, and F1 is quite unstable especially for rare class.",
    "431166": "Ah, I see now. \n\nI've actually been thinking about this recently. I think it's definitely a problem that you have a distribution that isn't properly bimodal - this is primarily why we have such sensitivity to the threshold...the model is unable to clearly distinguish between classes and so there is no obvious cut-off (or it's unstable). I suffer from this issue too. You are effectively trading precision and recall against each other in the middle of the chart. \n\nI thought about using a loss function (e.g. lovasz) which encourages a clear separation of the distributions but given the class imbalance I doubt it will be overly helpful.",
    "431862": "emm, actually, i found that sometimes larger validation loss can also lead to better LB. this is quite wired...",
    "433424": "Still, I don't see how he gets a monotonically increasing F1 if he is not using that metric to select snapshots (which I don't think he is).",
    "433598": "I guess the reason relates to the use of BCE as a loss function (if that is what is being used). @jakubczakon said 'monitoring' focal loss, implying it's not actually being used?\n\nOver-fitting BCE will mean you over-fit the popular classes but these are only 1/28 of the f1 and by over-fitting them you are probably capturing some structure that's helping predict the rarer classes better, thus boosting f1.",
    "433623": "Sorry for miscommunicating.\n\nI am using focal as my loss.\nI am monitoring both focal and f1 on the validation set.",
    "433627": "Hmmmmmmm, well I still think it's possibly related to this misalignment between the loss function being optimized and the f1 metric."
  },
  "source": "meta"
}