{
  "id": 551740,
  "title": "Overfitting or Something Else [Solved?]",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/551740",
  "author_name": "David List",
  "post_date": "2024-12-15T07:48:07.124000",
  "votes": 10,
  "comment_count": 26,
  "views": 0,
  "content": "<p>Anyone else seeing this?  Appears to happen every time if you train long enough.  Seems to be something more than your standard \"overfitting\".  Vanishing gradient maybe?</p>\n<p><strong>Edit:</strong>  Left chart is mislabeled.  Should be dice metric (which I believe is the same as f1.)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10704200%2Fd94e7df20771c673a1577793d32ac98b%2FOverfitting.jpg?generation=1734248765410721&amp;alt=media\" alt=\"\"></p>\n<p><strong>Solved?:</strong> This reddit post discusses the phenomenon which apparently occurs when you have a lot of negative (background only) samples which presumably we do:</p>\n<p><a href=\"https://www.reddit.com/r/MLQuestions/comments/1g0d2d9/comment/lt1l0if/\" target=\"_blank\">https://www.reddit.com/r/MLQuestions/comments/1g0d2d9/comment/lt1l0if/</a></p>",
  "messages": [
    {
      "id": 3072441,
      "postDate": "2024-12-15T07:48:07.123Z",
      "content": "<p>Anyone else seeing this?  Appears to happen every time if you train long enough.  Seems to be something more than your standard \"overfitting\".  Vanishing gradient maybe?</p>\n<p><strong>Edit:</strong>  Left chart is mislabeled.  Should be dice metric (which I believe is the same as f1.)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10704200%2Fd94e7df20771c673a1577793d32ac98b%2FOverfitting.jpg?generation=1734248765410721&amp;alt=media\" alt=\"\"></p>\n<p><strong>Solved?:</strong> This reddit post discusses the phenomenon which apparently occurs when you have a lot of negative (background only) samples which presumably we do:</p>\n<p><a href=\"https://www.reddit.com/r/MLQuestions/comments/1g0d2d9/comment/lt1l0if/\" target=\"_blank\">https://www.reddit.com/r/MLQuestions/comments/1g0d2d9/comment/lt1l0if/</a></p>",
      "rawMarkdown": "Anyone else seeing this?  Appears to happen every time if you train long enough.  Seems to be something more than your standard \"overfitting\".  Vanishing gradient maybe?\n\n**Edit:**  Left chart is mislabeled.  Should be dice metric (which I believe is the same as f1.)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10704200%2Fd94e7df20771c673a1577793d32ac98b%2FOverfitting.jpg?generation=1734248765410721&alt=media)\n\n**Solved?:** This reddit post discusses the phenomenon which apparently occurs when you have a lot of negative (background only) samples which presumably we do:\n\n[https://www.reddit.com/r/MLQuestions/comments/1g0d2d9/comment/lt1l0if/](https://www.reddit.com/r/MLQuestions/comments/1g0d2d9/comment/lt1l0if/)",
      "votes": 10
    },
    {
      "id": 3072837,
      "postDate": "2024-12-15T17:07:51.803Z",
      "content": "<p>Let me guess. You use Dice loss or similar from Monai? Your model cheats and predicts all ones in channel 0 (background) and all 0 in other channels thus minimising the loss.</p>",
      "rawMarkdown": "Let me guess. You use Dice loss or similar from Monai? Your model cheats and predicts all ones in channel 0 (background) and all 0 in other channels thus minimising the loss.",
      "votes": 3,
      "replies": [
        {
          "id": 3072947,
          "postDate": "2024-12-15T20:29:40.637Z",
          "content": "<p>This the UNet notebook from the competition github account (with slight modifications like increasing the epochs and getting it to work with the competition data.)</p>\n<p>The loss is TverskyLoss.</p>",
          "rawMarkdown": "This the UNet notebook from the competition github account (with slight modifications like increasing the epochs and getting it to work with the competition data.)\n\nThe loss is TverskyLoss.",
          "replies": [
            {
              "id": 3072950,
              "postDate": "2024-12-15T20:31:03.390Z",
              "content": "<p>They are all the same. Channel 0 is background and it overpowers everything else.</p>",
              "rawMarkdown": "They are all the same. Channel 0 is background and it overpowers everything else.",
              "votes": 2
            },
            {
              "id": 3072971,
              "postDate": "2024-12-15T21:21:27.620Z",
              "content": "<p>Makes sense, and actually kind of expected.  What is unexpected though is that it works for the first 900 or so epochs.  😀</p>",
              "rawMarkdown": "Makes sense, and actually kind of expected.  What is unexpected though is that it works for the first 900 or so epochs.  😀"
            },
            {
              "id": 3072993,
              "postDate": "2024-12-15T22:24:00.433Z",
              "content": "<p>Tversky loss with default alpha and beta (0.5 both) is identical to Dice Loss.</p>",
              "rawMarkdown": "Tversky loss with default alpha and beta (0.5 both) is identical to Dice Loss.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3072590,
      "postDate": "2024-12-15T12:22:16.760Z",
      "content": "<p>I have still a lot to see and learn but it looks like a overfitting to me. There is few real samples available and the targets are very small (very few pixels involved). So is not strange that at some point the parameters on the model starts to simply remember the targets that have seen ~900 steps. And starts to learn specific non general features from them.</p>",
      "rawMarkdown": "I have still a lot to see and learn but it looks like a overfitting to me. There is few real samples available and the targets are very small (very few pixels involved). So is not strange that at some point the parameters on the model starts to simply remember the targets that have seen ~900 steps. And starts to learn specific non general features from them.",
      "votes": 1,
      "replies": [
        {
          "id": 3072952,
          "postDate": "2024-12-15T20:36:03.447Z",
          "content": "<p>This is actually not the best example.  In most cases both the loss and the recall are essentially discontinuous down to almost zero.  True, could still be overfitting, but just seems a little more dramatic than I'm used to.</p>",
          "rawMarkdown": "This is actually not the best example.  In most cases both the loss and the recall are essentially discontinuous down to almost zero.  True, could still be overfitting, but just seems a little more dramatic than I'm used to.",
          "votes": 1,
          "replies": [
            {
              "id": 3073061,
              "postDate": "2024-12-16T02:13:37.673Z",
              "content": "<p>True. I haven't seen such decay in my experiments. I haven't experimented with Monai neither.</p>",
              "rawMarkdown": "True. I haven't seen such decay in my experiments. I haven't experimented with Monai neither."
            },
            {
              "id": 3073127,
              "postDate": "2024-12-16T04:48:59.653Z",
              "content": "<p>\"I have still a lot to see and learn but it looks like a overfitting to me. \"</p>\n<p>no it is NOT overfitting.</p>\n<p>overfitting refers to good results on train set but poor results on validation set.</p>",
              "rawMarkdown": "\"I have still a lot to see and learn but it looks like a overfitting to me. \"\n\nno it is NOT overfitting.\n\noverfitting refers to good results on train set but poor results on validation set.",
              "votes": 1
            },
            {
              "id": 3073271,
              "postDate": "2024-12-16T08:52:08.720Z",
              "content": "<p>I thought that recall was the metric for validation. How can loss and metric  decay toguether for train?</p>",
              "rawMarkdown": "I thought that recall was the metric for validation. How can loss and metric  decay toguether for train?",
              "votes": 1
            },
            {
              "id": 3073697,
              "postDate": "2024-12-16T18:56:45.010Z",
              "content": "<p>A problem here is my graph only shows validation recall.  If it also showed training recall the answer here would be a little more obvious.  I think <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> is suggesting that both training and validation recall collapse.  If, on the other hand, training recall continued to go up while validation recall collapses then that would be a strong case for overfitting.  If I have a chance I'll try to plot training recall as well.</p>",
              "rawMarkdown": "A problem here is my graph only shows validation recall.  If it also showed training recall the answer here would be a little more obvious.  I think @hengck23 is suggesting that both training and validation recall collapse.  If, on the other hand, training recall continued to go up while validation recall collapses then that would be a strong case for overfitting.  If I have a chance I'll try to plot training recall as well.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3074725,
      "postDate": "2024-12-18T01:40:06.433Z",
      "content": "<p>Seems like this is a common occurrence when you have data with a lot of negative (background only) samples which I'm sure we do.  Scroll to the end of the comments:</p>\n<p><a href=\"https://www.reddit.com/r/MLQuestions/comments/1g0d2d9/comment/lt1l0if/\" target=\"_blank\">https://www.reddit.com/r/MLQuestions/comments/1g0d2d9/comment/lt1l0if/</a></p>",
      "rawMarkdown": "Seems like this is a common occurrence when you have data with a lot of negative (background only) samples which I'm sure we do.  Scroll to the end of the comments:\n\n[https://www.reddit.com/r/MLQuestions/comments/1g0d2d9/comment/lt1l0if/](https://www.reddit.com/r/MLQuestions/comments/1g0d2d9/comment/lt1l0if/)",
      "votes": 2,
      "replies": [
        {
          "id": 3074842,
          "postDate": "2024-12-18T05:51:20.223Z",
          "content": "<p>If NN can cheat it will eventually 😉</p>",
          "rawMarkdown": "If NN can cheat it will eventually 😉",
          "votes": 1,
          "replies": [
            {
              "id": 3074900,
              "postDate": "2024-12-18T07:28:49.083Z",
              "content": "<p>Seems like I <strong>might</strong> be able to tune the Tversky Loss to at least delay this?  Something to try at least.</p>\n<p><a href=\"https://arxiv.org/pdf/1706.05721\" target=\"_blank\">https://arxiv.org/pdf/1706.05721</a></p>",
              "rawMarkdown": "Seems like I **might** be able to tune the Tversky Loss to at least delay this?  Something to try at least.\n\n[https://arxiv.org/pdf/1706.05721](https://arxiv.org/pdf/1706.05721)"
            }
          ]
        }
      ]
    },
    {
      "id": 3072804,
      "postDate": "2024-12-15T16:40:02.047Z",
      "content": "<p>To add on to what <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> said. In early experiments my models were predicting all 0s for thyroglobulin and beta-galactosidase. The same might be happening for you!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F40c264c7e04be383b43113aea5a0c4d4%2F1.JPG?generation=1734280704305634&amp;alt=media\" alt=\"Cropper\" style=\"max-width: 90%\"></p>",
      "rawMarkdown": "To add on to what @hengck23 said. In early experiments my models were predicting all 0s for thyroglobulin and beta-galactosidase. The same might be happening for you!\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F40c264c7e04be383b43113aea5a0c4d4%2F1.JPG?generation=1734280704305634&alt=media\" alt=\"Cropper\" style=\"max-width: 90%;\">",
      "votes": 2,
      "replies": [
        {
          "id": 3072948,
          "postDate": "2024-12-15T20:30:34.937Z",
          "content": "<p>Yeah, I think it starts with one and then that typically starts spreading to the others.</p>",
          "rawMarkdown": "Yeah, I think it starts with one and then that typically starts spreading to the others.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3072455,
      "postDate": "2024-12-15T08:18:43.833Z",
      "content": "<p>make a map where each pixel indicates the loss value for your training (and validation) images.<br>\nyou should see why.</p>\n<p>assume your code doesn't have bugs and there are no issues with numerical issues and the results are correct, it just tells you that :<br>\nthere exists a solution such that if the recall approaches zero, the loss approaches zeros too.<br>\n(i.e. detect everything as negative with high confidence)</p>\n<p>just choose two models (one at 600 and another at 900). compute loss and recall and verify</p>",
      "rawMarkdown": "make a map where each pixel indicates the loss value for your training (and validation) images.\nyou should see why.\n\nassume your code doesn't have bugs and there are no issues with numerical issues and the results are correct, it just tells you that :\nthere exists a solution such that if the recall approaches zero, the loss approaches zeros too.\n(i.e. detect everything as negative with high confidence)\n\njust choose two models (one at 600 and another at 900). compute loss and recall and verify\n",
      "votes": 2,
      "replies": [
        {
          "id": 3072461,
          "postDate": "2024-12-15T08:30:15.837Z",
          "content": "<p>Thanks!  I'll give it a try.</p>",
          "rawMarkdown": "Thanks!  I'll give it a try."
        }
      ]
    },
    {
      "id": 3093101,
      "postDate": "2025-01-10T13:37:13.307Z",
      "content": "<p>Hi David, I'm facing the same problem using DiceFocalLoss from MONAI.<br>\nDid you find a solution?<br>\nI tried training with just FocalLoss, but the results were worse compared to DiceFocalLoss (before overfitting happens). </p>\n<p>I'm thinking about combining more losses like Dice + Focal + others to see if it decrease the Dice importance. If you already solved it and can share, I'd really appreciate it!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2F41af716a87ef99f8e190016bb7e1e1af%2F__results___15_2.png?generation=1736515587722525&amp;alt=media\" alt=\"dice per epoch\"></p>",
      "rawMarkdown": "Hi David, I'm facing the same problem using DiceFocalLoss from MONAI.\nDid you find a solution?\nI tried training with just FocalLoss, but the results were worse compared to DiceFocalLoss (before overfitting happens). \n\nI'm thinking about combining more losses like Dice + Focal + others to see if it decrease the Dice importance. If you already solved it and can share, I'd really appreciate it!\n\n![dice per epoch](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2F41af716a87ef99f8e190016bb7e1e1af%2F__results___15_2.png?generation=1736515587722525&alt=media)\n\n",
      "replies": [
        {
          "id": 3093113,
          "postDate": "2025-01-10T13:46:12.337Z",
          "content": "<p>From <a href=\"https://www.kaggle.com/DennisSakva\" target=\"_blank\">@DennisSakva</a>:</p>\n<p>\"Let me guess. You use Dice loss or similar from Monai? Your model cheats and predicts all ones in channel 0 (background) and all 0 in other channels thus minimising the loss.\"</p>\n<p>So I think you should try reduce the weight of background.</p>",
          "rawMarkdown": "From @DennisSakva:\n\n\"Let me guess. You use Dice loss or similar from Monai? Your model cheats and predicts all ones in channel 0 (background) and all 0 in other channels thus minimising the loss.\"\n\nSo I think you should try reduce the weight of background.",
          "votes": 1
        },
        {
          "id": 3093128,
          "postDate": "2025-01-10T14:00:00.160Z",
          "content": "<p>An easy way to observe this phenomenon would be if you also had the validation dice for the background class plotted as well.</p>",
          "rawMarkdown": "An easy way to observe this phenomenon would be if you also had the validation dice for the background class plotted as well.",
          "votes": 1
        },
        {
          "id": 3093393,
          "postDate": "2025-01-10T20:39:05.157Z",
          "content": "<p><a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a> the ways I've found to \"fix\" this are to increase the label size or to increase the value of beta using Tversky loss.  When I was looking at this earlier neither translated into better overall model performance, though some of that could have been related to bugs in my model.  I can confirm, however, that you can get to at least 0.750 without fixing this.</p>",
          "rawMarkdown": "@sersasj the ways I've found to \"fix\" this are to increase the label size or to increase the value of beta using Tversky loss.  When I was looking at this earlier neither translated into better overall model performance, though some of that could have been related to bugs in my model.  I can confirm, however, that you can get to at least 0.750 without fixing this.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3076379,
      "postDate": "2024-12-20T00:15:19.190Z",
      "content": "<p>I believe the most direct solution is to increase the training data and the model layers. Additionally, I added gradient clipping, which also yielded some positive feedback.</p>",
      "rawMarkdown": "I believe the most direct solution is to increase the training data and the model layers. Additionally, I added gradient clipping, which also yielded some positive feedback.",
      "replies": [
        {
          "id": 3076380,
          "postDate": "2024-12-20T00:16:54.900Z",
          "content": "<p>Furthermore, I believe that optimizing the learning rate schedule and choosing the appropriate batch size are also important factors that influence the performance.&gt; I believe the most direct solution is to increase the training data and the model layers. Additionally, I added gradient clipping, which also yielded some positive feedback.</p>",
          "rawMarkdown": "Furthermore, I believe that optimizing the learning rate schedule and choosing the appropriate batch size are also important factors that influence the performance.> I believe the most direct solution is to increase the training data and the model layers. Additionally, I added gradient clipping, which also yielded some positive feedback.\n\n"
        }
      ]
    },
    {
      "id": 3074656,
      "postDate": "2024-12-17T22:39:03.933Z",
      "content": "<p>Just an FYI…  I misread the code.  Should be Dice Metric (which I believe is really just F1 score) instead of recall on the left chart.</p>",
      "rawMarkdown": "Just an FYI...  I misread the code.  Should be Dice Metric (which I believe is really just F1 score) instead of recall on the left chart."
    },
    {
      "id": 3098021,
      "postDate": "2025-01-16T01:20:36.313Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3072837,
      "author_name": "DennisSakva",
      "author_url": "",
      "post_date": "2024-12-15T17:07:51.803000",
      "content": "<p>Let me guess. You use Dice loss or similar from Monai? Your model cheats and predicts all ones in channel 0 (background) and all 0 in other channels thus minimising the loss.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3072947,
          "author_name": "David List",
          "author_url": "",
          "post_date": "2024-12-15T20:29:40.637000",
          "content": "<p>This the UNet notebook from the competition github account (with slight modifications like increasing the epochs and getting it to work with the competition data.)</p>\n<p>The loss is TverskyLoss.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3072950,
              "author_name": "DennisSakva",
              "author_url": "",
              "post_date": "2024-12-15T20:31:03.390000",
              "content": "<p>They are all the same. Channel 0 is background and it overpowers everything else.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3072971,
              "author_name": "David List",
              "author_url": "",
              "post_date": "2024-12-15T21:21:27.620000",
              "content": "<p>Makes sense, and actually kind of expected.  What is unexpected though is that it works for the first 900 or so epochs.  😀</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3072993,
              "author_name": "Andrei Zamfir",
              "author_url": "",
              "post_date": "2024-12-15T22:24:00.433000",
              "content": "<p>Tversky loss with default alpha and beta (0.5 both) is identical to Dice Loss.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3072590,
      "author_name": "Ángel Jacinto Sánchez Ruiz",
      "author_url": "",
      "post_date": "2024-12-15T12:22:16.760000",
      "content": "<p>I have still a lot to see and learn but it looks like a overfitting to me. There is few real samples available and the targets are very small (very few pixels involved). So is not strange that at some point the parameters on the model starts to simply remember the targets that have seen ~900 steps. And starts to learn specific non general features from them.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3072952,
          "author_name": "David List",
          "author_url": "",
          "post_date": "2024-12-15T20:36:03.447000",
          "content": "<p>This is actually not the best example.  In most cases both the loss and the recall are essentially discontinuous down to almost zero.  True, could still be overfitting, but just seems a little more dramatic than I'm used to.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3073061,
              "author_name": "Ángel Jacinto Sánchez Ruiz",
              "author_url": "",
              "post_date": "2024-12-16T02:13:37.673000",
              "content": "<p>True. I haven't seen such decay in my experiments. I haven't experimented with Monai neither.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3073127,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-12-16T04:48:59.653000",
              "content": "<p>\"I have still a lot to see and learn but it looks like a overfitting to me. \"</p>\n<p>no it is NOT overfitting.</p>\n<p>overfitting refers to good results on train set but poor results on validation set.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3073271,
              "author_name": "Ángel Jacinto Sánchez Ruiz",
              "author_url": "",
              "post_date": "2024-12-16T08:52:08.720000",
              "content": "<p>I thought that recall was the metric for validation. How can loss and metric  decay toguether for train?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3073697,
              "author_name": "David List",
              "author_url": "",
              "post_date": "2024-12-16T18:56:45.010000",
              "content": "<p>A problem here is my graph only shows validation recall.  If it also showed training recall the answer here would be a little more obvious.  I think <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> is suggesting that both training and validation recall collapse.  If, on the other hand, training recall continued to go up while validation recall collapses then that would be a strong case for overfitting.  If I have a chance I'll try to plot training recall as well.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3074725,
      "author_name": "David List",
      "author_url": "",
      "post_date": "2024-12-18T01:40:06.433000",
      "content": "<p>Seems like this is a common occurrence when you have data with a lot of negative (background only) samples which I'm sure we do.  Scroll to the end of the comments:</p>\n<p><a href=\"https://www.reddit.com/r/MLQuestions/comments/1g0d2d9/comment/lt1l0if/\" target=\"_blank\">https://www.reddit.com/r/MLQuestions/comments/1g0d2d9/comment/lt1l0if/</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 3074842,
          "author_name": "DennisSakva",
          "author_url": "",
          "post_date": "2024-12-18T05:51:20.223000",
          "content": "<p>If NN can cheat it will eventually 😉</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3074900,
              "author_name": "David List",
              "author_url": "",
              "post_date": "2024-12-18T07:28:49.083000",
              "content": "<p>Seems like I <strong>might</strong> be able to tune the Tversky Loss to at least delay this?  Something to try at least.</p>\n<p><a href=\"https://arxiv.org/pdf/1706.05721\" target=\"_blank\">https://arxiv.org/pdf/1706.05721</a></p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3072804,
      "author_name": "Bartley",
      "author_url": "",
      "post_date": "2024-12-15T16:40:02.047000",
      "content": "<p>To add on to what <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> said. In early experiments my models were predicting all 0s for thyroglobulin and beta-galactosidase. The same might be happening for you!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F40c264c7e04be383b43113aea5a0c4d4%2F1.JPG?generation=1734280704305634&amp;alt=media\" alt=\"Cropper\" style=\"max-width: 90%\"></p>",
      "votes": 2,
      "replies": [
        {
          "id": 3072948,
          "author_name": "David List",
          "author_url": "",
          "post_date": "2024-12-15T20:30:34.937000",
          "content": "<p>Yeah, I think it starts with one and then that typically starts spreading to the others.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3072455,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-12-15T08:18:43.833000",
      "content": "<p>make a map where each pixel indicates the loss value for your training (and validation) images.<br>\nyou should see why.</p>\n<p>assume your code doesn't have bugs and there are no issues with numerical issues and the results are correct, it just tells you that :<br>\nthere exists a solution such that if the recall approaches zero, the loss approaches zeros too.<br>\n(i.e. detect everything as negative with high confidence)</p>\n<p>just choose two models (one at 600 and another at 900). compute loss and recall and verify</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3072461,
          "author_name": "David List",
          "author_url": "",
          "post_date": "2024-12-15T08:30:15.837000",
          "content": "<p>Thanks!  I'll give it a try.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3093101,
      "author_name": "Sergio Alvarez",
      "author_url": "",
      "post_date": "2025-01-10T13:37:13.307000",
      "content": "<p>Hi David, I'm facing the same problem using DiceFocalLoss from MONAI.<br>\nDid you find a solution?<br>\nI tried training with just FocalLoss, but the results were worse compared to DiceFocalLoss (before overfitting happens). </p>\n<p>I'm thinking about combining more losses like Dice + Focal + others to see if it decrease the Dice importance. If you already solved it and can share, I'd really appreciate it!</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2F41af716a87ef99f8e190016bb7e1e1af%2F__results___15_2.png?generation=1736515587722525&amp;alt=media\" alt=\"dice per epoch\"></p>",
      "votes": 0,
      "replies": [
        {
          "id": 3093113,
          "author_name": "Ángel Jacinto Sánchez Ruiz",
          "author_url": "",
          "post_date": "2025-01-10T13:46:12.337000",
          "content": "<p>From <a href=\"https://www.kaggle.com/DennisSakva\" target=\"_blank\">@DennisSakva</a>:</p>\n<p>\"Let me guess. You use Dice loss or similar from Monai? Your model cheats and predicts all ones in channel 0 (background) and all 0 in other channels thus minimising the loss.\"</p>\n<p>So I think you should try reduce the weight of background.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3093128,
          "author_name": "Andrei Zamfir",
          "author_url": "",
          "post_date": "2025-01-10T14:00:00.160000",
          "content": "<p>An easy way to observe this phenomenon would be if you also had the validation dice for the background class plotted as well.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3093393,
          "author_name": "David List",
          "author_url": "",
          "post_date": "2025-01-10T20:39:05.157000",
          "content": "<p><a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a> the ways I've found to \"fix\" this are to increase the label size or to increase the value of beta using Tversky loss.  When I was looking at this earlier neither translated into better overall model performance, though some of that could have been related to bugs in my model.  I can confirm, however, that you can get to at least 0.750 without fixing this.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3076379,
      "author_name": "Lonelycrab",
      "author_url": "",
      "post_date": "2024-12-20T00:15:19.190000",
      "content": "<p>I believe the most direct solution is to increase the training data and the model layers. Additionally, I added gradient clipping, which also yielded some positive feedback.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3076380,
          "author_name": "Lonelycrab",
          "author_url": "",
          "post_date": "2024-12-20T00:16:54.900000",
          "content": "<p>Furthermore, I believe that optimizing the learning rate schedule and choosing the appropriate batch size are also important factors that influence the performance.&gt; I believe the most direct solution is to increase the training data and the model layers. Additionally, I added gradient clipping, which also yielded some positive feedback.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3074656,
      "author_name": "David List",
      "author_url": "",
      "post_date": "2024-12-17T22:39:03.933000",
      "content": "<p>Just an FYI…  I misread the code.  Should be Dice Metric (which I believe is really just F1 score) instead of recall on the left chart.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3098021,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-01-16T01:20:36.313000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3072441": "Anyone else seeing this?  Appears to happen every time if you train long enough.  Seems to be something more than your standard \"overfitting\".  Vanishing gradient maybe?\n\n**Edit:**  Left chart is mislabeled.  Should be dice metric (which I believe is the same as f1.)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10704200%2Fd94e7df20771c673a1577793d32ac98b%2FOverfitting.jpg?generation=1734248765410721&alt=media)\n\n**Solved?:** This reddit post discusses the phenomenon which apparently occurs when you have a lot of negative (background only) samples which presumably we do:\n\n[https://www.reddit.com/r/MLQuestions/comments/1g0d2d9/comment/lt1l0if/](https://www.reddit.com/r/MLQuestions/comments/1g0d2d9/comment/lt1l0if/)",
    "3072837": "Let me guess. You use Dice loss or similar from Monai? Your model cheats and predicts all ones in channel 0 (background) and all 0 in other channels thus minimising the loss.",
    "3072590": "I have still a lot to see and learn but it looks like a overfitting to me. There is few real samples available and the targets are very small (very few pixels involved). So is not strange that at some point the parameters on the model starts to simply remember the targets that have seen ~900 steps. And starts to learn specific non general features from them.",
    "3074725": "Seems like this is a common occurrence when you have data with a lot of negative (background only) samples which I'm sure we do.  Scroll to the end of the comments:\n\n[https://www.reddit.com/r/MLQuestions/comments/1g0d2d9/comment/lt1l0if/](https://www.reddit.com/r/MLQuestions/comments/1g0d2d9/comment/lt1l0if/)",
    "3072804": "To add on to what @hengck23 said. In early experiments my models were predicting all 0s for thyroglobulin and beta-galactosidase. The same might be happening for you!\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5570735%2F40c264c7e04be383b43113aea5a0c4d4%2F1.JPG?generation=1734280704305634&alt=media\" alt=\"Cropper\" style=\"max-width: 90%;\">",
    "3072455": "make a map where each pixel indicates the loss value for your training (and validation) images.\nyou should see why.\n\nassume your code doesn't have bugs and there are no issues with numerical issues and the results are correct, it just tells you that :\nthere exists a solution such that if the recall approaches zero, the loss approaches zeros too.\n(i.e. detect everything as negative with high confidence)\n\njust choose two models (one at 600 and another at 900). compute loss and recall and verify\n",
    "3093101": "Hi David, I'm facing the same problem using DiceFocalLoss from MONAI.\nDid you find a solution?\nI tried training with just FocalLoss, but the results were worse compared to DiceFocalLoss (before overfitting happens). \n\nI'm thinking about combining more losses like Dice + Focal + others to see if it decrease the Dice importance. If you already solved it and can share, I'd really appreciate it!\n\n![dice per epoch](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2F41af716a87ef99f8e190016bb7e1e1af%2F__results___15_2.png?generation=1736515587722525&alt=media)\n\n",
    "3076379": "I believe the most direct solution is to increase the training data and the model layers. Additionally, I added gradient clipping, which also yielded some positive feedback.",
    "3074656": "Just an FYI...  I misread the code.  Should be Dice Metric (which I believe is really just F1 score) instead of recall on the left chart.",
    "3098021": ""
  }
}