{
  "id": 400221,
  "title": "Why BCE Loss function performs much worse?",
  "url": "/competitions/birdclef-2023/discussion/400221",
  "author_name": "tfts",
  "post_date": "2023-04-07T09:37:17.587000",
  "votes": 12,
  "comment_count": 13,
  "views": 0,
  "content": "<p>I tried some experiments on loss function of Cross Entropy and Binary Cross Entropy in PyTorch.</p>\n<p>I once thought, they should be equal for multi-class classification problem. But the experiment say’s it’s not.</p>\n<ul>\n<li>The Cross Entropy loss function can reach 0.78 as the nice public notebook</li>\n<li>The BCE loss can’t learn anything, just get same score 0.47 like random guess</li>\n</ul>\n<p>I updated my local Pytorch from 1.9 to  1.13，it’s nice that the new version can accept the multl-column target data now. But I’m still confused. </p>\n<p>Does anyone like to share some idea about this?</p>",
  "messages": [
    {
      "id": 2213079,
      "postDate": "2023-04-07T09:37:17.587Z",
      "content": "<p>I tried some experiments on loss function of Cross Entropy and Binary Cross Entropy in PyTorch.</p>\n<p>I once thought, they should be equal for multi-class classification problem. But the experiment say’s it’s not.</p>\n<ul>\n<li>The Cross Entropy loss function can reach 0.78 as the nice public notebook</li>\n<li>The BCE loss can’t learn anything, just get same score 0.47 like random guess</li>\n</ul>\n<p>I updated my local Pytorch from 1.9 to  1.13，it’s nice that the new version can accept the multl-column target data now. But I’m still confused. </p>\n<p>Does anyone like to share some idea about this?</p>",
      "rawMarkdown": "I tried some experiments on loss function of Cross Entropy and Binary Cross Entropy in PyTorch.\n\nI once thought, they should be equal for multi-class classification problem. But the experiment say’s it’s not.\n\n- The Cross Entropy loss function can reach 0.78 as the nice public notebook\n- The BCE loss can’t learn anything, just get same score 0.47 like random guess\n\nI updated my local Pytorch from 1.9 to  1.13，it’s nice that the new version can accept the multl-column target data now. But I’m still confused. \n\nDoes anyone like to share some idea about this?\n",
      "votes": 11
    },
    {
      "id": 2216712,
      "postDate": "2023-04-10T09:40:22.057Z",
      "content": "<p>I'm not entirely sure what your process is currently so I would like to clarify some terminology first:</p>\n<ul>\n<li>Multi-class: means that there is only one correct label for a sample. So it means that we're expecting only one type of bird call within a audio sample. Categorical cross entropy(CCE) is a common loss function for this type of task,</li>\n<li>Multi-label: means that there could be none to multiple correct labels for a sample. BCEloss and Focal loss is a common loss function for this type of task.<br>\nTechnically speaking. this competition should be a multi-label classification task, but since for the majority of the samples there are only one to zero birds at the same time, this makes categorical cross entropy loss a viable loss function.</li>\n</ul>\n<p><em>However, <a href=\"https://research.facebook.com/publications/exploring-the-limits-of-weakly-supervised-pretraining/\" target=\"_blank\">in this paper</a> it does claim that using CCE for a multi-label task pretraining performed better than BCE.</em></p>\n<p>In the comment below you mentioned that when selecting the right label you chose the one with the max probability. The possible reason that doesn't work is that, assume that you have two classes that both output \"0.8\" probability, they don't actually mean the same \"0.8\". This also means that if one is 0.7 and the other is 0.8, there is a possibility the \"0.7\" actually is more likely to be true than the \"0.8\".  </p>\n<p><strong>Why is that?</strong> BCELoss treats every label output as an independent task, so you can think of it as having a separate model for each one of the label for now. In the ideal case I should be able to compare each of the models output probability, but what happens realistically is that some models may be overconfident (outputs 90% all the time) and some models will be underconfident (outputs 60% at max), causing the comparison of probabilities to not make any sense and leads to bad scores if you did so. There is actually an entire field dedicated to model calibration if you're interested in that.  <br>\nCCE on the other case, went through a softmax layer so it compares different label outputs by default. When one of the label has high probability, the other labels will be suppressed.</p>\n<p>But I am somewhat surprised about how bad it is performing. Did you remember to remove the softmax layer in the inference notebook?</p>",
      "rawMarkdown": "I'm not entirely sure what your process is currently so I would like to clarify some terminology first:\n- Multi-class: means that there is only one correct label for a sample. So it means that we're expecting only one type of bird call within a audio sample. Categorical cross entropy(CCE) is a common loss function for this type of task,\n- Multi-label: means that there could be none to multiple correct labels for a sample. BCEloss and Focal loss is a common loss function for this type of task.\nTechnically speaking. this competition should be a multi-label classification task, but since for the majority of the samples there are only one to zero birds at the same time, this makes categorical cross entropy loss a viable loss function.\n\n*However, [in this paper](https://research.facebook.com/publications/exploring-the-limits-of-weakly-supervised-pretraining/) it does claim that using CCE for a multi-label task pretraining performed better than BCE.*\n\nIn the comment below you mentioned that when selecting the right label you chose the one with the max probability. The possible reason that doesn't work is that, assume that you have two classes that both output \"0.8\" probability, they don't actually mean the same \"0.8\". This also means that if one is 0.7 and the other is 0.8, there is a possibility the \"0.7\" actually is more likely to be true than the \"0.8\".  \n\n**Why is that?** BCELoss treats every label output as an independent task, so you can think of it as having a separate model for each one of the label for now. In the ideal case I should be able to compare each of the models output probability, but what happens realistically is that some models may be overconfident (outputs 90% all the time) and some models will be underconfident (outputs 60% at max), causing the comparison of probabilities to not make any sense and leads to bad scores if you did so. There is actually an entire field dedicated to model calibration if you're interested in that.  \nCCE on the other case, went through a softmax layer so it compares different label outputs by default. When one of the label has high probability, the other labels will be suppressed.\n\nBut I am somewhat surprised about how bad it is performing. Did you remember to remove the softmax layer in the inference notebook?",
      "votes": 7
    },
    {
      "id": 2218164,
      "postDate": "2023-04-11T13:27:13.907Z",
      "content": "<p>Hi Diogenes,</p>\n<p>I had a similar problem, I solved it by changing:<br>\n<code>nn.BCEWithLogitsLoss(reduction='mean')</code><br>\nto <br>\n<code>nn.BCEWithLogitsLoss(reduction='sum')</code></p>\n<p>I hope this also solves it for you! Let me know if it works.</p>",
      "rawMarkdown": "Hi Diogenes,\n\nI had a similar problem, I solved it by changing:\n`nn.BCEWithLogitsLoss(reduction='mean')`\nto \n`nn.BCEWithLogitsLoss(reduction='sum')`\n\nI hope this also solves it for you! Let me know if it works.",
      "votes": 3,
      "replies": [
        {
          "id": 2218250,
          "postDate": "2023-04-11T14:22:26.907Z",
          "content": "<p>Interesting! I will try it, Thanks</p>",
          "rawMarkdown": "Interesting! I will try it, Thanks",
          "votes": 1
        },
        {
          "id": 2255744,
          "postDate": "2023-05-12T01:01:09.950Z",
          "content": "<p>Thats what did the trick for me thx ! <a href=\"https://www.kaggle.com/menno1111\" target=\"_blank\">@menno1111</a> </p>",
          "rawMarkdown": "Thats what did the trick for me thx ! @menno1111 ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2216131,
      "postDate": "2023-04-09T19:49:12.023Z",
      "content": "<p>Hey Diogenes,<br>\nInteresting question, nn.BCEWithLogitsLoss combines a sigmoid layer with BCE Loss. For the way this problem is set up, applying sigmoid means multiple classes can get a positive hit per predcition with sigmoid. nn.CrossEntropyLoss on the other hand applies a logsoftmax. With softmax, the predictions are forced towards only one positive hit per prediction. I'm not an expert, but let me know if this tracks.</p>",
      "rawMarkdown": "Hey Diogenes,\nInteresting question, nn.BCEWithLogitsLoss combines a sigmoid layer with BCE Loss. For the way this problem is set up, applying sigmoid means multiple classes can get a positive hit per predcition with sigmoid. nn.CrossEntropyLoss on the other hand applies a logsoftmax. With softmax, the predictions are forced towards only one positive hit per prediction. I'm not an expert, but let me know if this tracks.",
      "votes": 1
    },
    {
      "id": 2213885,
      "postDate": "2023-04-07T23:39:24.377Z",
      "content": "<p>Are you using BCELoss in PyTorch? If so, please note that this loss function does not include an activation function. To address this, you may want to consider using BCEWithLogitsLoss, which combines a Sigmoid layer and the BCELoss.</p>",
      "rawMarkdown": "Are you using BCELoss in PyTorch? If so, please note that this loss function does not include an activation function. To address this, you may want to consider using BCEWithLogitsLoss, which combines a Sigmoid layer and the BCELoss.",
      "votes": 1,
      "replies": [
        {
          "id": 2213938,
          "postDate": "2023-04-08T01:08:47.160Z",
          "content": "<p>Thanks for the answer. Actually I use nn.BCEWithLogitsLoss. I change only one line of code from <code>nn.CrossEntropyLoss(reduction=\"mean\")</code> to <code>nn.BCEWithLogitsLoss(reduction='mean')</code>, the performance changed a lot</p>",
          "rawMarkdown": "Thanks for the answer. Actually I use nn.BCEWithLogitsLoss. I change only one line of code from `nn.CrossEntropyLoss(reduction=\"mean\")` to `nn.BCEWithLogitsLoss(reduction='mean') `, the performance changed a lot",
          "votes": 2
        }
      ]
    },
    {
      "id": 2213852,
      "postDate": "2023-04-07T22:31:02.653Z",
      "content": "<p>Hi Diogenes,</p>\n<p>The binary cross entropy loss should be used when there are 2 possible classes (classes 0 or 1)</p>\n<p>I haven't looked properly at the data in this competition, but looks like there are more then 2 possible classes. This means the BCE is not appropriate.</p>\n<p>I hope I have understood the problem correctly and have helped in some way.</p>",
      "rawMarkdown": "Hi Diogenes,\n\nThe binary cross entropy loss should be used when there are 2 possible classes (classes 0 or 1)\n\nI haven't looked properly at the data in this competition, but looks like there are more then 2 possible classes. This means the BCE is not appropriate.\n\nI hope I have understood the problem correctly and have helped in some way.",
      "votes": -2,
      "replies": [
        {
          "id": 2213940,
          "postDate": "2023-04-08T01:14:04.013Z",
          "content": "<p>Hey Toni,</p>\n<p>For multi-class classification like this competition, we have 264 class totally. I think we could treat it like 264 binary-class classification. So, for each class, the model predict its probability yes or no, we can choose the largest prob to decide which class it is finally. theoretically, I once thought I could tune the parameters to similar result.</p>",
          "rawMarkdown": "Hey Toni,\n\nFor multi-class classification like this competition, we have 264 class totally. I think we could treat it like 264 binary-class classification. So, for each class, the model predict its probability yes or no, we can choose the largest prob to decide which class it is finally. theoretically, I once thought I could tune the parameters to similar result.",
          "votes": 2,
          "replies": [
            {
              "id": 2214095,
              "postDate": "2023-04-08T06:20:13.010Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        },
        {
          "id": 2214097,
          "postDate": "2023-04-08T06:20:54.480Z",
          "content": "<p>OK great. I have never used it like that. I will probably do the same experiment you have done, to get a better feeling for it. This is exacty I have come back to kaggle. Thanks :)</p>",
          "rawMarkdown": "OK great. I have never used it like that. I will probably do the same experiment you have done, to get a better feeling for it. This is exacty I have come back to kaggle. Thanks :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 2215052,
      "postDate": "2023-04-09T04:10:19.437Z",
      "content": "<p>Hi,</p>\n<p>I am pretty new to data science and I searched around for the BCE Loss function. I read that they are used to measure predicted values and actual values, but how wonder how does this actually apply to in this particular situation?</p>\n<p>Thank you</p>",
      "rawMarkdown": "Hi,\n\nI am pretty new to data science and I searched around for the BCE Loss function. I read that they are used to measure predicted values and actual values, but how wonder how does this actually apply to in this particular situation?\n\nThank you",
      "replies": [
        {
          "id": 2215441,
          "postDate": "2023-04-09T09:43:00.010Z",
          "content": "<p>you can treat this competition as a multi-class problem.</p>\n<p>the ground truth is a tensor with shape (examples, num_classes)<br>\nthe logits from the model is also a tensor with same shape.</p>\n<p>for BCE loss, the loss function will calculate BCE loss for each columns: t<em>log(p) + (1-t)</em>log(1-p)</p>\n<p>my loss function is like this</p>\n<pre><code>class CustomLoss(nn.Module):\n    def __init__(self) -&gt; None:\n        super().__init__()\n        # self.criterion = nn.CrossEntropyLoss(reduction=\"mean\")\n        self.criterion = nn.BCEWithLogitsLoss(reduction='mean') \n\n    def forward(self, y_pred, y_true):\n        return self.criterion(y_pred, y_true)\n</code></pre>",
          "rawMarkdown": "you can treat this competition as a multi-class problem.\n\nthe ground truth is a tensor with shape (examples, num_classes)\nthe logits from the model is also a tensor with same shape.\n\nfor BCE loss, the loss function will calculate BCE loss for each columns: t*log(p) + (1-t)*log(1-p)\n\nmy loss function is like this\n\n```\nclass CustomLoss(nn.Module):\n    def __init__(self) -> None:\n        super().__init__()\n        # self.criterion = nn.CrossEntropyLoss(reduction=\"mean\")\n        self.criterion = nn.BCEWithLogitsLoss(reduction='mean') \n    \n    def forward(self, y_pred, y_true):\n        return self.criterion(y_pred, y_true)\n```"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2216712,
      "author_name": "lhanhsin",
      "author_url": "",
      "post_date": "2023-04-10T09:40:22.057000",
      "content": "<p>I'm not entirely sure what your process is currently so I would like to clarify some terminology first:</p>\n<ul>\n<li>Multi-class: means that there is only one correct label for a sample. So it means that we're expecting only one type of bird call within a audio sample. Categorical cross entropy(CCE) is a common loss function for this type of task,</li>\n<li>Multi-label: means that there could be none to multiple correct labels for a sample. BCEloss and Focal loss is a common loss function for this type of task.<br>\nTechnically speaking. this competition should be a multi-label classification task, but since for the majority of the samples there are only one to zero birds at the same time, this makes categorical cross entropy loss a viable loss function.</li>\n</ul>\n<p><em>However, <a href=\"https://research.facebook.com/publications/exploring-the-limits-of-weakly-supervised-pretraining/\" target=\"_blank\">in this paper</a> it does claim that using CCE for a multi-label task pretraining performed better than BCE.</em></p>\n<p>In the comment below you mentioned that when selecting the right label you chose the one with the max probability. The possible reason that doesn't work is that, assume that you have two classes that both output \"0.8\" probability, they don't actually mean the same \"0.8\". This also means that if one is 0.7 and the other is 0.8, there is a possibility the \"0.7\" actually is more likely to be true than the \"0.8\".  </p>\n<p><strong>Why is that?</strong> BCELoss treats every label output as an independent task, so you can think of it as having a separate model for each one of the label for now. In the ideal case I should be able to compare each of the models output probability, but what happens realistically is that some models may be overconfident (outputs 90% all the time) and some models will be underconfident (outputs 60% at max), causing the comparison of probabilities to not make any sense and leads to bad scores if you did so. There is actually an entire field dedicated to model calibration if you're interested in that.  <br>\nCCE on the other case, went through a softmax layer so it compares different label outputs by default. When one of the label has high probability, the other labels will be suppressed.</p>\n<p>But I am somewhat surprised about how bad it is performing. Did you remember to remove the softmax layer in the inference notebook?</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 2218164,
      "author_name": "menno",
      "author_url": "",
      "post_date": "2023-04-11T13:27:13.907000",
      "content": "<p>Hi Diogenes,</p>\n<p>I had a similar problem, I solved it by changing:<br>\n<code>nn.BCEWithLogitsLoss(reduction='mean')</code><br>\nto <br>\n<code>nn.BCEWithLogitsLoss(reduction='sum')</code></p>\n<p>I hope this also solves it for you! Let me know if it works.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2218250,
          "author_name": "tfts",
          "author_url": "",
          "post_date": "2023-04-11T14:22:26.907000",
          "content": "<p>Interesting! I will try it, Thanks</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2255744,
          "author_name": "JEANMPIA",
          "author_url": "",
          "post_date": "2023-05-12T01:01:09.950000",
          "content": "<p>Thats what did the trick for me thx ! <a href=\"https://www.kaggle.com/menno1111\" target=\"_blank\">@menno1111</a> </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2216131,
      "author_name": "Michael Bolton",
      "author_url": "",
      "post_date": "2023-04-09T19:49:12.023000",
      "content": "<p>Hey Diogenes,<br>\nInteresting question, nn.BCEWithLogitsLoss combines a sigmoid layer with BCE Loss. For the way this problem is set up, applying sigmoid means multiple classes can get a positive hit per predcition with sigmoid. nn.CrossEntropyLoss on the other hand applies a logsoftmax. With softmax, the predictions are forced towards only one positive hit per prediction. I'm not an expert, but let me know if this tracks.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2213885,
      "author_name": "MYSO",
      "author_url": "",
      "post_date": "2023-04-07T23:39:24.377000",
      "content": "<p>Are you using BCELoss in PyTorch? If so, please note that this loss function does not include an activation function. To address this, you may want to consider using BCEWithLogitsLoss, which combines a Sigmoid layer and the BCELoss.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2213938,
          "author_name": "tfts",
          "author_url": "",
          "post_date": "2023-04-08T01:08:47.160000",
          "content": "<p>Thanks for the answer. Actually I use nn.BCEWithLogitsLoss. I change only one line of code from <code>nn.CrossEntropyLoss(reduction=\"mean\")</code> to <code>nn.BCEWithLogitsLoss(reduction='mean')</code>, the performance changed a lot</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2213852,
      "author_name": "Tone",
      "author_url": "",
      "post_date": "2023-04-07T22:31:02.653000",
      "content": "<p>Hi Diogenes,</p>\n<p>The binary cross entropy loss should be used when there are 2 possible classes (classes 0 or 1)</p>\n<p>I haven't looked properly at the data in this competition, but looks like there are more then 2 possible classes. This means the BCE is not appropriate.</p>\n<p>I hope I have understood the problem correctly and have helped in some way.</p>",
      "votes": -2,
      "replies": [
        {
          "id": 2213940,
          "author_name": "tfts",
          "author_url": "",
          "post_date": "2023-04-08T01:14:04.013000",
          "content": "<p>Hey Toni,</p>\n<p>For multi-class classification like this competition, we have 264 class totally. I think we could treat it like 264 binary-class classification. So, for each class, the model predict its probability yes or no, we can choose the largest prob to decide which class it is finally. theoretically, I once thought I could tune the parameters to similar result.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2214095,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-04-08T06:20:13.010000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2214097,
          "author_name": "Tone",
          "author_url": "",
          "post_date": "2023-04-08T06:20:54.480000",
          "content": "<p>OK great. I have never used it like that. I will probably do the same experiment you have done, to get a better feeling for it. This is exacty I have come back to kaggle. Thanks :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2215052,
      "author_name": "Cheng Hua",
      "author_url": "",
      "post_date": "2023-04-09T04:10:19.437000",
      "content": "<p>Hi,</p>\n<p>I am pretty new to data science and I searched around for the BCE Loss function. I read that they are used to measure predicted values and actual values, but how wonder how does this actually apply to in this particular situation?</p>\n<p>Thank you</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2215441,
          "author_name": "tfts",
          "author_url": "",
          "post_date": "2023-04-09T09:43:00.010000",
          "content": "<p>you can treat this competition as a multi-class problem.</p>\n<p>the ground truth is a tensor with shape (examples, num_classes)<br>\nthe logits from the model is also a tensor with same shape.</p>\n<p>for BCE loss, the loss function will calculate BCE loss for each columns: t<em>log(p) + (1-t)</em>log(1-p)</p>\n<p>my loss function is like this</p>\n<pre><code>class CustomLoss(nn.Module):\n    def __init__(self) -&gt; None:\n        super().__init__()\n        # self.criterion = nn.CrossEntropyLoss(reduction=\"mean\")\n        self.criterion = nn.BCEWithLogitsLoss(reduction='mean') \n\n    def forward(self, y_pred, y_true):\n        return self.criterion(y_pred, y_true)\n</code></pre>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2213079": "I tried some experiments on loss function of Cross Entropy and Binary Cross Entropy in PyTorch.\n\nI once thought, they should be equal for multi-class classification problem. But the experiment say’s it’s not.\n\n- The Cross Entropy loss function can reach 0.78 as the nice public notebook\n- The BCE loss can’t learn anything, just get same score 0.47 like random guess\n\nI updated my local Pytorch from 1.9 to  1.13，it’s nice that the new version can accept the multl-column target data now. But I’m still confused. \n\nDoes anyone like to share some idea about this?\n",
    "2216712": "I'm not entirely sure what your process is currently so I would like to clarify some terminology first:\n- Multi-class: means that there is only one correct label for a sample. So it means that we're expecting only one type of bird call within a audio sample. Categorical cross entropy(CCE) is a common loss function for this type of task,\n- Multi-label: means that there could be none to multiple correct labels for a sample. BCEloss and Focal loss is a common loss function for this type of task.\nTechnically speaking. this competition should be a multi-label classification task, but since for the majority of the samples there are only one to zero birds at the same time, this makes categorical cross entropy loss a viable loss function.\n\n*However, [in this paper](https://research.facebook.com/publications/exploring-the-limits-of-weakly-supervised-pretraining/) it does claim that using CCE for a multi-label task pretraining performed better than BCE.*\n\nIn the comment below you mentioned that when selecting the right label you chose the one with the max probability. The possible reason that doesn't work is that, assume that you have two classes that both output \"0.8\" probability, they don't actually mean the same \"0.8\". This also means that if one is 0.7 and the other is 0.8, there is a possibility the \"0.7\" actually is more likely to be true than the \"0.8\".  \n\n**Why is that?** BCELoss treats every label output as an independent task, so you can think of it as having a separate model for each one of the label for now. In the ideal case I should be able to compare each of the models output probability, but what happens realistically is that some models may be overconfident (outputs 90% all the time) and some models will be underconfident (outputs 60% at max), causing the comparison of probabilities to not make any sense and leads to bad scores if you did so. There is actually an entire field dedicated to model calibration if you're interested in that.  \nCCE on the other case, went through a softmax layer so it compares different label outputs by default. When one of the label has high probability, the other labels will be suppressed.\n\nBut I am somewhat surprised about how bad it is performing. Did you remember to remove the softmax layer in the inference notebook?",
    "2218164": "Hi Diogenes,\n\nI had a similar problem, I solved it by changing:\n`nn.BCEWithLogitsLoss(reduction='mean')`\nto \n`nn.BCEWithLogitsLoss(reduction='sum')`\n\nI hope this also solves it for you! Let me know if it works.",
    "2216131": "Hey Diogenes,\nInteresting question, nn.BCEWithLogitsLoss combines a sigmoid layer with BCE Loss. For the way this problem is set up, applying sigmoid means multiple classes can get a positive hit per predcition with sigmoid. nn.CrossEntropyLoss on the other hand applies a logsoftmax. With softmax, the predictions are forced towards only one positive hit per prediction. I'm not an expert, but let me know if this tracks.",
    "2213885": "Are you using BCELoss in PyTorch? If so, please note that this loss function does not include an activation function. To address this, you may want to consider using BCEWithLogitsLoss, which combines a Sigmoid layer and the BCELoss.",
    "2213852": "Hi Diogenes,\n\nThe binary cross entropy loss should be used when there are 2 possible classes (classes 0 or 1)\n\nI haven't looked properly at the data in this competition, but looks like there are more then 2 possible classes. This means the BCE is not appropriate.\n\nI hope I have understood the problem correctly and have helped in some way.",
    "2215052": "Hi,\n\nI am pretty new to data science and I searched around for the BCE Loss function. I read that they are used to measure predicted values and actual values, but how wonder how does this actually apply to in this particular situation?\n\nThank you"
  }
}