{
  "id": 244735,
  "title": "Why does not sigmoid equals softmax cross entropy?",
  "url": "/competitions/seti-breakthrough-listen/discussion/244735",
  "author_name": "Coin",
  "post_date": "2021-06-08T04:40:22.201000",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>I used same model together with same training strategy. Just by replacing pytorch nn.BCEWithLogits with nn.CrossEntropyLoss, the acc gets higher(from 0.5 to 0.6). I just found that for nn.BCEWithLogits, the scores computed by sigmoid function is smaller than that computed by softmax. Does anyone else observed this phenomenon? If so, how do you solve this problem, merely by using softmax and its associated cross entropy ?</p>",
  "messages": [
    {
      "id": 1341160,
      "postDate": "2021-06-08T13:14:14.747Z",
      "content": "<p>Let me chime in something that came off the top of my head, I may be wrong as I didn't go and research on both functionalities.</p>\n<p>One key difference is how you defined the number of classes in your model. In this case, did your model output 2 classes, or 1? If 1, then by all means, use <code>nn.BCEWithLogits</code>. According to docs, </p>\n<blockquote>\n  <blockquote>\n    <p>This loss combines a Sigmoid layer and the BCELoss in one single class.<br>\n    That means, it takes your raw logits from the output of the <code>head (classifier)</code>, and performs <code>sigmoid</code> and <code>BCE</code> loss on it. However, it does not work if you defined your model to have 2 classes, because then you should use <code>softmax</code> to transform the 2 outputs accordingly and calibrate them into probabilities.</p>\n  </blockquote>\n</blockquote>\n<p>So one red flag is when you say you are using the same model. You should really check if your model is outputting one or two classes here. In binary classification, I think there should be no difference when you use 1 class + BCE and 2 class + CE, given that BCE is a special case of CE in this case.</p>",
      "rawMarkdown": "Let me chime in something that came off the top of my head, I may be wrong as I didn't go and research on both functionalities.\n\nOne key difference is how you defined the number of classes in your model. In this case, did your model output 2 classes, or 1? If 1, then by all means, use `nn.BCEWithLogits`. According to docs, \n>>This loss combines a Sigmoid layer and the BCELoss in one single class.\nThat means, it takes your raw logits from the output of the `head (classifier)`, and performs `sigmoid` and `BCE` loss on it. However, it does not work if you defined your model to have 2 classes, because then you should use `softmax` to transform the 2 outputs accordingly and calibrate them into probabilities.\n\nSo one red flag is when you say you are using the same model. You should really check if your model is outputting one or two classes here. In binary classification, I think there should be no difference when you use 1 class + BCE and 2 class + CE, given that BCE is a special case of CE in this case.",
      "votes": 3
    },
    {
      "id": 1340544,
      "postDate": "2021-06-08T04:40:22.200Z",
      "content": "<p>Hi,</p>\n<p>I used same model together with same training strategy. Just by replacing pytorch nn.BCEWithLogits with nn.CrossEntropyLoss, the acc gets higher(from 0.5 to 0.6). I just found that for nn.BCEWithLogits, the scores computed by sigmoid function is smaller than that computed by softmax. Does anyone else observed this phenomenon? If so, how do you solve this problem, merely by using softmax and its associated cross entropy ?</p>",
      "rawMarkdown": "Hi,\n\nI used same model together with same training strategy. Just by replacing pytorch nn.BCEWithLogits with nn.CrossEntropyLoss, the acc gets higher(from 0.5 to 0.6). I just found that for nn.BCEWithLogits, the scores computed by sigmoid function is smaller than that computed by softmax. Does anyone else observed this phenomenon? If so, how do you solve this problem, merely by using softmax and its associated cross entropy ?",
      "votes": 1
    },
    {
      "id": 1353518,
      "postDate": "2021-06-17T06:37:36.177Z",
      "content": "<p>I make warmup longer and the model starts to converge now</p>",
      "rawMarkdown": "I make warmup longer and the model starts to converge now"
    },
    {
      "id": 1340706,
      "postDate": "2021-06-08T07:47:30.737Z",
      "content": "<p>I am not sure about the reason. <br>\nOne thing I can consider is that the distribution is different,<br>\nSo it affect on the gradient of loss function.</p>\n<p>Maybe gradient is not stable (?) the convergence is hard in beginning.</p>\n<p>Both gradient bce loss with sigmoid and softmax  have similar value around x = 0.5.<br>\nBut in forward calculation of the model, sigmoid value is easily changed near 0.5.<br>\nLike 0.3 - 0.7, the gradient range is more larger in sigmoid. It vanishing or exploding.<br>\nSo convergence in sigmoid loss is not making near 0.5 acc.</p>\n<p>maybe increasing batch size and reducing learning rate, gradient clipping things works.</p>",
      "rawMarkdown": "I am not sure about the reason. \nOne thing I can consider is that the distribution is different,\nSo it affect on the gradient of loss function.\n\nMaybe gradient is not stable (?) the convergence is hard in beginning.\n\nBoth gradient bce loss with sigmoid and softmax  have similar value around x = 0.5.\nBut in forward calculation of the model, sigmoid value is easily changed near 0.5.\nLike 0.3 - 0.7, the gradient range is more larger in sigmoid. It vanishing or exploding.\nSo convergence in sigmoid loss is not making near 0.5 acc.\n\nmaybe increasing batch size and reducing learning rate, gradient clipping things works.\n\n"
    }
  ],
  "comments": [
    {
      "id": 1341160,
      "author_name": "gao-hongnan",
      "author_url": "",
      "post_date": "2021-06-08T13:14:14.747000",
      "content": "<p>Let me chime in something that came off the top of my head, I may be wrong as I didn't go and research on both functionalities.</p>\n<p>One key difference is how you defined the number of classes in your model. In this case, did your model output 2 classes, or 1? If 1, then by all means, use <code>nn.BCEWithLogits</code>. According to docs, </p>\n<blockquote>\n  <blockquote>\n    <p>This loss combines a Sigmoid layer and the BCELoss in one single class.<br>\n    That means, it takes your raw logits from the output of the <code>head (classifier)</code>, and performs <code>sigmoid</code> and <code>BCE</code> loss on it. However, it does not work if you defined your model to have 2 classes, because then you should use <code>softmax</code> to transform the 2 outputs accordingly and calibrate them into probabilities.</p>\n  </blockquote>\n</blockquote>\n<p>So one red flag is when you say you are using the same model. You should really check if your model is outputting one or two classes here. In binary classification, I think there should be no difference when you use 1 class + BCE and 2 class + CE, given that BCE is a special case of CE in this case.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1353518,
      "author_name": "Coin",
      "author_url": "",
      "post_date": "2021-06-17T06:37:36.177000",
      "content": "<p>I make warmup longer and the model starts to converge now</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1340706,
      "author_name": "WOOSUNG YOON",
      "author_url": "",
      "post_date": "2021-06-08T07:47:30.737000",
      "content": "<p>I am not sure about the reason. <br>\nOne thing I can consider is that the distribution is different,<br>\nSo it affect on the gradient of loss function.</p>\n<p>Maybe gradient is not stable (?) the convergence is hard in beginning.</p>\n<p>Both gradient bce loss with sigmoid and softmax  have similar value around x = 0.5.<br>\nBut in forward calculation of the model, sigmoid value is easily changed near 0.5.<br>\nLike 0.3 - 0.7, the gradient range is more larger in sigmoid. It vanishing or exploding.<br>\nSo convergence in sigmoid loss is not making near 0.5 acc.</p>\n<p>maybe increasing batch size and reducing learning rate, gradient clipping things works.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1341160": "Let me chime in something that came off the top of my head, I may be wrong as I didn't go and research on both functionalities.\n\nOne key difference is how you defined the number of classes in your model. In this case, did your model output 2 classes, or 1? If 1, then by all means, use `nn.BCEWithLogits`. According to docs, \n>>This loss combines a Sigmoid layer and the BCELoss in one single class.\nThat means, it takes your raw logits from the output of the `head (classifier)`, and performs `sigmoid` and `BCE` loss on it. However, it does not work if you defined your model to have 2 classes, because then you should use `softmax` to transform the 2 outputs accordingly and calibrate them into probabilities.\n\nSo one red flag is when you say you are using the same model. You should really check if your model is outputting one or two classes here. In binary classification, I think there should be no difference when you use 1 class + BCE and 2 class + CE, given that BCE is a special case of CE in this case.",
    "1340544": "Hi,\n\nI used same model together with same training strategy. Just by replacing pytorch nn.BCEWithLogits with nn.CrossEntropyLoss, the acc gets higher(from 0.5 to 0.6). I just found that for nn.BCEWithLogits, the scores computed by sigmoid function is smaller than that computed by softmax. Does anyone else observed this phenomenon? If so, how do you solve this problem, merely by using softmax and its associated cross entropy ?",
    "1353518": "I make warmup longer and the model starts to converge now",
    "1340706": "I am not sure about the reason. \nOne thing I can consider is that the distribution is different,\nSo it affect on the gradient of loss function.\n\nMaybe gradient is not stable (?) the convergence is hard in beginning.\n\nBoth gradient bce loss with sigmoid and softmax  have similar value around x = 0.5.\nBut in forward calculation of the model, sigmoid value is easily changed near 0.5.\nLike 0.3 - 0.7, the gradient range is more larger in sigmoid. It vanishing or exploding.\nSo convergence in sigmoid loss is not making near 0.5 acc.\n\nmaybe increasing batch size and reducing learning rate, gradient clipping things works.\n\n"
  }
}