{
  "id": 101066,
  "title": "Training with softmax but inference with sigmoid?",
  "url": "/competitions/recursion-cellular-image-classification/discussion/101066",
  "author_name": "",
  "post_date": "2019-07-23T00:30:49.406799Z",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Is this ever done?</p>\n\n<p>I have been trying to train a model with sigmoid activation and cross entropy loss however none of my models train well at all. They just end up sitting with a steady high loss and zero accuracy.\nWhen I use softmax activation, they perform much better.</p>\n\n<p>However, for some of the challenge specific test time tricks I want to apply, I think I need a sigmoid activation so that I know the probability of each siRNA independently.</p>\n\n<p>So, is it possible for me to train the model with a softmax activation, however change it to sigmoid activation for inference? Is this ever done, will it work?</p>\n\n<p>Cheers,\nC</p>",
  "messages": [
    {
      "id": "582249",
      "postDate": "07/23/2019 00:30:49",
      "content": "<p>Is this ever done?</p>\n\n<p>I have been trying to train a model with sigmoid activation and cross entropy loss however none of my models train well at all. They just end up sitting with a steady high loss and zero accuracy.\nWhen I use softmax activation, they perform much better.</p>\n\n<p>However, for some of the challenge specific test time tricks I want to apply, I think I need a sigmoid activation so that I know the probability of each siRNA independently.</p>\n\n<p>So, is it possible for me to train the model with a softmax activation, however change it to sigmoid activation for inference? Is this ever done, will it work?</p>\n\n<p>Cheers,\nC</p>",
      "rawMarkdown": "Is this ever done?\n\nI have been trying to train a model with sigmoid activation and cross entropy loss however none of my models train well at all. They just end up sitting with a steady high loss and zero accuracy.\nWhen I use softmax activation, they perform much better.\n\nHowever, for some of the challenge specific test time tricks I want to apply, I think I need a sigmoid activation so that I know the probability of each siRNA independently.\n\nSo, is it possible for me to train the model with a softmax activation, however change it to sigmoid activation for inference? Is this ever done, will it work?\n\nCheers,\nC",
      "votes": null
    },
    {
      "id": "582264",
      "postDate": "07/23/2019 01:27:14",
      "content": "<p>Change the activation function is easy enough, I don't see why you can't change it in few lines of code and test it out.  </p>",
      "rawMarkdown": "Change the activation function is easy enough, I don't see why you can't change it in few lines of code and test it out.",
      "votes": null
    },
    {
      "id": "582267",
      "postDate": "07/23/2019 01:43:01",
      "content": "<p>Yup training some now to test it out.\nJust figured I would ask in case someone would jump in and say it wont work for some good reason (I haven't seen it be done in some extensive googling, so maybe there is a reason). \neg I am thinking maybe that since we typically want the model to converge to predict one thing very strongly and everything else really weakly - then when replacing softmax with sigmoid and expecting the result to reflect a confidence of each class independently, then all of the lower confidence classes will just be noise because they have been penalized really weakly during training.</p>",
      "rawMarkdown": "Yup training some now to test it out.\nJust figured I would ask in case someone would jump in and say it wont work for some good reason (I haven't seen it be done in some extensive googling, so maybe there is a reason). \neg I am thinking maybe that since we typically want the model to converge to predict one thing very strongly and everything else really weakly - then when replacing softmax with sigmoid and expecting the result to reflect a confidence of each class independently, then all of the lower confidence classes will just be noise because they have been penalized really weakly during training.",
      "votes": null
    },
    {
      "id": "582517",
      "postDate": "07/23/2019 09:16:22",
      "content": "<p>You probably don't have to. What's your idea with getting the probability of each siRNA independently?</p>\n\n<p>[multi-class] cross entropy loss generally assumes you are using softmax. Actually generally the implementations want their input in log space, just in case you're applying the activation by hand, please check your ml library doc, you might not have to do so.\nBinary cross entropy (BCE) is meant to use sigmoids, and is better to use a logit-space implementation to avoid numeric precision loss (which also doesn't require a sigmoid function before hand).</p>\n\n<p>Having said that, for both cases during inference you can compare the values before the non-linear activation, do you agree that:\nif     : x &gt; y\nthen: sigmoid(x) &gt; sigmoid(y) ?\nsame for exp(x) &gt; exp(y) as these functions are both monotonic.\nSo applying a sigmoid to the output won't change your predictions, right?</p>\n\n<p>For ensembling multiple model predictions, if you are using softmax you might need to run the activation before hand as different models would have different denominators: sum( [ exp(x) for x in preds ] )\nbut logits normally are added directly...</p>\n\n<p>Sorry if you already know all that (maybe I just didn't get your idea), but hopefully it helps someone else.</p>\n\n<p>Now learning slow might be another problem. Which optimiser are you using?</p>",
      "rawMarkdown": "You probably don't have to. What's your idea with getting the probability of each siRNA independently?\n\n[multi-class] cross entropy loss generally assumes you are using softmax. Actually generally the implementations want their input in log space, just in case you're applying the activation by hand, please check your ml library doc, you might not have to do so.\nBinary cross entropy (BCE) is meant to use sigmoids, and is better to use a logit-space implementation to avoid numeric precision loss (which also doesn't require a sigmoid function before hand).\n\nHaving said that, for both cases during inference you can compare the values before the non-linear activation, do you agree that:\nif     : x &gt; y\nthen: sigmoid(x) &gt; sigmoid(y) ?\nsame for exp(x) &gt; exp(y) as these functions are both monotonic.\nSo applying a sigmoid to the output won't change your predictions, right?\n\nFor ensembling multiple model predictions, if you are using softmax you might need to run the activation before hand as different models would have different denominators: sum( [ exp(x) for x in preds ] )\nbut logits normally are added directly...\n\nSorry if you already know all that (maybe I just didn't get your idea), but hopefully it helps someone else.\n\nNow learning slow might be another problem. Which optimiser are you using?",
      "votes": null
    },
    {
      "id": "583054",
      "postDate": "07/24/2019 00:47:46",
      "content": "<p>My idea to use each siRNA independently was to be able to average siRNA probabilities over each site and take the highest resulting probability; and also to leverage the observation that each plate only contains one of each siRNA - so when I predict a well with one siRNA more confidently than any other, then I can exclude it from any other well. So I would be kind of treating it as a multi-label problem.</p>\n\n<p>I am using Keras (tf backend) categorical_crossentropy, this has an input <code>from_logits</code> which is set to False. From reading the code, it looks like this will expect a tensor of probabilities (either sigmoid or softmax), divide those probabilities by the sum of the tensor (to turn it into a probability distribution, if it wasn't already) and then calculate cross entropy. So I 'think' this is doing what I want it to...</p>\n\n<p>My choice to use multiclass cross-entropy here is that I thought it would be best to try and train this as close to a single label problem as possible - since that is what it is. \nBut perhaps BCE may work better? penalizing each class independently..</p>\n\n<p>I certainly agree with all of your logic there for inference. However I was unsure how best to train the model so that looking at the logits without activation for the lower probability classes would actually be representative. I was concerned that since they are not penalized hard, the model would not really be trying to put the lower probability classes into the correct order and noise would dominate.</p>\n\n<p>Your final point - I think might be right on the money. Perhaps I have focused too early on how to set up my problem and my issue is hyper parameters and optimizers. At the time of writing my OP, I was using keras.optimizers.Adadelta(). However I have now changed to SGD with all default parameters and I am actually getting some training performance now with both activation functions. Problem is it is severely overfitting, training loss goes down but validation loss does not :( \nBut that is OK at least I am getting something, now I just need to pick an activation and loss and stick with it - so many choices now :(</p>",
      "rawMarkdown": "My idea to use each siRNA independently was to be able to average siRNA probabilities over each site and take the highest resulting probability; and also to leverage the observation that each plate only contains one of each siRNA - so when I predict a well with one siRNA more confidently than any other, then I can exclude it from any other well. So I would be kind of treating it as a multi-label problem.\n\nI am using Keras (tf backend) categorical_crossentropy, this has an input `from_logits` which is set to False. From reading the code, it looks like this will expect a tensor of probabilities (either sigmoid or softmax), divide those probabilities by the sum of the tensor (to turn it into a probability distribution, if it wasn't already) and then calculate cross entropy. So I 'think' this is doing what I want it to...\n\nMy choice to use multiclass cross-entropy here is that I thought it would be best to try and train this as close to a single label problem as possible - since that is what it is. \nBut perhaps BCE may work better? penalizing each class independently..\n\nI certainly agree with all of your logic there for inference. However I was unsure how best to train the model so that looking at the logits without activation for the lower probability classes would actually be representative. I was concerned that since they are not penalized hard, the model would not really be trying to put the lower probability classes into the correct order and noise would dominate.\n\nYour final point - I think might be right on the money. Perhaps I have focused too early on how to set up my problem and my issue is hyper parameters and optimizers. At the time of writing my OP, I was using keras.optimizers.Adadelta(). However I have now changed to SGD with all default parameters and I am actually getting some training performance now with both activation functions. Problem is it is severely overfitting, training loss goes down but validation loss does not :( \nBut that is OK at least I am getting something, now I just need to pick an activation and loss and stick with it - so many choices now :(",
      "votes": null
    },
    {
      "id": "583256",
      "postDate": "07/24/2019 08:29:21",
      "content": "<p>I still haven't been able to crack the secret of this competition yet, but just want to share funny experience.</p>\n\n<p>Accidentally, I mistook <code>binary_crossentropy</code> and <code>sigmoid</code> for the more appropriate  <code>categorical_crossentropy</code> and <code>softmax</code> which should be my bug ... However, the former version got a better LB score.</p>",
      "rawMarkdown": "I still haven't been able to crack the secret of this competition yet, but just want to share funny experience.\n\nAccidentally, I mistook `binary_crossentropy` and `sigmoid` for the more appropriate  `categorical_crossentropy` and `softmax` which should be my bug ... However, the former version got a better LB score.",
      "votes": null
    },
    {
      "id": "583263",
      "postDate": "07/24/2019 08:40:29",
      "content": "<p>Ha really, interesting. I just tried training a ResNet50 with BCE and sigmoid and it very, very quickly just learned to predict 0 for all classes. There are a couple other posts where they have success using BCE, although I don't know what activation - I assumed softmax, but I wonder how sigmoid can work.</p>",
      "rawMarkdown": "Ha really, interesting. I just tried training a ResNet50 with BCE and sigmoid and it very, very quickly just learned to predict 0 for all classes. There are a couple other posts where they have success using BCE, although I don't know what activation - I assumed softmax, but I wonder how sigmoid can work.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 582264,
      "author_name": "ryanzhang",
      "author_url": "",
      "post_date": "07/23/2019 01:27:14",
      "content": "<p>Change the activation function is easy enough, I don't see why you can't change it in few lines of code and test it out.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 582267,
          "author_name": "cherring",
          "author_url": "",
          "post_date": "07/23/2019 01:43:01",
          "content": "<p>Yup training some now to test it out.\nJust figured I would ask in case someone would jump in and say it wont work for some good reason (I haven't seen it be done in some extensive googling, so maybe there is a reason). \neg I am thinking maybe that since we typically want the model to converge to predict one thing very strongly and everything else really weakly - then when replacing softmax with sigmoid and expecting the result to reflect a confidence of each class independently, then all of the lower confidence classes will just be noise because they have been penalized really weakly during training.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 582517,
      "author_name": "hmendonca",
      "author_url": "",
      "post_date": "07/23/2019 09:16:22",
      "content": "<p>You probably don't have to. What's your idea with getting the probability of each siRNA independently?</p>\n\n<p>[multi-class] cross entropy loss generally assumes you are using softmax. Actually generally the implementations want their input in log space, just in case you're applying the activation by hand, please check your ml library doc, you might not have to do so.\nBinary cross entropy (BCE) is meant to use sigmoids, and is better to use a logit-space implementation to avoid numeric precision loss (which also doesn't require a sigmoid function before hand).</p>\n\n<p>Having said that, for both cases during inference you can compare the values before the non-linear activation, do you agree that:\nif     : x &gt; y\nthen: sigmoid(x) &gt; sigmoid(y) ?\nsame for exp(x) &gt; exp(y) as these functions are both monotonic.\nSo applying a sigmoid to the output won't change your predictions, right?</p>\n\n<p>For ensembling multiple model predictions, if you are using softmax you might need to run the activation before hand as different models would have different denominators: sum( [ exp(x) for x in preds ] )\nbut logits normally are added directly...</p>\n\n<p>Sorry if you already know all that (maybe I just didn't get your idea), but hopefully it helps someone else.</p>\n\n<p>Now learning slow might be another problem. Which optimiser are you using?</p>",
      "votes": null,
      "replies": [
        {
          "id": 583054,
          "author_name": "cherring",
          "author_url": "",
          "post_date": "07/24/2019 00:47:46",
          "content": "<p>My idea to use each siRNA independently was to be able to average siRNA probabilities over each site and take the highest resulting probability; and also to leverage the observation that each plate only contains one of each siRNA - so when I predict a well with one siRNA more confidently than any other, then I can exclude it from any other well. So I would be kind of treating it as a multi-label problem.</p>\n\n<p>I am using Keras (tf backend) categorical_crossentropy, this has an input <code>from_logits</code> which is set to False. From reading the code, it looks like this will expect a tensor of probabilities (either sigmoid or softmax), divide those probabilities by the sum of the tensor (to turn it into a probability distribution, if it wasn't already) and then calculate cross entropy. So I 'think' this is doing what I want it to...</p>\n\n<p>My choice to use multiclass cross-entropy here is that I thought it would be best to try and train this as close to a single label problem as possible - since that is what it is. \nBut perhaps BCE may work better? penalizing each class independently..</p>\n\n<p>I certainly agree with all of your logic there for inference. However I was unsure how best to train the model so that looking at the logits without activation for the lower probability classes would actually be representative. I was concerned that since they are not penalized hard, the model would not really be trying to put the lower probability classes into the correct order and noise would dominate.</p>\n\n<p>Your final point - I think might be right on the money. Perhaps I have focused too early on how to set up my problem and my issue is hyper parameters and optimizers. At the time of writing my OP, I was using keras.optimizers.Adadelta(). However I have now changed to SGD with all default parameters and I am actually getting some training performance now with both activation functions. Problem is it is severely overfitting, training loss goes down but validation loss does not :( \nBut that is OK at least I am getting something, now I just need to pick an activation and loss and stick with it - so many choices now :(</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 583256,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "07/24/2019 08:29:21",
      "content": "<p>I still haven't been able to crack the secret of this competition yet, but just want to share funny experience.</p>\n\n<p>Accidentally, I mistook <code>binary_crossentropy</code> and <code>sigmoid</code> for the more appropriate  <code>categorical_crossentropy</code> and <code>softmax</code> which should be my bug ... However, the former version got a better LB score.</p>",
      "votes": null,
      "replies": [
        {
          "id": 583263,
          "author_name": "cherring",
          "author_url": "",
          "post_date": "07/24/2019 08:40:29",
          "content": "<p>Ha really, interesting. I just tried training a ResNet50 with BCE and sigmoid and it very, very quickly just learned to predict 0 for all classes. There are a couple other posts where they have success using BCE, although I don't know what activation - I assumed softmax, but I wonder how sigmoid can work.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "582249": "Is this ever done?\n\nI have been trying to train a model with sigmoid activation and cross entropy loss however none of my models train well at all. They just end up sitting with a steady high loss and zero accuracy.\nWhen I use softmax activation, they perform much better.\n\nHowever, for some of the challenge specific test time tricks I want to apply, I think I need a sigmoid activation so that I know the probability of each siRNA independently.\n\nSo, is it possible for me to train the model with a softmax activation, however change it to sigmoid activation for inference? Is this ever done, will it work?\n\nCheers,\nC",
    "582264": "Change the activation function is easy enough, I don't see why you can't change it in few lines of code and test it out.",
    "582267": "Yup training some now to test it out.\nJust figured I would ask in case someone would jump in and say it wont work for some good reason (I haven't seen it be done in some extensive googling, so maybe there is a reason). \neg I am thinking maybe that since we typically want the model to converge to predict one thing very strongly and everything else really weakly - then when replacing softmax with sigmoid and expecting the result to reflect a confidence of each class independently, then all of the lower confidence classes will just be noise because they have been penalized really weakly during training.",
    "582517": "You probably don't have to. What's your idea with getting the probability of each siRNA independently?\n\n[multi-class] cross entropy loss generally assumes you are using softmax. Actually generally the implementations want their input in log space, just in case you're applying the activation by hand, please check your ml library doc, you might not have to do so.\nBinary cross entropy (BCE) is meant to use sigmoids, and is better to use a logit-space implementation to avoid numeric precision loss (which also doesn't require a sigmoid function before hand).\n\nHaving said that, for both cases during inference you can compare the values before the non-linear activation, do you agree that:\nif     : x &gt; y\nthen: sigmoid(x) &gt; sigmoid(y) ?\nsame for exp(x) &gt; exp(y) as these functions are both monotonic.\nSo applying a sigmoid to the output won't change your predictions, right?\n\nFor ensembling multiple model predictions, if you are using softmax you might need to run the activation before hand as different models would have different denominators: sum( [ exp(x) for x in preds ] )\nbut logits normally are added directly...\n\nSorry if you already know all that (maybe I just didn't get your idea), but hopefully it helps someone else.\n\nNow learning slow might be another problem. Which optimiser are you using?",
    "583054": "My idea to use each siRNA independently was to be able to average siRNA probabilities over each site and take the highest resulting probability; and also to leverage the observation that each plate only contains one of each siRNA - so when I predict a well with one siRNA more confidently than any other, then I can exclude it from any other well. So I would be kind of treating it as a multi-label problem.\n\nI am using Keras (tf backend) categorical_crossentropy, this has an input `from_logits` which is set to False. From reading the code, it looks like this will expect a tensor of probabilities (either sigmoid or softmax), divide those probabilities by the sum of the tensor (to turn it into a probability distribution, if it wasn't already) and then calculate cross entropy. So I 'think' this is doing what I want it to...\n\nMy choice to use multiclass cross-entropy here is that I thought it would be best to try and train this as close to a single label problem as possible - since that is what it is. \nBut perhaps BCE may work better? penalizing each class independently..\n\nI certainly agree with all of your logic there for inference. However I was unsure how best to train the model so that looking at the logits without activation for the lower probability classes would actually be representative. I was concerned that since they are not penalized hard, the model would not really be trying to put the lower probability classes into the correct order and noise would dominate.\n\nYour final point - I think might be right on the money. Perhaps I have focused too early on how to set up my problem and my issue is hyper parameters and optimizers. At the time of writing my OP, I was using keras.optimizers.Adadelta(). However I have now changed to SGD with all default parameters and I am actually getting some training performance now with both activation functions. Problem is it is severely overfitting, training loss goes down but validation loss does not :( \nBut that is OK at least I am getting something, now I just need to pick an activation and loss and stick with it - so many choices now :(",
    "583256": "I still haven't been able to crack the secret of this competition yet, but just want to share funny experience.\n\nAccidentally, I mistook `binary_crossentropy` and `sigmoid` for the more appropriate  `categorical_crossentropy` and `softmax` which should be my bug ... However, the former version got a better LB score.",
    "583263": "Ha really, interesting. I just tried training a ResNet50 with BCE and sigmoid and it very, very quickly just learned to predict 0 for all classes. There are a couple other posts where they have success using BCE, although I don't know what activation - I assumed softmax, but I wonder how sigmoid can work."
  },
  "source": "meta"
}