{
  "id": 171471,
  "title": "Should we use sigmoid or softmax for inference? We trained with Softmax but use sigmoid for inference???",
  "url": "/competitions/birdsong-recognition/discussion/171471",
  "author_name": "",
  "post_date": "2020-07-31T22:48:36.254876Z",
  "votes": 6,
  "comment_count": 3,
  "views": 0,
  "content": "<p>The last activation on our models when we train is Softmax. This is because we train with just 1 label. However, it is possible for the test data to have multiple birds, so I see some kernels which use softmax while training, but use sigmoid for inference.</p>\n\n<p>This is the list of kernels that were trained with Softmax, but make inference with Sigmoid:\n1. <a href=\"https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast\">https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast</a>\n2. <a href=\"https://www.kaggle.com/hidehisaarai1213/inference-pytorch-birdcall-resnet-baseline\">https://www.kaggle.com/hidehisaarai1213/inference-pytorch-birdcall-resnet-baseline</a> (model returns a softmax and sigmoid; softmax was used for training and sigmoid is using for inference)\n3. <a href=\"https://www.kaggle.com/radek1/esp-starter-pack-v3-res34-minmax\">https://www.kaggle.com/radek1/esp-starter-pack-v3-res34-minmax</a> (this uses <code>lme_pool</code> but at its core it is still 1/(1+exp(-x))</p>\n\n<p>Maybe somebody who is experienced with multilabel problems can tell me if this is allowed? It feels kind of dangerous to use sigmoid output if we never trained with it. </p>",
  "messages": [
    {
      "id": "953595",
      "postDate": "07/31/2020 22:48:36",
      "content": "<p>The last activation on our models when we train is Softmax. This is because we train with just 1 label. However, it is possible for the test data to have multiple birds, so I see some kernels which use softmax while training, but use sigmoid for inference.</p>\n\n<p>This is the list of kernels that were trained with Softmax, but make inference with Sigmoid:\n1. <a href=\"https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast\">https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast</a>\n2. <a href=\"https://www.kaggle.com/hidehisaarai1213/inference-pytorch-birdcall-resnet-baseline\">https://www.kaggle.com/hidehisaarai1213/inference-pytorch-birdcall-resnet-baseline</a> (model returns a softmax and sigmoid; softmax was used for training and sigmoid is using for inference)\n3. <a href=\"https://www.kaggle.com/radek1/esp-starter-pack-v3-res34-minmax\">https://www.kaggle.com/radek1/esp-starter-pack-v3-res34-minmax</a> (this uses <code>lme_pool</code> but at its core it is still 1/(1+exp(-x))</p>\n\n<p>Maybe somebody who is experienced with multilabel problems can tell me if this is allowed? It feels kind of dangerous to use sigmoid output if we never trained with it. </p>",
      "rawMarkdown": "The last activation on our models when we train is Softmax. This is because we train with just 1 label. However, it is possible for the test data to have multiple birds, so I see some kernels which use softmax while training, but use sigmoid for inference.\n\nThis is the list of kernels that were trained with Softmax, but make inference with Sigmoid:\n1. https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast\n2. https://www.kaggle.com/hidehisaarai1213/inference-pytorch-birdcall-resnet-baseline (model returns a softmax and sigmoid; softmax was used for training and sigmoid is using for inference)\n3. https://www.kaggle.com/radek1/esp-starter-pack-v3-res34-minmax (this uses `lme_pool` but at its core it is still 1/(1+exp(-x))\n\nMaybe somebody who is experienced with multilabel problems can tell me if this is allowed? It feels kind of dangerous to use sigmoid output if we never trained with it.",
      "votes": null
    },
    {
      "id": "953639",
      "postDate": "08/01/2020 00:12:07",
      "content": "<p>For 2, I used sigmoid for training. I'm not sure whether it is good to train with softmax and infer with sigmoid.</p>",
      "rawMarkdown": "For 2, I used sigmoid for training. I'm not sure whether it is good to train with softmax and infer with sigmoid.",
      "votes": null
    },
    {
      "id": "955741",
      "postDate": "08/02/2020 22:02:14",
      "content": "<p>Softmax slightly increases the weighting of the highest output. Sigmoid is probably preferable because it makes finding a reasonable threshold much easier. softmax will slightly change the calibration on each prediction based on distribution of other classes which is probably not what we want at test time. Softmax might be helpful to converge during training though. I think the method of training with softmax and swapping with sigmoid at the end is reasonable. </p>",
      "rawMarkdown": "Softmax slightly increases the weighting of the highest output. Sigmoid is probably preferable because it makes finding a reasonable threshold much easier. softmax will slightly change the calibration on each prediction based on distribution of other classes which is probably not what we want at test time. Softmax might be helpful to converge during training though. I think the method of training with softmax and swapping with sigmoid at the end is reasonable.",
      "votes": null
    },
    {
      "id": "958109",
      "postDate": "08/04/2020 19:18:25",
      "content": "<p><code>TIL</code></p>",
      "rawMarkdown": "`TIL`",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 953639,
      "author_name": "hidehisaarai1213",
      "author_url": "",
      "post_date": "08/01/2020 00:12:07",
      "content": "<p>For 2, I used sigmoid for training. I'm not sure whether it is good to train with softmax and infer with sigmoid.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 955741,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "08/02/2020 22:02:14",
      "content": "<p>Softmax slightly increases the weighting of the highest output. Sigmoid is probably preferable because it makes finding a reasonable threshold much easier. softmax will slightly change the calibration on each prediction based on distribution of other classes which is probably not what we want at test time. Softmax might be helpful to converge during training though. I think the method of training with softmax and swapping with sigmoid at the end is reasonable. </p>",
      "votes": null,
      "replies": [
        {
          "id": 958109,
          "author_name": "authman",
          "author_url": "",
          "post_date": "08/04/2020 19:18:25",
          "content": "<p><code>TIL</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "953595": "The last activation on our models when we train is Softmax. This is because we train with just 1 label. However, it is possible for the test data to have multiple birds, so I see some kernels which use softmax while training, but use sigmoid for inference.\n\nThis is the list of kernels that were trained with Softmax, but make inference with Sigmoid:\n1. https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast\n2. https://www.kaggle.com/hidehisaarai1213/inference-pytorch-birdcall-resnet-baseline (model returns a softmax and sigmoid; softmax was used for training and sigmoid is using for inference)\n3. https://www.kaggle.com/radek1/esp-starter-pack-v3-res34-minmax (this uses `lme_pool` but at its core it is still 1/(1+exp(-x))\n\nMaybe somebody who is experienced with multilabel problems can tell me if this is allowed? It feels kind of dangerous to use sigmoid output if we never trained with it.",
    "953639": "For 2, I used sigmoid for training. I'm not sure whether it is good to train with softmax and infer with sigmoid.",
    "955741": "Softmax slightly increases the weighting of the highest output. Sigmoid is probably preferable because it makes finding a reasonable threshold much easier. softmax will slightly change the calibration on each prediction based on distribution of other classes which is probably not what we want at test time. Softmax might be helpful to converge during training though. I think the method of training with softmax and swapping with sigmoid at the end is reasonable.",
    "958109": "`TIL`"
  },
  "source": "meta"
}