{
  "id": 170856,
  "title": "Output layer activation functions for training/predicting",
  "url": "/competitions/birdsong-recognition/discussion/170856",
  "author_name": "",
  "post_date": "2020-07-29T10:27:35.987004200Z",
  "votes": 3,
  "comment_count": 1,
  "views": 0,
  "content": "<p>The training data consists of audio files which belong to one class, however in the prediction audio files multiple birds can occur in a 5 second clip. Thus we are training with a single label, but predicting with multiple labels.\nIn some notebooks, i.e. <a href=\"https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast\">this</a> one, it looks like <em>softmax</em> is used for training and <em>sigmoid</em> for predicting. Would that be the best practice for this competition?</p>\n\n<p>In general, what activation function for the output layer should be used for training/predicting and is it advisable to use different activation functions for training and predicting?</p>\n\n<p>I am quite new in the data science world so any background information is also welcome 😃 </p>",
  "messages": [
    {
      "id": "950304",
      "postDate": "07/29/2020 10:27:35",
      "content": "<p>The training data consists of audio files which belong to one class, however in the prediction audio files multiple birds can occur in a 5 second clip. Thus we are training with a single label, but predicting with multiple labels.\nIn some notebooks, i.e. <a href=\"https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast\">this</a> one, it looks like <em>softmax</em> is used for training and <em>sigmoid</em> for predicting. Would that be the best practice for this competition?</p>\n\n<p>In general, what activation function for the output layer should be used for training/predicting and is it advisable to use different activation functions for training and predicting?</p>\n\n<p>I am quite new in the data science world so any background information is also welcome 😃 </p>",
      "rawMarkdown": "The training data consists of audio files which belong to one class, however in the prediction audio files multiple birds can occur in a 5 second clip. Thus we are training with a single label, but predicting with multiple labels.\nIn some notebooks, i.e. [this](https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast) one, it looks like *softmax* is used for training and *sigmoid* for predicting. Would that be the best practice for this competition?\n\nIn general, what activation function for the output layer should be used for training/predicting and is it advisable to use different activation functions for training and predicting?\n\nI am quite new in the data science world so any background information is also welcome 😃",
      "votes": null
    },
    {
      "id": "957368",
      "postDate": "08/04/2020 08:48:09",
      "content": "<p>Using Softmax in training and changing it to Sigmoid during inference is a bad idea (I tried doing it and realized the maths later),\nWhen you declare a softmax activation in training, you are asking the model to just maximize the score of one class, doesn't matter how much. \nWhile in Sigmoid, you force the model to give values close to 1 for TRUE classes. </p>\n\n<p>I hope you are getting the picture. </p>",
      "rawMarkdown": "Using Softmax in training and changing it to Sigmoid during inference is a bad idea (I tried doing it and realized the maths later),\nWhen you declare a softmax activation in training, you are asking the model to just maximize the score of one class, doesn't matter how much. \nWhile in Sigmoid, you force the model to give values close to 1 for TRUE classes. \n\nI hope you are getting the picture.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 957368,
      "author_name": "humblediscipulus",
      "author_url": "",
      "post_date": "08/04/2020 08:48:09",
      "content": "<p>Using Softmax in training and changing it to Sigmoid during inference is a bad idea (I tried doing it and realized the maths later),\nWhen you declare a softmax activation in training, you are asking the model to just maximize the score of one class, doesn't matter how much. \nWhile in Sigmoid, you force the model to give values close to 1 for TRUE classes. </p>\n\n<p>I hope you are getting the picture. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "950304": "The training data consists of audio files which belong to one class, however in the prediction audio files multiple birds can occur in a 5 second clip. Thus we are training with a single label, but predicting with multiple labels.\nIn some notebooks, i.e. [this](https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast) one, it looks like *softmax* is used for training and *sigmoid* for predicting. Would that be the best practice for this competition?\n\nIn general, what activation function for the output layer should be used for training/predicting and is it advisable to use different activation functions for training and predicting?\n\nI am quite new in the data science world so any background information is also welcome 😃",
    "957368": "Using Softmax in training and changing it to Sigmoid during inference is a bad idea (I tried doing it and realized the maths later),\nWhen you declare a softmax activation in training, you are asking the model to just maximize the score of one class, doesn't matter how much. \nWhile in Sigmoid, you force the model to give values close to 1 for TRUE classes. \n\nI hope you are getting the picture."
  },
  "source": "meta"
}