{
  "id": 89169,
  "title": "Last layer activation (Sigmoid or softmax)",
  "url": "/competitions/freesound-audio-tagging-2019/discussion/89169",
  "author_name": "",
  "post_date": "2019-04-11T23:46:25.061120100Z",
  "votes": 1,
  "comment_count": 10,
  "views": 0,
  "content": "<p>This competition is multi label but not so much. There is less than 10% with more than one label in the train set. I have better results with the softmax final classifier activation than with sigmoid. I am using binary cross entropy for the loss. What final activation are you using?</p>",
  "messages": [
    {
      "id": "514802",
      "postDate": "04/11/2019 23:46:25",
      "content": "<p>This competition is multi label but not so much. There is less than 10% with more than one label in the train set. I have better results with the softmax final classifier activation than with sigmoid. I am using binary cross entropy for the loss. What final activation are you using?</p>",
      "rawMarkdown": "This competition is multi label but not so much. There is less than 10% with more than one label in the train set. I have better results with the softmax final classifier activation than with sigmoid. I am using binary cross entropy for the loss. What final activation are you using?",
      "votes": null
    },
    {
      "id": "514804",
      "postDate": "04/11/2019 23:51:47",
      "content": "<p>I have tried the same (just flattening all the multi labels into single labeled samples, then apply conventional clf), and got similar a little less score. I'd like to follow the original purpose of this competition to seek for possibility of multi label classification, now trying multi-label clf only...</p>",
      "rawMarkdown": "I have tried the same (just flattening all the multi labels into single labeled samples, then apply conventional clf), and got similar a little less score. I'd like to follow the original purpose of this competition to seek for possibility of multi label classification, now trying multi-label clf only...",
      "votes": null
    },
    {
      "id": "514818",
      "postDate": "04/12/2019 00:13:39",
      "content": "<p>I am using sigmoid in final activation layer w/ BCE loss. When I used softmax I got a lower score.</p>",
      "rawMarkdown": "I am using sigmoid in final activation layer w/ BCE loss. When I used softmax I got a lower score.",
      "votes": null
    },
    {
      "id": "514891",
      "postDate": "04/12/2019 02:29:39",
      "content": "<p>Softmax with KL divergence as the loss.</p>",
      "rawMarkdown": "Softmax with KL divergence as the loss.",
      "votes": null
    },
    {
      "id": "515684",
      "postDate": "04/13/2019 00:30:50",
      "content": "<p>i do believe sigmoid and bce is the way to go but right now i have not been able to get a good score with the approach. I am getting better scores with categorical CE and softmax. I would keep experimenting. </p>",
      "rawMarkdown": "i do believe sigmoid and bce is the way to go but right now i have not been able to get a good score with the approach. I am getting better scores with categorical CE and softmax. I would keep experimenting.",
      "votes": null
    },
    {
      "id": "515689",
      "postDate": "04/13/2019 00:39:13",
      "content": "<p>It could be because  many of data are single labeled...</p>",
      "rawMarkdown": "It could be because  many of data are single labeled...",
      "votes": null
    },
    {
      "id": "535471",
      "postDate": "05/23/2019 01:52:03",
      "content": "<p>I'm also curious about this. Did you come to any conclusion?\nFrom what I've read so far, multilabel should be sigmoid with bce, but I've read other stackoverflow posts debating otherwise.\nHave you looked at linear activation and BCEwithLogits loss?\n<a href=\"https://www.kaggle.com/ratthachat/fat19-mixup-keras-on-preprocesseddata-lb632\">https://www.kaggle.com/ratthachat/fat19-mixup-keras-on-preprocesseddata-lb632</a></p>",
      "rawMarkdown": "I'm also curious about this. Did you come to any conclusion?\nFrom what I've read so far, multilabel should be sigmoid with bce, but I've read other stackoverflow posts debating otherwise.\nHave you looked at linear activation and BCEwithLogits loss?\nhttps://www.kaggle.com/ratthachat/fat19-mixup-keras-on-preprocesseddata-lb632",
      "votes": null
    },
    {
      "id": "535496",
      "postDate": "05/23/2019 03:15:42",
      "content": "<p>Are you asking to me? Any conclusion depends on each conditions.\nLet's see what performs better... :)</p>",
      "rawMarkdown": "Are you asking to me? Any conclusion depends on each conditions.\nLet's see what performs better... :)",
      "votes": null
    },
    {
      "id": "535531",
      "postDate": "05/23/2019 05:11:56",
      "content": "<p>Sigmoid + BCEloss is the same as linear + bcewithlogitloss. The only difference is the later is considered more numerically stable.</p>",
      "rawMarkdown": "Sigmoid + BCEloss is the same as linear + bcewithlogitloss. The only difference is the later is considered more numerically stable.",
      "votes": null
    },
    {
      "id": "535687",
      "postDate": "05/23/2019 10:49:05",
      "content": "<p>How do you interpret the output of linear activation and BCEwithLogits loss? The output range is very wide. Do you just set a manual threshold and just clip them to 0 and 1?</p>",
      "rawMarkdown": "How do you interpret the output of linear activation and BCEwithLogits loss? The output range is very wide. Do you just set a manual threshold and just clip them to 0 and 1?",
      "votes": null
    },
    {
      "id": "535692",
      "postDate": "05/23/2019 10:55:33",
      "content": "<p>Even though I've read that Sigmoid+BCE is the way to tackle multilabel problems, softmax+CcE gives me most consistent and best result for some reason. Also my output score and LB score is within 1% difference. I just started deep learning, so there are so many things I'm trying to understand at the moment. :)</p>",
      "rawMarkdown": "Even though I've read that Sigmoid+BCE is the way to tackle multilabel problems, softmax+CcE gives me most consistent and best result for some reason. Also my output score and LB score is within 1% difference. I just started deep learning, so there are so many things I'm trying to understand at the moment. :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 514804,
      "author_name": "daisukelab",
      "author_url": "",
      "post_date": "04/11/2019 23:51:47",
      "content": "<p>I have tried the same (just flattening all the multi labels into single labeled samples, then apply conventional clf), and got similar a little less score. I'd like to follow the original purpose of this competition to seek for possibility of multi label classification, now trying multi-label clf only...</p>",
      "votes": null,
      "replies": [
        {
          "id": 535471,
          "author_name": "chikim",
          "author_url": "",
          "post_date": "05/23/2019 01:52:03",
          "content": "<p>I'm also curious about this. Did you come to any conclusion?\nFrom what I've read so far, multilabel should be sigmoid with bce, but I've read other stackoverflow posts debating otherwise.\nHave you looked at linear activation and BCEwithLogits loss?\n<a href=\"https://www.kaggle.com/ratthachat/fat19-mixup-keras-on-preprocesseddata-lb632\">https://www.kaggle.com/ratthachat/fat19-mixup-keras-on-preprocesseddata-lb632</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 535496,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "05/23/2019 03:15:42",
          "content": "<p>Are you asking to me? Any conclusion depends on each conditions.\nLet's see what performs better... :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 535531,
          "author_name": "ebouteillon",
          "author_url": "",
          "post_date": "05/23/2019 05:11:56",
          "content": "<p>Sigmoid + BCEloss is the same as linear + bcewithlogitloss. The only difference is the later is considered more numerically stable.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 535687,
          "author_name": "chikim",
          "author_url": "",
          "post_date": "05/23/2019 10:49:05",
          "content": "<p>How do you interpret the output of linear activation and BCEwithLogits loss? The output range is very wide. Do you just set a manual threshold and just clip them to 0 and 1?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 535692,
          "author_name": "chikim",
          "author_url": "",
          "post_date": "05/23/2019 10:55:33",
          "content": "<p>Even though I've read that Sigmoid+BCE is the way to tackle multilabel problems, softmax+CcE gives me most consistent and best result for some reason. Also my output score and LB score is within 1% difference. I just started deep learning, so there are so many things I'm trying to understand at the moment. :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 514818,
      "author_name": "jamesrequa",
      "author_url": "",
      "post_date": "04/12/2019 00:13:39",
      "content": "<p>I am using sigmoid in final activation layer w/ BCE loss. When I used softmax I got a lower score.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 514891,
      "author_name": "andris",
      "author_url": "",
      "post_date": "04/12/2019 02:29:39",
      "content": "<p>Softmax with KL divergence as the loss.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 515684,
      "author_name": "ahmedalesh",
      "author_url": "",
      "post_date": "04/13/2019 00:30:50",
      "content": "<p>i do believe sigmoid and bce is the way to go but right now i have not been able to get a good score with the approach. I am getting better scores with categorical CE and softmax. I would keep experimenting. </p>",
      "votes": null,
      "replies": [
        {
          "id": 515689,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "04/13/2019 00:39:13",
          "content": "<p>It could be because  many of data are single labeled...</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "514802": "This competition is multi label but not so much. There is less than 10% with more than one label in the train set. I have better results with the softmax final classifier activation than with sigmoid. I am using binary cross entropy for the loss. What final activation are you using?",
    "514804": "I have tried the same (just flattening all the multi labels into single labeled samples, then apply conventional clf), and got similar a little less score. I'd like to follow the original purpose of this competition to seek for possibility of multi label classification, now trying multi-label clf only...",
    "514818": "I am using sigmoid in final activation layer w/ BCE loss. When I used softmax I got a lower score.",
    "514891": "Softmax with KL divergence as the loss.",
    "515684": "i do believe sigmoid and bce is the way to go but right now i have not been able to get a good score with the approach. I am getting better scores with categorical CE and softmax. I would keep experimenting.",
    "515689": "It could be because  many of data are single labeled...",
    "535471": "I'm also curious about this. Did you come to any conclusion?\nFrom what I've read so far, multilabel should be sigmoid with bce, but I've read other stackoverflow posts debating otherwise.\nHave you looked at linear activation and BCEwithLogits loss?\nhttps://www.kaggle.com/ratthachat/fat19-mixup-keras-on-preprocesseddata-lb632",
    "535496": "Are you asking to me? Any conclusion depends on each conditions.\nLet's see what performs better... :)",
    "535531": "Sigmoid + BCEloss is the same as linear + bcewithlogitloss. The only difference is the later is considered more numerically stable.",
    "535687": "How do you interpret the output of linear activation and BCEwithLogits loss? The output range is very wide. Do you just set a manual threshold and just clip them to 0 and 1?",
    "535692": "Even though I've read that Sigmoid+BCE is the way to tackle multilabel problems, softmax+CcE gives me most consistent and best result for some reason. Also my output score and LB score is within 1% difference. I just started deep learning, so there are so many things I'm trying to understand at the moment. :)"
  },
  "source": "meta"
}