{
  "id": 500256,
  "title": "Why would multiclass produce a better LB score than multilabel?",
  "url": "/competitions/birdclef-2024/discussion/500256",
  "author_name": "",
  "post_date": "2024-05-04T20:46:37.661868500Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>My highest ranking LB submission was my 26th one using softmax 0.64. The problem is multilabel, but I've yet to be able to replicate my multiclass LB with mutlilabel losses (focal, binary cross entropy). My best scoring LB was submission 26 and I'm now on 130 so I've tried a lot of experiments.</p>\n<p>I'm having trouble understanding why this would be the case. Does it sound like I have something fundamentally wrong in my code?</p>",
  "messages": [
    {
      "id": "2793644",
      "postDate": "05/04/2024 20:46:37",
      "content": "<p>My highest ranking LB submission was my 26th one using softmax 0.64. The problem is multilabel, but I've yet to be able to replicate my multiclass LB with mutlilabel losses (focal, binary cross entropy). My best scoring LB was submission 26 and I'm now on 130 so I've tried a lot of experiments.</p>\n<p>I'm having trouble understanding why this would be the case. Does it sound like I have something fundamentally wrong in my code?</p>",
      "rawMarkdown": "My highest ranking LB submission was my 26th one using softmax 0.64. The problem is multilabel, but I've yet to be able to replicate my multiclass LB with mutlilabel losses (focal, binary cross entropy). My best scoring LB was submission 26 and I'm now on 130 so I've tried a lot of experiments.\n\n\nI'm having trouble understanding why this would be the case. Does it sound like I have something fundamentally wrong in my code?",
      "votes": null
    },
    {
      "id": "2794302",
      "postDate": "05/05/2024 08:00:30",
      "content": "<p>As far as I know this competition have very unreliable LB, because soundscapes are from sources the model have never seen, so ~0.01 could be just the randomness of data (my exact same model with different random seed can have score diff of 0.03). However, using sigmoid rather than softmax pretty much always gave me a higher score for my codes. For me tuning up mixup helped because, like you said, it's a multi label problem but the training data are mostly single label.</p>",
      "rawMarkdown": "As far as I know this competition have very unreliable LB, because soundscapes are from sources the model have never seen, so ~0.01 could be just the randomness of data (my exact same model with different random seed can have score diff of 0.03). However, using sigmoid rather than softmax pretty much always gave me a higher score for my codes. For me tuning up mixup helped because, like you said, it's a multi label problem but the training data are mostly single label.",
      "votes": null
    },
    {
      "id": "2797589",
      "postDate": "05/06/2024 19:45:59",
      "content": "<p>I found a bug where I wasn't actually using secondary labels when I thought I was so that could be part of the problem.</p>",
      "rawMarkdown": "I found a bug where I wasn't actually using secondary labels when I thought I was so that could be part of the problem.",
      "votes": null
    },
    {
      "id": "2805842",
      "postDate": "05/10/2024 19:00:31",
      "content": "<p>Also going to put out here that I had a bug in my submission script where I wasn't using all of the audio, but 48 overlapping chunks.. So I think the multiclass loss inadvertently allowed me to score better. I've only been able to do one submission since the fix and the focal loss matched my previous best multiclass loss with the same model configuration.</p>",
      "rawMarkdown": "Also going to put out here that I had a bug in my submission script where I wasn't using all of the audio, but 48 overlapping chunks.. So I think the multiclass loss inadvertently allowed me to score better. I've only been able to do one submission since the fix and the focal loss matched my previous best multiclass loss with the same model configuration.",
      "votes": null
    },
    {
      "id": "2824646",
      "postDate": "05/20/2024 00:19:08",
      "content": "<p>In my case, both sigmoid and softmax produces 0.66. The reason of the effectiveness of softmax is maybe 1 or 2:</p>\n<ol>\n<li><p>the model trained with usual Cross Entropy has an ability to catch the primary label with a high confidence at the cost of other labels. </p></li>\n<li><p>a particular label has a correlation with other labels, such that not only primary labels but also another label is captured by training on the cross entropy loss that uses softmax. This maybe lead to the ability to predict multi-labels.</p></li>\n</ol>\n<p>I’m sorry if I’m missing something.</p>",
      "rawMarkdown": "In my case, both sigmoid and softmax produces 0.66. The reason of the effectiveness of softmax is maybe 1 or 2:\n\n1. the model trained with usual Cross Entropy has an ability to catch the primary label with a high confidence at the cost of other labels. \n\n2. a particular label has a correlation with other labels, such that not only primary labels but also another label is captured by training on the cross entropy loss that uses softmax. This maybe lead to the ability to predict multi-labels.\n\nI’m sorry if I’m missing something.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2794302,
      "author_name": "llleeeoooh",
      "author_url": "",
      "post_date": "05/05/2024 08:00:30",
      "content": "<p>As far as I know this competition have very unreliable LB, because soundscapes are from sources the model have never seen, so ~0.01 could be just the randomness of data (my exact same model with different random seed can have score diff of 0.03). However, using sigmoid rather than softmax pretty much always gave me a higher score for my codes. For me tuning up mixup helped because, like you said, it's a multi label problem but the training data are mostly single label.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2797589,
      "author_name": "willrice",
      "author_url": "",
      "post_date": "05/06/2024 19:45:59",
      "content": "<p>I found a bug where I wasn't actually using secondary labels when I thought I was so that could be part of the problem.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2805842,
      "author_name": "willrice",
      "author_url": "",
      "post_date": "05/10/2024 19:00:31",
      "content": "<p>Also going to put out here that I had a bug in my submission script where I wasn't using all of the audio, but 48 overlapping chunks.. So I think the multiclass loss inadvertently allowed me to score better. I've only been able to do one submission since the fix and the focal loss matched my previous best multiclass loss with the same model configuration.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2824646,
      "author_name": "johnlemon3",
      "author_url": "",
      "post_date": "05/20/2024 00:19:08",
      "content": "<p>In my case, both sigmoid and softmax produces 0.66. The reason of the effectiveness of softmax is maybe 1 or 2:</p>\n<ol>\n<li><p>the model trained with usual Cross Entropy has an ability to catch the primary label with a high confidence at the cost of other labels. </p></li>\n<li><p>a particular label has a correlation with other labels, such that not only primary labels but also another label is captured by training on the cross entropy loss that uses softmax. This maybe lead to the ability to predict multi-labels.</p></li>\n</ol>\n<p>I’m sorry if I’m missing something.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2793644": "My highest ranking LB submission was my 26th one using softmax 0.64. The problem is multilabel, but I've yet to be able to replicate my multiclass LB with mutlilabel losses (focal, binary cross entropy). My best scoring LB was submission 26 and I'm now on 130 so I've tried a lot of experiments.\n\n\nI'm having trouble understanding why this would be the case. Does it sound like I have something fundamentally wrong in my code?",
    "2794302": "As far as I know this competition have very unreliable LB, because soundscapes are from sources the model have never seen, so ~0.01 could be just the randomness of data (my exact same model with different random seed can have score diff of 0.03). However, using sigmoid rather than softmax pretty much always gave me a higher score for my codes. For me tuning up mixup helped because, like you said, it's a multi label problem but the training data are mostly single label.",
    "2797589": "I found a bug where I wasn't actually using secondary labels when I thought I was so that could be part of the problem.",
    "2805842": "Also going to put out here that I had a bug in my submission script where I wasn't using all of the audio, but 48 overlapping chunks.. So I think the multiclass loss inadvertently allowed me to score better. I've only been able to do one submission since the fix and the focal loss matched my previous best multiclass loss with the same model configuration.",
    "2824646": "In my case, both sigmoid and softmax produces 0.66. The reason of the effectiveness of softmax is maybe 1 or 2:\n\n1. the model trained with usual Cross Entropy has an ability to catch the primary label with a high confidence at the cost of other labels. \n\n2. a particular label has a correlation with other labels, such that not only primary labels but also another label is captured by training on the cross entropy loss that uses softmax. This maybe lead to the ability to predict multi-labels.\n\nI’m sorry if I’m missing something."
  },
  "source": "meta"
}