{
  "id": 240559,
  "title": "Bi-Tempered Logistic Loss",
  "url": "/competitions/birdclef-2021/discussion/240559",
  "author_name": "",
  "post_date": "2021-05-20T10:47:16.327288300Z",
  "votes": 14,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Maybe this loss function can be of use here, anyone tried it with success or can rule it out?</p>\n<p>The idea is that, when during training, a random ~5s clip is picked, that sound may not represent the label call well, it become mislabeled in the training. The Bi-Tempered Logistic Loss is created to handle this kind of problem.</p>\n<p><a href=\"https://ai.googleblog.com/2019/08/bi-tempered-logistic-loss-for-training.html\" target=\"_blank\">https://ai.googleblog.com/2019/08/bi-tempered-logistic-loss-for-training.html</a><br>\n<a href=\"https://arxiv.org/pdf/1906.03361.pdf\" target=\"_blank\">https://arxiv.org/pdf/1906.03361.pdf</a></p>\n<p>Update I: <br>\nAfter a full train run - it works better(F1 validation score) than BCE(WL), and that with only a first test fold run without tuning the loss settings.</p>\n<p>Update II:<br>\nDid a new run, with less randomness, given that mostly of the sounds are cut in random, so kept that same during both testing.<br>\nHad better validation score again and also a better LB score vs BCE loss, a small gain though but still, the val score also increases with some tuning of parameter settings, so maybe there are more gain to find. Or maybe this is all in the margin of error and just luck, but this is what I found within this loss.<br>\nI will keep the BiTLoss for now in my model training, maybe ensemble it with others, time will tell. Too much tuning, testing and training in other parts of this competition, so I'll leave further loss tuning for now :)</p>",
  "messages": [
    {
      "id": "1316157",
      "postDate": "05/20/2021 10:47:16",
      "content": "<p>Maybe this loss function can be of use here, anyone tried it with success or can rule it out?</p>\n<p>The idea is that, when during training, a random ~5s clip is picked, that sound may not represent the label call well, it become mislabeled in the training. The Bi-Tempered Logistic Loss is created to handle this kind of problem.</p>\n<p><a href=\"https://ai.googleblog.com/2019/08/bi-tempered-logistic-loss-for-training.html\" target=\"_blank\">https://ai.googleblog.com/2019/08/bi-tempered-logistic-loss-for-training.html</a><br>\n<a href=\"https://arxiv.org/pdf/1906.03361.pdf\" target=\"_blank\">https://arxiv.org/pdf/1906.03361.pdf</a></p>\n<p>Update I: <br>\nAfter a full train run - it works better(F1 validation score) than BCE(WL), and that with only a first test fold run without tuning the loss settings.</p>\n<p>Update II:<br>\nDid a new run, with less randomness, given that mostly of the sounds are cut in random, so kept that same during both testing.<br>\nHad better validation score again and also a better LB score vs BCE loss, a small gain though but still, the val score also increases with some tuning of parameter settings, so maybe there are more gain to find. Or maybe this is all in the margin of error and just luck, but this is what I found within this loss.<br>\nI will keep the BiTLoss for now in my model training, maybe ensemble it with others, time will tell. Too much tuning, testing and training in other parts of this competition, so I'll leave further loss tuning for now :)</p>",
      "rawMarkdown": "Maybe this loss function can be of use here, anyone tried it with success or can rule it out?\n\nThe idea is that, when during training, a random ~5s clip is picked, that sound may not represent the label call well, it become mislabeled in the training. The Bi-Tempered Logistic Loss is created to handle this kind of problem.\n\nhttps://ai.googleblog.com/2019/08/bi-tempered-logistic-loss-for-training.html\nhttps://arxiv.org/pdf/1906.03361.pdf\n\nUpdate I: \nAfter a full train run - it works better(F1 validation score) than BCE(WL), and that with only a first test fold run without tuning the loss settings.\n\nUpdate II:\nDid a new run, with less randomness, given that mostly of the sounds are cut in random, so kept that same during both testing.\nHad better validation score again and also a better LB score vs BCE loss, a small gain though but still, the val score also increases with some tuning of parameter settings, so maybe there are more gain to find. Or maybe this is all in the margin of error and just luck, but this is what I found within this loss.\nI will keep the BiTLoss for now in my model training, maybe ensemble it with others, time will tell. Too much tuning, testing and training in other parts of this competition, so I'll leave further loss tuning for now :)",
      "votes": null
    },
    {
      "id": "1316645",
      "postDate": "05/20/2021 18:32:52",
      "content": "<p>In other competition, I used the Bi-Tempered Logistic Loss function.<br>\nPlease use it if necessary. :)</p>\n<p><a href=\"https://www.kaggle.com/piantic/train-cassava-starter-using-various-loss-funcs?scriptVersionId=52092486&amp;cellId=37\" target=\"_blank\">link</a></p>",
      "rawMarkdown": "In other competition, I used the Bi-Tempered Logistic Loss function.\nPlease use it if necessary. :)\n\n[link](https://www.kaggle.com/piantic/train-cassava-starter-using-various-loss-funcs?scriptVersionId=52092486&cellId=37)",
      "votes": null
    },
    {
      "id": "1316660",
      "postDate": "05/20/2021 18:54:51",
      "content": "<p>Thanks. Yes I also had the same version saved from the same competition :) trying a custom version now for a fold run, with different settings, see how it goes.</p>",
      "rawMarkdown": "Thanks. Yes I also had the same version saved from the same competition :) trying a custom version now for a fold run, with different settings, see how it goes.",
      "votes": null
    },
    {
      "id": "1316667",
      "postDate": "05/20/2021 19:02:46",
      "content": "<p>I hadn't seen this before, thanks! Just quickly skimming so far… The paper seems focused on softmax loss, but we're generally allowed to have multiple birds in a segment (so, treating each label as an independent binary classification problem is usually seen as preferable to softmax'ing). Do you make any modifications for allowing multiple labels?</p>",
      "rawMarkdown": "I hadn't seen this before, thanks! Just quickly skimming so far... The paper seems focused on softmax loss, but we're generally allowed to have multiple birds in a segment (so, treating each label as an independent binary classification problem is usually seen as preferable to softmax'ing). Do you make any modifications for allowing multiple labels?",
      "votes": null
    },
    {
      "id": "1317620",
      "postDate": "05/21/2021 14:52:49",
      "content": "<p>Yes it's a great loss function in general.</p>\n<p>It’s just an idea to the problem with the mislabeled croped samples and maybe find a solution for it, which I wanted to share the community and take feedback on, if some already had tested this or other similar techniques. 🙂 This loss function was just a quick idea that crossed my mind, there are many losses out there for handle mislabeled images to test on.</p>\n<p>I'm still testing it and I'm passing it through a sigmoid part for the probabilities instead of directly using the softmax part of it, and with some other changes to the default settings - it works, if it's better than say bce, the full run will tell.</p>",
      "rawMarkdown": "Yes it's a great loss function in general.\n\nIt’s just an idea to the problem with the mislabeled croped samples and maybe find a solution for it, which I wanted to share the community and take feedback on, if some already had tested this or other similar techniques. 🙂 This loss function was just a quick idea that crossed my mind, there are many losses out there for handle mislabeled images to test on.\n\nI'm still testing it and I'm passing it through a sigmoid part for the probabilities instead of directly using the softmax part of it, and with some other changes to the default settings - it works, if it's better than say bce, the full run will tell.",
      "votes": null
    },
    {
      "id": "1318451",
      "postDate": "05/22/2021 10:45:35",
      "content": "<p>Update: After a full train run - it works better(F1 validation score) than BCE(WL), and that with only a first test fold run without tuning the loss settings.</p>",
      "rawMarkdown": "Update: After a full train run - it works better(F1 validation score) than BCE(WL), and that with only a first test fold run without tuning the loss settings.",
      "votes": null
    },
    {
      "id": "1319222",
      "postDate": "05/23/2021 03:20:30",
      "content": "<p>great works,thanks</p>",
      "rawMarkdown": "great works,thanks",
      "votes": null
    },
    {
      "id": "1319397",
      "postDate": "05/23/2021 06:45:36",
      "content": "<p><a href=\"https://www.kaggle.com/kirderf\" target=\"_blank\">@kirderf</a>  thanks  you get a better lB  also with it ? </p>",
      "rawMarkdown": "kirderf  thanks  you get a better lB  also with it ?",
      "votes": null
    },
    {
      "id": "1319771",
      "postDate": "05/23/2021 13:39:49",
      "content": "<p><a href=\"https://www.kaggle.com/kirderf\" target=\"_blank\">@kirderf</a> thanks, good idea. what did you use for loss settings for the first run t1 and t2?</p>",
      "rawMarkdown": "kirderf thanks, good idea. what did you use for loss settings for the first run t1 and t2?",
      "votes": null
    },
    {
      "id": "1321600",
      "postDate": "05/24/2021 19:35:06",
      "content": "<p>Yes slightly better LB score as well</p>",
      "rawMarkdown": "Yes slightly better LB score as well",
      "votes": null
    },
    {
      "id": "1321603",
      "postDate": "05/24/2021 19:48:52",
      "content": "<p>Update II: <br>\nDid a new run, with less randomness, given that mostly of the sounds are cut in random, so kept that same during both testing. <br>\nHad better validation score again and also a better LB score vs BCE loss, a small gain though but still, the val score also increases with some tuning of parameter settings, so maybe there are more gain to find. Or maybe this is all in the margin of error and just luck, but this is what I found within this loss. <br>\nI will keep the BiTLoss for now in my model training, maybe ensemble it with others, time will tell. Too much tuning, testing and training in other parts of this competition, so I'll leave further loss tuning for now :)</p>",
      "rawMarkdown": "Update II: \nDid a new run, with less randomness, given that mostly of the sounds are cut in random, so kept that same during both testing. \nHad better validation score again and also a better LB score vs BCE loss, a small gain though but still, the val score also increases with some tuning of parameter settings, so maybe there are more gain to find. Or maybe this is all in the margin of error and just luck, but this is what I found within this loss. \nI will keep the BiTLoss for now in my model training, maybe ensemble it with others, time will tell. Too much tuning, testing and training in other parts of this competition, so I'll leave further loss tuning for now :)",
      "votes": null
    },
    {
      "id": "1322800",
      "postDate": "05/25/2021 17:29:59",
      "content": "<p>Thanks for looking into this, Kerderf!</p>\n<p>I've got a number of 'good heuristic' methods for getting much-better-than-random training segments from the Xeno-Canto recordings. It seems like one would want to do this kind of loss tuning on top of easy heuristics, to reduce the amount+impact of 'bad' training segments in the first place.</p>\n<p>I also found this interesting paper on label smoothing with weak labels, after wondering why there wasn't some comparison to label smoothing in the bitempered loss paper: <br>\n<a href=\"http://www.sanjivk.com/LabelSmoothing_ICML20.pdf\" target=\"_blank\">http://www.sanjivk.com/LabelSmoothing_ICML20.pdf</a></p>",
      "rawMarkdown": "Thanks for looking into this, Kerderf!\n\nI've got a number of 'good heuristic' methods for getting much-better-than-random training segments from the Xeno-Canto recordings. It seems like one would want to do this kind of loss tuning on top of easy heuristics, to reduce the amount+impact of 'bad' training segments in the first place.\n\nI also found this interesting paper on label smoothing with weak labels, after wondering why there wasn't some comparison to label smoothing in the bitempered loss paper: \nhttp://www.sanjivk.com/LabelSmoothing_ICML20.pdf",
      "votes": null
    },
    {
      "id": "1323378",
      "postDate": "05/26/2021 07:25:23",
      "content": "<p>Thanks 👍 Yes I'm also using label smoothing in the BiTLoss. Also testing other methods like majority/full mixup/cutmix training with extra 'no call/noisy label' to help(?) the classifing, and then find a good PP to the prob..Have seen in a paper that fmix and cutmix is slighly better than mixup, link below.  Though I'm only using cutmix/mixup so far. Best is of course is to try only segment true calls before training.<br>\nTo small variations between the result to be groundbreaking but maybe in a final ensembling it can help. Just ideas, fun to evaluate …🙂</p>\n<blockquote>\n  <p>B. Audio Classification<br>\n  We now perform experiments on the Google Commands<br>\n  data set, which was created to promote deep learning research<br>\n  on speech recognition problems. It is comprised of 65,000 one<br>\n  second utterances of one of 30 words, with 10 of those words<br>\n  being the target classes and the rest considered unrelated or<br>\n  background noise. We perform MSDA on a Mel-frequency<br>\n  spectrogram of each utterance. The results for a PreAct<br>\n  ResNet-18 are given in Table II. We evaluate FMix, MixUp,<br>\n  and CutMix for the standard α = 1 used for the majority of<br>\n  our experiments and α = 0.2 recommended by Zhang et al.<br>\n  [2] for MixUp. We see in both cases that FMix and CutMix<br>\n  improve performance over MixUp outside the margin of error,<br>\n  with the best result achieved by FMix with α = 1.</p>\n</blockquote>\n<p><a href=\"https://arxiv.org/abs/2002.12047\" target=\"_blank\">https://arxiv.org/abs/2002.12047</a></p>",
      "rawMarkdown": "Thanks 👍 Yes I'm also using label smoothing in the BiTLoss. Also testing other methods like majority/full mixup/cutmix training with extra 'no call/noisy label' to help(?) the classifing, and then find a good PP to the prob..Have seen in a paper that fmix and cutmix is slighly better than mixup, link below.  Though I'm only using cutmix/mixup so far. Best is of course is to try only segment true calls before training.\nTo small variations between the result to be groundbreaking but maybe in a final ensembling it can help. Just ideas, fun to evaluate …🙂\n\n> B. Audio Classification\nWe now perform experiments on the Google Commands\ndata set, which was created to promote deep learning research\non speech recognition problems. It is comprised of 65,000 one\nsecond utterances of one of 30 words, with 10 of those words\nbeing the target classes and the rest considered unrelated or\nbackground noise. We perform MSDA on a Mel-frequency\nspectrogram of each utterance. The results for a PreAct\nResNet-18 are given in Table II. We evaluate FMix, MixUp,\nand CutMix for the standard α = 1 used for the majority of\nour experiments and α = 0.2 recommended by Zhang et al.\n[2] for MixUp. We see in both cases that FMix and CutMix\nimprove performance over MixUp outside the margin of error,\nwith the best result achieved by FMix with α = 1.\n\nhttps://arxiv.org/abs/2002.12047",
      "votes": null
    },
    {
      "id": "1329130",
      "postDate": "05/30/2021 21:38:17",
      "content": "<p>Hmm, I don't seem to be doing it right. The blog says, \"Setting both t1 and t2 to 1.0 recovers the logistic loss function,\" but, when I try those settings, I don't get the same results as BCEWithLogitsLoss.</p>\n<pre><code>import torch\nfrom torch import nn\n\ndevice = \"cpu\"\n\nactivations = torch.FloatTensor([[-0.5,  0.1,  2.0]]).to(device)\nlabels = torch.FloatTensor([[0.2, 0.5, 0.3]]).to(device)\n\nloss = bi_tempered_logistic_loss(activations=activations, labels=labels, t1=1.0, t2=1.0)\nprint(\"Loss, t1=1.0, t2=1.0: \", loss)\n\ncriterion = nn.BCEWithLogitsLoss()\nloss = criterion(activations, labels).item()\nprint(\"Loss BCEWithLogitLoss: \", loss)\n</code></pre>\n<p>These numbers should be equal, right?<br>\n<a href=\"https://github.com/mlpanda/bi-tempered-loss-pytorch\" target=\"_blank\">Original Github Repot</a></p>",
      "rawMarkdown": "Hmm, I don't seem to be doing it right. The blog says, \"Setting both t1 and t2 to 1.0 recovers the logistic loss function,\" but, when I try those settings, I don't get the same results as BCEWithLogitsLoss.\n\n```\nimport torch\nfrom torch import nn\n\ndevice = \"cpu\"\n\nactivations = torch.FloatTensor([[-0.5,  0.1,  2.0]]).to(device)\nlabels = torch.FloatTensor([[0.2, 0.5, 0.3]]).to(device)\n\nloss = bi_tempered_logistic_loss(activations=activations, labels=labels, t1=1.0, t2=1.0)\nprint(\"Loss, t1=1.0, t2=1.0: \", loss)\n\ncriterion = nn.BCEWithLogitsLoss()\nloss = criterion(activations, labels).item()\nprint(\"Loss BCEWithLogitLoss: \", loss)\n```\nThese numbers should be equal, right?\n[Original Github Repot](https://github.com/mlpanda/bi-tempered-loss-pytorch)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1316645,
      "author_name": "piantic",
      "author_url": "",
      "post_date": "05/20/2021 18:32:52",
      "content": "<p>In other competition, I used the Bi-Tempered Logistic Loss function.<br>\nPlease use it if necessary. :)</p>\n<p><a href=\"https://www.kaggle.com/piantic/train-cassava-starter-using-various-loss-funcs?scriptVersionId=52092486&amp;cellId=37\" target=\"_blank\">link</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1316660,
          "author_name": "kirderf",
          "author_url": "",
          "post_date": "05/20/2021 18:54:51",
          "content": "<p>Thanks. Yes I also had the same version saved from the same competition :) trying a custom version now for a fold run, with different settings, see how it goes.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1319222,
          "author_name": "hanson0910",
          "author_url": "",
          "post_date": "05/23/2021 03:20:30",
          "content": "<p>great works,thanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1316667,
      "author_name": "tomdenton",
      "author_url": "",
      "post_date": "05/20/2021 19:02:46",
      "content": "<p>I hadn't seen this before, thanks! Just quickly skimming so far… The paper seems focused on softmax loss, but we're generally allowed to have multiple birds in a segment (so, treating each label as an independent binary classification problem is usually seen as preferable to softmax'ing). Do you make any modifications for allowing multiple labels?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1317620,
          "author_name": "kirderf",
          "author_url": "",
          "post_date": "05/21/2021 14:52:49",
          "content": "<p>Yes it's a great loss function in general.</p>\n<p>It’s just an idea to the problem with the mislabeled croped samples and maybe find a solution for it, which I wanted to share the community and take feedback on, if some already had tested this or other similar techniques. 🙂 This loss function was just a quick idea that crossed my mind, there are many losses out there for handle mislabeled images to test on.</p>\n<p>I'm still testing it and I'm passing it through a sigmoid part for the probabilities instead of directly using the softmax part of it, and with some other changes to the default settings - it works, if it's better than say bce, the full run will tell.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1318451,
          "author_name": "kirderf",
          "author_url": "",
          "post_date": "05/22/2021 10:45:35",
          "content": "<p>Update: After a full train run - it works better(F1 validation score) than BCE(WL), and that with only a first test fold run without tuning the loss settings.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1319397,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "05/23/2021 06:45:36",
          "content": "<p><a href=\"https://www.kaggle.com/kirderf\" target=\"_blank\">@kirderf</a>  thanks  you get a better lB  also with it ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1319771,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "05/23/2021 13:39:49",
          "content": "<p><a href=\"https://www.kaggle.com/kirderf\" target=\"_blank\">@kirderf</a> thanks, good idea. what did you use for loss settings for the first run t1 and t2?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1321600,
          "author_name": "kirderf",
          "author_url": "",
          "post_date": "05/24/2021 19:35:06",
          "content": "<p>Yes slightly better LB score as well</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1321603,
      "author_name": "kirderf",
      "author_url": "",
      "post_date": "05/24/2021 19:48:52",
      "content": "<p>Update II: <br>\nDid a new run, with less randomness, given that mostly of the sounds are cut in random, so kept that same during both testing. <br>\nHad better validation score again and also a better LB score vs BCE loss, a small gain though but still, the val score also increases with some tuning of parameter settings, so maybe there are more gain to find. Or maybe this is all in the margin of error and just luck, but this is what I found within this loss. <br>\nI will keep the BiTLoss for now in my model training, maybe ensemble it with others, time will tell. Too much tuning, testing and training in other parts of this competition, so I'll leave further loss tuning for now :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1322800,
          "author_name": "tomdenton",
          "author_url": "",
          "post_date": "05/25/2021 17:29:59",
          "content": "<p>Thanks for looking into this, Kerderf!</p>\n<p>I've got a number of 'good heuristic' methods for getting much-better-than-random training segments from the Xeno-Canto recordings. It seems like one would want to do this kind of loss tuning on top of easy heuristics, to reduce the amount+impact of 'bad' training segments in the first place.</p>\n<p>I also found this interesting paper on label smoothing with weak labels, after wondering why there wasn't some comparison to label smoothing in the bitempered loss paper: <br>\n<a href=\"http://www.sanjivk.com/LabelSmoothing_ICML20.pdf\" target=\"_blank\">http://www.sanjivk.com/LabelSmoothing_ICML20.pdf</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1323378,
          "author_name": "kirderf",
          "author_url": "",
          "post_date": "05/26/2021 07:25:23",
          "content": "<p>Thanks 👍 Yes I'm also using label smoothing in the BiTLoss. Also testing other methods like majority/full mixup/cutmix training with extra 'no call/noisy label' to help(?) the classifing, and then find a good PP to the prob..Have seen in a paper that fmix and cutmix is slighly better than mixup, link below.  Though I'm only using cutmix/mixup so far. Best is of course is to try only segment true calls before training.<br>\nTo small variations between the result to be groundbreaking but maybe in a final ensembling it can help. Just ideas, fun to evaluate …🙂</p>\n<blockquote>\n  <p>B. Audio Classification<br>\n  We now perform experiments on the Google Commands<br>\n  data set, which was created to promote deep learning research<br>\n  on speech recognition problems. It is comprised of 65,000 one<br>\n  second utterances of one of 30 words, with 10 of those words<br>\n  being the target classes and the rest considered unrelated or<br>\n  background noise. We perform MSDA on a Mel-frequency<br>\n  spectrogram of each utterance. The results for a PreAct<br>\n  ResNet-18 are given in Table II. We evaluate FMix, MixUp,<br>\n  and CutMix for the standard α = 1 used for the majority of<br>\n  our experiments and α = 0.2 recommended by Zhang et al.<br>\n  [2] for MixUp. We see in both cases that FMix and CutMix<br>\n  improve performance over MixUp outside the margin of error,<br>\n  with the best result achieved by FMix with α = 1.</p>\n</blockquote>\n<p><a href=\"https://arxiv.org/abs/2002.12047\" target=\"_blank\">https://arxiv.org/abs/2002.12047</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1329130,
      "author_name": "philipkd",
      "author_url": "",
      "post_date": "05/30/2021 21:38:17",
      "content": "<p>Hmm, I don't seem to be doing it right. The blog says, \"Setting both t1 and t2 to 1.0 recovers the logistic loss function,\" but, when I try those settings, I don't get the same results as BCEWithLogitsLoss.</p>\n<pre><code>import torch\nfrom torch import nn\n\ndevice = \"cpu\"\n\nactivations = torch.FloatTensor([[-0.5,  0.1,  2.0]]).to(device)\nlabels = torch.FloatTensor([[0.2, 0.5, 0.3]]).to(device)\n\nloss = bi_tempered_logistic_loss(activations=activations, labels=labels, t1=1.0, t2=1.0)\nprint(\"Loss, t1=1.0, t2=1.0: \", loss)\n\ncriterion = nn.BCEWithLogitsLoss()\nloss = criterion(activations, labels).item()\nprint(\"Loss BCEWithLogitLoss: \", loss)\n</code></pre>\n<p>These numbers should be equal, right?<br>\n<a href=\"https://github.com/mlpanda/bi-tempered-loss-pytorch\" target=\"_blank\">Original Github Repot</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1316157": "Maybe this loss function can be of use here, anyone tried it with success or can rule it out?\n\nThe idea is that, when during training, a random ~5s clip is picked, that sound may not represent the label call well, it become mislabeled in the training. The Bi-Tempered Logistic Loss is created to handle this kind of problem.\n\nhttps://ai.googleblog.com/2019/08/bi-tempered-logistic-loss-for-training.html\nhttps://arxiv.org/pdf/1906.03361.pdf\n\nUpdate I: \nAfter a full train run - it works better(F1 validation score) than BCE(WL), and that with only a first test fold run without tuning the loss settings.\n\nUpdate II:\nDid a new run, with less randomness, given that mostly of the sounds are cut in random, so kept that same during both testing.\nHad better validation score again and also a better LB score vs BCE loss, a small gain though but still, the val score also increases with some tuning of parameter settings, so maybe there are more gain to find. Or maybe this is all in the margin of error and just luck, but this is what I found within this loss.\nI will keep the BiTLoss for now in my model training, maybe ensemble it with others, time will tell. Too much tuning, testing and training in other parts of this competition, so I'll leave further loss tuning for now :)",
    "1316645": "In other competition, I used the Bi-Tempered Logistic Loss function.\nPlease use it if necessary. :)\n\n[link](https://www.kaggle.com/piantic/train-cassava-starter-using-various-loss-funcs?scriptVersionId=52092486&cellId=37)",
    "1316660": "Thanks. Yes I also had the same version saved from the same competition :) trying a custom version now for a fold run, with different settings, see how it goes.",
    "1316667": "I hadn't seen this before, thanks! Just quickly skimming so far... The paper seems focused on softmax loss, but we're generally allowed to have multiple birds in a segment (so, treating each label as an independent binary classification problem is usually seen as preferable to softmax'ing). Do you make any modifications for allowing multiple labels?",
    "1317620": "Yes it's a great loss function in general.\n\nIt’s just an idea to the problem with the mislabeled croped samples and maybe find a solution for it, which I wanted to share the community and take feedback on, if some already had tested this or other similar techniques. 🙂 This loss function was just a quick idea that crossed my mind, there are many losses out there for handle mislabeled images to test on.\n\nI'm still testing it and I'm passing it through a sigmoid part for the probabilities instead of directly using the softmax part of it, and with some other changes to the default settings - it works, if it's better than say bce, the full run will tell.",
    "1318451": "Update: After a full train run - it works better(F1 validation score) than BCE(WL), and that with only a first test fold run without tuning the loss settings.",
    "1319222": "great works,thanks",
    "1319397": "kirderf  thanks  you get a better lB  also with it ?",
    "1319771": "kirderf thanks, good idea. what did you use for loss settings for the first run t1 and t2?",
    "1321600": "Yes slightly better LB score as well",
    "1321603": "Update II: \nDid a new run, with less randomness, given that mostly of the sounds are cut in random, so kept that same during both testing. \nHad better validation score again and also a better LB score vs BCE loss, a small gain though but still, the val score also increases with some tuning of parameter settings, so maybe there are more gain to find. Or maybe this is all in the margin of error and just luck, but this is what I found within this loss. \nI will keep the BiTLoss for now in my model training, maybe ensemble it with others, time will tell. Too much tuning, testing and training in other parts of this competition, so I'll leave further loss tuning for now :)",
    "1322800": "Thanks for looking into this, Kerderf!\n\nI've got a number of 'good heuristic' methods for getting much-better-than-random training segments from the Xeno-Canto recordings. It seems like one would want to do this kind of loss tuning on top of easy heuristics, to reduce the amount+impact of 'bad' training segments in the first place.\n\nI also found this interesting paper on label smoothing with weak labels, after wondering why there wasn't some comparison to label smoothing in the bitempered loss paper: \nhttp://www.sanjivk.com/LabelSmoothing_ICML20.pdf",
    "1323378": "Thanks 👍 Yes I'm also using label smoothing in the BiTLoss. Also testing other methods like majority/full mixup/cutmix training with extra 'no call/noisy label' to help(?) the classifing, and then find a good PP to the prob..Have seen in a paper that fmix and cutmix is slighly better than mixup, link below.  Though I'm only using cutmix/mixup so far. Best is of course is to try only segment true calls before training.\nTo small variations between the result to be groundbreaking but maybe in a final ensembling it can help. Just ideas, fun to evaluate …🙂\n\n> B. Audio Classification\nWe now perform experiments on the Google Commands\ndata set, which was created to promote deep learning research\non speech recognition problems. It is comprised of 65,000 one\nsecond utterances of one of 30 words, with 10 of those words\nbeing the target classes and the rest considered unrelated or\nbackground noise. We perform MSDA on a Mel-frequency\nspectrogram of each utterance. The results for a PreAct\nResNet-18 are given in Table II. We evaluate FMix, MixUp,\nand CutMix for the standard α = 1 used for the majority of\nour experiments and α = 0.2 recommended by Zhang et al.\n[2] for MixUp. We see in both cases that FMix and CutMix\nimprove performance over MixUp outside the margin of error,\nwith the best result achieved by FMix with α = 1.\n\nhttps://arxiv.org/abs/2002.12047",
    "1329130": "Hmm, I don't seem to be doing it right. The blog says, \"Setting both t1 and t2 to 1.0 recovers the logistic loss function,\" but, when I try those settings, I don't get the same results as BCEWithLogitsLoss.\n\n```\nimport torch\nfrom torch import nn\n\ndevice = \"cpu\"\n\nactivations = torch.FloatTensor([[-0.5,  0.1,  2.0]]).to(device)\nlabels = torch.FloatTensor([[0.2, 0.5, 0.3]]).to(device)\n\nloss = bi_tempered_logistic_loss(activations=activations, labels=labels, t1=1.0, t2=1.0)\nprint(\"Loss, t1=1.0, t2=1.0: \", loss)\n\ncriterion = nn.BCEWithLogitsLoss()\nloss = criterion(activations, labels).item()\nprint(\"Loss BCEWithLogitLoss: \", loss)\n```\nThese numbers should be equal, right?\n[Original Github Repot](https://github.com/mlpanda/bi-tempered-loss-pytorch)"
  },
  "source": "meta"
}