{
  "id": 318252,
  "title": "Is there a massive data shift between training and test audio files?",
  "url": "/competitions/birdclef-2022/discussion/318252",
  "author_name": "",
  "post_date": "2022-04-11T10:36:00.716899900Z",
  "votes": 32,
  "comment_count": 23,
  "views": 0,
  "content": "<p>Hello everyone,</p>\n<p>This is my first audio competition and I'm really puzzled by something and I'd like to share it with everyone, hoping to understand better what is happening here.</p>\n<p>I have been struggling a lot to beat random scores so far (&gt;0.51).</p>\n<p>I tried different approaches (all of them pretty basic using melspectrograms and CV models but inspired from the previous winning solutions):</p>\n<ul>\n<li>22 classes in a multi class settings : I predict the most likely bird present each audio file among the 21 scored classes, the 22th classes representing all the other classes of birds during training.</li>\n<li>22 classes in a multi label settings : each class has a binary score and I compute the competition metric with a threshold of 0.5</li>\n<li>152 classes in a multi label settings</li>\n</ul>\n<p>I have a 5 fold cross-validation pipeline training on 5s audio clips (randomly sampled) where I compute the competition metric according to this post <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/314999\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/314999</a> from <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a>.</p>\n<p>For all of those approaches, I can get CV scores &gt;0.70 without tweaking anything about the prediction thresholds. However, when submitting to the public LB I've always ended up with random scores (0.49-0.51).</p>\n<p>I spent a lot of time trying to find a bug in my submission pipeline but I actually think everything works as expected, the only problem is that the basic thresholds I can use during CV just does not work at all for the leaderboard.</p>\n<p>I just managed to reach 0.6 (single fold just to try beating randomness) by doing something I would consider quite absurd: with the 152 multi-label setting I only look at the 21 scored species predictions, normalize them so that they sum to 1, and then predict any bird which has a probability &gt;0.25.</p>\n<p>So here are a few questions:</p>\n<ul>\n<li>Is there something fundamentally different between the training set and the test set?</li>\n<li>Am I the only one having this issue?</li>\n<li>If no : how do you deal with choosing thresholds that don't match your cross validation strategy ?</li>\n<li>if yes : how did you solve the problem? any idea what could be going on?</li>\n</ul>\n<p>Cheers!</p>",
  "messages": [
    {
      "id": "1752028",
      "postDate": "04/11/2022 10:36:00",
      "content": "<p>Hello everyone,</p>\n<p>This is my first audio competition and I'm really puzzled by something and I'd like to share it with everyone, hoping to understand better what is happening here.</p>\n<p>I have been struggling a lot to beat random scores so far (&gt;0.51).</p>\n<p>I tried different approaches (all of them pretty basic using melspectrograms and CV models but inspired from the previous winning solutions):</p>\n<ul>\n<li>22 classes in a multi class settings : I predict the most likely bird present each audio file among the 21 scored classes, the 22th classes representing all the other classes of birds during training.</li>\n<li>22 classes in a multi label settings : each class has a binary score and I compute the competition metric with a threshold of 0.5</li>\n<li>152 classes in a multi label settings</li>\n</ul>\n<p>I have a 5 fold cross-validation pipeline training on 5s audio clips (randomly sampled) where I compute the competition metric according to this post <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/314999\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/314999</a> from <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a>.</p>\n<p>For all of those approaches, I can get CV scores &gt;0.70 without tweaking anything about the prediction thresholds. However, when submitting to the public LB I've always ended up with random scores (0.49-0.51).</p>\n<p>I spent a lot of time trying to find a bug in my submission pipeline but I actually think everything works as expected, the only problem is that the basic thresholds I can use during CV just does not work at all for the leaderboard.</p>\n<p>I just managed to reach 0.6 (single fold just to try beating randomness) by doing something I would consider quite absurd: with the 152 multi-label setting I only look at the 21 scored species predictions, normalize them so that they sum to 1, and then predict any bird which has a probability &gt;0.25.</p>\n<p>So here are a few questions:</p>\n<ul>\n<li>Is there something fundamentally different between the training set and the test set?</li>\n<li>Am I the only one having this issue?</li>\n<li>If no : how do you deal with choosing thresholds that don't match your cross validation strategy ?</li>\n<li>if yes : how did you solve the problem? any idea what could be going on?</li>\n</ul>\n<p>Cheers!</p>",
      "rawMarkdown": "Hello everyone,\n\nThis is my first audio competition and I'm really puzzled by something and I'd like to share it with everyone, hoping to understand better what is happening here.\n\nI have been struggling a lot to beat random scores so far (>0.51).\n\nI tried different approaches (all of them pretty basic using melspectrograms and CV models but inspired from the previous winning solutions):\n- 22 classes in a multi class settings : I predict the most likely bird present each audio file among the 21 scored classes, the 22th classes representing all the other classes of birds during training.\n-  22 classes in a multi label settings : each class has a binary score and I compute the competition metric with a threshold of 0.5\n- 152 classes in a multi label settings\n\nI have a 5 fold cross-validation pipeline training on 5s audio clips (randomly sampled) where I compute the competition metric according to this post https://www.kaggle.com/competitions/birdclef-2022/discussion/314999 from @dschettler8845.\n\nFor all of those approaches, I can get CV scores >0.70 without tweaking anything about the prediction thresholds. However, when submitting to the public LB I've always ended up with random scores (0.49-0.51).\n\nI spent a lot of time trying to find a bug in my submission pipeline but I actually think everything works as expected, the only problem is that the basic thresholds I can use during CV just does not work at all for the leaderboard.\n\nI just managed to reach 0.6 (single fold just to try beating randomness) by doing something I would consider quite absurd: with the 152 multi-label setting I only look at the 21 scored species predictions, normalize them so that they sum to 1, and then predict any bird which has a probability >0.25.\n\nSo here are a few questions:\n- Is there something fundamentally different between the training set and the test set?\n- Am I the only one having this issue?\n- If no : how do you deal with choosing thresholds that don't match your cross validation strategy ?\n- if yes : how did you solve the problem? any idea what could be going on?\n\nCheers!",
      "votes": null
    },
    {
      "id": "1752115",
      "postDate": "04/11/2022 12:31:32",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a>,</p>\n<p>I started a few days ago and I probably observe the same problem - LB score is around 0.5 when the same threshold is used during the evaluation of CV in training and prediction for submission. Decreasing the threshold during submission might increase the LB to 0.69. I am still trying to understand what happens to build an adequate CV strategy.</p>",
      "rawMarkdown": "Hi @optimo,\n\nI started a few days ago and I probably observe the same problem - LB score is around 0.5 when the same threshold is used during the evaluation of CV in training and prediction for submission. Decreasing the threshold during submission might increase the LB to 0.69. I am still trying to understand what happens to build an adequate CV strategy.",
      "votes": null
    },
    {
      "id": "1752300",
      "postDate": "04/11/2022 15:41:39",
      "content": "<p>There's a large domain shift between training (Xeno-Canto targeted human recordings of specific birds) and test (automated soundscape recordings). </p>\n<p>Mario Lasseck's writeup from BirdClef 2019 has some good summary of effective augmentation strategies which have helped in the past for overcoming the train/test domain shift:<br>\n<a href=\"https://www.semanticscholar.org/paper/Bird-Species-Identification-in-Soundscapes-Lasseck/39c2a49662ed9912842b9e82db1b40c65145eb86\" target=\"_blank\">https://www.semanticscholar.org/paper/Bird-Species-Identification-in-Soundscapes-Lasseck/39c2a49662ed9912842b9e82db1b40c65145eb86</a></p>",
      "rawMarkdown": "There's a large domain shift between training (Xeno-Canto targeted human recordings of specific birds) and test (automated soundscape recordings). \n\nMario Lasseck's writeup from BirdClef 2019 has some good summary of effective augmentation strategies which have helped in the past for overcoming the train/test domain shift:\nhttps://www.semanticscholar.org/paper/Bird-Species-Identification-in-Soundscapes-Lasseck/39c2a49662ed9912842b9e82db1b40c65145eb86",
      "votes": null
    },
    {
      "id": "1752305",
      "postDate": "04/11/2022 15:46:13",
      "content": "<p>Thank you very much, I'll have a look!</p>",
      "rawMarkdown": "Thank you very much, I'll have a look!",
      "votes": null
    },
    {
      "id": "1753403",
      "postDate": "04/12/2022 19:10:46",
      "content": "<p>Yes, I had the same problem !<br>\nI came to the conclusion to trust my CV score because public lb is only <strong>16%</strong> </p>",
      "rawMarkdown": "Yes, I had the same problem !\nI came to the conclusion to trust my CV score because public lb is only **16%**",
      "votes": null
    },
    {
      "id": "1753406",
      "postDate": "04/12/2022 19:12:36",
      "content": "<p>I also found that it is very easy to overfit on the lb by being aggressive on the threshold.</p>",
      "rawMarkdown": "I also found that it is very easy to overfit on the lb by being aggressive on the threshold.",
      "votes": null
    },
    {
      "id": "1753415",
      "postDate": "04/12/2022 19:27:31",
      "content": "<p>Well, the domain shift will still exist on the private LB so you can't \"simply\" trust your local CV (or I think I'll end up with a 0.49 private too 😫).</p>\n<p>I agree with you that simply relying on \"what has been my best threshold so far\" based on LB is probably a bad solution. The hardest thing here is to emulate in someway a cross validation strategy which allows you to get similar performances in CV and LB.</p>\n<p>Have you managed to do so? I'm experimenting with background noises at the moment, I think it's one part of the solution but I'm not quite there yet.</p>",
      "rawMarkdown": "Well, the domain shift will still exist on the private LB so you can't \"simply\" trust your local CV (or I think I'll end up with a 0.49 private too 😫).\n\nI agree with you that simply relying on \"what has been my best threshold so far\" based on LB is probably a bad solution. The hardest thing here is to emulate in someway a cross validation strategy which allows you to get similar performances in CV and LB.\n\nHave you managed to do so? I'm experimenting with background noises at the moment, I think it's one part of the solution but I'm not quite there yet.",
      "votes": null
    },
    {
      "id": "1754172",
      "postDate": "04/13/2022 12:36:29",
      "content": "<p>What do you mean by domain shift ?<br>\nI only tested one resnet34. I didn't observe a correlation between cv and lb.<br>\nAnd also the 5 fold blend always score lower than single fold in my case.</p>",
      "rawMarkdown": "What do you mean by domain shift ?\nI only tested one resnet34. I didn't observe a correlation between cv and lb.\nAnd also the 5 fold blend always score lower than single fold in my case.",
      "votes": null
    },
    {
      "id": "1754205",
      "postDate": "04/13/2022 13:17:26",
      "content": "<p>The authors of the challenge specified that the test set contains ~5000 1-minute long soundscapes compared to 80 10-minute long soundscapes during previous year competition. </p>\n<p>Looking at birdclef-2021 private LB, it seems to me more or less stable. </p>\n<p>However, we shouldn't underestimate the shakeup power of macro-F1 score :D</p>\n<p>Just to mention, this year public LB contains 5000 * 0.16 = 800 minutes soundscapes in contrast to 80 * 10 * 0.35 = 280 minutes of soundscapes during last year challenge, which makes me believe that current LB should be even more stable. </p>\n<p>Maybe authors decided to increased test set size to mitigate transition from F1-micro to F1-macro metric.</p>",
      "rawMarkdown": "The authors of the challenge specified that the test set contains ~5000 1-minute long soundscapes compared to 80 10-minute long soundscapes during previous year competition. \n\nLooking at birdclef-2021 private LB, it seems to me more or less stable. \n\nHowever, we shouldn't underestimate the shakeup power of macro-F1 score :D\n\nJust to mention, this year public LB contains 5000 * 0.16 = 800 minutes soundscapes in contrast to 80 * 10 * 0.35 = 280 minutes of soundscapes during last year challenge, which makes me believe that current LB should be even more stable. \n\nMaybe authors decided to increased test set size to mitigate transition from F1-micro to F1-macro metric.",
      "votes": null
    },
    {
      "id": "1754231",
      "postDate": "04/13/2022 13:39:09",
      "content": "<p>By domain shift I mean that the audios from test files are fundamentally different that the one given for training as <a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a> confirmed (I'm not expert but I guess: background noise, number of species at the same time etc…) which explains why a model that works fine in CV with a threshold of 0.5 will never detect any bird in the test data. When lowering the thresholds then you end up with correct detection.</p>\n<p><a href=\"https://www.kaggle.com/amedprof\" target=\"_blank\">@amedprof</a> did you have to lower your threshold to reach 0.70?</p>\n<p><a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> without revealing any secret, can you confirm that your models do not behave similarly on training and testing data ?</p>",
      "rawMarkdown": "By domain shift I mean that the audios from test files are fundamentally different that the one given for training as @tomdenton confirmed (I'm not expert but I guess: background noise, number of species at the same time etc...) which explains why a model that works fine in CV with a threshold of 0.5 will never detect any bird in the test data. When lowering the thresholds then you end up with correct detection.\n\n@amedprof did you have to lower your threshold to reach 0.70?\n\n@martynoveduard without revealing any secret, can you confirm that your models do not behave similarly on training and testing data ?",
      "votes": null
    },
    {
      "id": "1755248",
      "postDate": "04/14/2022 13:03:29",
      "content": "<p>How did you estimate your CV? we don't have 5 seconds audio files.</p>",
      "rawMarkdown": "How did you estimate your CV? we don't have 5 seconds audio files.",
      "votes": null
    },
    {
      "id": "1755261",
      "postDate": "04/14/2022 13:12:26",
      "content": "<p>I select a random 5 seconds crop for each file to compute my CV.<br>\nSo there is randomness in this approach, but I can estimate the standard deviation by computing different times the CV with different crops.</p>\n<p>The limitation here is that 5 second crops in the training set might not contain the desired label, however this should underestimate the real score (CV should be less since ground truth is blurry), in practice the domain shift make this score uncorrelated to the LB.</p>\n<p>What is your approach ?</p>",
      "rawMarkdown": "I select a random 5 seconds crop for each file to compute my CV.\nSo there is randomness in this approach, but I can estimate the standard deviation by computing different times the CV with different crops.\n\nThe limitation here is that 5 second crops in the training set might not contain the desired label, however this should underestimate the real score (CV should be less since ground truth is blurry), in practice the domain shift make this score uncorrelated to the LB.\n\nWhat is your approach ?",
      "votes": null
    },
    {
      "id": "1755305",
      "postDate": "04/14/2022 13:51:26",
      "content": "<p>When I validate, i chose 5-second clip starting from 0 seconds. Then 0~5 seconds interval is chosen for validation. I think this is not good. With this setting, my cv &gt; 0.8</p>",
      "rawMarkdown": "When I validate, i chose 5-second clip starting from 0 seconds. Then 0~5 seconds interval is chosen for validation. I think this is not good. With this setting, my cv > 0.8",
      "votes": null
    },
    {
      "id": "1755318",
      "postDate": "04/14/2022 13:55:57",
      "content": "<p>Do you need to lower your thresholds during inference ? Or is your pipeline robust to domain shift ?</p>",
      "rawMarkdown": "Do you need to lower your thresholds during inference ? Or is your pipeline robust to domain shift ?",
      "votes": null
    },
    {
      "id": "1755334",
      "postDate": "04/14/2022 14:03:22",
      "content": "<p>while I can't say details fully, my cv score was better when i lowered the threshold for my checking micro f1 score. <br>\nat this time, i don't consider the domain shift you mentioned. i don't think my models are robust to domain shift</p>",
      "rawMarkdown": "while I can't say details fully, my cv score was better when i lowered the threshold for my checking micro f1 score. \nat this time, i don't consider the domain shift you mentioned. i don't think my models are robust to domain shift",
      "votes": null
    },
    {
      "id": "1755370",
      "postDate": "04/14/2022 14:39:04",
      "content": "<p>Fair enough, I actually see the opposite, I monitor the competition metric (not micro f1 score) with different thresholds 0.5, 0.2 and 0.1 and my scores are systematically better with a threshold of 0.5.</p>",
      "rawMarkdown": "Fair enough, I actually see the opposite, I monitor the competition metric (not micro f1 score) with different thresholds 0.5, 0.2 and 0.1 and my scores are systematically better with a threshold of 0.5.",
      "votes": null
    },
    {
      "id": "1755382",
      "postDate": "04/14/2022 15:02:16",
      "content": "<p>In my case cv score at  f1 - 0.3 &gt; f1 - 0.2 &gt; f1 - 0.5 &gt; f1 - 0.1 </p>",
      "rawMarkdown": "In my case cv score at  f1 - 0.3 > f1 - 0.2 > f1 - 0.5 > f1 - 0.1",
      "votes": null
    },
    {
      "id": "1755866",
      "postDate": "04/15/2022 04:21:49",
      "content": "<p>yep, I can confirm that my current validation scheme gives the wrong threshold estimation, however I think this is because I'm using wrong proxy-metric to competition metric and we also don't have data with hard labels, so validation is kinda noisy. </p>\n<p>however, I suppose that it is possible to establish good CV/LB correlation in this comp, we'll see.</p>",
      "rawMarkdown": "yep, I can confirm that my current validation scheme gives the wrong threshold estimation, however I think this is because I'm using wrong proxy-metric to competition metric and we also don't have data with hard labels, so validation is kinda noisy. \n\nhowever, I suppose that it is possible to establish good CV/LB correlation in this comp, we'll see.",
      "votes": null
    },
    {
      "id": "1763682",
      "postDate": "04/21/2022 17:51:28",
      "content": "<p><a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> Hello. Congrats on your great place on the leaderboard. I am dealing now with finding a proper metric. So, I am wondering if you let me know which metric you finally chose for the cross validation, f1-macro or you made your own?</p>",
      "rawMarkdown": "optimo Hello. Congrats on your great place on the leaderboard. I am dealing now with finding a proper metric. So, I am wondering if you let me know which metric you finally chose for the cross validation, f1-macro or you made your own?",
      "votes": null
    },
    {
      "id": "1763734",
      "postDate": "04/21/2022 19:01:37",
      "content": "<p>This is a very common issue in audio recognition competitions. There is a large domain shift between the training data and the test data. You need to be very careful in choosing the thresholds that you use in your submission pipeline.</p>",
      "rawMarkdown": "This is a very common issue in audio recognition competitions. There is a large domain shift between the training data and the test data. You need to be very careful in choosing the thresholds that you use in your submission pipeline.",
      "votes": null
    },
    {
      "id": "1763793",
      "postDate": "04/21/2022 20:21:30",
      "content": "<p>I’m monitoring the competition loss following this : <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/314999\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/314999</a></p>",
      "rawMarkdown": "I’m monitoring the competition loss following this : https://www.kaggle.com/competitions/birdclef-2022/discussion/314999",
      "votes": null
    },
    {
      "id": "1764466",
      "postDate": "04/22/2022 13:59:29",
      "content": "<p>Thank you for your reply. Oh, I saw this post. However, I think the accuracy defined there:</p>\n<blockquote>\n  <p>'Accuracy' is given as an equally weighted average between the True Positive score and True Negative score)</p>\n</blockquote>\n<p>is not a correct metric that satisfies what the Competition Host <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/311493#1716290\" target=\"_blank\">commented</a> : </p>\n<blockquote>\n  <p>Hi! There's typically much more 'negative' audio for a given species than positive audio. So we chose a weighting which equalizes the impact of the positive/negative labels for each species. (and also equalizing the total weight for each species.)</p>\n</blockquote>",
      "rawMarkdown": "Thank you for your reply. Oh, I saw this post. However, I think the accuracy defined there:\n>'Accuracy' is given as an equally weighted average between the True Positive score and True Negative score)\n\nis not a correct metric that satisfies what the Competition Host [commented](https://www.kaggle.com/competitions/birdclef-2022/discussion/311493#1716290) : \n> Hi! There's typically much more 'negative' audio for a given species than positive audio. So we chose a weighting which equalizes the impact of the positive/negative labels for each species. (and also equalizing the total weight for each species.)",
      "votes": null
    },
    {
      "id": "1764476",
      "postDate": "04/22/2022 14:18:22",
      "content": "<p>I think overall it’s similar to the completion metric, the only difference is that the host can decide to remove some rows from the final submission. No idea how they decide to remove a row however.</p>",
      "rawMarkdown": "I think overall it’s similar to the completion metric, the only difference is that the host can decide to remove some rows from the final submission. No idea how they decide to remove a row however.",
      "votes": null
    },
    {
      "id": "1764679",
      "postDate": "04/22/2022 17:31:05",
      "content": "<p>They mentioned that they only keep the rows including scored birds and remove the rest.</p>",
      "rawMarkdown": "They mentioned that they only keep the rows including scored birds and remove the rest.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1752115,
      "author_name": "egortrushin",
      "author_url": "",
      "post_date": "04/11/2022 12:31:32",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a>,</p>\n<p>I started a few days ago and I probably observe the same problem - LB score is around 0.5 when the same threshold is used during the evaluation of CV in training and prediction for submission. Decreasing the threshold during submission might increase the LB to 0.69. I am still trying to understand what happens to build an adequate CV strategy.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1752300,
      "author_name": "tomdenton",
      "author_url": "",
      "post_date": "04/11/2022 15:41:39",
      "content": "<p>There's a large domain shift between training (Xeno-Canto targeted human recordings of specific birds) and test (automated soundscape recordings). </p>\n<p>Mario Lasseck's writeup from BirdClef 2019 has some good summary of effective augmentation strategies which have helped in the past for overcoming the train/test domain shift:<br>\n<a href=\"https://www.semanticscholar.org/paper/Bird-Species-Identification-in-Soundscapes-Lasseck/39c2a49662ed9912842b9e82db1b40c65145eb86\" target=\"_blank\">https://www.semanticscholar.org/paper/Bird-Species-Identification-in-Soundscapes-Lasseck/39c2a49662ed9912842b9e82db1b40c65145eb86</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1752305,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "04/11/2022 15:46:13",
          "content": "<p>Thank you very much, I'll have a look!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1753403,
      "author_name": "amedprof",
      "author_url": "",
      "post_date": "04/12/2022 19:10:46",
      "content": "<p>Yes, I had the same problem !<br>\nI came to the conclusion to trust my CV score because public lb is only <strong>16%</strong> </p>",
      "votes": null,
      "replies": [
        {
          "id": 1753406,
          "author_name": "amedprof",
          "author_url": "",
          "post_date": "04/12/2022 19:12:36",
          "content": "<p>I also found that it is very easy to overfit on the lb by being aggressive on the threshold.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1753415,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "04/12/2022 19:27:31",
          "content": "<p>Well, the domain shift will still exist on the private LB so you can't \"simply\" trust your local CV (or I think I'll end up with a 0.49 private too 😫).</p>\n<p>I agree with you that simply relying on \"what has been my best threshold so far\" based on LB is probably a bad solution. The hardest thing here is to emulate in someway a cross validation strategy which allows you to get similar performances in CV and LB.</p>\n<p>Have you managed to do so? I'm experimenting with background noises at the moment, I think it's one part of the solution but I'm not quite there yet.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1754172,
          "author_name": "amedprof",
          "author_url": "",
          "post_date": "04/13/2022 12:36:29",
          "content": "<p>What do you mean by domain shift ?<br>\nI only tested one resnet34. I didn't observe a correlation between cv and lb.<br>\nAnd also the 5 fold blend always score lower than single fold in my case.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1754205,
          "author_name": "martynoveduard",
          "author_url": "",
          "post_date": "04/13/2022 13:17:26",
          "content": "<p>The authors of the challenge specified that the test set contains ~5000 1-minute long soundscapes compared to 80 10-minute long soundscapes during previous year competition. </p>\n<p>Looking at birdclef-2021 private LB, it seems to me more or less stable. </p>\n<p>However, we shouldn't underestimate the shakeup power of macro-F1 score :D</p>\n<p>Just to mention, this year public LB contains 5000 * 0.16 = 800 minutes soundscapes in contrast to 80 * 10 * 0.35 = 280 minutes of soundscapes during last year challenge, which makes me believe that current LB should be even more stable. </p>\n<p>Maybe authors decided to increased test set size to mitigate transition from F1-micro to F1-macro metric.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1754231,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "04/13/2022 13:39:09",
          "content": "<p>By domain shift I mean that the audios from test files are fundamentally different that the one given for training as <a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a> confirmed (I'm not expert but I guess: background noise, number of species at the same time etc…) which explains why a model that works fine in CV with a threshold of 0.5 will never detect any bird in the test data. When lowering the thresholds then you end up with correct detection.</p>\n<p><a href=\"https://www.kaggle.com/amedprof\" target=\"_blank\">@amedprof</a> did you have to lower your threshold to reach 0.70?</p>\n<p><a href=\"https://www.kaggle.com/martynoveduard\" target=\"_blank\">@martynoveduard</a> without revealing any secret, can you confirm that your models do not behave similarly on training and testing data ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1755866,
          "author_name": "martynoveduard",
          "author_url": "",
          "post_date": "04/15/2022 04:21:49",
          "content": "<p>yep, I can confirm that my current validation scheme gives the wrong threshold estimation, however I think this is because I'm using wrong proxy-metric to competition metric and we also don't have data with hard labels, so validation is kinda noisy. </p>\n<p>however, I suppose that it is possible to establish good CV/LB correlation in this comp, we'll see.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1755248,
      "author_name": "deepkim",
      "author_url": "",
      "post_date": "04/14/2022 13:03:29",
      "content": "<p>How did you estimate your CV? we don't have 5 seconds audio files.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1755261,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "04/14/2022 13:12:26",
          "content": "<p>I select a random 5 seconds crop for each file to compute my CV.<br>\nSo there is randomness in this approach, but I can estimate the standard deviation by computing different times the CV with different crops.</p>\n<p>The limitation here is that 5 second crops in the training set might not contain the desired label, however this should underestimate the real score (CV should be less since ground truth is blurry), in practice the domain shift make this score uncorrelated to the LB.</p>\n<p>What is your approach ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1755305,
          "author_name": "deepkim",
          "author_url": "",
          "post_date": "04/14/2022 13:51:26",
          "content": "<p>When I validate, i chose 5-second clip starting from 0 seconds. Then 0~5 seconds interval is chosen for validation. I think this is not good. With this setting, my cv &gt; 0.8</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1755318,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "04/14/2022 13:55:57",
          "content": "<p>Do you need to lower your thresholds during inference ? Or is your pipeline robust to domain shift ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1755334,
          "author_name": "deepkim",
          "author_url": "",
          "post_date": "04/14/2022 14:03:22",
          "content": "<p>while I can't say details fully, my cv score was better when i lowered the threshold for my checking micro f1 score. <br>\nat this time, i don't consider the domain shift you mentioned. i don't think my models are robust to domain shift</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1755370,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "04/14/2022 14:39:04",
          "content": "<p>Fair enough, I actually see the opposite, I monitor the competition metric (not micro f1 score) with different thresholds 0.5, 0.2 and 0.1 and my scores are systematically better with a threshold of 0.5.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1755382,
          "author_name": "deepkim",
          "author_url": "",
          "post_date": "04/14/2022 15:02:16",
          "content": "<p>In my case cv score at  f1 - 0.3 &gt; f1 - 0.2 &gt; f1 - 0.5 &gt; f1 - 0.1 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1763682,
      "author_name": "tjamali",
      "author_url": "",
      "post_date": "04/21/2022 17:51:28",
      "content": "<p><a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> Hello. Congrats on your great place on the leaderboard. I am dealing now with finding a proper metric. So, I am wondering if you let me know which metric you finally chose for the cross validation, f1-macro or you made your own?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1763793,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "04/21/2022 20:21:30",
          "content": "<p>I’m monitoring the competition loss following this : <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/314999\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/314999</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1764466,
          "author_name": "tjamali",
          "author_url": "",
          "post_date": "04/22/2022 13:59:29",
          "content": "<p>Thank you for your reply. Oh, I saw this post. However, I think the accuracy defined there:</p>\n<blockquote>\n  <p>'Accuracy' is given as an equally weighted average between the True Positive score and True Negative score)</p>\n</blockquote>\n<p>is not a correct metric that satisfies what the Competition Host <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/311493#1716290\" target=\"_blank\">commented</a> : </p>\n<blockquote>\n  <p>Hi! There's typically much more 'negative' audio for a given species than positive audio. So we chose a weighting which equalizes the impact of the positive/negative labels for each species. (and also equalizing the total weight for each species.)</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1764476,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "04/22/2022 14:18:22",
          "content": "<p>I think overall it’s similar to the completion metric, the only difference is that the host can decide to remove some rows from the final submission. No idea how they decide to remove a row however.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1764679,
          "author_name": "tjamali",
          "author_url": "",
          "post_date": "04/22/2022 17:31:05",
          "content": "<p>They mentioned that they only keep the rows including scored birds and remove the rest.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1763734,
      "author_name": "",
      "author_url": "",
      "post_date": "04/21/2022 19:01:37",
      "content": "<p>This is a very common issue in audio recognition competitions. There is a large domain shift between the training data and the test data. You need to be very careful in choosing the thresholds that you use in your submission pipeline.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1752028": "Hello everyone,\n\nThis is my first audio competition and I'm really puzzled by something and I'd like to share it with everyone, hoping to understand better what is happening here.\n\nI have been struggling a lot to beat random scores so far (>0.51).\n\nI tried different approaches (all of them pretty basic using melspectrograms and CV models but inspired from the previous winning solutions):\n- 22 classes in a multi class settings : I predict the most likely bird present each audio file among the 21 scored classes, the 22th classes representing all the other classes of birds during training.\n-  22 classes in a multi label settings : each class has a binary score and I compute the competition metric with a threshold of 0.5\n- 152 classes in a multi label settings\n\nI have a 5 fold cross-validation pipeline training on 5s audio clips (randomly sampled) where I compute the competition metric according to this post https://www.kaggle.com/competitions/birdclef-2022/discussion/314999 from @dschettler8845.\n\nFor all of those approaches, I can get CV scores >0.70 without tweaking anything about the prediction thresholds. However, when submitting to the public LB I've always ended up with random scores (0.49-0.51).\n\nI spent a lot of time trying to find a bug in my submission pipeline but I actually think everything works as expected, the only problem is that the basic thresholds I can use during CV just does not work at all for the leaderboard.\n\nI just managed to reach 0.6 (single fold just to try beating randomness) by doing something I would consider quite absurd: with the 152 multi-label setting I only look at the 21 scored species predictions, normalize them so that they sum to 1, and then predict any bird which has a probability >0.25.\n\nSo here are a few questions:\n- Is there something fundamentally different between the training set and the test set?\n- Am I the only one having this issue?\n- If no : how do you deal with choosing thresholds that don't match your cross validation strategy ?\n- if yes : how did you solve the problem? any idea what could be going on?\n\nCheers!",
    "1752115": "Hi @optimo,\n\nI started a few days ago and I probably observe the same problem - LB score is around 0.5 when the same threshold is used during the evaluation of CV in training and prediction for submission. Decreasing the threshold during submission might increase the LB to 0.69. I am still trying to understand what happens to build an adequate CV strategy.",
    "1752300": "There's a large domain shift between training (Xeno-Canto targeted human recordings of specific birds) and test (automated soundscape recordings). \n\nMario Lasseck's writeup from BirdClef 2019 has some good summary of effective augmentation strategies which have helped in the past for overcoming the train/test domain shift:\nhttps://www.semanticscholar.org/paper/Bird-Species-Identification-in-Soundscapes-Lasseck/39c2a49662ed9912842b9e82db1b40c65145eb86",
    "1752305": "Thank you very much, I'll have a look!",
    "1753403": "Yes, I had the same problem !\nI came to the conclusion to trust my CV score because public lb is only **16%**",
    "1753406": "I also found that it is very easy to overfit on the lb by being aggressive on the threshold.",
    "1753415": "Well, the domain shift will still exist on the private LB so you can't \"simply\" trust your local CV (or I think I'll end up with a 0.49 private too 😫).\n\nI agree with you that simply relying on \"what has been my best threshold so far\" based on LB is probably a bad solution. The hardest thing here is to emulate in someway a cross validation strategy which allows you to get similar performances in CV and LB.\n\nHave you managed to do so? I'm experimenting with background noises at the moment, I think it's one part of the solution but I'm not quite there yet.",
    "1754172": "What do you mean by domain shift ?\nI only tested one resnet34. I didn't observe a correlation between cv and lb.\nAnd also the 5 fold blend always score lower than single fold in my case.",
    "1754205": "The authors of the challenge specified that the test set contains ~5000 1-minute long soundscapes compared to 80 10-minute long soundscapes during previous year competition. \n\nLooking at birdclef-2021 private LB, it seems to me more or less stable. \n\nHowever, we shouldn't underestimate the shakeup power of macro-F1 score :D\n\nJust to mention, this year public LB contains 5000 * 0.16 = 800 minutes soundscapes in contrast to 80 * 10 * 0.35 = 280 minutes of soundscapes during last year challenge, which makes me believe that current LB should be even more stable. \n\nMaybe authors decided to increased test set size to mitigate transition from F1-micro to F1-macro metric.",
    "1754231": "By domain shift I mean that the audios from test files are fundamentally different that the one given for training as @tomdenton confirmed (I'm not expert but I guess: background noise, number of species at the same time etc...) which explains why a model that works fine in CV with a threshold of 0.5 will never detect any bird in the test data. When lowering the thresholds then you end up with correct detection.\n\n@amedprof did you have to lower your threshold to reach 0.70?\n\n@martynoveduard without revealing any secret, can you confirm that your models do not behave similarly on training and testing data ?",
    "1755248": "How did you estimate your CV? we don't have 5 seconds audio files.",
    "1755261": "I select a random 5 seconds crop for each file to compute my CV.\nSo there is randomness in this approach, but I can estimate the standard deviation by computing different times the CV with different crops.\n\nThe limitation here is that 5 second crops in the training set might not contain the desired label, however this should underestimate the real score (CV should be less since ground truth is blurry), in practice the domain shift make this score uncorrelated to the LB.\n\nWhat is your approach ?",
    "1755305": "When I validate, i chose 5-second clip starting from 0 seconds. Then 0~5 seconds interval is chosen for validation. I think this is not good. With this setting, my cv > 0.8",
    "1755318": "Do you need to lower your thresholds during inference ? Or is your pipeline robust to domain shift ?",
    "1755334": "while I can't say details fully, my cv score was better when i lowered the threshold for my checking micro f1 score. \nat this time, i don't consider the domain shift you mentioned. i don't think my models are robust to domain shift",
    "1755370": "Fair enough, I actually see the opposite, I monitor the competition metric (not micro f1 score) with different thresholds 0.5, 0.2 and 0.1 and my scores are systematically better with a threshold of 0.5.",
    "1755382": "In my case cv score at  f1 - 0.3 > f1 - 0.2 > f1 - 0.5 > f1 - 0.1",
    "1755866": "yep, I can confirm that my current validation scheme gives the wrong threshold estimation, however I think this is because I'm using wrong proxy-metric to competition metric and we also don't have data with hard labels, so validation is kinda noisy. \n\nhowever, I suppose that it is possible to establish good CV/LB correlation in this comp, we'll see.",
    "1763682": "optimo Hello. Congrats on your great place on the leaderboard. I am dealing now with finding a proper metric. So, I am wondering if you let me know which metric you finally chose for the cross validation, f1-macro or you made your own?",
    "1763734": "This is a very common issue in audio recognition competitions. There is a large domain shift between the training data and the test data. You need to be very careful in choosing the thresholds that you use in your submission pipeline.",
    "1763793": "I’m monitoring the competition loss following this : https://www.kaggle.com/competitions/birdclef-2022/discussion/314999",
    "1764466": "Thank you for your reply. Oh, I saw this post. However, I think the accuracy defined there:\n>'Accuracy' is given as an equally weighted average between the True Positive score and True Negative score)\n\nis not a correct metric that satisfies what the Competition Host [commented](https://www.kaggle.com/competitions/birdclef-2022/discussion/311493#1716290) : \n> Hi! There's typically much more 'negative' audio for a given species than positive audio. So we chose a weighting which equalizes the impact of the positive/negative labels for each species. (and also equalizing the total weight for each species.)",
    "1764476": "I think overall it’s similar to the completion metric, the only difference is that the host can decide to remove some rows from the final submission. No idea how they decide to remove a row however.",
    "1764679": "They mentioned that they only keep the rows including scored birds and remove the rest."
  },
  "source": "meta"
}