{
  "id": 201920,
  "title": "Have anyone observed instability between CV and LB?",
  "url": "/competitions/rfcx-species-audio-detection/discussion/201920",
  "author_name": "",
  "post_date": "2020-12-07T09:46:29.051089500Z",
  "votes": 32,
  "comment_count": 21,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>We have been confronting a problem of uncorrelated CV - LB recently. There are several possible explanations like below:</p>\n<ul>\n<li>Bad validation strategy</li>\n</ul>\n<p>This may well be the case, and until recently we were doing validation with LWLRAP in random crop windows of n seconds, rather than with LWLRAP at the clip level, which is the correct metric used in testing as <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/198132#1101848\" target=\"_blank\">pointed out in the CV - LB thread</a>, but even after moving to the correct metric, we still have the problem above.</p>\n<ul>\n<li>Test set is too small</li>\n</ul>\n<p>As you know, public test set is only 21% of the test data, which means 1992 * 0.21 ≒ 418 clips. It is possible that this number is too small to eliminate the randomness. In that case, solution is simple - stick to the CV - but there could be one problem in this policy as I describe below.</p>\n<ul>\n<li>Test annotation can be noisy</li>\n</ul>\n<p>Some of you have already noticed this, but train annotation is not perfect. There are some (call events of) species that do exist, but not just be annotated. This is okay for train data, because we can at least use annotated part of the data to train our model, but not very good if the annotation is in the same level in test dataset. If annotation in the test dataset has some missing labels, it would be quite problematic because even if our models correctly find some call events (with respect to the <strong>perfect</strong> annotation), it would be treated as false positives. I just asked to the host about this concern in <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197782#1101468\" target=\"_blank\">a different thread</a> recently, but no answer has been provided so far.</p>\n<p>Do you have similar problems? What do you think about my concerns? I'll be waiting for you comment :)</p>",
  "messages": [
    {
      "id": "1104865",
      "postDate": "12/07/2020 09:46:29",
      "content": "<p>Hi everyone,</p>\n<p>We have been confronting a problem of uncorrelated CV - LB recently. There are several possible explanations like below:</p>\n<ul>\n<li>Bad validation strategy</li>\n</ul>\n<p>This may well be the case, and until recently we were doing validation with LWLRAP in random crop windows of n seconds, rather than with LWLRAP at the clip level, which is the correct metric used in testing as <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/198132#1101848\" target=\"_blank\">pointed out in the CV - LB thread</a>, but even after moving to the correct metric, we still have the problem above.</p>\n<ul>\n<li>Test set is too small</li>\n</ul>\n<p>As you know, public test set is only 21% of the test data, which means 1992 * 0.21 ≒ 418 clips. It is possible that this number is too small to eliminate the randomness. In that case, solution is simple - stick to the CV - but there could be one problem in this policy as I describe below.</p>\n<ul>\n<li>Test annotation can be noisy</li>\n</ul>\n<p>Some of you have already noticed this, but train annotation is not perfect. There are some (call events of) species that do exist, but not just be annotated. This is okay for train data, because we can at least use annotated part of the data to train our model, but not very good if the annotation is in the same level in test dataset. If annotation in the test dataset has some missing labels, it would be quite problematic because even if our models correctly find some call events (with respect to the <strong>perfect</strong> annotation), it would be treated as false positives. I just asked to the host about this concern in <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197782#1101468\" target=\"_blank\">a different thread</a> recently, but no answer has been provided so far.</p>\n<p>Do you have similar problems? What do you think about my concerns? I'll be waiting for you comment :)</p>",
      "rawMarkdown": "Hi everyone,\n\nWe have been confronting a problem of uncorrelated CV - LB recently. There are several possible explanations like below:\n\n* Bad validation strategy\n\nThis may well be the case, and until recently we were doing validation with LWLRAP in random crop windows of n seconds, rather than with LWLRAP at the clip level, which is the correct metric used in testing as [pointed out in the CV - LB thread](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/198132#1101848), but even after moving to the correct metric, we still have the problem above.\n\n* Test set is too small\n\nAs you know, public test set is only 21% of the test data, which means 1992 \\* 0.21 ≒ 418 clips. It is possible that this number is too small to eliminate the randomness. In that case, solution is simple - stick to the CV - but there could be one problem in this policy as I describe below.\n\n* Test annotation can be noisy\n\nSome of you have already noticed this, but train annotation is not perfect. There are some (call events of) species that do exist, but not just be annotated. This is okay for train data, because we can at least use annotated part of the data to train our model, but not very good if the annotation is in the same level in test dataset. If annotation in the test dataset has some missing labels, it would be quite problematic because even if our models correctly find some call events (with respect to the **perfect** annotation), it would be treated as false positives. I just asked to the host about this concern in [a different thread](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197782#1101468) recently, but no answer has been provided so far.\n\nDo you have similar problems? What do you think about my concerns? I'll be waiting for you comment :)",
      "votes": null
    },
    {
      "id": "1104879",
      "postDate": "12/07/2020 10:07:26",
      "content": "<p>I have seen high LB randomness from beginning of this competition. With mitigations i can clamp it to ~0.02 range, but it still makes testing of little things quite problematic. For me validation accuracy at certain point of the training gives the best indication of LB performance.</p>\n<p>I thought that it was due to my usage of BCELoss and it being bad approximation for competition metric, but if you have the same problem with LWLRAP, then it must be something else.</p>\n<blockquote>\n  <p>Test set is too small<br>\n  As you know, public test set is only 21% of the test data, which means 1992 * 0.21 ≒ 418 clips</p>\n</blockquote>\n<p>This is probably it. Amount of birds in public test data could be in ~1000 events range or even lower.</p>\n<blockquote>\n  <p>If annotation in the test dataset has some missing labels, it would be quite problematic because even if our models correctly find some call events (with respect to the perfect annotation), it would be treated as false positives.</p>\n</blockquote>\n<p>In my experience, this is not a problem. Accuracy correlates with LB score nicely (with caveat of LB randomness). It would be a big issue if competition metric would be per-event, but with such low number of birds to label and per-clip labels i do not expect significant amount of missing annotations.</p>",
      "rawMarkdown": "I have seen high LB randomness from beginning of this competition. With mitigations i can clamp it to ~0.02 range, but it still makes testing of little things quite problematic. For me validation accuracy at certain point of the training gives the best indication of LB performance.\n\nI thought that it was due to my usage of BCELoss and it being bad approximation for competition metric, but if you have the same problem with LWLRAP, then it must be something else.\n\n> Test set is too small\n> As you know, public test set is only 21% of the test data, which means 1992 * 0.21 ≒ 418 clips\n\nThis is probably it. Amount of birds in public test data could be in ~1000 events range or even lower.\n\n> If annotation in the test dataset has some missing labels, it would be quite problematic because even if our models correctly find some call events (with respect to the perfect annotation), it would be treated as false positives.\n\nIn my experience, this is not a problem. Accuracy correlates with LB score nicely (with caveat of LB randomness). It would be a big issue if competition metric would be per-event, but with such low number of birds to label and per-clip labels i do not expect significant amount of missing annotations.",
      "votes": null
    },
    {
      "id": "1105190",
      "postDate": "12/07/2020 15:49:22",
      "content": "<p>I have asked some questions to the host in the discussion train_tp vs. train_fp but the host doesn't seem to respond to me during Dec.4- Dec.7(Today). I doubt the test ground truth is 100% correct.</p>\n<p>Also, I agree that the host didn't respond these days as well as in the discussion: Welcome to the Rainforest Connection challenge! 💔</p>",
      "rawMarkdown": "I have asked some questions to the host in the discussion train_tp vs. train_fp but the host doesn't seem to respond to me during Dec.4- Dec.7(Today). I doubt the test ground truth is 100% correct.\n\nAlso, I agree that the host didn't respond these days as well as in the discussion: Welcome to the Rainforest Connection challenge! 💔",
      "votes": null
    },
    {
      "id": "1105506",
      "postDate": "12/08/2020 00:05:56",
      "content": "<p>I also noticed inconsistent CV - LB results. How the test set is annotated is an important question, I hope the host can provide some answer. </p>",
      "rawMarkdown": "I also noticed inconsistent CV - LB results. How the test set is annotated is an important question, I hope the host can provide some answer.",
      "votes": null
    },
    {
      "id": "1105509",
      "postDate": "12/08/2020 00:14:56",
      "content": "<blockquote>\n  <p>For me validation accuracy at certain point of the training gives the best indication of LB performance.</p>\n</blockquote>\n<p>May I ask why do you use accuracy instead of LWLRAP?</p>\n<p>Seems you've somehow constructed a good validation strategy…good to know, thank you for sharing!</p>",
      "rawMarkdown": "> For me validation accuracy at certain point of the training gives the best indication of LB performance.\n\nMay I ask why do you use accuracy instead of LWLRAP?\n\nSeems you've somehow constructed a good validation strategy...good to know, thank you for sharing!",
      "votes": null
    },
    {
      "id": "1105517",
      "postDate": "12/08/2020 00:25:02",
      "content": "<p><a href=\"https://www.kaggle.com/fffrrt\" target=\"_blank\">@fffrrt</a> can I assume, from your answer and my intuition, that predicting each test .flac one single label correct would give us good enough result?</p>",
      "rawMarkdown": "fffrrt can I assume, from your answer and my intuition, that predicting each test .flac one single label correct would give us good enough result?",
      "votes": null
    },
    {
      "id": "1105859",
      "postDate": "12/08/2020 09:10:56",
      "content": "<blockquote>\n  <p>May I ask why do you use accuracy instead of LWLRAP?</p>\n</blockquote>\n<p>I used accuracy when i finally managed to reduce LB randomness to acceptable levels. Given how much headache it was, i did not touch that part of code until now.</p>\n<p>LWLRAP might be better, but at the time i settled for first working solution.</p>",
      "rawMarkdown": "> May I ask why do you use accuracy instead of LWLRAP?\n\nI used accuracy when i finally managed to reduce LB randomness to acceptable levels. Given how much headache it was, i did not touch that part of code until now.\n\nLWLRAP might be better, but at the time i settled for first working solution.",
      "votes": null
    },
    {
      "id": "1107042",
      "postDate": "12/09/2020 10:40:57",
      "content": "<p><a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> Thanks for your sharing! May I ask what degree of CV / LB inconsistency are you guys experiencing?<br>\nI did a resnest50 baseline, training on 15s crops and doing validation on whole clip using LWLRAP. When I train 5fold for 30 epochs, 5fold CV is .809 and LB is .824. However, when I train 5fold for 35 epochs, 5fold CV is .820 while LB is .760. LB is indeed very strange here…</p>",
      "rawMarkdown": "hidehisaarai1213 Thanks for your sharing! May I ask what degree of CV / LB inconsistency are you guys experiencing?\nI did a resnest50 baseline, training on 15s crops and doing validation on whole clip using LWLRAP. When I train 5fold for 30 epochs, 5fold CV is .809 and LB is .824. However, when I train 5fold for 35 epochs, 5fold CV is .820 while LB is .760. LB is indeed very strange here...",
      "votes": null
    },
    {
      "id": "1107124",
      "postDate": "12/09/2020 12:14:09",
      "content": "<p>We observe the same level of inconsistency as yours.</p>\n<blockquote>\n  <p>When I train 5fold for 30 epochs, 5fold CV is .809 and LB is .824. However, when I train 5fold for 35 epochs, 5fold CV is .820 while LB is .760.</p>\n</blockquote>\n<p>We also see this kind of things - still working on figuring out why…</p>",
      "rawMarkdown": "We observe the same level of inconsistency as yours.\n\n> When I train 5fold for 30 epochs, 5fold CV is .809 and LB is .824. However, when I train 5fold for 35 epochs, 5fold CV is .820 while LB is .760.\n\nWe also see this kind of things - still working on figuring out why...",
      "votes": null
    },
    {
      "id": "1107273",
      "postDate": "12/09/2020 14:48:55",
      "content": "<p>Interesting. thanks for sharing.</p>\n<p>I reached 0.86 CV with resnest50, 70 epochs and extensive augmentation but LB is in 0.6 range. I was convinced I messed up somewhere but now you got me thinking. will report when/if i escape from this pothole</p>",
      "rawMarkdown": "Interesting. thanks for sharing.\n\nI reached 0.86 CV with resnest50, 70 epochs and extensive augmentation but LB is in 0.6 range. I was convinced I messed up somewhere but now you got me thinking. will report when/if i escape from this pothole",
      "votes": null
    },
    {
      "id": "1107371",
      "postDate": "12/09/2020 16:22:28",
      "content": "<blockquote>\n  <p>until recently we were doing validation with LWLRAP in random crop windows of n seconds, rather than with LWLRAP at the clip level</p>\n</blockquote>\n<p>Do you split train and validation by file/clip, train on n second chunks then validate on full 60 second clips?</p>",
      "rawMarkdown": "> until recently we were doing validation with LWLRAP in random crop windows of n seconds, rather than with LWLRAP at the clip level\n\nDo you split train and validation by file/clip, train on n second chunks then validate on full 60 second clips?",
      "votes": null
    },
    {
      "id": "1107534",
      "postDate": "12/09/2020 18:57:27",
      "content": "<p><a href=\"https://www.kaggle.com/hi170840\" target=\"_blank\">@hi170840</a> <a href=\"https://www.kaggle.com/roguekk007\" target=\"_blank\">@roguekk007</a> Have you noticed this on models without SED? I am not using SED at the moment so maybe that's the reason for my huge gap. </p>",
      "rawMarkdown": "hi170840 @roguekk007 Have you noticed this on models without SED? I am not using SED at the moment so maybe that's the reason for my huge gap.",
      "votes": null
    },
    {
      "id": "1107557",
      "postDate": "12/09/2020 19:15:44",
      "content": "<p>without SED i get really nice correlation. For example below is one experiments (using correct metric):</p>\n<pre><code>1 FOLD - 0.836\n2 FOLD - 0.778\n3 FOLD - 0.835\n4 FOLD - 0.817\n5 FOLD - 0.801\n</code></pre>\n<p>I calculated this 5 fold by combining all validation prediction with targets (instead of taking average of CV scores) </p>\n<pre><code>5 FOLD CV score: 0.8142345929598827\nLB        score: 0.815\n</code></pre>\n<p>now there are other experiments which are done without SED, they are all nicely correlated and follow similar trends as score above … </p>\n<p>I think the CV LB correlation come due to how training is done (since they are multiple ways to train on this dataset). </p>",
      "rawMarkdown": "without SED i get really nice correlation. For example below is one experiments (using correct metric):\n\n```\n1 FOLD - 0.836\n2 FOLD - 0.778\n3 FOLD - 0.835\n4 FOLD - 0.817\n5 FOLD - 0.801\n```\n\nI calculated this 5 fold by combining all validation prediction with targets (instead of taking average of CV scores) \n\n```\n5 FOLD CV score: 0.8142345929598827\nLB        score: 0.815\n```\n\nnow there are other experiments which are done without SED, they are all nicely correlated and follow similar trends as score above ... \n\nI think the CV LB correlation come due to how training is done (since they are multiple ways to train on this dataset).",
      "votes": null
    },
    {
      "id": "1107593",
      "postDate": "12/09/2020 19:51:03",
      "content": "<p>thx <a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> and good luck </p>",
      "rawMarkdown": "thx @drhabib and good luck",
      "votes": null
    },
    {
      "id": "1107868",
      "postDate": "12/10/2020 02:56:48",
      "content": "<p>I split by clip and train on n second chunks. In the validation phase, I first cut 60 second clip to n second chunks and then get the prediction for each chunk. After prediction, I aggregate the prediction by taking max probability of each class in each clip.</p>",
      "rawMarkdown": "I split by clip and train on n second chunks. In the validation phase, I first cut 60 second clip to n second chunks and then get the prediction for each chunk. After prediction, I aggregate the prediction by taking max probability of each class in each clip.",
      "votes": null
    },
    {
      "id": "1108337",
      "postDate": "12/10/2020 14:56:02",
      "content": "<p>Thank you! I aggregate by taking the max of the n second clip when creating a submission so makes sense to also do the same for validation. </p>",
      "rawMarkdown": "Thank you! I aggregate by taking the max of the n second clip when creating a submission so makes sense to also do the same for validation.",
      "votes": null
    },
    {
      "id": "1108342",
      "postDate": "12/10/2020 14:57:57",
      "content": "<p>I also make the prediction for the test in the same way :)</p>",
      "rawMarkdown": "I also make the prediction for the test in the same way :)",
      "votes": null
    },
    {
      "id": "1110885",
      "postDate": "12/13/2020 07:10:51",
      "content": "<p><a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> Hi! I wonder if there's any reason behind aggregating predictions for several chunks instead of just feeding a larger image into CNN? (which I believe you did using SED in Cornell birdcall, correct me if I'm wrong</p>",
      "rawMarkdown": "hidehisaarai1213 Hi! I wonder if there's any reason behind aggregating predictions for several chunks instead of just feeding a larger image into CNN? (which I believe you did using SED in Cornell birdcall, correct me if I'm wrong",
      "votes": null
    },
    {
      "id": "1110926",
      "postDate": "12/13/2020 07:58:06",
      "content": "<p>Well there aren't so much concrete reason behind this 😅 I just felt like to do so, but maybe the other way is better (I haven't tried yet)</p>",
      "rawMarkdown": "Well there aren't so much concrete reason behind this 😅 I just felt like to do so, but maybe the other way is better (I haven't tried yet)",
      "votes": null
    },
    {
      "id": "1136270",
      "postDate": "01/02/2021 22:40:06",
      "content": "<p>I have also seen this kind of behavior. The same CV can give me a difference of few points in LB … <br>\nMoreover, I have also seen that a small change in my training such as modifying the weight decay value of my optimizer can change the LB of few points which is quite weird to me. The training seems quite unstable in this competition, at least with the naive approach (cropping few seconds, doing a prediction on it and evaluate at the segment level)</p>",
      "rawMarkdown": "I have also seen this kind of behavior. The same CV can give me a difference of few points in LB ... \nMoreover, I have also seen that a small change in my training such as modifying the weight decay value of my optimizer can change the LB of few points which is quite weird to me. The training seems quite unstable in this competition, at least with the naive approach (cropping few seconds, doing a prediction on it and evaluate at the segment level)",
      "votes": null
    },
    {
      "id": "1136271",
      "postDate": "01/02/2021 22:41:40",
      "content": "<p>May I ask what do you mean when you are talking about SED training ?</p>",
      "rawMarkdown": "May I ask what do you mean when you are talking about SED training ?",
      "votes": null
    },
    {
      "id": "1166563",
      "postDate": "01/23/2021 17:29:01",
      "content": "<p><a href=\"https://www.kaggle.com/fffrrt\" target=\"_blank\">@fffrrt</a> <br>\nBeing driven crazy by LB - CV randomness  I am using LWLRAP and doing validation on clip level.</p>\n<p>By using accuracy as CV metrics, are you using multi-label accuracy or single label accuracy ?   If it is multi-label accuracy, what sigmoid probs threshold do you use ?</p>",
      "rawMarkdown": "fffrrt \nBeing driven crazy by LB - CV randomness  I am using LWLRAP and doing validation on clip level.\n\nBy using accuracy as CV metrics, are you using multi-label accuracy or single label accuracy ?   If it is multi-label accuracy, what sigmoid probs threshold do you use ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1104879,
      "author_name": "fffrrt",
      "author_url": "",
      "post_date": "12/07/2020 10:07:26",
      "content": "<p>I have seen high LB randomness from beginning of this competition. With mitigations i can clamp it to ~0.02 range, but it still makes testing of little things quite problematic. For me validation accuracy at certain point of the training gives the best indication of LB performance.</p>\n<p>I thought that it was due to my usage of BCELoss and it being bad approximation for competition metric, but if you have the same problem with LWLRAP, then it must be something else.</p>\n<blockquote>\n  <p>Test set is too small<br>\n  As you know, public test set is only 21% of the test data, which means 1992 * 0.21 ≒ 418 clips</p>\n</blockquote>\n<p>This is probably it. Amount of birds in public test data could be in ~1000 events range or even lower.</p>\n<blockquote>\n  <p>If annotation in the test dataset has some missing labels, it would be quite problematic because even if our models correctly find some call events (with respect to the perfect annotation), it would be treated as false positives.</p>\n</blockquote>\n<p>In my experience, this is not a problem. Accuracy correlates with LB score nicely (with caveat of LB randomness). It would be a big issue if competition metric would be per-event, but with such low number of birds to label and per-clip labels i do not expect significant amount of missing annotations.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1105509,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "12/08/2020 00:14:56",
          "content": "<blockquote>\n  <p>For me validation accuracy at certain point of the training gives the best indication of LB performance.</p>\n</blockquote>\n<p>May I ask why do you use accuracy instead of LWLRAP?</p>\n<p>Seems you've somehow constructed a good validation strategy…good to know, thank you for sharing!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1105517,
          "author_name": "wubinbai",
          "author_url": "",
          "post_date": "12/08/2020 00:25:02",
          "content": "<p><a href=\"https://www.kaggle.com/fffrrt\" target=\"_blank\">@fffrrt</a> can I assume, from your answer and my intuition, that predicting each test .flac one single label correct would give us good enough result?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1105859,
          "author_name": "fffrrt",
          "author_url": "",
          "post_date": "12/08/2020 09:10:56",
          "content": "<blockquote>\n  <p>May I ask why do you use accuracy instead of LWLRAP?</p>\n</blockquote>\n<p>I used accuracy when i finally managed to reduce LB randomness to acceptable levels. Given how much headache it was, i did not touch that part of code until now.</p>\n<p>LWLRAP might be better, but at the time i settled for first working solution.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1166563,
          "author_name": "garfieldchh",
          "author_url": "",
          "post_date": "01/23/2021 17:29:01",
          "content": "<p><a href=\"https://www.kaggle.com/fffrrt\" target=\"_blank\">@fffrrt</a> <br>\nBeing driven crazy by LB - CV randomness  I am using LWLRAP and doing validation on clip level.</p>\n<p>By using accuracy as CV metrics, are you using multi-label accuracy or single label accuracy ?   If it is multi-label accuracy, what sigmoid probs threshold do you use ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1105190,
      "author_name": "wubinbai",
      "author_url": "",
      "post_date": "12/07/2020 15:49:22",
      "content": "<p>I have asked some questions to the host in the discussion train_tp vs. train_fp but the host doesn't seem to respond to me during Dec.4- Dec.7(Today). I doubt the test ground truth is 100% correct.</p>\n<p>Also, I agree that the host didn't respond these days as well as in the discussion: Welcome to the Rainforest Connection challenge! 💔</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1105506,
      "author_name": "mnpinto",
      "author_url": "",
      "post_date": "12/08/2020 00:05:56",
      "content": "<p>I also noticed inconsistent CV - LB results. How the test set is annotated is an important question, I hope the host can provide some answer. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1107042,
      "author_name": "roguekk007",
      "author_url": "",
      "post_date": "12/09/2020 10:40:57",
      "content": "<p><a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> Thanks for your sharing! May I ask what degree of CV / LB inconsistency are you guys experiencing?<br>\nI did a resnest50 baseline, training on 15s crops and doing validation on whole clip using LWLRAP. When I train 5fold for 30 epochs, 5fold CV is .809 and LB is .824. However, when I train 5fold for 35 epochs, 5fold CV is .820 while LB is .760. LB is indeed very strange here…</p>",
      "votes": null,
      "replies": [
        {
          "id": 1107124,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "12/09/2020 12:14:09",
          "content": "<p>We observe the same level of inconsistency as yours.</p>\n<blockquote>\n  <p>When I train 5fold for 30 epochs, 5fold CV is .809 and LB is .824. However, when I train 5fold for 35 epochs, 5fold CV is .820 while LB is .760.</p>\n</blockquote>\n<p>We also see this kind of things - still working on figuring out why…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1107273,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "12/09/2020 14:48:55",
          "content": "<p>Interesting. thanks for sharing.</p>\n<p>I reached 0.86 CV with resnest50, 70 epochs and extensive augmentation but LB is in 0.6 range. I was convinced I messed up somewhere but now you got me thinking. will report when/if i escape from this pothole</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1107534,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "12/09/2020 18:57:27",
          "content": "<p><a href=\"https://www.kaggle.com/hi170840\" target=\"_blank\">@hi170840</a> <a href=\"https://www.kaggle.com/roguekk007\" target=\"_blank\">@roguekk007</a> Have you noticed this on models without SED? I am not using SED at the moment so maybe that's the reason for my huge gap. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1107557,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "12/09/2020 19:15:44",
          "content": "<p>without SED i get really nice correlation. For example below is one experiments (using correct metric):</p>\n<pre><code>1 FOLD - 0.836\n2 FOLD - 0.778\n3 FOLD - 0.835\n4 FOLD - 0.817\n5 FOLD - 0.801\n</code></pre>\n<p>I calculated this 5 fold by combining all validation prediction with targets (instead of taking average of CV scores) </p>\n<pre><code>5 FOLD CV score: 0.8142345929598827\nLB        score: 0.815\n</code></pre>\n<p>now there are other experiments which are done without SED, they are all nicely correlated and follow similar trends as score above … </p>\n<p>I think the CV LB correlation come due to how training is done (since they are multiple ways to train on this dataset). </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1107593,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "12/09/2020 19:51:03",
          "content": "<p>thx <a href=\"https://www.kaggle.com/drhabib\" target=\"_blank\">@drhabib</a> and good luck </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1136271,
          "author_name": "ludovick",
          "author_url": "",
          "post_date": "01/02/2021 22:41:40",
          "content": "<p>May I ask what do you mean when you are talking about SED training ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1107371,
      "author_name": "jackvial",
      "author_url": "",
      "post_date": "12/09/2020 16:22:28",
      "content": "<blockquote>\n  <p>until recently we were doing validation with LWLRAP in random crop windows of n seconds, rather than with LWLRAP at the clip level</p>\n</blockquote>\n<p>Do you split train and validation by file/clip, train on n second chunks then validate on full 60 second clips?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1107868,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "12/10/2020 02:56:48",
          "content": "<p>I split by clip and train on n second chunks. In the validation phase, I first cut 60 second clip to n second chunks and then get the prediction for each chunk. After prediction, I aggregate the prediction by taking max probability of each class in each clip.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1108337,
          "author_name": "jackvial",
          "author_url": "",
          "post_date": "12/10/2020 14:56:02",
          "content": "<p>Thank you! I aggregate by taking the max of the n second clip when creating a submission so makes sense to also do the same for validation. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1108342,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "12/10/2020 14:57:57",
          "content": "<p>I also make the prediction for the test in the same way :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1110885,
          "author_name": "roguekk007",
          "author_url": "",
          "post_date": "12/13/2020 07:10:51",
          "content": "<p><a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> Hi! I wonder if there's any reason behind aggregating predictions for several chunks instead of just feeding a larger image into CNN? (which I believe you did using SED in Cornell birdcall, correct me if I'm wrong</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1110926,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "12/13/2020 07:58:06",
          "content": "<p>Well there aren't so much concrete reason behind this 😅 I just felt like to do so, but maybe the other way is better (I haven't tried yet)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1136270,
      "author_name": "ludovick",
      "author_url": "",
      "post_date": "01/02/2021 22:40:06",
      "content": "<p>I have also seen this kind of behavior. The same CV can give me a difference of few points in LB … <br>\nMoreover, I have also seen that a small change in my training such as modifying the weight decay value of my optimizer can change the LB of few points which is quite weird to me. The training seems quite unstable in this competition, at least with the naive approach (cropping few seconds, doing a prediction on it and evaluate at the segment level)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1104865": "Hi everyone,\n\nWe have been confronting a problem of uncorrelated CV - LB recently. There are several possible explanations like below:\n\n* Bad validation strategy\n\nThis may well be the case, and until recently we were doing validation with LWLRAP in random crop windows of n seconds, rather than with LWLRAP at the clip level, which is the correct metric used in testing as [pointed out in the CV - LB thread](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/198132#1101848), but even after moving to the correct metric, we still have the problem above.\n\n* Test set is too small\n\nAs you know, public test set is only 21% of the test data, which means 1992 \\* 0.21 ≒ 418 clips. It is possible that this number is too small to eliminate the randomness. In that case, solution is simple - stick to the CV - but there could be one problem in this policy as I describe below.\n\n* Test annotation can be noisy\n\nSome of you have already noticed this, but train annotation is not perfect. There are some (call events of) species that do exist, but not just be annotated. This is okay for train data, because we can at least use annotated part of the data to train our model, but not very good if the annotation is in the same level in test dataset. If annotation in the test dataset has some missing labels, it would be quite problematic because even if our models correctly find some call events (with respect to the **perfect** annotation), it would be treated as false positives. I just asked to the host about this concern in [a different thread](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/197782#1101468) recently, but no answer has been provided so far.\n\nDo you have similar problems? What do you think about my concerns? I'll be waiting for you comment :)",
    "1104879": "I have seen high LB randomness from beginning of this competition. With mitigations i can clamp it to ~0.02 range, but it still makes testing of little things quite problematic. For me validation accuracy at certain point of the training gives the best indication of LB performance.\n\nI thought that it was due to my usage of BCELoss and it being bad approximation for competition metric, but if you have the same problem with LWLRAP, then it must be something else.\n\n> Test set is too small\n> As you know, public test set is only 21% of the test data, which means 1992 * 0.21 ≒ 418 clips\n\nThis is probably it. Amount of birds in public test data could be in ~1000 events range or even lower.\n\n> If annotation in the test dataset has some missing labels, it would be quite problematic because even if our models correctly find some call events (with respect to the perfect annotation), it would be treated as false positives.\n\nIn my experience, this is not a problem. Accuracy correlates with LB score nicely (with caveat of LB randomness). It would be a big issue if competition metric would be per-event, but with such low number of birds to label and per-clip labels i do not expect significant amount of missing annotations.",
    "1105190": "I have asked some questions to the host in the discussion train_tp vs. train_fp but the host doesn't seem to respond to me during Dec.4- Dec.7(Today). I doubt the test ground truth is 100% correct.\n\nAlso, I agree that the host didn't respond these days as well as in the discussion: Welcome to the Rainforest Connection challenge! 💔",
    "1105506": "I also noticed inconsistent CV - LB results. How the test set is annotated is an important question, I hope the host can provide some answer.",
    "1105509": "> For me validation accuracy at certain point of the training gives the best indication of LB performance.\n\nMay I ask why do you use accuracy instead of LWLRAP?\n\nSeems you've somehow constructed a good validation strategy...good to know, thank you for sharing!",
    "1105517": "fffrrt can I assume, from your answer and my intuition, that predicting each test .flac one single label correct would give us good enough result?",
    "1105859": "> May I ask why do you use accuracy instead of LWLRAP?\n\nI used accuracy when i finally managed to reduce LB randomness to acceptable levels. Given how much headache it was, i did not touch that part of code until now.\n\nLWLRAP might be better, but at the time i settled for first working solution.",
    "1107042": "hidehisaarai1213 Thanks for your sharing! May I ask what degree of CV / LB inconsistency are you guys experiencing?\nI did a resnest50 baseline, training on 15s crops and doing validation on whole clip using LWLRAP. When I train 5fold for 30 epochs, 5fold CV is .809 and LB is .824. However, when I train 5fold for 35 epochs, 5fold CV is .820 while LB is .760. LB is indeed very strange here...",
    "1107124": "We observe the same level of inconsistency as yours.\n\n> When I train 5fold for 30 epochs, 5fold CV is .809 and LB is .824. However, when I train 5fold for 35 epochs, 5fold CV is .820 while LB is .760.\n\nWe also see this kind of things - still working on figuring out why...",
    "1107273": "Interesting. thanks for sharing.\n\nI reached 0.86 CV with resnest50, 70 epochs and extensive augmentation but LB is in 0.6 range. I was convinced I messed up somewhere but now you got me thinking. will report when/if i escape from this pothole",
    "1107371": "> until recently we were doing validation with LWLRAP in random crop windows of n seconds, rather than with LWLRAP at the clip level\n\nDo you split train and validation by file/clip, train on n second chunks then validate on full 60 second clips?",
    "1107534": "hi170840 @roguekk007 Have you noticed this on models without SED? I am not using SED at the moment so maybe that's the reason for my huge gap.",
    "1107557": "without SED i get really nice correlation. For example below is one experiments (using correct metric):\n\n```\n1 FOLD - 0.836\n2 FOLD - 0.778\n3 FOLD - 0.835\n4 FOLD - 0.817\n5 FOLD - 0.801\n```\n\nI calculated this 5 fold by combining all validation prediction with targets (instead of taking average of CV scores) \n\n```\n5 FOLD CV score: 0.8142345929598827\nLB        score: 0.815\n```\n\nnow there are other experiments which are done without SED, they are all nicely correlated and follow similar trends as score above ... \n\nI think the CV LB correlation come due to how training is done (since they are multiple ways to train on this dataset).",
    "1107593": "thx @drhabib and good luck",
    "1107868": "I split by clip and train on n second chunks. In the validation phase, I first cut 60 second clip to n second chunks and then get the prediction for each chunk. After prediction, I aggregate the prediction by taking max probability of each class in each clip.",
    "1108337": "Thank you! I aggregate by taking the max of the n second clip when creating a submission so makes sense to also do the same for validation.",
    "1108342": "I also make the prediction for the test in the same way :)",
    "1110885": "hidehisaarai1213 Hi! I wonder if there's any reason behind aggregating predictions for several chunks instead of just feeding a larger image into CNN? (which I believe you did using SED in Cornell birdcall, correct me if I'm wrong",
    "1110926": "Well there aren't so much concrete reason behind this 😅 I just felt like to do so, but maybe the other way is better (I haven't tried yet)",
    "1136270": "I have also seen this kind of behavior. The same CV can give me a difference of few points in LB ... \nMoreover, I have also seen that a small change in my training such as modifying the weight decay value of my optimizer can change the LB of few points which is quite weird to me. The training seems quite unstable in this competition, at least with the naive approach (cropping few seconds, doing a prediction on it and evaluate at the segment level)",
    "1136271": "May I ask what do you mean when you are talking about SED training ?",
    "1166563": "fffrrt \nBeing driven crazy by LB - CV randomness  I am using LWLRAP and doing validation on clip level.\n\nBy using accuracy as CV metrics, are you using multi-label accuracy or single label accuracy ?   If it is multi-label accuracy, what sigmoid probs threshold do you use ?"
  },
  "source": "meta"
}