{
  "id": 209684,
  "title": "Why are CV and LB different?",
  "url": "/competitions/rfcx-species-audio-detection/discussion/209684",
  "author_name": "",
  "post_date": "2021-01-08T08:17:50.576686900Z",
  "votes": 30,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Why are CV and LB different?<br>\nI've come up with three reasons I think. But I may be wrong.</p>\n<h1>1. Domain shift</h1>\n<p>In the <a href=\"https://www.kaggle.com/c/birdsong-recognition\" target=\"_blank\">last competition</a>, main theme is <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183204\" target=\"_blank\">domain shift and missing labels</a>.<br>\nIn this point, this competition is similar.</p>\n<p>In this competition, I think domain shift exist. But this is not a big problem.<br>\nIn training sound, it contains noisy sound (white noise and pink noise).<br>\nBut test sound is clean(or a little noisy sound). </p>\n<p>But I think this domain shift has <strong>no negative impact.</strong><br>\nBecause it's a domain shift from noisy to clean.</p>\n<h1>2. Missing labels</h1>\n<p>I <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/209040\" target=\"_blank\">reported</a> that missing labels exist in tp sound. <br>\nIf validation data contains missing labels, your CV is different from LB.</p>\n<p>For example, y_label is not correct(contains missing label).</p>\n<ul>\n<li>y_label = [1,0,0]</li>\n<li>y_true = [1,1,0]</li>\n</ul>\n<p>2nd annotation is missed in y_label. Then y_predict = [1,0,0] is best CV in your local. But  y_predict=[1,1,0] is better in the test. In other words, no matter how much you improve your CV, it will not always work in a testing environment. <strong>The validation data is not reliable.</strong></p>\n<h1>3. A gap training and prediction</h1>\n<p>Many participants seem to use <a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection?scriptVersionId=40731755\" target=\"_blank\">SED</a>. And I also use SED.<br>\nIf you use SED(PANNs architecture), you should be careful.</p>\n<p>If you use SED like below, there is a gap training and prediction.</p>\n<ul>\n<li>training with weak label</li>\n<li>prediction with <code>framewise_output</code><br>\n(What is framewise_output?(Now I'm writing \"How to use SED\". Coming soon.))</li>\n</ul>\n<p>Weak label training is fitted <code>clipwise_output</code> <strong>not <code>framewise_output</code>.</strong> <code>clipwise_output</code> is a time-compressed version of <code>framewise_output</code>. <code>clipwise_output</code> is correlated with <code>framewise_output</code>, but not equal. </p>\n<p>Comparatively, <code>framewise_output</code> prediction is good at short sound event. In this competition, there is many short sound event. Therefore I use <code>framewise_output</code> for prediction. But there is a gap training and prediction. Then CV(<code>clipwise_output</code> training) is different from LB(<code>framewise_output</code> prediction).</p>\n<p>If you change validation strategy, you may avoid this problem. For example, in validation phase you use  np.max(<code>framewise_output</code>, axis=time) instead of  <code>clipwise_output</code>. This strategy may improve a gap CV and LB.</p>\n<h1>4. Conclusion</h1>\n<p>Finally, I don't have any above countermeasures. Because I cannot fix <strong>2. Missing labels</strong>.</p>\n<p>My strategy is \"trust LB\" not \"trust CV\". Fortunately, <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/207901#1134198\" target=\"_blank\">public LB is almost the same as private LB</a>. I believe shaking almost never happens.</p>",
  "messages": [
    {
      "id": "1144085",
      "postDate": "01/08/2021 08:17:50",
      "content": "<p>Why are CV and LB different?<br>\nI've come up with three reasons I think. But I may be wrong.</p>\n<h1>1. Domain shift</h1>\n<p>In the <a href=\"https://www.kaggle.com/c/birdsong-recognition\" target=\"_blank\">last competition</a>, main theme is <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/183204\" target=\"_blank\">domain shift and missing labels</a>.<br>\nIn this point, this competition is similar.</p>\n<p>In this competition, I think domain shift exist. But this is not a big problem.<br>\nIn training sound, it contains noisy sound (white noise and pink noise).<br>\nBut test sound is clean(or a little noisy sound). </p>\n<p>But I think this domain shift has <strong>no negative impact.</strong><br>\nBecause it's a domain shift from noisy to clean.</p>\n<h1>2. Missing labels</h1>\n<p>I <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/209040\" target=\"_blank\">reported</a> that missing labels exist in tp sound. <br>\nIf validation data contains missing labels, your CV is different from LB.</p>\n<p>For example, y_label is not correct(contains missing label).</p>\n<ul>\n<li>y_label = [1,0,0]</li>\n<li>y_true = [1,1,0]</li>\n</ul>\n<p>2nd annotation is missed in y_label. Then y_predict = [1,0,0] is best CV in your local. But  y_predict=[1,1,0] is better in the test. In other words, no matter how much you improve your CV, it will not always work in a testing environment. <strong>The validation data is not reliable.</strong></p>\n<h1>3. A gap training and prediction</h1>\n<p>Many participants seem to use <a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection?scriptVersionId=40731755\" target=\"_blank\">SED</a>. And I also use SED.<br>\nIf you use SED(PANNs architecture), you should be careful.</p>\n<p>If you use SED like below, there is a gap training and prediction.</p>\n<ul>\n<li>training with weak label</li>\n<li>prediction with <code>framewise_output</code><br>\n(What is framewise_output?(Now I'm writing \"How to use SED\". Coming soon.))</li>\n</ul>\n<p>Weak label training is fitted <code>clipwise_output</code> <strong>not <code>framewise_output</code>.</strong> <code>clipwise_output</code> is a time-compressed version of <code>framewise_output</code>. <code>clipwise_output</code> is correlated with <code>framewise_output</code>, but not equal. </p>\n<p>Comparatively, <code>framewise_output</code> prediction is good at short sound event. In this competition, there is many short sound event. Therefore I use <code>framewise_output</code> for prediction. But there is a gap training and prediction. Then CV(<code>clipwise_output</code> training) is different from LB(<code>framewise_output</code> prediction).</p>\n<p>If you change validation strategy, you may avoid this problem. For example, in validation phase you use  np.max(<code>framewise_output</code>, axis=time) instead of  <code>clipwise_output</code>. This strategy may improve a gap CV and LB.</p>\n<h1>4. Conclusion</h1>\n<p>Finally, I don't have any above countermeasures. Because I cannot fix <strong>2. Missing labels</strong>.</p>\n<p>My strategy is \"trust LB\" not \"trust CV\". Fortunately, <a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/207901#1134198\" target=\"_blank\">public LB is almost the same as private LB</a>. I believe shaking almost never happens.</p>",
      "rawMarkdown": "Why are CV and LB different?\nI've come up with three reasons I think. But I may be wrong.\n\n# 1. Domain shift\nIn the [last competition](https://www.kaggle.com/c/birdsong-recognition), main theme is [domain shift and missing labels](https://www.kaggle.com/c/birdsong-recognition/discussion/183204).\nIn this point, this competition is similar.\n\nIn this competition, I think domain shift exist. But this is not a big problem.\nIn training sound, it contains noisy sound (white noise and pink noise).\nBut test sound is clean(or a little noisy sound). \n\nBut I think this domain shift has **no negative impact.**\nBecause it's a domain shift from noisy to clean.\n\n# 2. Missing labels\nI [reported](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/209040) that missing labels exist in tp sound. \nIf validation data contains missing labels, your CV is different from LB.\n\nFor example, y_label is not correct(contains missing label).\n+ y_label = [1,0,0]\n+ y_true = [1,1,0]\n\n2nd annotation is missed in y_label. Then y_predict = [1,0,0] is best CV in your local. But  y_predict=[1,1,0] is better in the test. In other words, no matter how much you improve your CV, it will not always work in a testing environment. **The validation data is not reliable.**\n\n# 3. A gap training and prediction\nMany participants seem to use [SED](https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection?scriptVersionId=40731755). And I also use SED.\nIf you use SED(PANNs architecture), you should be careful.\n\nIf you use SED like below, there is a gap training and prediction.\n+ training with weak label\n+ prediction with ```framewise_output```\n(What is framewise_output?(Now I'm writing \"How to use SED\". Coming soon.))\n\nWeak label training is fitted ```clipwise_output``` **not ```framewise_output```.** ```clipwise_output``` is a time-compressed version of ```framewise_output```. ```clipwise_output``` is correlated with ```framewise_output```, but not equal. \n\nComparatively, ```framewise_output``` prediction is good at short sound event. In this competition, there is many short sound event. Therefore I use ```framewise_output``` for prediction. But there is a gap training and prediction. Then CV(```clipwise_output``` training) is different from LB(```framewise_output``` prediction).\n\nIf you change validation strategy, you may avoid this problem. For example, in validation phase you use  np.max(```framewise_output```, axis=time) instead of  ```clipwise_output```. This strategy may improve a gap CV and LB.\n\n# 4. Conclusion\nFinally, I don't have any above countermeasures. Because I cannot fix **2. Missing labels**.\n\nMy strategy is \"trust LB\" not \"trust CV\". Fortunately, [public LB is almost the same as private LB](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/207901#1134198). I believe shaking almost never happens.",
      "votes": null
    },
    {
      "id": "1144216",
      "postDate": "01/08/2021 10:08:07",
      "content": "<p><a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> Thanks.</p>\n<p>how do you predict at inference time</p>",
      "rawMarkdown": "shinmurashinmura Thanks.\n\nhow do you predict at inference time",
      "votes": null
    },
    {
      "id": "1144242",
      "postDate": "01/08/2021 10:44:15",
      "content": "<p>Thank you. Are you train on clipwise_output?</p>",
      "rawMarkdown": "Thank you. Are you train on clipwise_output?",
      "votes": null
    },
    {
      "id": "1144429",
      "postDate": "01/08/2021 13:09:38",
      "content": "<p>Yes, I trained with clipwise_output and prediction is \"framewise_outuput.</p>",
      "rawMarkdown": "Yes, I trained with clipwise_output and prediction is \"framewise_outuput.",
      "votes": null
    },
    {
      "id": "1144435",
      "postDate": "01/08/2021 13:11:53",
      "content": "<p><a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/208830#1139552\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/208830#1139552</a></p>",
      "rawMarkdown": "https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/208830#1139552",
      "votes": null
    },
    {
      "id": "1144445",
      "postDate": "01/08/2021 13:18:39",
      "content": "<p><a href=\"https://www.kaggle.com/shinmura0\" target=\"_blank\">@shinmura0</a> how is your individual fold CV?</p>",
      "rawMarkdown": "shinmura0 how is your individual fold CV?",
      "votes": null
    },
    {
      "id": "1149876",
      "postDate": "01/12/2021 08:16:19",
      "content": "<p>I used StratifiedKFold(sklearn).<br>\nAnd n_splits=5.</p>",
      "rawMarkdown": "I used StratifiedKFold(sklearn).\nAnd n_splits=5.",
      "votes": null
    },
    {
      "id": "1155222",
      "postDate": "01/16/2021 09:59:42",
      "content": "<p>Makes sense <a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> </p>",
      "rawMarkdown": "Makes sense @shinmurashinmura",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1144216,
      "author_name": "gopidurgaprasad",
      "author_url": "",
      "post_date": "01/08/2021 10:08:07",
      "content": "<p><a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> Thanks.</p>\n<p>how do you predict at inference time</p>",
      "votes": null,
      "replies": [
        {
          "id": 1144435,
          "author_name": "shinmurashinmura",
          "author_url": "",
          "post_date": "01/08/2021 13:11:53",
          "content": "<p><a href=\"https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/208830#1139552\" target=\"_blank\">https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/208830#1139552</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1144242,
      "author_name": "vlomme",
      "author_url": "",
      "post_date": "01/08/2021 10:44:15",
      "content": "<p>Thank you. Are you train on clipwise_output?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1144429,
          "author_name": "shinmurashinmura",
          "author_url": "",
          "post_date": "01/08/2021 13:09:38",
          "content": "<p>Yes, I trained with clipwise_output and prediction is \"framewise_outuput.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1144445,
      "author_name": "tugstugi",
      "author_url": "",
      "post_date": "01/08/2021 13:18:39",
      "content": "<p><a href=\"https://www.kaggle.com/shinmura0\" target=\"_blank\">@shinmura0</a> how is your individual fold CV?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1149876,
          "author_name": "shinmurashinmura",
          "author_url": "",
          "post_date": "01/12/2021 08:16:19",
          "content": "<p>I used StratifiedKFold(sklearn).<br>\nAnd n_splits=5.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1155222,
      "author_name": "pattnaiksatyajit",
      "author_url": "",
      "post_date": "01/16/2021 09:59:42",
      "content": "<p>Makes sense <a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1144085": "Why are CV and LB different?\nI've come up with three reasons I think. But I may be wrong.\n\n# 1. Domain shift\nIn the [last competition](https://www.kaggle.com/c/birdsong-recognition), main theme is [domain shift and missing labels](https://www.kaggle.com/c/birdsong-recognition/discussion/183204).\nIn this point, this competition is similar.\n\nIn this competition, I think domain shift exist. But this is not a big problem.\nIn training sound, it contains noisy sound (white noise and pink noise).\nBut test sound is clean(or a little noisy sound). \n\nBut I think this domain shift has **no negative impact.**\nBecause it's a domain shift from noisy to clean.\n\n# 2. Missing labels\nI [reported](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/209040) that missing labels exist in tp sound. \nIf validation data contains missing labels, your CV is different from LB.\n\nFor example, y_label is not correct(contains missing label).\n+ y_label = [1,0,0]\n+ y_true = [1,1,0]\n\n2nd annotation is missed in y_label. Then y_predict = [1,0,0] is best CV in your local. But  y_predict=[1,1,0] is better in the test. In other words, no matter how much you improve your CV, it will not always work in a testing environment. **The validation data is not reliable.**\n\n# 3. A gap training and prediction\nMany participants seem to use [SED](https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection?scriptVersionId=40731755). And I also use SED.\nIf you use SED(PANNs architecture), you should be careful.\n\nIf you use SED like below, there is a gap training and prediction.\n+ training with weak label\n+ prediction with ```framewise_output```\n(What is framewise_output?(Now I'm writing \"How to use SED\". Coming soon.))\n\nWeak label training is fitted ```clipwise_output``` **not ```framewise_output```.** ```clipwise_output``` is a time-compressed version of ```framewise_output```. ```clipwise_output``` is correlated with ```framewise_output```, but not equal. \n\nComparatively, ```framewise_output``` prediction is good at short sound event. In this competition, there is many short sound event. Therefore I use ```framewise_output``` for prediction. But there is a gap training and prediction. Then CV(```clipwise_output``` training) is different from LB(```framewise_output``` prediction).\n\nIf you change validation strategy, you may avoid this problem. For example, in validation phase you use  np.max(```framewise_output```, axis=time) instead of  ```clipwise_output```. This strategy may improve a gap CV and LB.\n\n# 4. Conclusion\nFinally, I don't have any above countermeasures. Because I cannot fix **2. Missing labels**.\n\nMy strategy is \"trust LB\" not \"trust CV\". Fortunately, [public LB is almost the same as private LB](https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/207901#1134198). I believe shaking almost never happens.",
    "1144216": "shinmurashinmura Thanks.\n\nhow do you predict at inference time",
    "1144242": "Thank you. Are you train on clipwise_output?",
    "1144429": "Yes, I trained with clipwise_output and prediction is \"framewise_outuput.",
    "1144435": "https://www.kaggle.com/c/rfcx-species-audio-detection/discussion/208830#1139552",
    "1144445": "shinmura0 how is your individual fold CV?",
    "1149876": "I used StratifiedKFold(sklearn).\nAnd n_splits=5.",
    "1155222": "Makes sense @shinmurashinmura"
  },
  "source": "meta"
}