{
  "id": 171421,
  "title": "Is Resnet alone giving better score",
  "url": "/competitions/birdsong-recognition/discussion/171421",
  "author_name": "",
  "post_date": "2020-07-31T17:50:15.019649700Z",
  "votes": 4,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Did any models giving better score other than Resnet50, i can able to see many notebooks has been published based on resnet Architecture.\nDoes any other Model giving more accuracy or better results?</p>",
  "messages": [
    {
      "id": "953344",
      "postDate": "07/31/2020 17:50:15",
      "content": "<p>Did any models giving better score other than Resnet50, i can able to see many notebooks has been published based on resnet Architecture.\nDoes any other Model giving more accuracy or better results?</p>",
      "rawMarkdown": "Did any models giving better score other than Resnet50, i can able to see many notebooks has been published based on resnet Architecture.\nDoes any other Model giving more accuracy or better results?",
      "votes": null
    },
    {
      "id": "955747",
      "postDate": "08/02/2020 22:15:31",
      "content": "<p>I have used DenseNet.\nIt got better local fold-0 score, but LB score was low than ResNet.</p>\n\n<p>| |fold-0|LB|\n| -- | -- | -- |\n|ResNet50|0.6489874215|0.555|\n|DenseNet121|0.6738970588|0.506|</p>",
      "rawMarkdown": "I have used DenseNet.\nIt got better local fold-0 score, but LB score was low than ResNet.\n\n| |fold-0|LB|\n| -- | -- | -- |\n|ResNet50|0.6489874215|0.555|\n|DenseNet121|0.6738970588|0.506|",
      "votes": null
    },
    {
      "id": "957370",
      "postDate": "08/04/2020 08:49:21",
      "content": "<p>Yup... I am also getting similar trends here. The deeper the network, poorer the LB score.</p>",
      "rawMarkdown": "Yup... I am also getting similar trends here. The deeper the network, poorer the LB score.",
      "votes": null
    },
    {
      "id": "958421",
      "postDate": "08/05/2020 01:22:46",
      "content": "<p>One of the difference is test data has \"no call\".\nIf I set high threshold after sigmoid, I may get better score.\n(I ref <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/167263\">this discussion</a>.)</p>\n\n<p>Other way, to add \"no call\" to training data.\nI'm trying now.</p>",
      "rawMarkdown": "One of the difference is test data has \"no call\".\nIf I set high threshold after sigmoid, I may get better score.\n(I ref [this discussion](https://www.kaggle.com/c/birdsong-recognition/discussion/167263).)\n\nOther way, to add \"no call\" to training data.\nI'm trying now.",
      "votes": null
    },
    {
      "id": "958499",
      "postDate": "08/05/2020 02:56:07",
      "content": "<p>Did you use the same threshold?\nI found that optimal threshold (for pubLB) is different between different models.</p>",
      "rawMarkdown": "Did you use the same threshold?\nI found that optimal threshold (for pubLB) is different between different models.",
      "votes": null
    },
    {
      "id": "958579",
      "postDate": "08/05/2020 04:09:47",
      "content": "<p>Oh, didn't consider that. I wish we could have a threshold independent metric to compare models. The 2 submissions per day rule is making it really hard to know what works and what doesn't. </p>\n\n<p>Seems like everyone is using resnet following ur public kernel and have got results ~0.56. I tried efficient net b5, which on train/val set was performing way better than resnet but when I submitted it, it ended up in ~0.48 (way worse than the baseline)</p>\n\n<p>My intuition was that it is way too deep and overfitting on the training data, which led to poor generalization on the test set. As the distribution of train and test set are quite different, I strongly believe that to be true. </p>\n\n<p>Your opinion on this?\n<a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> \n<a href=\"/takamichitoda\">@takamichitoda</a> </p>",
      "rawMarkdown": "Oh, didn't consider that. I wish we could have a threshold independent metric to compare models. The 2 submissions per day rule is making it really hard to know what works and what doesn't. \n\nSeems like everyone is using resnet following ur public kernel and have got results ~0.56. I tried efficient net b5, which on train/val set was performing way better than resnet but when I submitted it, it ended up in ~0.48 (way worse than the baseline)\n\nMy intuition was that it is way too deep and overfitting on the training data, which led to poor generalization on the test set. As the distribution of train and test set are quite different, I strongly believe that to be true. \n\nYour opinion on this?\n@hidehisaarai1213 \n@takamichitoda",
      "votes": null
    },
    {
      "id": "958629",
      "postDate": "08/05/2020 04:59:23",
      "content": "<p>I agree with what Hidehisa said. In addition to different thresholds based on model architecture, I find different thresholds are needed, even for the same model architecture, if you modify it somewhat (augmentation, external data, etc.)</p>\n\n<p>We are somewhat flying blind because we do not reliably know the distribution of classes in test set, since test set does not come from the same distribution as the train set. That is one of the biggest challenges in this competition.</p>\n\n<p>I hope to overcome that by having a really robust solution (I haven't beaten the top public kernel, but I am improving every day)</p>",
      "rawMarkdown": "I agree with what Hidehisa said. In addition to different thresholds based on model architecture, I find different thresholds are needed, even for the same model architecture, if you modify it somewhat (augmentation, external data, etc.)\n\nWe are somewhat flying blind because we do not reliably know the distribution of classes in test set, since test set does not come from the same distribution as the train set. That is one of the biggest challenges in this competition.\n\nI hope to overcome that by having a really robust solution (I haven't beaten the top public kernel, but I am improving every day)",
      "votes": null
    },
    {
      "id": "959132",
      "postDate": "08/05/2020 11:23:35",
      "content": "<p>i suggest that if we want to compare different kaggler results, someone should upload some \"standard validation set\" (and used in additional to LB public score)</p>\n\n<p>Current public LB is difficult to use because there is also domain shift.</p>\n\n<p>there is external \"Xeno-Canto Bird Recordings\" in another post, maybe we can use that for kagglers comparison of different models. someone needs to make some evaluation script (e.g. chunck lever F1 score and file F1 level score, etc)</p>",
      "rawMarkdown": "i suggest that if we want to compare different kaggler results, someone should upload some \"standard validation set\" (and used in additional to LB public score)\n\nCurrent public LB is difficult to use because there is also domain shift.\n\nthere is external \"Xeno-Canto Bird Recordings\" in another post, maybe we can use that for kagglers comparison of different models. someone needs to make some evaluation script (e.g. chunck lever F1 score and file F1 level score, etc)",
      "votes": null
    },
    {
      "id": "959172",
      "postDate": "08/05/2020 11:56:36",
      "content": "<p>Sounds interesting, but is that trust worthy? What we are going to predict is rather similar to public LB if public LB and private LB is similar.</p>",
      "rawMarkdown": "Sounds interesting, but is that trust worthy? What we are going to predict is rather similar to public LB if public LB and private LB is similar.",
      "votes": null
    },
    {
      "id": "959310",
      "postDate": "08/05/2020 13:42:30",
      "content": "<p>there are two parts to it:\n1. good design of model\n2. good domain adaption</p>\n\n<p>i would prefer to decouple the two parts. For domain adaption (which is this kaggle challenge is about), the correct way to test could be:</p>\n\n<ol>\n<li><p>assume you have a good model on clean data. say it achieve F1 0.90 on your local cross-validation soundscape file.</p></li>\n<li><p>you submit and has public LB = 0.60</p></li>\n<li><p>then you should add more noise (or add other disturbance) to your local cross-validation soundscape file until your local CV is about 0.60.</p></li>\n<li><p>improve your model and inference pipeline, say your local CV improves to 0.65.</p></li>\n<li><p>test your improved model and hope public LB also shows same improvement.</p></li>\n<li><p>if local CV and public LB correlates, it means your validation data should have capture some data characteristics of the public data.</p></li>\n<li><p>for private LB data, there is no correlation of results. we can only assume that public/private has some common distribution and make our model as robust as possible.</p></li>\n</ol>\n\n<hr>\n\n<p>in previous kaggle, we adjust our model to improve public LB.</p>\n\n<p>in this kaggle, due to \"complete black-box' (different test data), we need to adjust both our model and local validation data. we need to predict the public (and private) test data as well.</p>",
      "rawMarkdown": "there are two parts to it:\n1. good design of model\n2. good domain adaption\n\ni would prefer to decouple the two parts. For domain adaption (which is this kaggle challenge is about), the correct way to test could be:\n\n1. assume you have a good model on clean data. say it achieve F1 0.90 on your local cross-validation soundscape file.\n\n2. you submit and has public LB = 0.60\n\n3. then you should add more noise (or add other disturbance) to your local cross-validation soundscape file until your local CV is about 0.60.\n\n4. improve your model and inference pipeline, say your local CV improves to 0.65.\n\n5. test your improved model and hope public LB also shows same improvement.\n\n6. if local CV and public LB correlates, it means your validation data should have capture some data characteristics of the public data.\n\n7. for private LB data, there is no correlation of results. we can only assume that public/private has some common distribution and make our model as robust as possible.\n\n---\n\nin previous kaggle, we adjust our model to improve public LB.\n\nin this kaggle, due to \"complete black-box' (different test data), we need to adjust both our model and local validation data. we need to predict the public (and private) test data as well.",
      "votes": null
    },
    {
      "id": "959845",
      "postDate": "08/06/2020 00:51:44",
      "content": "<p>I have used same threshold.\nI didn't care threshold because the expected score increase is about 0.02 from this discussion. (I thought I must got 0.56 over in threshold=0.5 at least)</p>\n\n<blockquote>\n  <p>As the distribution of train and test set are quite different,</p>\n</blockquote>\n\n<p>I think so too. And more, we must predict multi label and \"no call\" in test set.</p>\n\n<p>The important thing I think is make good validation set than modeling.</p>",
      "rawMarkdown": "I have used same threshold.\nI didn't care threshold because the expected score increase is about 0.02 from this discussion. (I thought I must got 0.56 over in threshold=0.5 at least)\n\n&gt; As the distribution of train and test set are quite different,\n\nI think so too. And more, we must predict multi label and \"no call\" in test set.\n\nThe important thing I think is make good validation set than modeling.",
      "votes": null
    },
    {
      "id": "959953",
      "postDate": "08/06/2020 03:57:28",
      "content": "<blockquote>\n  <p>My intuition was that it is way too deep and overfitting on the training data, which led to poor generalization on the test set. As the distribution of train and test set are quite different, I strongly believe that to be true.</p>\n</blockquote>\n\n<p>Agree. We must make models in good balance.</p>",
      "rawMarkdown": "&gt; My intuition was that it is way too deep and overfitting on the training data, which led to poor generalization on the test set. As the distribution of train and test set are quite different, I strongly believe that to be true.\n\nAgree. We must make models in good balance.",
      "votes": null
    },
    {
      "id": "959959",
      "postDate": "08/06/2020 04:02:29",
      "content": "<blockquote>\n  <p>i would prefer to decouple the two parts.</p>\n</blockquote>\n\n<p>Got it.</p>\n\n<blockquote>\n  <p>the correct way to test could be:</p>\n</blockquote>\n\n<p>this is quite interesting. I think that what kind of \"noise\" we add to make our local CV match public LB matters a lot. It's not just additive gaussian noise, but also label noise, unrelated class existence, call-type distribution change, etc.</p>",
      "rawMarkdown": "&gt; i would prefer to decouple the two parts.\n\nGot it.\n\n&gt; the correct way to test could be:\n\nthis is quite interesting. I think that what kind of \"noise\" we add to make our local CV match public LB matters a lot. It's not just additive gaussian noise, but also label noise, unrelated class existence, call-type distribution change, etc.",
      "votes": null
    },
    {
      "id": "960711",
      "postDate": "08/06/2020 16:02:06",
      "content": "<p><a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> \nthere is one trick. this is a \"thought experiment\", i haven't tried it yet. you can try this in local cross validation or submission LB if you have free submission slots left.</p>\n\n<p>in your submission kernel, measure snr for events in the 5 sec window. you ca do the following probing:\n1. standard submission, say you get LB score1\n2. if there is no signal with snr &gt; threshold1 in the 5 sec window, set prediction to \"nocall\". Submit and get LB score2\n3. if there is no signal with snr &gt; threshold2, set prediction to \"nocall\". Submit and get LB score3\netc ...</p>\n\n<p>Then you can roughly measure the snr of the public test set, by comparing the changes in LB scores.</p>\n\n<p>Now snr is only one of the properties of audio. you can think of other spectrum/audio properties to do \"probing\" . Once you know the properties of the test set, you can simulate same properties from your train set.</p>",
      "rawMarkdown": "hidehisaarai1213 \nthere is one trick. this is a \"thought experiment\", i haven't tried it yet. you can try this in local cross validation or submission LB if you have free submission slots left.\n\nin your submission kernel, measure snr for events in the 5 sec window. you ca do the following probing:\n1. standard submission, say you get LB score1\n2. if there is no signal with snr &gt; threshold1 in the 5 sec window, set prediction to \"nocall\". Submit and get LB score2\n3. if there is no signal with snr &gt; threshold2, set prediction to \"nocall\". Submit and get LB score3\netc ...\n\nThen you can roughly measure the snr of the public test set, by comparing the changes in LB scores.\n\nNow snr is only one of the properties of audio. you can think of other spectrum/audio properties to do \"probing\" . Once you know the properties of the test set, you can simulate same properties from your train set.",
      "votes": null
    },
    {
      "id": "961541",
      "postDate": "08/07/2020 09:18:07",
      "content": "<p>I have never tried probing in any kind of competition, but sounds nice to try when I have nothing to submit except for this one, thanks!</p>",
      "rawMarkdown": "I have never tried probing in any kind of competition, but sounds nice to try when I have nothing to submit except for this one, thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 955747,
      "author_name": "takamichitoda",
      "author_url": "",
      "post_date": "08/02/2020 22:15:31",
      "content": "<p>I have used DenseNet.\nIt got better local fold-0 score, but LB score was low than ResNet.</p>\n\n<p>| |fold-0|LB|\n| -- | -- | -- |\n|ResNet50|0.6489874215|0.555|\n|DenseNet121|0.6738970588|0.506|</p>",
      "votes": null,
      "replies": [
        {
          "id": 957370,
          "author_name": "humblediscipulus",
          "author_url": "",
          "post_date": "08/04/2020 08:49:21",
          "content": "<p>Yup... I am also getting similar trends here. The deeper the network, poorer the LB score.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 958421,
          "author_name": "takamichitoda",
          "author_url": "",
          "post_date": "08/05/2020 01:22:46",
          "content": "<p>One of the difference is test data has \"no call\".\nIf I set high threshold after sigmoid, I may get better score.\n(I ref <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/167263\">this discussion</a>.)</p>\n\n<p>Other way, to add \"no call\" to training data.\nI'm trying now.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 958499,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "08/05/2020 02:56:07",
          "content": "<p>Did you use the same threshold?\nI found that optimal threshold (for pubLB) is different between different models.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 958579,
          "author_name": "humblediscipulus",
          "author_url": "",
          "post_date": "08/05/2020 04:09:47",
          "content": "<p>Oh, didn't consider that. I wish we could have a threshold independent metric to compare models. The 2 submissions per day rule is making it really hard to know what works and what doesn't. </p>\n\n<p>Seems like everyone is using resnet following ur public kernel and have got results ~0.56. I tried efficient net b5, which on train/val set was performing way better than resnet but when I submitted it, it ended up in ~0.48 (way worse than the baseline)</p>\n\n<p>My intuition was that it is way too deep and overfitting on the training data, which led to poor generalization on the test set. As the distribution of train and test set are quite different, I strongly believe that to be true. </p>\n\n<p>Your opinion on this?\n<a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> \n<a href=\"/takamichitoda\">@takamichitoda</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 958629,
          "author_name": "returnofsputnik",
          "author_url": "",
          "post_date": "08/05/2020 04:59:23",
          "content": "<p>I agree with what Hidehisa said. In addition to different thresholds based on model architecture, I find different thresholds are needed, even for the same model architecture, if you modify it somewhat (augmentation, external data, etc.)</p>\n\n<p>We are somewhat flying blind because we do not reliably know the distribution of classes in test set, since test set does not come from the same distribution as the train set. That is one of the biggest challenges in this competition.</p>\n\n<p>I hope to overcome that by having a really robust solution (I haven't beaten the top public kernel, but I am improving every day)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 959845,
          "author_name": "takamichitoda",
          "author_url": "",
          "post_date": "08/06/2020 00:51:44",
          "content": "<p>I have used same threshold.\nI didn't care threshold because the expected score increase is about 0.02 from this discussion. (I thought I must got 0.56 over in threshold=0.5 at least)</p>\n\n<blockquote>\n  <p>As the distribution of train and test set are quite different,</p>\n</blockquote>\n\n<p>I think so too. And more, we must predict multi label and \"no call\" in test set.</p>\n\n<p>The important thing I think is make good validation set than modeling.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 959953,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "08/06/2020 03:57:28",
          "content": "<blockquote>\n  <p>My intuition was that it is way too deep and overfitting on the training data, which led to poor generalization on the test set. As the distribution of train and test set are quite different, I strongly believe that to be true.</p>\n</blockquote>\n\n<p>Agree. We must make models in good balance.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 959132,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/05/2020 11:23:35",
      "content": "<p>i suggest that if we want to compare different kaggler results, someone should upload some \"standard validation set\" (and used in additional to LB public score)</p>\n\n<p>Current public LB is difficult to use because there is also domain shift.</p>\n\n<p>there is external \"Xeno-Canto Bird Recordings\" in another post, maybe we can use that for kagglers comparison of different models. someone needs to make some evaluation script (e.g. chunck lever F1 score and file F1 level score, etc)</p>",
      "votes": null,
      "replies": [
        {
          "id": 959172,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "08/05/2020 11:56:36",
          "content": "<p>Sounds interesting, but is that trust worthy? What we are going to predict is rather similar to public LB if public LB and private LB is similar.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 959310,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/05/2020 13:42:30",
          "content": "<p>there are two parts to it:\n1. good design of model\n2. good domain adaption</p>\n\n<p>i would prefer to decouple the two parts. For domain adaption (which is this kaggle challenge is about), the correct way to test could be:</p>\n\n<ol>\n<li><p>assume you have a good model on clean data. say it achieve F1 0.90 on your local cross-validation soundscape file.</p></li>\n<li><p>you submit and has public LB = 0.60</p></li>\n<li><p>then you should add more noise (or add other disturbance) to your local cross-validation soundscape file until your local CV is about 0.60.</p></li>\n<li><p>improve your model and inference pipeline, say your local CV improves to 0.65.</p></li>\n<li><p>test your improved model and hope public LB also shows same improvement.</p></li>\n<li><p>if local CV and public LB correlates, it means your validation data should have capture some data characteristics of the public data.</p></li>\n<li><p>for private LB data, there is no correlation of results. we can only assume that public/private has some common distribution and make our model as robust as possible.</p></li>\n</ol>\n\n<hr>\n\n<p>in previous kaggle, we adjust our model to improve public LB.</p>\n\n<p>in this kaggle, due to \"complete black-box' (different test data), we need to adjust both our model and local validation data. we need to predict the public (and private) test data as well.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 959959,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "08/06/2020 04:02:29",
          "content": "<blockquote>\n  <p>i would prefer to decouple the two parts.</p>\n</blockquote>\n\n<p>Got it.</p>\n\n<blockquote>\n  <p>the correct way to test could be:</p>\n</blockquote>\n\n<p>this is quite interesting. I think that what kind of \"noise\" we add to make our local CV match public LB matters a lot. It's not just additive gaussian noise, but also label noise, unrelated class existence, call-type distribution change, etc.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 960711,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/06/2020 16:02:06",
          "content": "<p><a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> \nthere is one trick. this is a \"thought experiment\", i haven't tried it yet. you can try this in local cross validation or submission LB if you have free submission slots left.</p>\n\n<p>in your submission kernel, measure snr for events in the 5 sec window. you ca do the following probing:\n1. standard submission, say you get LB score1\n2. if there is no signal with snr &gt; threshold1 in the 5 sec window, set prediction to \"nocall\". Submit and get LB score2\n3. if there is no signal with snr &gt; threshold2, set prediction to \"nocall\". Submit and get LB score3\netc ...</p>\n\n<p>Then you can roughly measure the snr of the public test set, by comparing the changes in LB scores.</p>\n\n<p>Now snr is only one of the properties of audio. you can think of other spectrum/audio properties to do \"probing\" . Once you know the properties of the test set, you can simulate same properties from your train set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961541,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "08/07/2020 09:18:07",
          "content": "<p>I have never tried probing in any kind of competition, but sounds nice to try when I have nothing to submit except for this one, thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "953344": "Did any models giving better score other than Resnet50, i can able to see many notebooks has been published based on resnet Architecture.\nDoes any other Model giving more accuracy or better results?",
    "955747": "I have used DenseNet.\nIt got better local fold-0 score, but LB score was low than ResNet.\n\n| |fold-0|LB|\n| -- | -- | -- |\n|ResNet50|0.6489874215|0.555|\n|DenseNet121|0.6738970588|0.506|",
    "957370": "Yup... I am also getting similar trends here. The deeper the network, poorer the LB score.",
    "958421": "One of the difference is test data has \"no call\".\nIf I set high threshold after sigmoid, I may get better score.\n(I ref [this discussion](https://www.kaggle.com/c/birdsong-recognition/discussion/167263).)\n\nOther way, to add \"no call\" to training data.\nI'm trying now.",
    "958499": "Did you use the same threshold?\nI found that optimal threshold (for pubLB) is different between different models.",
    "958579": "Oh, didn't consider that. I wish we could have a threshold independent metric to compare models. The 2 submissions per day rule is making it really hard to know what works and what doesn't. \n\nSeems like everyone is using resnet following ur public kernel and have got results ~0.56. I tried efficient net b5, which on train/val set was performing way better than resnet but when I submitted it, it ended up in ~0.48 (way worse than the baseline)\n\nMy intuition was that it is way too deep and overfitting on the training data, which led to poor generalization on the test set. As the distribution of train and test set are quite different, I strongly believe that to be true. \n\nYour opinion on this?\n@hidehisaarai1213 \n@takamichitoda",
    "958629": "I agree with what Hidehisa said. In addition to different thresholds based on model architecture, I find different thresholds are needed, even for the same model architecture, if you modify it somewhat (augmentation, external data, etc.)\n\nWe are somewhat flying blind because we do not reliably know the distribution of classes in test set, since test set does not come from the same distribution as the train set. That is one of the biggest challenges in this competition.\n\nI hope to overcome that by having a really robust solution (I haven't beaten the top public kernel, but I am improving every day)",
    "959132": "i suggest that if we want to compare different kaggler results, someone should upload some \"standard validation set\" (and used in additional to LB public score)\n\nCurrent public LB is difficult to use because there is also domain shift.\n\nthere is external \"Xeno-Canto Bird Recordings\" in another post, maybe we can use that for kagglers comparison of different models. someone needs to make some evaluation script (e.g. chunck lever F1 score and file F1 level score, etc)",
    "959172": "Sounds interesting, but is that trust worthy? What we are going to predict is rather similar to public LB if public LB and private LB is similar.",
    "959310": "there are two parts to it:\n1. good design of model\n2. good domain adaption\n\ni would prefer to decouple the two parts. For domain adaption (which is this kaggle challenge is about), the correct way to test could be:\n\n1. assume you have a good model on clean data. say it achieve F1 0.90 on your local cross-validation soundscape file.\n\n2. you submit and has public LB = 0.60\n\n3. then you should add more noise (or add other disturbance) to your local cross-validation soundscape file until your local CV is about 0.60.\n\n4. improve your model and inference pipeline, say your local CV improves to 0.65.\n\n5. test your improved model and hope public LB also shows same improvement.\n\n6. if local CV and public LB correlates, it means your validation data should have capture some data characteristics of the public data.\n\n7. for private LB data, there is no correlation of results. we can only assume that public/private has some common distribution and make our model as robust as possible.\n\n---\n\nin previous kaggle, we adjust our model to improve public LB.\n\nin this kaggle, due to \"complete black-box' (different test data), we need to adjust both our model and local validation data. we need to predict the public (and private) test data as well.",
    "959845": "I have used same threshold.\nI didn't care threshold because the expected score increase is about 0.02 from this discussion. (I thought I must got 0.56 over in threshold=0.5 at least)\n\n&gt; As the distribution of train and test set are quite different,\n\nI think so too. And more, we must predict multi label and \"no call\" in test set.\n\nThe important thing I think is make good validation set than modeling.",
    "959953": "&gt; My intuition was that it is way too deep and overfitting on the training data, which led to poor generalization on the test set. As the distribution of train and test set are quite different, I strongly believe that to be true.\n\nAgree. We must make models in good balance.",
    "959959": "&gt; i would prefer to decouple the two parts.\n\nGot it.\n\n&gt; the correct way to test could be:\n\nthis is quite interesting. I think that what kind of \"noise\" we add to make our local CV match public LB matters a lot. It's not just additive gaussian noise, but also label noise, unrelated class existence, call-type distribution change, etc.",
    "960711": "hidehisaarai1213 \nthere is one trick. this is a \"thought experiment\", i haven't tried it yet. you can try this in local cross validation or submission LB if you have free submission slots left.\n\nin your submission kernel, measure snr for events in the 5 sec window. you ca do the following probing:\n1. standard submission, say you get LB score1\n2. if there is no signal with snr &gt; threshold1 in the 5 sec window, set prediction to \"nocall\". Submit and get LB score2\n3. if there is no signal with snr &gt; threshold2, set prediction to \"nocall\". Submit and get LB score3\netc ...\n\nThen you can roughly measure the snr of the public test set, by comparing the changes in LB scores.\n\nNow snr is only one of the properties of audio. you can think of other spectrum/audio properties to do \"probing\" . Once you know the properties of the test set, you can simulate same properties from your train set.",
    "961541": "I have never tried probing in any kind of competition, but sounds nice to try when I have nothing to submit except for this one, thanks!"
  },
  "source": "meta"
}