{
  "id": 170104,
  "title": "Low F1-Score with \"example_test_audio\" using popular notebook",
  "url": "/competitions/birdsong-recognition/discussion/170104",
  "author_name": "",
  "post_date": "2020-07-26T12:50:39.017049200Z",
  "votes": 8,
  "comment_count": 22,
  "views": 0,
  "content": "<p>In this <a href=\"https://www.kaggle.com/jpison/inference-resnest50-fast-with-example-test\">notebook</a>, I tried to validate with the official 'example_test_audio' data, the most popular <a href=\"https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast\">inference notebook</a>, shared by <a href=\"/ttahara\">@ttahara</a>.</p>\n\n<p>However, the obtained micro F1-Score is too low:  F1-Score=0.0180</p>\n\n<p>What is wrong?</p>\n\n<p>Thanks in advance</p>",
  "messages": [
    {
      "id": "946221",
      "postDate": "07/26/2020 12:50:39",
      "content": "<p>In this <a href=\"https://www.kaggle.com/jpison/inference-resnest50-fast-with-example-test\">notebook</a>, I tried to validate with the official 'example_test_audio' data, the most popular <a href=\"https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast\">inference notebook</a>, shared by <a href=\"/ttahara\">@ttahara</a>.</p>\n\n<p>However, the obtained micro F1-Score is too low:  F1-Score=0.0180</p>\n\n<p>What is wrong?</p>\n\n<p>Thanks in advance</p>",
      "rawMarkdown": "In this [notebook](https://www.kaggle.com/jpison/inference-resnest50-fast-with-example-test), I tried to validate with the official 'example\\_test\\_audio' data, the most popular [inference notebook](https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast), shared by @ttahara.\n\nHowever, the obtained micro F1-Score is too low:  F1-Score=0.0180\n\nWhat is wrong?\n\nThanks in advance",
      "votes": null
    },
    {
      "id": "946346",
      "postDate": "07/26/2020 14:18:48",
      "content": "<p>Hi,</p>\n\n<p>You can click on the right \"Version 2 of 2\" then press \"Compare Diff\" button. You can view the difference between your code and Tawara's code. For example, you just have <code>if site==\"site_3\"</code> but Tawara has <code>if site == \"site_3\" or site == \"example\":</code></p>\n\n<p>There are a few other differences. Maybe you can find it better than me, since it is your code.</p>",
      "rawMarkdown": "Hi,\n\nYou can click on the right \"Version 2 of 2\" then press \"Compare Diff\" button. You can view the difference between your code and Tawara's code. For example, you just have `if site==\"site_3\"` but Tawara has `if site == \"site_3\" or site == \"example\":`\n\nThere are a few other differences. Maybe you can find it better than me, since it is your code.",
      "votes": null
    },
    {
      "id": "946563",
      "postDate": "07/26/2020 16:42:27",
      "content": "<p>Hi,</p>\n\n<p>It is the opposite. </p>\n\n<p>I included <code>if site == \"site_3\" or site == \"example\":</code>  to preprocess the two .mp3 files in the 'example_test_audio' folder.</p>",
      "rawMarkdown": "Hi,\n\nIt is the opposite. \n\nI included `if site == \"site_3\" or site == \"example\":`  to preprocess the two .mp3 files in the 'example\\_test\\_audio' folder.",
      "votes": null
    },
    {
      "id": "947455",
      "postDate": "07/27/2020 09:27:34",
      "content": "<p>Thanks for sharing. I compared it with <a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> 's baseline resnet50 model and got a micro f1-score of 0.0166. Am really curious on why that's the case. From inspection, it seems that the model is overpredicting a lot of nocalls. This can really cause a massive shakeup, if the rest of the test data set has a lot of labelled birdcalls</p>",
      "rawMarkdown": "Thanks for sharing. I compared it with @hidehisaarai1213 's baseline resnet50 model and got a micro f1-score of 0.0166. Am really curious on why that's the case. From inspection, it seems that the model is overpredicting a lot of nocalls. This can really cause a massive shakeup, if the rest of the test data set has a lot of labelled birdcalls",
      "votes": null
    },
    {
      "id": "947471",
      "postDate": "07/27/2020 09:43:03",
      "content": "<p>Here is the notebook I forked from yours, then applied to <a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> 's baseline model if anyone is interested:\n<a href=\"https://www.kaggle.com/alanchn31/forked-resnet50-with-example-test\">https://www.kaggle.com/alanchn31/forked-resnet50-with-example-test</a></p>",
      "rawMarkdown": "Here is the notebook I forked from yours, then applied to @hidehisaarai1213 's baseline model if anyone is interested:\nhttps://www.kaggle.com/alanchn31/forked-resnet50-with-example-test",
      "votes": null
    },
    {
      "id": "947573",
      "postDate": "07/27/2020 10:58:33",
      "content": "<p>Yes, but the model obtain a 0.568 in LB.\nReducing the threshold I got a local micro f1-score of 0.058, but it is too far from LB! I have been checking the code and I cannot find errors in it. If the code is fine, then I think the files in the directory 'example_test_audio' may not be representative of the testing database.</p>",
      "rawMarkdown": "Yes, but the model obtain a 0.568 in LB.\nReducing the threshold I got a local micro f1-score of 0.058, but it is too far from LB! I have been checking the code and I cannot find errors in it. If the code is fine, then I think the files in the directory 'example_test_audio' may not be representative of the testing database.",
      "votes": null
    },
    {
      "id": "947574",
      "postDate": "07/27/2020 10:58:47",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "947588",
      "postDate": "07/27/2020 11:11:02",
      "content": "<p>Yes, agree. We need to see if exampletestaudio is representative. Another point to consider is that, the current LB is only calculated on 27% of the test set. Furthermore, in the test set used to evaluate public LB now, a lot of recordings are nocall. I wonder how different the 73% will be from this 27% used to evaluate our standings</p>",
      "rawMarkdown": "Yes, agree. We need to see if exampletestaudio is representative. Another point to consider is that, the current LB is only calculated on 27% of the test set. Furthermore, in the test set used to evaluate public LB now, a lot of recordings are nocall. I wonder how different the 73% will be from this 27% used to evaluate our standings",
      "votes": null
    },
    {
      "id": "947664",
      "postDate": "07/27/2020 12:16:03",
      "content": "<p>1) the records(BLKFR-10 and ORANGE-7) contain labels not from the training sample, such as mouqua and squirrel \n2) the Public test contains a lot of nocall(54%), so the current best solutions are not very good\n3) nocall is also calculated in f1</p>",
      "rawMarkdown": "1) the records(BLKFR-10 and ORANGE-7) contain labels not from the training sample, such as mouqua and squirrel \n2) the Public test contains a lot of nocall(54%), so the current best solutions are not very good\n3) nocall is also calculated in f1",
      "votes": null
    },
    {
      "id": "947709",
      "postDate": "07/27/2020 12:44:06",
      "content": "<p>Can you explain how do you know test set contains 54% nocall?</p>",
      "rawMarkdown": "Can you explain how do you know test set contains 54% nocall?",
      "votes": null
    },
    {
      "id": "947719",
      "postDate": "07/27/2020 12:52:32",
      "content": "<p>If you send only nocall everywhere, the rating is 0.544</p>",
      "rawMarkdown": "If you send only nocall everywhere, the rating is 0.544",
      "votes": null
    },
    {
      "id": "947725",
      "postDate": "07/27/2020 12:56:43",
      "content": "<p>ruh roh</p>",
      "rawMarkdown": "ruh roh",
      "votes": null
    },
    {
      "id": "948131",
      "postDate": "07/27/2020 17:25:04",
      "content": "<p>Thank you <a href=\"/vlomme\">@vlomme</a> for your response! I didn't realise that.\nNew <a href=\"https://www.kaggle.com/jpison/inference-resnest50-fast-with-example-test-audio\">version 4</a>, includes 'nocall' label into the f1-score calculation and considers only the labels of the competition.\nNow, the search of the best local micro f1-score=0.455090 and LB=0.566 is obtained with a threshold=0.70. Close to the <a href=\"/ttahara\">@ttahara</a>'s solution (threshold=0.60, LB=0.568).</p>",
      "rawMarkdown": "Thank you @vlomme for your response! I didn't realise that.\nNew [version 4](https://www.kaggle.com/jpison/inference-resnest50-fast-with-example-test-audio), includes 'nocall' label into the f1-score calculation and considers only the labels of the competition.\nNow, the search of the best local micro f1-score=0.455090 and LB=0.566 is obtained with a threshold=0.70. Close to the @ttahara's solution (threshold=0.60, LB=0.568).",
      "votes": null
    },
    {
      "id": "948151",
      "postDate": "07/27/2020 17:35:35",
      "content": "<p>I think they should have similar distribution.</p>",
      "rawMarkdown": "I think they should have similar distribution.",
      "votes": null
    },
    {
      "id": "957581",
      "postDate": "08/04/2020 12:23:54",
      "content": "<p>If you run the same with BirdCLEF 2020 SSW*.wav audio provided at <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/158877#911336\">https://www.kaggle.com/c/birdsong-recognition/discussion/158877#911336</a> (17 species) then you will notice very low F1 score due to non-nocall in test set. It's the opposite for best threshold (more random for low values).</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F698363%2F0607d9fe7840d738bf53c2b2c070419f%2FSW.png?generation=1596543515489705&amp;alt=media\" alt=\"\"></p>\n\n<p>I'm wondering if nocall distribution from public set will be similar to private set ...</p>",
      "rawMarkdown": "If you run the same with BirdCLEF 2020 SSW*.wav audio provided at https://www.kaggle.com/c/birdsong-recognition/discussion/158877#911336 (17 species) then you will notice very low F1 score due to non-nocall in test set. It's the opposite for best threshold (more random for low values).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F698363%2F0607d9fe7840d738bf53c2b2c070419f%2FSW.png?generation=1596543515489705&amp;alt=media)\n\nI'm wondering if nocall distribution from public set will be similar to private set ...",
      "votes": null
    },
    {
      "id": "959300",
      "postDate": "08/05/2020 13:33:59",
      "content": "<p>But, to calculate f1-score, did you remove the samples with birds out of the competition?</p>",
      "rawMarkdown": "But, to calculate f1-score, did you remove the samples with birds out of the competition?",
      "votes": null
    },
    {
      "id": "959540",
      "postDate": "08/05/2020 17:16:33",
      "content": "<p><a href=\"/mpware\">@mpware</a> Hi, I have a little question:\nWhy did you say there are only 17 species match this competition?\nI have checked the <code>gt_birdclef2020_validation_data.csv</code> file and found that all 76 species belong to the SSW audios are included in the training classes in this comp.\nLike the first row in the form in your picture, why <code>cangoo</code> is <code>nocall</code>? It is a species in training data.</p>",
      "rawMarkdown": "mpware Hi, I have a little question:\nWhy did you say there are only 17 species match this competition?\nI have checked the `gt_birdclef2020_validation_data.csv` file and found that all 76 species belong to the SSW audios are included in the training classes in this comp.\nLike the first row in the form in your picture, why `cangoo` is `nocall`? It is a species in training data.",
      "votes": null
    },
    {
      "id": "960191",
      "postDate": "08/06/2020 07:54:42",
      "content": "<p>Yes removed. \nBTW: For BLKFR-10 + ORANGE-7, predicting nocall for all gives a F1 score around 0.7.</p>",
      "rawMarkdown": "Yes removed. \nBTW: For BLKFR-10 + ORANGE-7, predicting nocall for all gives a F1 score around 0.7.",
      "votes": null
    },
    {
      "id": "960194",
      "postDate": "08/06/2020 07:57:43",
      "content": "<p>The ground truth matching to SSW* files contains only 17 species in this competition. The screenshot I provided is a ground truth + inference (cangoo not detected so it becomes nocall)</p>",
      "rawMarkdown": "The ground truth matching to SSW* files contains only 17 species in this competition. The screenshot I provided is a ground truth + inference (cangoo not detected so it becomes nocall)",
      "votes": null
    },
    {
      "id": "960310",
      "postDate": "08/06/2020 09:54:39",
      "content": "<p>I will check again. Is there any other file need to be used except <code>gt_birdclef2020_validation_data.csv</code>? I may miss some information.</p>",
      "rawMarkdown": "I will check again. Is there any other file need to be used except `gt_birdclef2020_validation_data.csv`? I may miss some information.",
      "votes": null
    },
    {
      "id": "960346",
      "postDate": "08/06/2020 10:19:08",
      "content": "<p>The files you need are under <code>BirdCLEF2020__Validation_Audio_and_Ground_Truth/gt</code> folder.\nHere is an example:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F698363%2Fc35c92f736d0c7ec7975794e756870fd%2FSSWexample.png?generation=1596709110319567&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "The files you need are under `BirdCLEF2020__Validation_Audio_and_Ground_Truth/gt` folder.\nHere is an example:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F698363%2Fc35c92f736d0c7ec7975794e756870fd%2FSSWexample.png?generation=1596709110319567&amp;alt=media)",
      "votes": null
    },
    {
      "id": "960360",
      "postDate": "08/06/2020 10:37:57",
      "content": "<p>Oh! I totally use the wrong file. Thanks for helping.\nBut what is <code>gt_birdclef2020_validation_data.csv</code> under <code>BirdCLEF2020__Validation_Audio_and_Ground_Truth/scripts</code>?</p>",
      "rawMarkdown": "Oh! I totally use the wrong file. Thanks for helping.\nBut what is `gt_birdclef2020_validation_data.csv` under `BirdCLEF2020__Validation_Audio_and_Ground_Truth/scripts`?",
      "votes": null
    },
    {
      "id": "1003851",
      "postDate": "09/09/2020 09:54:43",
      "content": "<p>I think the file <em>gt_birdclef2020_validation_data.csv</em> and the files under  <em>BirdCLEF2020__Validation_Audio_and_Ground_Truth/gt</em> folder contain the same information.</p>",
      "rawMarkdown": "I think the file *gt_birdclef2020_validation_data.csv* and the files under  *BirdCLEF2020__Validation_Audio_and_Ground_Truth/gt* folder contain the same information.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 946346,
      "author_name": "returnofsputnik",
      "author_url": "",
      "post_date": "07/26/2020 14:18:48",
      "content": "<p>Hi,</p>\n\n<p>You can click on the right \"Version 2 of 2\" then press \"Compare Diff\" button. You can view the difference between your code and Tawara's code. For example, you just have <code>if site==\"site_3\"</code> but Tawara has <code>if site == \"site_3\" or site == \"example\":</code></p>\n\n<p>There are a few other differences. Maybe you can find it better than me, since it is your code.</p>",
      "votes": null,
      "replies": [
        {
          "id": 946563,
          "author_name": "jpison",
          "author_url": "",
          "post_date": "07/26/2020 16:42:27",
          "content": "<p>Hi,</p>\n\n<p>It is the opposite. </p>\n\n<p>I included <code>if site == \"site_3\" or site == \"example\":</code>  to preprocess the two .mp3 files in the 'example_test_audio' folder.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 947455,
      "author_name": "alanchn31",
      "author_url": "",
      "post_date": "07/27/2020 09:27:34",
      "content": "<p>Thanks for sharing. I compared it with <a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> 's baseline resnet50 model and got a micro f1-score of 0.0166. Am really curious on why that's the case. From inspection, it seems that the model is overpredicting a lot of nocalls. This can really cause a massive shakeup, if the rest of the test data set has a lot of labelled birdcalls</p>",
      "votes": null,
      "replies": [
        {
          "id": 947573,
          "author_name": "jpison",
          "author_url": "",
          "post_date": "07/27/2020 10:58:33",
          "content": "<p>Yes, but the model obtain a 0.568 in LB.\nReducing the threshold I got a local micro f1-score of 0.058, but it is too far from LB! I have been checking the code and I cannot find errors in it. If the code is fine, then I think the files in the directory 'example_test_audio' may not be representative of the testing database.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 947588,
          "author_name": "alanchn31",
          "author_url": "",
          "post_date": "07/27/2020 11:11:02",
          "content": "<p>Yes, agree. We need to see if exampletestaudio is representative. Another point to consider is that, the current LB is only calculated on 27% of the test set. Furthermore, in the test set used to evaluate public LB now, a lot of recordings are nocall. I wonder how different the 73% will be from this 27% used to evaluate our standings</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 948151,
          "author_name": "jpison",
          "author_url": "",
          "post_date": "07/27/2020 17:35:35",
          "content": "<p>I think they should have similar distribution.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 947471,
      "author_name": "alanchn31",
      "author_url": "",
      "post_date": "07/27/2020 09:43:03",
      "content": "<p>Here is the notebook I forked from yours, then applied to <a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> 's baseline model if anyone is interested:\n<a href=\"https://www.kaggle.com/alanchn31/forked-resnet50-with-example-test\">https://www.kaggle.com/alanchn31/forked-resnet50-with-example-test</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 947574,
          "author_name": "jpison",
          "author_url": "",
          "post_date": "07/27/2020 10:58:47",
          "content": "<p>Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 947664,
      "author_name": "vlomme",
      "author_url": "",
      "post_date": "07/27/2020 12:16:03",
      "content": "<p>1) the records(BLKFR-10 and ORANGE-7) contain labels not from the training sample, such as mouqua and squirrel \n2) the Public test contains a lot of nocall(54%), so the current best solutions are not very good\n3) nocall is also calculated in f1</p>",
      "votes": null,
      "replies": [
        {
          "id": 947709,
          "author_name": "returnofsputnik",
          "author_url": "",
          "post_date": "07/27/2020 12:44:06",
          "content": "<p>Can you explain how do you know test set contains 54% nocall?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 947719,
          "author_name": "vlomme",
          "author_url": "",
          "post_date": "07/27/2020 12:52:32",
          "content": "<p>If you send only nocall everywhere, the rating is 0.544</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 947725,
          "author_name": "returnofsputnik",
          "author_url": "",
          "post_date": "07/27/2020 12:56:43",
          "content": "<p>ruh roh</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 948131,
          "author_name": "jpison",
          "author_url": "",
          "post_date": "07/27/2020 17:25:04",
          "content": "<p>Thank you <a href=\"/vlomme\">@vlomme</a> for your response! I didn't realise that.\nNew <a href=\"https://www.kaggle.com/jpison/inference-resnest50-fast-with-example-test-audio\">version 4</a>, includes 'nocall' label into the f1-score calculation and considers only the labels of the competition.\nNow, the search of the best local micro f1-score=0.455090 and LB=0.566 is obtained with a threshold=0.70. Close to the <a href=\"/ttahara\">@ttahara</a>'s solution (threshold=0.60, LB=0.568).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 957581,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "08/04/2020 12:23:54",
          "content": "<p>If you run the same with BirdCLEF 2020 SSW*.wav audio provided at <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/158877#911336\">https://www.kaggle.com/c/birdsong-recognition/discussion/158877#911336</a> (17 species) then you will notice very low F1 score due to non-nocall in test set. It's the opposite for best threshold (more random for low values).</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F698363%2F0607d9fe7840d738bf53c2b2c070419f%2FSW.png?generation=1596543515489705&amp;alt=media\" alt=\"\"></p>\n\n<p>I'm wondering if nocall distribution from public set will be similar to private set ...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 959300,
          "author_name": "jpison",
          "author_url": "",
          "post_date": "08/05/2020 13:33:59",
          "content": "<p>But, to calculate f1-score, did you remove the samples with birds out of the competition?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 959540,
          "author_name": "karlyukang",
          "author_url": "",
          "post_date": "08/05/2020 17:16:33",
          "content": "<p><a href=\"/mpware\">@mpware</a> Hi, I have a little question:\nWhy did you say there are only 17 species match this competition?\nI have checked the <code>gt_birdclef2020_validation_data.csv</code> file and found that all 76 species belong to the SSW audios are included in the training classes in this comp.\nLike the first row in the form in your picture, why <code>cangoo</code> is <code>nocall</code>? It is a species in training data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 960191,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "08/06/2020 07:54:42",
          "content": "<p>Yes removed. \nBTW: For BLKFR-10 + ORANGE-7, predicting nocall for all gives a F1 score around 0.7.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 960194,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "08/06/2020 07:57:43",
          "content": "<p>The ground truth matching to SSW* files contains only 17 species in this competition. The screenshot I provided is a ground truth + inference (cangoo not detected so it becomes nocall)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 960310,
          "author_name": "karlyukang",
          "author_url": "",
          "post_date": "08/06/2020 09:54:39",
          "content": "<p>I will check again. Is there any other file need to be used except <code>gt_birdclef2020_validation_data.csv</code>? I may miss some information.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 960346,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "08/06/2020 10:19:08",
          "content": "<p>The files you need are under <code>BirdCLEF2020__Validation_Audio_and_Ground_Truth/gt</code> folder.\nHere is an example:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F698363%2Fc35c92f736d0c7ec7975794e756870fd%2FSSWexample.png?generation=1596709110319567&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 960360,
          "author_name": "karlyukang",
          "author_url": "",
          "post_date": "08/06/2020 10:37:57",
          "content": "<p>Oh! I totally use the wrong file. Thanks for helping.\nBut what is <code>gt_birdclef2020_validation_data.csv</code> under <code>BirdCLEF2020__Validation_Audio_and_Ground_Truth/scripts</code>?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1003851,
          "author_name": "vzaguskin",
          "author_url": "",
          "post_date": "09/09/2020 09:54:43",
          "content": "<p>I think the file <em>gt_birdclef2020_validation_data.csv</em> and the files under  <em>BirdCLEF2020__Validation_Audio_and_Ground_Truth/gt</em> folder contain the same information.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "946221": "In this [notebook](https://www.kaggle.com/jpison/inference-resnest50-fast-with-example-test), I tried to validate with the official 'example\\_test\\_audio' data, the most popular [inference notebook](https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast), shared by @ttahara.\n\nHowever, the obtained micro F1-Score is too low:  F1-Score=0.0180\n\nWhat is wrong?\n\nThanks in advance",
    "946346": "Hi,\n\nYou can click on the right \"Version 2 of 2\" then press \"Compare Diff\" button. You can view the difference between your code and Tawara's code. For example, you just have `if site==\"site_3\"` but Tawara has `if site == \"site_3\" or site == \"example\":`\n\nThere are a few other differences. Maybe you can find it better than me, since it is your code.",
    "946563": "Hi,\n\nIt is the opposite. \n\nI included `if site == \"site_3\" or site == \"example\":`  to preprocess the two .mp3 files in the 'example\\_test\\_audio' folder.",
    "947455": "Thanks for sharing. I compared it with @hidehisaarai1213 's baseline resnet50 model and got a micro f1-score of 0.0166. Am really curious on why that's the case. From inspection, it seems that the model is overpredicting a lot of nocalls. This can really cause a massive shakeup, if the rest of the test data set has a lot of labelled birdcalls",
    "947471": "Here is the notebook I forked from yours, then applied to @hidehisaarai1213 's baseline model if anyone is interested:\nhttps://www.kaggle.com/alanchn31/forked-resnet50-with-example-test",
    "947573": "Yes, but the model obtain a 0.568 in LB.\nReducing the threshold I got a local micro f1-score of 0.058, but it is too far from LB! I have been checking the code and I cannot find errors in it. If the code is fine, then I think the files in the directory 'example_test_audio' may not be representative of the testing database.",
    "947574": "Thanks!",
    "947588": "Yes, agree. We need to see if exampletestaudio is representative. Another point to consider is that, the current LB is only calculated on 27% of the test set. Furthermore, in the test set used to evaluate public LB now, a lot of recordings are nocall. I wonder how different the 73% will be from this 27% used to evaluate our standings",
    "947664": "1) the records(BLKFR-10 and ORANGE-7) contain labels not from the training sample, such as mouqua and squirrel \n2) the Public test contains a lot of nocall(54%), so the current best solutions are not very good\n3) nocall is also calculated in f1",
    "947709": "Can you explain how do you know test set contains 54% nocall?",
    "947719": "If you send only nocall everywhere, the rating is 0.544",
    "947725": "ruh roh",
    "948131": "Thank you @vlomme for your response! I didn't realise that.\nNew [version 4](https://www.kaggle.com/jpison/inference-resnest50-fast-with-example-test-audio), includes 'nocall' label into the f1-score calculation and considers only the labels of the competition.\nNow, the search of the best local micro f1-score=0.455090 and LB=0.566 is obtained with a threshold=0.70. Close to the @ttahara's solution (threshold=0.60, LB=0.568).",
    "948151": "I think they should have similar distribution.",
    "957581": "If you run the same with BirdCLEF 2020 SSW*.wav audio provided at https://www.kaggle.com/c/birdsong-recognition/discussion/158877#911336 (17 species) then you will notice very low F1 score due to non-nocall in test set. It's the opposite for best threshold (more random for low values).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F698363%2F0607d9fe7840d738bf53c2b2c070419f%2FSW.png?generation=1596543515489705&amp;alt=media)\n\nI'm wondering if nocall distribution from public set will be similar to private set ...",
    "959300": "But, to calculate f1-score, did you remove the samples with birds out of the competition?",
    "959540": "mpware Hi, I have a little question:\nWhy did you say there are only 17 species match this competition?\nI have checked the `gt_birdclef2020_validation_data.csv` file and found that all 76 species belong to the SSW audios are included in the training classes in this comp.\nLike the first row in the form in your picture, why `cangoo` is `nocall`? It is a species in training data.",
    "960191": "Yes removed. \nBTW: For BLKFR-10 + ORANGE-7, predicting nocall for all gives a F1 score around 0.7.",
    "960194": "The ground truth matching to SSW* files contains only 17 species in this competition. The screenshot I provided is a ground truth + inference (cangoo not detected so it becomes nocall)",
    "960310": "I will check again. Is there any other file need to be used except `gt_birdclef2020_validation_data.csv`? I may miss some information.",
    "960346": "The files you need are under `BirdCLEF2020__Validation_Audio_and_Ground_Truth/gt` folder.\nHere is an example:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F698363%2Fc35c92f736d0c7ec7975794e756870fd%2FSSWexample.png?generation=1596709110319567&amp;alt=media)",
    "960360": "Oh! I totally use the wrong file. Thanks for helping.\nBut what is `gt_birdclef2020_validation_data.csv` under `BirdCLEF2020__Validation_Audio_and_Ground_Truth/scripts`?",
    "1003851": "I think the file *gt_birdclef2020_validation_data.csv* and the files under  *BirdCLEF2020__Validation_Audio_and_Ground_Truth/gt* folder contain the same information."
  },
  "source": "meta"
}