{
  "id": 318999,
  "title": "The lower threshhold has good score.",
  "url": "/competitions/birdclef-2022/discussion/318999",
  "author_name": "shinmura0",
  "post_date": "2022-04-15T03:03:46.057000",
  "votes": 42,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I am glad to participate in the bird identification competition again.</p>\n<p>I used <a href=\"https://www.kaggle.com/competitions/rfcx-species-audio-detection/discussion/211007#1151299\" target=\"_blank\">PANNs</a> for both <a href=\"https://www.kaggle.com/competitions/birdclef-2021\" target=\"_blank\">the last competition</a> and this one, and with PANNs, the threshold is important. This threshold means the threshold of the sigmoid function at the end of the model.</p>\n<p>My experiments is interesting. So I share them with you.</p>\n<table>\n<thead>\n<tr>\n<th>Threshold of the sigmoid function</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.5</td>\n<td>0.56</td>\n</tr>\n<tr>\n<td>0.3</td>\n<td>0.61</td>\n</tr>\n<tr>\n<td>0.2</td>\n<td>0.64</td>\n</tr>\n<tr>\n<td>0.1</td>\n<td>0.68</td>\n</tr>\n<tr>\n<td><strong>0.05</strong></td>\n<td><strong>0.71</strong></td>\n</tr>\n<tr>\n<td><strong>0.01</strong></td>\n<td><strong>0.71</strong></td>\n</tr>\n</tbody>\n</table>\n<p>These score is the same model result (only thresholds were changed). <a href=\"https://www.kaggle.com/competitions/birdclef-2021\" target=\"_blank\">In the previous competition</a>, the best threshold was <strong>0.6 or 0.7</strong>. However, in this competition, the threshold is much lower (0.01 or 0.05). I think the difference lies in whether the emphasis is on precision or recall.</p>\n<p>In the previous competition, about half of the cases were nocall (negative). Therefore, the emphasis was on precision, and <strong>a high threshold</strong> was appropriate. However, In this competition,  the emphasis is on positive, and recall is more important. and <strong>a low threshold</strong> was appropriate. Finally, there are many false positive in soundscape_453028782.ogg with the threshold of 0.01 like below. But it's good score (LB0.71).</p>\n<p><img src=\"https://user-images.githubusercontent.com/34497776/163506679-705a3030-5c07-4290-a658-d3daabed8f7b.png\" alt=\"image\"></p>\n<p>Do you have any idea?<br>\n(I don't completely understand <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/314999\" target=\"_blank\">the metrics for this competition</a>)</p>",
  "messages": [
    {
      "id": 1755821,
      "postDate": "2022-04-15T03:03:46.057Z",
      "content": "<p>I am glad to participate in the bird identification competition again.</p>\n<p>I used <a href=\"https://www.kaggle.com/competitions/rfcx-species-audio-detection/discussion/211007#1151299\" target=\"_blank\">PANNs</a> for both <a href=\"https://www.kaggle.com/competitions/birdclef-2021\" target=\"_blank\">the last competition</a> and this one, and with PANNs, the threshold is important. This threshold means the threshold of the sigmoid function at the end of the model.</p>\n<p>My experiments is interesting. So I share them with you.</p>\n<table>\n<thead>\n<tr>\n<th>Threshold of the sigmoid function</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.5</td>\n<td>0.56</td>\n</tr>\n<tr>\n<td>0.3</td>\n<td>0.61</td>\n</tr>\n<tr>\n<td>0.2</td>\n<td>0.64</td>\n</tr>\n<tr>\n<td>0.1</td>\n<td>0.68</td>\n</tr>\n<tr>\n<td><strong>0.05</strong></td>\n<td><strong>0.71</strong></td>\n</tr>\n<tr>\n<td><strong>0.01</strong></td>\n<td><strong>0.71</strong></td>\n</tr>\n</tbody>\n</table>\n<p>These score is the same model result (only thresholds were changed). <a href=\"https://www.kaggle.com/competitions/birdclef-2021\" target=\"_blank\">In the previous competition</a>, the best threshold was <strong>0.6 or 0.7</strong>. However, in this competition, the threshold is much lower (0.01 or 0.05). I think the difference lies in whether the emphasis is on precision or recall.</p>\n<p>In the previous competition, about half of the cases were nocall (negative). Therefore, the emphasis was on precision, and <strong>a high threshold</strong> was appropriate. However, In this competition,  the emphasis is on positive, and recall is more important. and <strong>a low threshold</strong> was appropriate. Finally, there are many false positive in soundscape_453028782.ogg with the threshold of 0.01 like below. But it's good score (LB0.71).</p>\n<p><img src=\"https://user-images.githubusercontent.com/34497776/163506679-705a3030-5c07-4290-a658-d3daabed8f7b.png\" alt=\"image\"></p>\n<p>Do you have any idea?<br>\n(I don't completely understand <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/314999\" target=\"_blank\">the metrics for this competition</a>)</p>",
      "rawMarkdown": "I am glad to participate in the bird identification competition again.\n\nI used [PANNs](https://www.kaggle.com/competitions/rfcx-species-audio-detection/discussion/211007#1151299) for both [the last competition](https://www.kaggle.com/competitions/birdclef-2021) and this one, and with PANNs, the threshold is important. This threshold means the threshold of the sigmoid function at the end of the model.\n\nMy experiments is interesting. So I share them with you.\n\n|Threshold of the sigmoid function|LB|\n|---|---|\n|0.5|0.56|\n|0.3|0.61|\n|0.2|0.64|\n|0.1|0.68|\n|**0.05**|**0.71**|\n|**0.01**|**0.71**|\n\nThese score is the same model result (only thresholds were changed). [In the previous competition](https://www.kaggle.com/competitions/birdclef-2021), the best threshold was **0.6 or 0.7**. However, in this competition, the threshold is much lower (0.01 or 0.05). I think the difference lies in whether the emphasis is on precision or recall.\n\nIn the previous competition, about half of the cases were nocall (negative). Therefore, the emphasis was on precision, and **a high threshold** was appropriate. However, In this competition,  the emphasis is on positive, and recall is more important. and **a low threshold** was appropriate. Finally, there are many false positive in soundscape_453028782.ogg with the threshold of 0.01 like below. But it's good score (LB0.71).\n\n![image](https://user-images.githubusercontent.com/34497776/163506679-705a3030-5c07-4290-a658-d3daabed8f7b.png)\n\nDo you have any idea?\n(I don't completely understand [the metrics for this competition](https://www.kaggle.com/competitions/birdclef-2022/discussion/314999))",
      "votes": 40
    },
    {
      "id": 1755842,
      "postDate": "2022-04-15T03:37:01.253Z",
      "content": "<p>Maybe this comment help you:<br>\n<a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/311493#1716290\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/311493#1716290</a></p>\n<blockquote>\n  <p>There's typically much more 'negative' audio for a given species than positive audio. So we chose a weighting which equalizes the impact of the positive/negative labels for each species. (and also equalizing the total weight for each species.)</p>\n</blockquote>\n<p>This indicates they leveled positive/negative class imbalance in metric calculation.</p>",
      "rawMarkdown": "Maybe this comment help you:\nhttps://www.kaggle.com/competitions/birdclef-2022/discussion/311493#1716290\n\n> There's typically much more 'negative' audio for a given species than positive audio. So we chose a weighting which equalizes the impact of the positive/negative labels for each species. (and also equalizing the total weight for each species.)\n\nThis indicates they leveled positive/negative class imbalance in metric calculation.",
      "votes": 5,
      "replies": [
        {
          "id": 1755847,
          "postDate": "2022-04-15T03:41:19.597Z",
          "content": "<p>Thank you.</p>",
          "rawMarkdown": "Thank you.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1774528,
      "postDate": "2022-05-02T08:41:43.337Z",
      "content": "<p><a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> </p>\n<blockquote>\n  <p>Finally, there are many false positive in soundscape_453028782.ogg with the threshold of 0.01 like below.</p>\n</blockquote>\n<p>If you are concerned about the ratio of positive predictions, you might try to output TPR and TNR on the evaluation set as well. My hypothesis is that LB's evaluation metrics use an algorithm similar to Balanced Accuracy[1]. If so, it would be more reasonable to set threshold with respect to TPR/TNR rather than to the number of positive/negative examples the model predicts.</p>\n<p>Also, the following is a quote from the competition's evaluation metrics description, which states that \"un-scored rows are dropped\" before the evaluation algorithm was applied. If we interpret this literally, it means that there were no \"no calls\" in the evaluation of this competition. If this hypothesis is correct, it is reasonable to select the model thresholds with the \"no call\" clips excluded from the validation set.</p>\n<blockquote>\n  <p>Given the amount of audio data used in this competition it wasn't feasible to label every single species found in every soundscape. Instead only a subset of species are actually scored for any given audio file. <strong>After dropping all of the un-scored rows</strong> we technically run a weighted classification accuracy with the weights set such that all of the species are assigned the same total weight and the true negatives and true positives for each species have the same weight.</p>\n</blockquote>\n<h2>Reference</h2>\n<ul>\n<li>[1] <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/321883\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/321883</a></li>\n</ul>",
      "rawMarkdown": "@shinmurashinmura \n\n> Finally, there are many false positive in soundscape_453028782.ogg with the threshold of 0.01 like below.\n\nIf you are concerned about the ratio of positive predictions, you might try to output TPR and TNR on the evaluation set as well. My hypothesis is that LB's evaluation metrics use an algorithm similar to Balanced Accuracy[1]. If so, it would be more reasonable to set threshold with respect to TPR/TNR rather than to the number of positive/negative examples the model predicts.\n\nAlso, the following is a quote from the competition's evaluation metrics description, which states that \"un-scored rows are dropped\" before the evaluation algorithm was applied. If we interpret this literally, it means that there were no \"no calls\" in the evaluation of this competition. If this hypothesis is correct, it is reasonable to select the model thresholds with the \"no call\" clips excluded from the validation set.\n\n> Given the amount of audio data used in this competition it wasn't feasible to label every single species found in every soundscape. Instead only a subset of species are actually scored for any given audio file. **After dropping all of the un-scored rows** we technically run a weighted classification accuracy with the weights set such that all of the species are assigned the same total weight and the true negatives and true positives for each species have the same weight.\n\n## Reference\n\n- [1] https://www.kaggle.com/competitions/birdclef-2022/discussion/321883",
      "votes": 4
    },
    {
      "id": 1784764,
      "postDate": "2022-05-11T12:55:45.920Z",
      "content": "<p>Just trust your local cv.</p>",
      "rawMarkdown": "Just trust your local cv."
    },
    {
      "id": 1784760,
      "postDate": "2022-05-11T12:53:49.660Z",
      "content": "<p>Trust your local cv.</p>",
      "rawMarkdown": "Trust your local cv."
    },
    {
      "id": 1766304,
      "postDate": "2022-04-24T12:18:03.833Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1755842,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2022-04-15T03:37:01.253000",
      "content": "<p>Maybe this comment help you:<br>\n<a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/311493#1716290\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/311493#1716290</a></p>\n<blockquote>\n  <p>There's typically much more 'negative' audio for a given species than positive audio. So we chose a weighting which equalizes the impact of the positive/negative labels for each species. (and also equalizing the total weight for each species.)</p>\n</blockquote>\n<p>This indicates they leveled positive/negative class imbalance in metric calculation.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1755847,
          "author_name": "shinmura0",
          "author_url": "",
          "post_date": "2022-04-15T03:41:19.597000",
          "content": "<p>Thank you.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1774528,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2022-05-02T08:41:43.337000",
      "content": "<p><a href=\"https://www.kaggle.com/shinmurashinmura\" target=\"_blank\">@shinmurashinmura</a> </p>\n<blockquote>\n  <p>Finally, there are many false positive in soundscape_453028782.ogg with the threshold of 0.01 like below.</p>\n</blockquote>\n<p>If you are concerned about the ratio of positive predictions, you might try to output TPR and TNR on the evaluation set as well. My hypothesis is that LB's evaluation metrics use an algorithm similar to Balanced Accuracy[1]. If so, it would be more reasonable to set threshold with respect to TPR/TNR rather than to the number of positive/negative examples the model predicts.</p>\n<p>Also, the following is a quote from the competition's evaluation metrics description, which states that \"un-scored rows are dropped\" before the evaluation algorithm was applied. If we interpret this literally, it means that there were no \"no calls\" in the evaluation of this competition. If this hypothesis is correct, it is reasonable to select the model thresholds with the \"no call\" clips excluded from the validation set.</p>\n<blockquote>\n  <p>Given the amount of audio data used in this competition it wasn't feasible to label every single species found in every soundscape. Instead only a subset of species are actually scored for any given audio file. <strong>After dropping all of the un-scored rows</strong> we technically run a weighted classification accuracy with the weights set such that all of the species are assigned the same total weight and the true negatives and true positives for each species have the same weight.</p>\n</blockquote>\n<h2>Reference</h2>\n<ul>\n<li>[1] <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/321883\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/321883</a></li>\n</ul>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1784764,
      "author_name": "Turbo",
      "author_url": "",
      "post_date": "2022-05-11T12:55:45.920000",
      "content": "<p>Just trust your local cv.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1784760,
      "author_name": "Turbo",
      "author_url": "",
      "post_date": "2022-05-11T12:53:49.660000",
      "content": "<p>Trust your local cv.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1766304,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-24T12:18:03.833000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1755821": "I am glad to participate in the bird identification competition again.\n\nI used [PANNs](https://www.kaggle.com/competitions/rfcx-species-audio-detection/discussion/211007#1151299) for both [the last competition](https://www.kaggle.com/competitions/birdclef-2021) and this one, and with PANNs, the threshold is important. This threshold means the threshold of the sigmoid function at the end of the model.\n\nMy experiments is interesting. So I share them with you.\n\n|Threshold of the sigmoid function|LB|\n|---|---|\n|0.5|0.56|\n|0.3|0.61|\n|0.2|0.64|\n|0.1|0.68|\n|**0.05**|**0.71**|\n|**0.01**|**0.71**|\n\nThese score is the same model result (only thresholds were changed). [In the previous competition](https://www.kaggle.com/competitions/birdclef-2021), the best threshold was **0.6 or 0.7**. However, in this competition, the threshold is much lower (0.01 or 0.05). I think the difference lies in whether the emphasis is on precision or recall.\n\nIn the previous competition, about half of the cases were nocall (negative). Therefore, the emphasis was on precision, and **a high threshold** was appropriate. However, In this competition,  the emphasis is on positive, and recall is more important. and **a low threshold** was appropriate. Finally, there are many false positive in soundscape_453028782.ogg with the threshold of 0.01 like below. But it's good score (LB0.71).\n\n![image](https://user-images.githubusercontent.com/34497776/163506679-705a3030-5c07-4290-a658-d3daabed8f7b.png)\n\nDo you have any idea?\n(I don't completely understand [the metrics for this competition](https://www.kaggle.com/competitions/birdclef-2022/discussion/314999))",
    "1755842": "Maybe this comment help you:\nhttps://www.kaggle.com/competitions/birdclef-2022/discussion/311493#1716290\n\n> There's typically much more 'negative' audio for a given species than positive audio. So we chose a weighting which equalizes the impact of the positive/negative labels for each species. (and also equalizing the total weight for each species.)\n\nThis indicates they leveled positive/negative class imbalance in metric calculation.",
    "1774528": "@shinmurashinmura \n\n> Finally, there are many false positive in soundscape_453028782.ogg with the threshold of 0.01 like below.\n\nIf you are concerned about the ratio of positive predictions, you might try to output TPR and TNR on the evaluation set as well. My hypothesis is that LB's evaluation metrics use an algorithm similar to Balanced Accuracy[1]. If so, it would be more reasonable to set threshold with respect to TPR/TNR rather than to the number of positive/negative examples the model predicts.\n\nAlso, the following is a quote from the competition's evaluation metrics description, which states that \"un-scored rows are dropped\" before the evaluation algorithm was applied. If we interpret this literally, it means that there were no \"no calls\" in the evaluation of this competition. If this hypothesis is correct, it is reasonable to select the model thresholds with the \"no call\" clips excluded from the validation set.\n\n> Given the amount of audio data used in this competition it wasn't feasible to label every single species found in every soundscape. Instead only a subset of species are actually scored for any given audio file. **After dropping all of the un-scored rows** we technically run a weighted classification accuracy with the weights set such that all of the species are assigned the same total weight and the true negatives and true positives for each species have the same weight.\n\n## Reference\n\n- [1] https://www.kaggle.com/competitions/birdclef-2022/discussion/321883",
    "1784764": "Just trust your local cv.",
    "1784760": "Trust your local cv.",
    "1766304": ""
  }
}