{
  "id": 323201,
  "title": "LB probe result: Estimated call rate for test soundscape is about 90%",
  "url": "/competitions/birdclef-2022/discussion/323201",
  "author_name": "Bilzard",
  "post_date": "2022-05-05T08:47:08.614000",
  "votes": 18,
  "comment_count": 7,
  "views": 0,
  "content": "<h1>TL;DR</h1>\n<p>I estimated call ratio of the test sound scapes using a pre-trained binary classifier on external data.<br>\nThe result shows that the estimated call ratio of the test sound scapes is <strong>90%</strong>. This is considerably more frequent than the frequency of labels in a single soundscape given as a sample.</p>\n<h1>Background</h1>\n<p>I hypothesized that <strong>no-calls were absent in the test soundscapes</strong> since I read the following explanation of the evaluation metrics.</p>\n<blockquote>\n  <p><strong>After dropping all of the un-scored rows</strong> we technically run a weighted classification accuracy with the weights set such that all of the species are assigned the same total weight and the true negatives and true positives for each species have the same weight.</p>\n</blockquote>\n<h1>Experimental Summary</h1>\n<ol>\n<li>train a binary classifier using the DCASE2018[1] public dataset (BirdVox-DCASE-20K, freefield1010).</li>\n<li>tabulate the percentage of calls in the test data using the binary classifier trained in 1.</li>\n<li>estimate the proportion of calls using the LB probing method (LB hack II [2])</li>\n</ol>\n<p>The performance of the binary classifier pre-trained in process 1 was AUC=0.92 and accuracy=0.82 for the BirdCLEF2021[3] training soundscape.</p>\n<h1>Experimental Results</h1>\n<p>The estimate of the proportion of calls was <strong>0.904</strong>.<br>\nEven discounting the fact that the accuracy of the binary classifier is not 100%, this indicates that there are quite a few call labels in the test soundscapes.<br>\nThis result provides some support for the hypothesis that no-call labels are not present in the test soundscape.</p>\n<h1>Disclaimer</h1>\n<p>I regret that, for various reasons, I am unable to publish the materials needed to reproduce this experiment. (However, if you have your own trained binary classifier and use LB hack II, it is possible to reproduce.)<br>\nAlso, please be aware that the results of this discussion may be incorrect due to bugs in the code, unexpected assumption errors, or assumptions in the estimation method that do not hold.</p>\n<h1>Reference</h1>\n<ul>\n<li>[1] <a href=\"https://dcase.community/challenge2018/task-bird-audio-detection\" target=\"_blank\">https://dcase.community/challenge2018/task-bird-audio-detection</a></li>\n<li>[2] <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/322606\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/322606</a></li>\n<li>[3] <a href=\"https://www.kaggle.com/competitions/birdclef-2021\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2021</a></li>\n</ul>",
  "messages": [
    {
      "id": 1778336,
      "postDate": "2022-05-05T08:47:08.613Z",
      "content": "<h1>TL;DR</h1>\n<p>I estimated call ratio of the test sound scapes using a pre-trained binary classifier on external data.<br>\nThe result shows that the estimated call ratio of the test sound scapes is <strong>90%</strong>. This is considerably more frequent than the frequency of labels in a single soundscape given as a sample.</p>\n<h1>Background</h1>\n<p>I hypothesized that <strong>no-calls were absent in the test soundscapes</strong> since I read the following explanation of the evaluation metrics.</p>\n<blockquote>\n  <p><strong>After dropping all of the un-scored rows</strong> we technically run a weighted classification accuracy with the weights set such that all of the species are assigned the same total weight and the true negatives and true positives for each species have the same weight.</p>\n</blockquote>\n<h1>Experimental Summary</h1>\n<ol>\n<li>train a binary classifier using the DCASE2018[1] public dataset (BirdVox-DCASE-20K, freefield1010).</li>\n<li>tabulate the percentage of calls in the test data using the binary classifier trained in 1.</li>\n<li>estimate the proportion of calls using the LB probing method (LB hack II [2])</li>\n</ol>\n<p>The performance of the binary classifier pre-trained in process 1 was AUC=0.92 and accuracy=0.82 for the BirdCLEF2021[3] training soundscape.</p>\n<h1>Experimental Results</h1>\n<p>The estimate of the proportion of calls was <strong>0.904</strong>.<br>\nEven discounting the fact that the accuracy of the binary classifier is not 100%, this indicates that there are quite a few call labels in the test soundscapes.<br>\nThis result provides some support for the hypothesis that no-call labels are not present in the test soundscape.</p>\n<h1>Disclaimer</h1>\n<p>I regret that, for various reasons, I am unable to publish the materials needed to reproduce this experiment. (However, if you have your own trained binary classifier and use LB hack II, it is possible to reproduce.)<br>\nAlso, please be aware that the results of this discussion may be incorrect due to bugs in the code, unexpected assumption errors, or assumptions in the estimation method that do not hold.</p>\n<h1>Reference</h1>\n<ul>\n<li>[1] <a href=\"https://dcase.community/challenge2018/task-bird-audio-detection\" target=\"_blank\">https://dcase.community/challenge2018/task-bird-audio-detection</a></li>\n<li>[2] <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/322606\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/322606</a></li>\n<li>[3] <a href=\"https://www.kaggle.com/competitions/birdclef-2021\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2021</a></li>\n</ul>",
      "rawMarkdown": "# TL;DR\n\nI estimated call ratio of the test sound scapes using a pre-trained binary classifier on external data.\nThe result shows that the estimated call ratio of the test sound scapes is **90%**. This is considerably more frequent than the frequency of labels in a single soundscape given as a sample.\n\n# Background\n\nI hypothesized that **no-calls were absent in the test soundscapes** since I read the following explanation of the evaluation metrics.\n\n> **After dropping all of the un-scored rows** we technically run a weighted classification accuracy with the weights set such that all of the species are assigned the same total weight and the true negatives and true positives for each species have the same weight.\n\n# Experimental Summary\n\n1. train a binary classifier using the DCASE2018[1] public dataset (BirdVox-DCASE-20K, freefield1010).\n2. tabulate the percentage of calls in the test data using the binary classifier trained in 1.\n3. estimate the proportion of calls using the LB probing method (LB hack II [2])\n\nThe performance of the binary classifier pre-trained in process 1 was AUC=0.92 and accuracy=0.82 for the BirdCLEF2021[3] training soundscape.\n\n# Experimental Results\n\nThe estimate of the proportion of calls was **0.904**.\nEven discounting the fact that the accuracy of the binary classifier is not 100%, this indicates that there are quite a few call labels in the test soundscapes.\nThis result provides some support for the hypothesis that no-call labels are not present in the test soundscape.\n\n# Disclaimer\n\nI regret that, for various reasons, I am unable to publish the materials needed to reproduce this experiment. (However, if you have your own trained binary classifier and use LB hack II, it is possible to reproduce.)\nAlso, please be aware that the results of this discussion may be incorrect due to bugs in the code, unexpected assumption errors, or assumptions in the estimation method that do not hold.\n\n# Reference\n\n- [1] https://dcase.community/challenge2018/task-bird-audio-detection\n- [2] https://www.kaggle.com/competitions/birdclef-2022/discussion/322606\n- [3] https://www.kaggle.com/competitions/birdclef-2021",
      "votes": 18
    },
    {
      "id": 1779033,
      "postDate": "2022-05-06T03:10:19.877Z",
      "content": "<p>I don't know why someone downvote this topic.<br>\nIt may certainly not be the method envisioned by the host, but I would think that if you find a hole in a competition like this, it would be better to share it so that all participants can use it rather than to let a few hackers monopolize it.<br>\nPlease understand that I am posting this in this context.</p>",
      "rawMarkdown": "I don't know why someone downvote this topic.\nIt may certainly not be the method envisioned by the host, but I would think that if you find a hole in a competition like this, it would be better to share it so that all participants can use it rather than to let a few hackers monopolize it.\nPlease understand that I am posting this in this context.",
      "votes": 4
    },
    {
      "id": 1778485,
      "postDate": "2022-05-05T11:06:43.860Z",
      "content": "<h2>How to gain more accuracy on estimate?</h2>\n<p>By increasing the sample of estimates with varying the magnification factor of p, we can increase the accuracy of the estimates.</p>\n<ol>\n<li>multiplying a factor a (0&lt; a &lt; 1) by p: p = p * a, and submit notebooks</li>\n<li>from observed notebook scores m_0 and m_1, calculate p_est</li>\n<li>recover original p_est :p_est = p_est / a</li>\n</ol>\n<h2>Results</h2>\n<table>\n<thead>\n<tr>\n<th>a</th>\n<th>m_0</th>\n<th>m_1</th>\n<th>p_est</th>\n<th>e_theory</th>\n<th>weight</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>0.71</td>\n<td>0.52</td>\n<td>0.905</td>\n<td>0.047</td>\n<td>4</td>\n</tr>\n<tr>\n<td>0.75</td>\n<td>0.71</td>\n<td>0.57</td>\n<td>0.889</td>\n<td>0.063</td>\n<td>3</td>\n</tr>\n<tr>\n<td>0.5</td>\n<td>0.71</td>\n<td>0.62</td>\n<td>0.857</td>\n<td>0.095</td>\n<td>2</td>\n</tr>\n<tr>\n<td>0.25</td>\n<td>0.71</td>\n<td>0.67</td>\n<td>0.762</td>\n<td>0.190</td>\n<td>1</td>\n</tr>\n</tbody>\n</table>\n<p>Note: the disadvantage of this method is that the error is proportional to 1/a.<br>\nE.g. for a=0.25, error(p)_theory = 0.19.<br>\nSo weighted mean is a rational way of averaging these estimates:</p>\n<p>$$<br>\n\\text{weighted_mean}(p_{\\text{est}}) = \\frac{\\sum_i{w_i p_{i}}}{\\sum_i{w_i}} = 0.876<br>\n$$</p>",
      "rawMarkdown": "## How to gain more accuracy on estimate?\n\nBy increasing the sample of estimates with varying the magnification factor of p, we can increase the accuracy of the estimates.\n\n1. multiplying a factor a (0< a < 1) by p: p = p * a, and submit notebooks\n2. from observed notebook scores m_0 and m_1, calculate p_est\n2. recover original p_est :p_est = p_est / a\n\n## Results\n\na|m_0|m_1|p_est|e_theory|weight\n--|--|--|--|--|--\n1|0.71|0.52|0.905|0.047|4\n0.75|0.71|0.57|0.889|0.063|3\n0.5|0.71|0.62|0.857|0.095|2\n0.25|0.71|0.67|0.762|0.190|1\n\nNote: the disadvantage of this method is that the error is proportional to 1/a.\nE.g. for a=0.25, error(p)_theory = 0.19.\nSo weighted mean is a rational way of averaging these estimates:\n\n$$\n\\text{weighted_mean}(p_{\\text{est}}) = \\frac{\\sum_i{w_i p_{i}}}{\\sum_i{w_i}} = 0.876\n$$",
      "votes": 1
    },
    {
      "id": 1778363,
      "postDate": "2022-05-05T09:01:37.053Z",
      "content": "<p>Note that using the same binary classifier, I also estimated the call ratio for a sample test soundscape; it was 1/12 = <strong>8.3%</strong>. Thus, I have already confirmed that the predictions of the binary classifier are not biased toward the positive values.</p>",
      "rawMarkdown": "Note that using the same binary classifier, I also estimated the call ratio for a sample test soundscape; it was 1/12 = **8.3%**. Thus, I have already confirmed that the predictions of the binary classifier are not biased toward the positive values.",
      "votes": 1,
      "replies": [
        {
          "id": 1778450,
          "postDate": "2022-05-05T10:04:21.483Z",
          "content": "<p>Using the same binary classifier used in this experiment, the following figures show the results of the predicted probabilities of call for a sample test soundscape. The upper figure represents the spectrogram of the soundscape. Although it is difficult to see from the spectrogram with the naked eye, if one actually listens to the audio, one can hear that a sound like a bird's call was actually recorded in the 25-30 second clip.</p>\n<p><a href=\"https://ibb.co/vzFmdW8\"><img src=\"https://i.ibb.co/bd0LvVn/Screen-Shot-2022-05-05-at-18-59-14.png\" alt=\"Screen-Shot-2022-05-05-at-18-59-14\"></a></p>",
          "rawMarkdown": "Using the same binary classifier used in this experiment, the following figures show the results of the predicted probabilities of call for a sample test soundscape. The upper figure represents the spectrogram of the soundscape. Although it is difficult to see from the spectrogram with the naked eye, if one actually listens to the audio, one can hear that a sound like a bird's call was actually recorded in the 25-30 second clip.\n\n<a href=\"https://ibb.co/vzFmdW8\"><img src=\"https://i.ibb.co/bd0LvVn/Screen-Shot-2022-05-05-at-18-59-14.png\" alt=\"Screen-Shot-2022-05-05-at-18-59-14\" border=\"0\"></a>",
          "votes": 1
        }
      ]
    },
    {
      "id": 1778422,
      "postDate": "2022-05-05T09:46:17.400Z",
      "content": "<p>If the consequences of this discussion are correct, it means that a binary classifier that classifies call/no-call will not help much in this competition (although it may improve slightly because of the inclusion of partial no-call labels).</p>\n<p>Conversely, it is possible to verify the correctness of the consequences of this discussion to some extent by actually using the binary classifier and seeing how much the scores improve.</p>",
      "rawMarkdown": "If the consequences of this discussion are correct, it means that a binary classifier that classifies call/no-call will not help much in this competition (although it may improve slightly because of the inclusion of partial no-call labels).\n\nConversely, it is possible to verify the correctness of the consequences of this discussion to some extent by actually using the binary classifier and seeing how much the scores improve.",
      "votes": 2
    },
    {
      "id": 1778915,
      "postDate": "2022-05-05T21:32:44.663Z",
      "content": "<p>Great Work !!!!</p>",
      "rawMarkdown": "Great Work !!!!"
    },
    {
      "id": 1778339,
      "postDate": "2022-05-05T08:51:01.137Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1779033,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2022-05-06T03:10:19.877000",
      "content": "<p>I don't know why someone downvote this topic.<br>\nIt may certainly not be the method envisioned by the host, but I would think that if you find a hole in a competition like this, it would be better to share it so that all participants can use it rather than to let a few hackers monopolize it.<br>\nPlease understand that I am posting this in this context.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1778485,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2022-05-05T11:06:43.860000",
      "content": "<h2>How to gain more accuracy on estimate?</h2>\n<p>By increasing the sample of estimates with varying the magnification factor of p, we can increase the accuracy of the estimates.</p>\n<ol>\n<li>multiplying a factor a (0&lt; a &lt; 1) by p: p = p * a, and submit notebooks</li>\n<li>from observed notebook scores m_0 and m_1, calculate p_est</li>\n<li>recover original p_est :p_est = p_est / a</li>\n</ol>\n<h2>Results</h2>\n<table>\n<thead>\n<tr>\n<th>a</th>\n<th>m_0</th>\n<th>m_1</th>\n<th>p_est</th>\n<th>e_theory</th>\n<th>weight</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>1</td>\n<td>0.71</td>\n<td>0.52</td>\n<td>0.905</td>\n<td>0.047</td>\n<td>4</td>\n</tr>\n<tr>\n<td>0.75</td>\n<td>0.71</td>\n<td>0.57</td>\n<td>0.889</td>\n<td>0.063</td>\n<td>3</td>\n</tr>\n<tr>\n<td>0.5</td>\n<td>0.71</td>\n<td>0.62</td>\n<td>0.857</td>\n<td>0.095</td>\n<td>2</td>\n</tr>\n<tr>\n<td>0.25</td>\n<td>0.71</td>\n<td>0.67</td>\n<td>0.762</td>\n<td>0.190</td>\n<td>1</td>\n</tr>\n</tbody>\n</table>\n<p>Note: the disadvantage of this method is that the error is proportional to 1/a.<br>\nE.g. for a=0.25, error(p)_theory = 0.19.<br>\nSo weighted mean is a rational way of averaging these estimates:</p>\n<p>$$<br>\n\\text{weighted_mean}(p_{\\text{est}}) = \\frac{\\sum_i{w_i p_{i}}}{\\sum_i{w_i}} = 0.876<br>\n$$</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1778363,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2022-05-05T09:01:37.053000",
      "content": "<p>Note that using the same binary classifier, I also estimated the call ratio for a sample test soundscape; it was 1/12 = <strong>8.3%</strong>. Thus, I have already confirmed that the predictions of the binary classifier are not biased toward the positive values.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1778450,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-05-05T10:04:21.483000",
          "content": "<p>Using the same binary classifier used in this experiment, the following figures show the results of the predicted probabilities of call for a sample test soundscape. The upper figure represents the spectrogram of the soundscape. Although it is difficult to see from the spectrogram with the naked eye, if one actually listens to the audio, one can hear that a sound like a bird's call was actually recorded in the 25-30 second clip.</p>\n<p><a href=\"https://ibb.co/vzFmdW8\"><img src=\"https://i.ibb.co/bd0LvVn/Screen-Shot-2022-05-05-at-18-59-14.png\" alt=\"Screen-Shot-2022-05-05-at-18-59-14\"></a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1778422,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2022-05-05T09:46:17.400000",
      "content": "<p>If the consequences of this discussion are correct, it means that a binary classifier that classifies call/no-call will not help much in this competition (although it may improve slightly because of the inclusion of partial no-call labels).</p>\n<p>Conversely, it is possible to verify the correctness of the consequences of this discussion to some extent by actually using the binary classifier and seeing how much the scores improve.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1778915,
      "author_name": "Arunodhayan",
      "author_url": "",
      "post_date": "2022-05-05T21:32:44.663000",
      "content": "<p>Great Work !!!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1778339,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-05T08:51:01.137000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1778336": "# TL;DR\n\nI estimated call ratio of the test sound scapes using a pre-trained binary classifier on external data.\nThe result shows that the estimated call ratio of the test sound scapes is **90%**. This is considerably more frequent than the frequency of labels in a single soundscape given as a sample.\n\n# Background\n\nI hypothesized that **no-calls were absent in the test soundscapes** since I read the following explanation of the evaluation metrics.\n\n> **After dropping all of the un-scored rows** we technically run a weighted classification accuracy with the weights set such that all of the species are assigned the same total weight and the true negatives and true positives for each species have the same weight.\n\n# Experimental Summary\n\n1. train a binary classifier using the DCASE2018[1] public dataset (BirdVox-DCASE-20K, freefield1010).\n2. tabulate the percentage of calls in the test data using the binary classifier trained in 1.\n3. estimate the proportion of calls using the LB probing method (LB hack II [2])\n\nThe performance of the binary classifier pre-trained in process 1 was AUC=0.92 and accuracy=0.82 for the BirdCLEF2021[3] training soundscape.\n\n# Experimental Results\n\nThe estimate of the proportion of calls was **0.904**.\nEven discounting the fact that the accuracy of the binary classifier is not 100%, this indicates that there are quite a few call labels in the test soundscapes.\nThis result provides some support for the hypothesis that no-call labels are not present in the test soundscape.\n\n# Disclaimer\n\nI regret that, for various reasons, I am unable to publish the materials needed to reproduce this experiment. (However, if you have your own trained binary classifier and use LB hack II, it is possible to reproduce.)\nAlso, please be aware that the results of this discussion may be incorrect due to bugs in the code, unexpected assumption errors, or assumptions in the estimation method that do not hold.\n\n# Reference\n\n- [1] https://dcase.community/challenge2018/task-bird-audio-detection\n- [2] https://www.kaggle.com/competitions/birdclef-2022/discussion/322606\n- [3] https://www.kaggle.com/competitions/birdclef-2021",
    "1779033": "I don't know why someone downvote this topic.\nIt may certainly not be the method envisioned by the host, but I would think that if you find a hole in a competition like this, it would be better to share it so that all participants can use it rather than to let a few hackers monopolize it.\nPlease understand that I am posting this in this context.",
    "1778485": "## How to gain more accuracy on estimate?\n\nBy increasing the sample of estimates with varying the magnification factor of p, we can increase the accuracy of the estimates.\n\n1. multiplying a factor a (0< a < 1) by p: p = p * a, and submit notebooks\n2. from observed notebook scores m_0 and m_1, calculate p_est\n2. recover original p_est :p_est = p_est / a\n\n## Results\n\na|m_0|m_1|p_est|e_theory|weight\n--|--|--|--|--|--\n1|0.71|0.52|0.905|0.047|4\n0.75|0.71|0.57|0.889|0.063|3\n0.5|0.71|0.62|0.857|0.095|2\n0.25|0.71|0.67|0.762|0.190|1\n\nNote: the disadvantage of this method is that the error is proportional to 1/a.\nE.g. for a=0.25, error(p)_theory = 0.19.\nSo weighted mean is a rational way of averaging these estimates:\n\n$$\n\\text{weighted_mean}(p_{\\text{est}}) = \\frac{\\sum_i{w_i p_{i}}}{\\sum_i{w_i}} = 0.876\n$$",
    "1778363": "Note that using the same binary classifier, I also estimated the call ratio for a sample test soundscape; it was 1/12 = **8.3%**. Thus, I have already confirmed that the predictions of the binary classifier are not biased toward the positive values.",
    "1778422": "If the consequences of this discussion are correct, it means that a binary classifier that classifies call/no-call will not help much in this competition (although it may improve slightly because of the inclusion of partial no-call labels).\n\nConversely, it is possible to verify the correctness of the consequences of this discussion to some extent by actually using the binary classifier and seeing how much the scores improve.",
    "1778915": "Great Work !!!!",
    "1778339": ""
  }
}