{
  "id": 322419,
  "title": "LB hack tip: How to estimate LB score of each species",
  "url": "/competitions/birdclef-2022/discussion/322419",
  "author_name": "",
  "post_date": "2022-05-02T03:37:08.467625700Z",
  "votes": 15,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I will share on how to estimate the contribution of each bird species to the LB score.</p>\n<h2>Calculation: how to estimate the score of each bird species</h2>\n<p>Here I make the following two assumptions.</p>\n<ol>\n<li>the LB score is calculated by the average of the scores of the individual bird species (macro-average score)</li>\n<li>if the predictions are reversed, the LB score is 1 minus the original score. That is, <code>score(y, 1-y_pred) = 1 - score(y, y_pred)</code>.</li>\n</ol>\n<p>(We can be sure about 2 by submitting all predictions inverted for any given model.)</p>\n<p>Under the above assumptions, the scores for individual bird species are s_1, . , s_21, the LB score is calculated as follows.</p>\n<p>$$<br>\nm_0 = \\frac{1}{21} (s_0 + … + s_i + … + s_{20}) \\tag{1}<br>\n$$</p>\n<p>Now, if we invert the predictions for only one bird species (let's say i), the score would be</p>\n<p>$$<br>\nm_1 = \\frac{1}{21} \\left(s_0 + … + (1 - s_i) + … + s_{20} \\right) \\tag{2}<br>\n$$</p>\n<p>Solving equations (1) and (2) for s_i yields</p>\n<p>$$<br>\ns_i = \\frac{21}{2} (m_0 - m_1) + 0.5 \\tag{3}<br>\n$$</p>\n<p>That is, we can estimate the score of bird i by substituting the difference between the score m_0 for a normal submission and the predicted inverted score m_1 for a certain bird species into equation (3).</p>\n<h2>Experiment 1: score estimation for \"skylar\"</h2>\n<p>I tried the above experiment on the most sampled bird species, \"skylar\", and found no difference in scores between m_0 and m_1. I assume this is due to either the score of \"skylar\" alone being less than 0.6, or the \"skylar\" not being used in the publicLB score calculation.</p>\n<h2>Experiment 2: Estimation of the average of the scores of the top 10 birds</h2>\n<p>In Experiment 1, we were unable to detect score differences due to prediction reversal due to the small number of bird species, but increasing the number of bird species makes it easier to detect score differences due to prediction reversal.</p>\n<p>Therefore, I decided to invert predictions for the top 10 bird species with the largest sample size and estimate the mean of the scores for these bird species.</p>\n<p>Letting s_top10 be the average of the scores of top 10 species and s_low11 be the average of the scores of the remaining 11 bird species, the LB score is calculated as follows.</p>\n<p>$$<br>\n21 m_0 = 10 s_\\text{top10} + 11 s_\\text{low11} \\tag{4}<br>\n$$</p>\n<p>The inverted predicted scores for the top 10 bird species can also be calculated as follows</p>\n<p>$$<br>\n21 m_1 = 10 (1 - s_\\text{top10}) + 11 s_\\text{low11} \\tag{5}<br>\n$$</p>\n<p>Solving equations (4) and (5) for s_top10 yields</p>\n<p>$$<br>\ns_\\text{top10} = \\frac{21}{20} (m_0 - m_1) + 0.5 \\tag{6}<br>\n$$</p>\n<p>I calculated m_0 and m_1 using the public Notebook model[1] and found them to be 0.71 and 0.49, respectively. Therefore, s_top10 can be calculated to be <strong>0.731</strong>.</p>\n<p>For the bottom 11 species, we can follow in the similar process, we estimated s_low11 to be <strong>0.700</strong>.</p>\n<p>These results indicate that the top 10 bird species with the largest sample size contribute more to the LB score.</p>\n<p>Although the low precision of the LB score does not allow us to predict exact values, these methods give us a rough idea of the degree to which your model's predictions are biased toward a particular species for public test data.</p>\n<p>Finally, I will share the source code I used for my experiments.</p>\n<p>Happy Kaggle.</p>\n<h2>Source code of score inversion</h2>\n<pre><code>def invert_target(input_df, birds_to_invert):\n    assert type(birds_to_invert) == list, type(birds)\n    tmp_df = input_df.copy()\n    prefixes, dates, birds, end_secs = zip(*tmp_df.row_id.str.split(\"_\"))\n    tmp_df[\"prefix\"] = prefixes\n    tmp_df[\"date\"] = dates\n    tmp_df[\"bird\"] = birds\n    tmp_df[\"end_sec\"] = end_secs\n\n    idxs = tmp_df.query(\"bird in @birds_to_invert\").index\n    tmp_df.loc[idxs, \"target\"] = tmp_df.loc[idxs, \"target\"].apply(lambda x: not(x))\n\n    return tmp_df[[\"row_id\", \"target\"]]\n</code></pre>\n<pre><code>sample_submission = ... # your model's prediction\n\ntrain = pd.read_csv(\"../input/birdclef-2022/train_metadata.csv\")\nscored = train.query(\"primary_label in @scored_birds\")\nscored_count = scored[\"primary_label\"].value_counts()\ntop10 = scored_count[:10].index.tolist()\nlow11 = scored_count[10:].index.tolist()\ntmp_df = invert_target(sample_submission, low11)\nsample_submission[\"target\"] = tmp_df[\"target\"]\n</code></pre>\n<h2>Reference</h2>\n<ul>\n<li>[1] <a href=\"https://www.kaggle.com/code/kaerunantoka/birdclef2022-ex005-f0-infer\" target=\"_blank\">https://www.kaggle.com/code/kaerunantoka/birdclef2022-ex005-f0-infer</a></li>\n</ul>",
  "messages": [
    {
      "id": "1774279",
      "postDate": "05/02/2022 03:37:08",
      "content": "<p>I will share on how to estimate the contribution of each bird species to the LB score.</p>\n<h2>Calculation: how to estimate the score of each bird species</h2>\n<p>Here I make the following two assumptions.</p>\n<ol>\n<li>the LB score is calculated by the average of the scores of the individual bird species (macro-average score)</li>\n<li>if the predictions are reversed, the LB score is 1 minus the original score. That is, <code>score(y, 1-y_pred) = 1 - score(y, y_pred)</code>.</li>\n</ol>\n<p>(We can be sure about 2 by submitting all predictions inverted for any given model.)</p>\n<p>Under the above assumptions, the scores for individual bird species are s_1, . , s_21, the LB score is calculated as follows.</p>\n<p>$$<br>\nm_0 = \\frac{1}{21} (s_0 + … + s_i + … + s_{20}) \\tag{1}<br>\n$$</p>\n<p>Now, if we invert the predictions for only one bird species (let's say i), the score would be</p>\n<p>$$<br>\nm_1 = \\frac{1}{21} \\left(s_0 + … + (1 - s_i) + … + s_{20} \\right) \\tag{2}<br>\n$$</p>\n<p>Solving equations (1) and (2) for s_i yields</p>\n<p>$$<br>\ns_i = \\frac{21}{2} (m_0 - m_1) + 0.5 \\tag{3}<br>\n$$</p>\n<p>That is, we can estimate the score of bird i by substituting the difference between the score m_0 for a normal submission and the predicted inverted score m_1 for a certain bird species into equation (3).</p>\n<h2>Experiment 1: score estimation for \"skylar\"</h2>\n<p>I tried the above experiment on the most sampled bird species, \"skylar\", and found no difference in scores between m_0 and m_1. I assume this is due to either the score of \"skylar\" alone being less than 0.6, or the \"skylar\" not being used in the publicLB score calculation.</p>\n<h2>Experiment 2: Estimation of the average of the scores of the top 10 birds</h2>\n<p>In Experiment 1, we were unable to detect score differences due to prediction reversal due to the small number of bird species, but increasing the number of bird species makes it easier to detect score differences due to prediction reversal.</p>\n<p>Therefore, I decided to invert predictions for the top 10 bird species with the largest sample size and estimate the mean of the scores for these bird species.</p>\n<p>Letting s_top10 be the average of the scores of top 10 species and s_low11 be the average of the scores of the remaining 11 bird species, the LB score is calculated as follows.</p>\n<p>$$<br>\n21 m_0 = 10 s_\\text{top10} + 11 s_\\text{low11} \\tag{4}<br>\n$$</p>\n<p>The inverted predicted scores for the top 10 bird species can also be calculated as follows</p>\n<p>$$<br>\n21 m_1 = 10 (1 - s_\\text{top10}) + 11 s_\\text{low11} \\tag{5}<br>\n$$</p>\n<p>Solving equations (4) and (5) for s_top10 yields</p>\n<p>$$<br>\ns_\\text{top10} = \\frac{21}{20} (m_0 - m_1) + 0.5 \\tag{6}<br>\n$$</p>\n<p>I calculated m_0 and m_1 using the public Notebook model[1] and found them to be 0.71 and 0.49, respectively. Therefore, s_top10 can be calculated to be <strong>0.731</strong>.</p>\n<p>For the bottom 11 species, we can follow in the similar process, we estimated s_low11 to be <strong>0.700</strong>.</p>\n<p>These results indicate that the top 10 bird species with the largest sample size contribute more to the LB score.</p>\n<p>Although the low precision of the LB score does not allow us to predict exact values, these methods give us a rough idea of the degree to which your model's predictions are biased toward a particular species for public test data.</p>\n<p>Finally, I will share the source code I used for my experiments.</p>\n<p>Happy Kaggle.</p>\n<h2>Source code of score inversion</h2>\n<pre><code>def invert_target(input_df, birds_to_invert):\n    assert type(birds_to_invert) == list, type(birds)\n    tmp_df = input_df.copy()\n    prefixes, dates, birds, end_secs = zip(*tmp_df.row_id.str.split(\"_\"))\n    tmp_df[\"prefix\"] = prefixes\n    tmp_df[\"date\"] = dates\n    tmp_df[\"bird\"] = birds\n    tmp_df[\"end_sec\"] = end_secs\n\n    idxs = tmp_df.query(\"bird in @birds_to_invert\").index\n    tmp_df.loc[idxs, \"target\"] = tmp_df.loc[idxs, \"target\"].apply(lambda x: not(x))\n\n    return tmp_df[[\"row_id\", \"target\"]]\n</code></pre>\n<pre><code>sample_submission = ... # your model's prediction\n\ntrain = pd.read_csv(\"../input/birdclef-2022/train_metadata.csv\")\nscored = train.query(\"primary_label in @scored_birds\")\nscored_count = scored[\"primary_label\"].value_counts()\ntop10 = scored_count[:10].index.tolist()\nlow11 = scored_count[10:].index.tolist()\ntmp_df = invert_target(sample_submission, low11)\nsample_submission[\"target\"] = tmp_df[\"target\"]\n</code></pre>\n<h2>Reference</h2>\n<ul>\n<li>[1] <a href=\"https://www.kaggle.com/code/kaerunantoka/birdclef2022-ex005-f0-infer\" target=\"_blank\">https://www.kaggle.com/code/kaerunantoka/birdclef2022-ex005-f0-infer</a></li>\n</ul>",
      "rawMarkdown": "I will share on how to estimate the contribution of each bird species to the LB score.\n\n## Calculation: how to estimate the score of each bird species\n\nHere I make the following two assumptions.\n\n1. the LB score is calculated by the average of the scores of the individual bird species (macro-average score)\n2. if the predictions are reversed, the LB score is 1 minus the original score. That is, `score(y, 1-y_pred) = 1 - score(y, y_pred)`.\n\n(We can be sure about 2 by submitting all predictions inverted for any given model.)\n\nUnder the above assumptions, the scores for individual bird species are s_1, . , s_21, the LB score is calculated as follows.\n\n$$\nm_0 = \\frac{1}{21} (s_0 + ... + s_i + ... + s_{20}) \\tag{1}\n$$\n\nNow, if we invert the predictions for only one bird species (let's say i), the score would be\n\n$$\nm_1 = \\frac{1}{21} \\left(s_0 + ... + (1 - s_i) + ... + s_{20} \\right) \\tag{2}\n$$\n\nSolving equations (1) and (2) for s_i yields\n\n$$\ns_i = \\frac{21}{2} (m_0 - m_1) + 0.5 \\tag{3}\n$$\n\nThat is, we can estimate the score of bird i by substituting the difference between the score m_0 for a normal submission and the predicted inverted score m_1 for a certain bird species into equation (3).\n\n## Experiment 1: score estimation for \"skylar\"\n\nI tried the above experiment on the most sampled bird species, \"skylar\", and found no difference in scores between m_0 and m_1. I assume this is due to either the score of \"skylar\" alone being less than 0.6, or the \"skylar\" not being used in the publicLB score calculation.\n\n## Experiment 2: Estimation of the average of the scores of the top 10 birds\n\nIn Experiment 1, we were unable to detect score differences due to prediction reversal due to the small number of bird species, but increasing the number of bird species makes it easier to detect score differences due to prediction reversal.\n\nTherefore, I decided to invert predictions for the top 10 bird species with the largest sample size and estimate the mean of the scores for these bird species.\n\nLetting s_top10 be the average of the scores of top 10 species and s_low11 be the average of the scores of the remaining 11 bird species, the LB score is calculated as follows.\n\n$$\n21 m_0 = 10 s_\\text{top10} + 11 s_\\text{low11} \\tag{4}\n$$\n\nThe inverted predicted scores for the top 10 bird species can also be calculated as follows\n\n$$\n21 m_1 = 10 (1 - s_\\text{top10}) + 11 s_\\text{low11} \\tag{5}\n$$\n\nSolving equations (4) and (5) for s_top10 yields\n\n$$\ns_\\text{top10} = \\frac{21}{20} (m_0 - m_1) + 0.5 \\tag{6}\n$$\n\nI calculated m_0 and m_1 using the public Notebook model[1] and found them to be 0.71 and 0.49, respectively. Therefore, s_top10 can be calculated to be **0.731**.\n\nFor the bottom 11 species, we can follow in the similar process, we estimated s_low11 to be **0.700**.\n\nThese results indicate that the top 10 bird species with the largest sample size contribute more to the LB score.\n\nAlthough the low precision of the LB score does not allow us to predict exact values, these methods give us a rough idea of the degree to which your model's predictions are biased toward a particular species for public test data.\n\nFinally, I will share the source code I used for my experiments.\n\nHappy Kaggle.\n\n## Source code of score inversion\n\n```python\ndef invert_target(input_df, birds_to_invert):\n    assert type(birds_to_invert) == list, type(birds)\n    tmp_df = input_df.copy()\n    prefixes, dates, birds, end_secs = zip(*tmp_df.row_id.str.split(\"_\"))\n    tmp_df[\"prefix\"] = prefixes\n    tmp_df[\"date\"] = dates\n    tmp_df[\"bird\"] = birds\n    tmp_df[\"end_sec\"] = end_secs\n    \n    idxs = tmp_df.query(\"bird in @birds_to_invert\").index\n    tmp_df.loc[idxs, \"target\"] = tmp_df.loc[idxs, \"target\"].apply(lambda x: not(x))\n    \n    return tmp_df[[\"row_id\", \"target\"]]\n```\n\n```python\nsample_submission = ... # your model's prediction\n\ntrain = pd.read_csv(\"../input/birdclef-2022/train_metadata.csv\")\nscored = train.query(\"primary_label in @scored_birds\")\nscored_count = scored[\"primary_label\"].value_counts()\ntop10 = scored_count[:10].index.tolist()\nlow11 = scored_count[10:].index.tolist()\ntmp_df = invert_target(sample_submission, low11)\nsample_submission[\"target\"] = tmp_df[\"target\"]\n```\n\n## Reference\n\n- [1] https://www.kaggle.com/code/kaerunantoka/birdclef2022-ex005-f0-infer",
      "votes": null
    },
    {
      "id": "1774281",
      "postDate": "05/02/2022 03:46:14",
      "content": "<p>Note: I confirmed assumption 2 seems to be correct, by obtaining below result:</p>\n<pre><code>score(y, y_pred) + score(y, 1 - y_pred) = 0.71 + 0.28 = 0.99\n</code></pre>\n<p>Strictly speaking, the sum of the scores is 0.99 instead of 1. This is probably because the LB scores are truncated to the third decimal place. In that case, the expected value of the sum of scores would be 0.99.</p>",
      "rawMarkdown": "Note: I confirmed assumption 2 seems to be correct, by obtaining below result:\n\n```\nscore(y, y_pred) + score(y, 1 - y_pred) = 0.71 + 0.28 = 0.99\n```\n\nStrictly speaking, the sum of the scores is 0.99 instead of 1. This is probably because the LB scores are truncated to the third decimal place. In that case, the expected value of the sum of scores would be 0.99.",
      "votes": null
    },
    {
      "id": "1774284",
      "postDate": "05/02/2022 03:49:35",
      "content": "<p>In Experiment 2, assuming that the sum of the scores with and without inverting predictions is 0.99, the estimated scores are as follows:</p>\n<pre><code>s_top10 = 0.726\ns_low11 = 0.695\n</code></pre>",
      "rawMarkdown": "In Experiment 2, assuming that the sum of the scores with and without inverting predictions is 0.99, the estimated scores are as follows:\n\n````\ns_top10 = 0.726\ns_low11 = 0.695\n````",
      "votes": null
    },
    {
      "id": "1774366",
      "postDate": "05/02/2022 06:03:34",
      "content": "<p>I have discovered one interesting fact.<br>\nIn the public notebook model (show reference [1] in the original post), the top 5 most frequent bird species in the training data contribute an estimated average of only 0.558 to the score. The top 6 to 10 each contribute the most to the score, with an estimated mean score of 0.894.<br>\nThis result indicates that a large frequency of occurrence in the training data does not necessarily mean a large contribution to the public LB score. The reason is unknown at this time, but it is possible that the distribution of species in the public test data is quite variable from species to species or that some species are excluded from the public test data calculations.</p>\n<pre><code>top5, mid_top5, mid_low5, low6 = \n(['skylar', 'houfin', 'jabwar', 'warwhe1', 'yefcan'],\n ['apapan', 'iiwi', 'omao', 'hawama', 'hawcre'],\n ['barpet', 'akiapo', 'elepai', 'aniani', 'hawgoo'],\n ['ercfra', 'hawpet1', 'puaioh', 'hawhaw', 'crehon', 'maupar'])\n\noriginal score: 0.71\nscore with inverted prediction of top5: 0.68\nscore with inverted prediction of mid_top5: 0.52\nscore with inverted prediction of mid_low5: 0.56\nscore with inverted prediction of top5: 0.65\n\nestimated mean score of top5: 0.558\nestimated mean score of mid_top5: 0.894\nestimated mean score of mid_low5: 0.810\nestimated mean score of low6: 0.600\n</code></pre>\n<p>I use below code to estimate above scores.</p>\n<pre><code>def est_mean_score(m0, m1, n, s=0.99):\n    \"\"\"\n    m0: original score\n    m1: score with predictions inverted in some species\n    n: number of species for which predictions were inverted\n    s: expected value of the sum of the original score and the predicted inverted score\n    \"\"\"\n    return 21 / (2 * n) * (m0 - m1) + s / 2\n</code></pre>",
      "rawMarkdown": "I have discovered one interesting fact.\nIn the public notebook model (show reference [1] in the original post), the top 5 most frequent bird species in the training data contribute an estimated average of only 0.558 to the score. The top 6 to 10 each contribute the most to the score, with an estimated mean score of 0.894.\nThis result indicates that a large frequency of occurrence in the training data does not necessarily mean a large contribution to the public LB score. The reason is unknown at this time, but it is possible that the distribution of species in the public test data is quite variable from species to species or that some species are excluded from the public test data calculations.\n\n```\ntop5, mid_top5, mid_low5, low6 = \n(['skylar', 'houfin', 'jabwar', 'warwhe1', 'yefcan'],\n ['apapan', 'iiwi', 'omao', 'hawama', 'hawcre'],\n ['barpet', 'akiapo', 'elepai', 'aniani', 'hawgoo'],\n ['ercfra', 'hawpet1', 'puaioh', 'hawhaw', 'crehon', 'maupar'])\n\noriginal score: 0.71\nscore with inverted prediction of top5: 0.68\nscore with inverted prediction of mid_top5: 0.52\nscore with inverted prediction of mid_low5: 0.56\nscore with inverted prediction of top5: 0.65\n\nestimated mean score of top5: 0.558\nestimated mean score of mid_top5: 0.894\nestimated mean score of mid_low5: 0.810\nestimated mean score of low6: 0.600\n```\n\nI use below code to estimate above scores.\n```python\ndef est_mean_score(m0, m1, n, s=0.99):\n    \"\"\"\n    m0: original score\n    m1: score with predictions inverted in some species\n    n: number of species for which predictions were inverted\n    s: expected value of the sum of the original score and the predicted inverted score\n    \"\"\"\n    return 21 / (2 * n) * (m0 - m1) + s / 2\n```",
      "votes": null
    },
    {
      "id": "1774394",
      "postDate": "05/02/2022 06:24:53",
      "content": "<p>It is surprising that the species in the top 11-15 sample size had the largest contribution to the score, even though they were only given less than 20 samples for training.</p>\n<p><a href=\"https://ibb.co/JqjhpP9\"><img src=\"https://i.ibb.co/1MGSXxk/Screen-Shot-2022-05-02-at-15-20-00.png\" alt=\"Screen-Shot-2022-05-02-at-15-20-00\"></a></p>\n<p>(source: <a href=\"https://www.kaggle.com/code/tatamikenn/birdclef22-eda-on-the-scored-species\" target=\"_blank\">https://www.kaggle.com/code/tatamikenn/birdclef22-eda-on-the-scored-species</a> )</p>",
      "rawMarkdown": "It is surprising that the species in the top 11-15 sample size had the largest contribution to the score, even though they were only given less than 20 samples for training.\n\n<a href=\"https://ibb.co/JqjhpP9\"><img src=\"https://i.ibb.co/1MGSXxk/Screen-Shot-2022-05-02-at-15-20-00.png\" alt=\"Screen-Shot-2022-05-02-at-15-20-00\" border=\"0\"></a>\n\n(source: https://www.kaggle.com/code/tatamikenn/birdclef22-eda-on-the-scored-species )",
      "votes": null
    },
    {
      "id": "1775366",
      "postDate": "05/03/2022 00:32:20",
      "content": "<blockquote>\n  <p>This result indicates that a large frequency of occurrence in the training data does not necessarily mean a large contribution to the public LB score.</p>\n</blockquote>\n<p>This may be due to the difference in the distribution of labels between the training dataset and the test soundscape.<br>\nEven if the frequency of occurrence of a label in the training dataset is high, it does not necessarily mean that the frequency in the test soundscape is high. What may be happening here is that the top five most frequent labels in the training dataset have a much lower frequency of occurrence in the test soundscape.<br>\nFor the labels that appear more frequently in training, the model is likely to output more positive predictions than for the other species. In contrast, in the test soundscape, the model generates more FPs when the frequency of occurrence is lower than expected. This would result in a smaller TNR and thus a smaller contribution to the score for these species.</p>",
      "rawMarkdown": "> This result indicates that a large frequency of occurrence in the training data does not necessarily mean a large contribution to the public LB score.\n\nThis may be due to the difference in the distribution of labels between the training dataset and the test soundscape.\nEven if the frequency of occurrence of a label in the training dataset is high, it does not necessarily mean that the frequency in the test soundscape is high. What may be happening here is that the top five most frequent labels in the training dataset have a much lower frequency of occurrence in the test soundscape.\nFor the labels that appear more frequently in training, the model is likely to output more positive predictions than for the other species. In contrast, in the test soundscape, the model generates more FPs when the frequency of occurrence is lower than expected. This would result in a smaller TNR and thus a smaller contribution to the score for these species.",
      "votes": null
    },
    {
      "id": "1775368",
      "postDate": "05/03/2022 00:41:16",
      "content": "<p>As evidence to support the above reasoning, I cite the following comment. According to this, 10 of the 21 species were initially targeted by the host, and the other 11 were accidentally recorded on the collection of target calls. Therefore, it is possible that samples of non-target calls in the test soundscape are less frequent than those of the target calls.</p>\n<blockquote>\n  <p>If I recall correctly, the other 11 were recorded incidentally during attempts to get audio for the 10 that are of primary interest.<br>\n  <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/307752#1730837\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/307752#1730837</a></p>\n</blockquote>\n<p>Also, to begin with, most of the sound sources of non-Hawaiian endemic species downloaded from xeno-canto were recorded outside of Hawaii, and it is natural to imagine that the frequency of labels in the training data set is considerably higher than the frequency of labels actually recorded in Hawaii .</p>",
      "rawMarkdown": "As evidence to support the above reasoning, I cite the following comment. According to this, 10 of the 21 species were initially targeted by the host, and the other 11 were accidentally recorded on the collection of target calls. Therefore, it is possible that samples of non-target calls in the test soundscape are less frequent than those of the target calls.\n\n> If I recall correctly, the other 11 were recorded incidentally during attempts to get audio for the 10 that are of primary interest.\nhttps://www.kaggle.com/competitions/birdclef-2022/discussion/307752#1730837\n\nAlso, to begin with, most of the sound sources of non-Hawaiian endemic species downloaded from xeno-canto were recorded outside of Hawaii, and it is natural to imagine that the frequency of labels in the training data set is considerably higher than the frequency of labels actually recorded in Hawaii .",
      "votes": null
    },
    {
      "id": "1775373",
      "postDate": "05/03/2022 00:47:50",
      "content": "<p>The implications derived from the above are as follows. The frequency of species labels in the training data could be adjusted to be closer to that in the test soundscape, thereby increasing the score contribution of species with low score contribution at this point in time.</p>",
      "rawMarkdown": "The implications derived from the above are as follows. The frequency of species labels in the training data could be adjusted to be closer to that in the test soundscape, thereby increasing the score contribution of species with low score contribution at this point in time.",
      "votes": null
    },
    {
      "id": "1776347",
      "postDate": "05/03/2022 23:49:51",
      "content": "<p>You've been publishing very useful information, thank you! 😄</p>\n<p>Do you think the most likely scenario is that the most frequent species regain their contribution to the scores in the private set?</p>",
      "rawMarkdown": "You've been publishing very useful information, thank you! 😄\n\nDo you think the most likely scenario is that the most frequent species regain their contribution to the scores in the private set?",
      "votes": null
    },
    {
      "id": "1776362",
      "postDate": "05/04/2022 00:11:06",
      "content": "<p><a href=\"https://www.kaggle.com/flrotm\" target=\"_blank\">@flrotm</a> While we cannot say for sure about the distribution of the private test set, the conclusion of this thread is that we need to be aware of the difference in species distribution between the training data set and the public (private) data set.</p>",
      "rawMarkdown": "flrotm While we cannot say for sure about the distribution of the private test set, the conclusion of this thread is that we need to be aware of the difference in species distribution between the training data set and the public (private) data set.",
      "votes": null
    },
    {
      "id": "1778251",
      "postDate": "05/05/2022 07:33:51",
      "content": "<h1>Theoretical error in scores</h1>\n<p>In general, let U be a subset of the species, and the average score s_U of U can be calculated as follows.</p>\n<p>$$<br>\ns_U = \\frac{21}{2|U|}(m_0 - m_1) + 0.99 \\tag{7}<br>\n$$</p>\n<p>In this case, since m_0 and m_1 have significant digits to the third decimal place, the theoretical upper bound of error for s_U can be calculated as follows.</p>\n<p>$$<br>\n\\text{error}(s_U) = \\frac{21}{2|U|}0.01 \\tag{8}<br>\n$$</p>\n<p>Since|S|=5 for top5, mid_top5 and mid_low_5, the error of the score is as follows.</p>\n<p>$$<br>\n\\text{error}(s_{\\text{top5}}) = 0.021 \\tag{9}<br>\n$$</p>\n<p>Also, since|S|=6 for low6, we obtain the following result.</p>\n<p>$$<br>\n\\text{error}(s_{\\text{low6}}) = 0.0175 \\tag{10}<br>\n$$</p>\n<p>Furthermore, for |S|=1, we get the following result.</p>\n<p>$$<br>\n\\text{error}(s_i) = 0.105 \\tag{11}<br>\n$$</p>",
      "rawMarkdown": "# Theoretical error in scores\n\nIn general, let U be a subset of the species, and the average score s_U of U can be calculated as follows.\n\n$$\ns_U = \\frac{21}{2\\|U\\|}(m_0 - m_1) + 0.99 \\tag{7}\n$$\n\nIn this case, since m_0 and m_1 have significant digits to the third decimal place, the theoretical upper bound of error for s_U can be calculated as follows.\n\n$$\n\\text{error}(s_U) = \\frac{21}{2\\|U\\|}0.01 \\tag{8}\n$$\n\nSince|S|=5 for top5, mid_top5 and mid_low_5, the error of the score is as follows.\n\n$$\n\\text{error}(s_{\\text{top5}}) = 0.021 \\tag{9}\n$$\n\nAlso, since|S|=6 for low6, we obtain the following result.\n\n$$\n\\text{error}(s_{\\text{low6}}) = 0.0175 \\tag{10}\n$$\n\nFurthermore, for |S|=1, we get the following result.\n\n$$\n\\text{error}(s_i) = 0.105 \\tag{11}\n$$",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1774281,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "05/02/2022 03:46:14",
      "content": "<p>Note: I confirmed assumption 2 seems to be correct, by obtaining below result:</p>\n<pre><code>score(y, y_pred) + score(y, 1 - y_pred) = 0.71 + 0.28 = 0.99\n</code></pre>\n<p>Strictly speaking, the sum of the scores is 0.99 instead of 1. This is probably because the LB scores are truncated to the third decimal place. In that case, the expected value of the sum of scores would be 0.99.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1774284,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "05/02/2022 03:49:35",
          "content": "<p>In Experiment 2, assuming that the sum of the scores with and without inverting predictions is 0.99, the estimated scores are as follows:</p>\n<pre><code>s_top10 = 0.726\ns_low11 = 0.695\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1774366,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "05/02/2022 06:03:34",
      "content": "<p>I have discovered one interesting fact.<br>\nIn the public notebook model (show reference [1] in the original post), the top 5 most frequent bird species in the training data contribute an estimated average of only 0.558 to the score. The top 6 to 10 each contribute the most to the score, with an estimated mean score of 0.894.<br>\nThis result indicates that a large frequency of occurrence in the training data does not necessarily mean a large contribution to the public LB score. The reason is unknown at this time, but it is possible that the distribution of species in the public test data is quite variable from species to species or that some species are excluded from the public test data calculations.</p>\n<pre><code>top5, mid_top5, mid_low5, low6 = \n(['skylar', 'houfin', 'jabwar', 'warwhe1', 'yefcan'],\n ['apapan', 'iiwi', 'omao', 'hawama', 'hawcre'],\n ['barpet', 'akiapo', 'elepai', 'aniani', 'hawgoo'],\n ['ercfra', 'hawpet1', 'puaioh', 'hawhaw', 'crehon', 'maupar'])\n\noriginal score: 0.71\nscore with inverted prediction of top5: 0.68\nscore with inverted prediction of mid_top5: 0.52\nscore with inverted prediction of mid_low5: 0.56\nscore with inverted prediction of top5: 0.65\n\nestimated mean score of top5: 0.558\nestimated mean score of mid_top5: 0.894\nestimated mean score of mid_low5: 0.810\nestimated mean score of low6: 0.600\n</code></pre>\n<p>I use below code to estimate above scores.</p>\n<pre><code>def est_mean_score(m0, m1, n, s=0.99):\n    \"\"\"\n    m0: original score\n    m1: score with predictions inverted in some species\n    n: number of species for which predictions were inverted\n    s: expected value of the sum of the original score and the predicted inverted score\n    \"\"\"\n    return 21 / (2 * n) * (m0 - m1) + s / 2\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1774394,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "05/02/2022 06:24:53",
          "content": "<p>It is surprising that the species in the top 11-15 sample size had the largest contribution to the score, even though they were only given less than 20 samples for training.</p>\n<p><a href=\"https://ibb.co/JqjhpP9\"><img src=\"https://i.ibb.co/1MGSXxk/Screen-Shot-2022-05-02-at-15-20-00.png\" alt=\"Screen-Shot-2022-05-02-at-15-20-00\"></a></p>\n<p>(source: <a href=\"https://www.kaggle.com/code/tatamikenn/birdclef22-eda-on-the-scored-species\" target=\"_blank\">https://www.kaggle.com/code/tatamikenn/birdclef22-eda-on-the-scored-species</a> )</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1775366,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "05/03/2022 00:32:20",
          "content": "<blockquote>\n  <p>This result indicates that a large frequency of occurrence in the training data does not necessarily mean a large contribution to the public LB score.</p>\n</blockquote>\n<p>This may be due to the difference in the distribution of labels between the training dataset and the test soundscape.<br>\nEven if the frequency of occurrence of a label in the training dataset is high, it does not necessarily mean that the frequency in the test soundscape is high. What may be happening here is that the top five most frequent labels in the training dataset have a much lower frequency of occurrence in the test soundscape.<br>\nFor the labels that appear more frequently in training, the model is likely to output more positive predictions than for the other species. In contrast, in the test soundscape, the model generates more FPs when the frequency of occurrence is lower than expected. This would result in a smaller TNR and thus a smaller contribution to the score for these species.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1775368,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "05/03/2022 00:41:16",
          "content": "<p>As evidence to support the above reasoning, I cite the following comment. According to this, 10 of the 21 species were initially targeted by the host, and the other 11 were accidentally recorded on the collection of target calls. Therefore, it is possible that samples of non-target calls in the test soundscape are less frequent than those of the target calls.</p>\n<blockquote>\n  <p>If I recall correctly, the other 11 were recorded incidentally during attempts to get audio for the 10 that are of primary interest.<br>\n  <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/307752#1730837\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/307752#1730837</a></p>\n</blockquote>\n<p>Also, to begin with, most of the sound sources of non-Hawaiian endemic species downloaded from xeno-canto were recorded outside of Hawaii, and it is natural to imagine that the frequency of labels in the training data set is considerably higher than the frequency of labels actually recorded in Hawaii .</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1775373,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "05/03/2022 00:47:50",
          "content": "<p>The implications derived from the above are as follows. The frequency of species labels in the training data could be adjusted to be closer to that in the test soundscape, thereby increasing the score contribution of species with low score contribution at this point in time.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1776347,
          "author_name": "flrotm",
          "author_url": "",
          "post_date": "05/03/2022 23:49:51",
          "content": "<p>You've been publishing very useful information, thank you! 😄</p>\n<p>Do you think the most likely scenario is that the most frequent species regain their contribution to the scores in the private set?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1776362,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "05/04/2022 00:11:06",
          "content": "<p><a href=\"https://www.kaggle.com/flrotm\" target=\"_blank\">@flrotm</a> While we cannot say for sure about the distribution of the private test set, the conclusion of this thread is that we need to be aware of the difference in species distribution between the training data set and the public (private) data set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1778251,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "05/05/2022 07:33:51",
      "content": "<h1>Theoretical error in scores</h1>\n<p>In general, let U be a subset of the species, and the average score s_U of U can be calculated as follows.</p>\n<p>$$<br>\ns_U = \\frac{21}{2|U|}(m_0 - m_1) + 0.99 \\tag{7}<br>\n$$</p>\n<p>In this case, since m_0 and m_1 have significant digits to the third decimal place, the theoretical upper bound of error for s_U can be calculated as follows.</p>\n<p>$$<br>\n\\text{error}(s_U) = \\frac{21}{2|U|}0.01 \\tag{8}<br>\n$$</p>\n<p>Since|S|=5 for top5, mid_top5 and mid_low_5, the error of the score is as follows.</p>\n<p>$$<br>\n\\text{error}(s_{\\text{top5}}) = 0.021 \\tag{9}<br>\n$$</p>\n<p>Also, since|S|=6 for low6, we obtain the following result.</p>\n<p>$$<br>\n\\text{error}(s_{\\text{low6}}) = 0.0175 \\tag{10}<br>\n$$</p>\n<p>Furthermore, for |S|=1, we get the following result.</p>\n<p>$$<br>\n\\text{error}(s_i) = 0.105 \\tag{11}<br>\n$$</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1774279": "I will share on how to estimate the contribution of each bird species to the LB score.\n\n## Calculation: how to estimate the score of each bird species\n\nHere I make the following two assumptions.\n\n1. the LB score is calculated by the average of the scores of the individual bird species (macro-average score)\n2. if the predictions are reversed, the LB score is 1 minus the original score. That is, `score(y, 1-y_pred) = 1 - score(y, y_pred)`.\n\n(We can be sure about 2 by submitting all predictions inverted for any given model.)\n\nUnder the above assumptions, the scores for individual bird species are s_1, . , s_21, the LB score is calculated as follows.\n\n$$\nm_0 = \\frac{1}{21} (s_0 + ... + s_i + ... + s_{20}) \\tag{1}\n$$\n\nNow, if we invert the predictions for only one bird species (let's say i), the score would be\n\n$$\nm_1 = \\frac{1}{21} \\left(s_0 + ... + (1 - s_i) + ... + s_{20} \\right) \\tag{2}\n$$\n\nSolving equations (1) and (2) for s_i yields\n\n$$\ns_i = \\frac{21}{2} (m_0 - m_1) + 0.5 \\tag{3}\n$$\n\nThat is, we can estimate the score of bird i by substituting the difference between the score m_0 for a normal submission and the predicted inverted score m_1 for a certain bird species into equation (3).\n\n## Experiment 1: score estimation for \"skylar\"\n\nI tried the above experiment on the most sampled bird species, \"skylar\", and found no difference in scores between m_0 and m_1. I assume this is due to either the score of \"skylar\" alone being less than 0.6, or the \"skylar\" not being used in the publicLB score calculation.\n\n## Experiment 2: Estimation of the average of the scores of the top 10 birds\n\nIn Experiment 1, we were unable to detect score differences due to prediction reversal due to the small number of bird species, but increasing the number of bird species makes it easier to detect score differences due to prediction reversal.\n\nTherefore, I decided to invert predictions for the top 10 bird species with the largest sample size and estimate the mean of the scores for these bird species.\n\nLetting s_top10 be the average of the scores of top 10 species and s_low11 be the average of the scores of the remaining 11 bird species, the LB score is calculated as follows.\n\n$$\n21 m_0 = 10 s_\\text{top10} + 11 s_\\text{low11} \\tag{4}\n$$\n\nThe inverted predicted scores for the top 10 bird species can also be calculated as follows\n\n$$\n21 m_1 = 10 (1 - s_\\text{top10}) + 11 s_\\text{low11} \\tag{5}\n$$\n\nSolving equations (4) and (5) for s_top10 yields\n\n$$\ns_\\text{top10} = \\frac{21}{20} (m_0 - m_1) + 0.5 \\tag{6}\n$$\n\nI calculated m_0 and m_1 using the public Notebook model[1] and found them to be 0.71 and 0.49, respectively. Therefore, s_top10 can be calculated to be **0.731**.\n\nFor the bottom 11 species, we can follow in the similar process, we estimated s_low11 to be **0.700**.\n\nThese results indicate that the top 10 bird species with the largest sample size contribute more to the LB score.\n\nAlthough the low precision of the LB score does not allow us to predict exact values, these methods give us a rough idea of the degree to which your model's predictions are biased toward a particular species for public test data.\n\nFinally, I will share the source code I used for my experiments.\n\nHappy Kaggle.\n\n## Source code of score inversion\n\n```python\ndef invert_target(input_df, birds_to_invert):\n    assert type(birds_to_invert) == list, type(birds)\n    tmp_df = input_df.copy()\n    prefixes, dates, birds, end_secs = zip(*tmp_df.row_id.str.split(\"_\"))\n    tmp_df[\"prefix\"] = prefixes\n    tmp_df[\"date\"] = dates\n    tmp_df[\"bird\"] = birds\n    tmp_df[\"end_sec\"] = end_secs\n    \n    idxs = tmp_df.query(\"bird in @birds_to_invert\").index\n    tmp_df.loc[idxs, \"target\"] = tmp_df.loc[idxs, \"target\"].apply(lambda x: not(x))\n    \n    return tmp_df[[\"row_id\", \"target\"]]\n```\n\n```python\nsample_submission = ... # your model's prediction\n\ntrain = pd.read_csv(\"../input/birdclef-2022/train_metadata.csv\")\nscored = train.query(\"primary_label in @scored_birds\")\nscored_count = scored[\"primary_label\"].value_counts()\ntop10 = scored_count[:10].index.tolist()\nlow11 = scored_count[10:].index.tolist()\ntmp_df = invert_target(sample_submission, low11)\nsample_submission[\"target\"] = tmp_df[\"target\"]\n```\n\n## Reference\n\n- [1] https://www.kaggle.com/code/kaerunantoka/birdclef2022-ex005-f0-infer",
    "1774281": "Note: I confirmed assumption 2 seems to be correct, by obtaining below result:\n\n```\nscore(y, y_pred) + score(y, 1 - y_pred) = 0.71 + 0.28 = 0.99\n```\n\nStrictly speaking, the sum of the scores is 0.99 instead of 1. This is probably because the LB scores are truncated to the third decimal place. In that case, the expected value of the sum of scores would be 0.99.",
    "1774284": "In Experiment 2, assuming that the sum of the scores with and without inverting predictions is 0.99, the estimated scores are as follows:\n\n````\ns_top10 = 0.726\ns_low11 = 0.695\n````",
    "1774366": "I have discovered one interesting fact.\nIn the public notebook model (show reference [1] in the original post), the top 5 most frequent bird species in the training data contribute an estimated average of only 0.558 to the score. The top 6 to 10 each contribute the most to the score, with an estimated mean score of 0.894.\nThis result indicates that a large frequency of occurrence in the training data does not necessarily mean a large contribution to the public LB score. The reason is unknown at this time, but it is possible that the distribution of species in the public test data is quite variable from species to species or that some species are excluded from the public test data calculations.\n\n```\ntop5, mid_top5, mid_low5, low6 = \n(['skylar', 'houfin', 'jabwar', 'warwhe1', 'yefcan'],\n ['apapan', 'iiwi', 'omao', 'hawama', 'hawcre'],\n ['barpet', 'akiapo', 'elepai', 'aniani', 'hawgoo'],\n ['ercfra', 'hawpet1', 'puaioh', 'hawhaw', 'crehon', 'maupar'])\n\noriginal score: 0.71\nscore with inverted prediction of top5: 0.68\nscore with inverted prediction of mid_top5: 0.52\nscore with inverted prediction of mid_low5: 0.56\nscore with inverted prediction of top5: 0.65\n\nestimated mean score of top5: 0.558\nestimated mean score of mid_top5: 0.894\nestimated mean score of mid_low5: 0.810\nestimated mean score of low6: 0.600\n```\n\nI use below code to estimate above scores.\n```python\ndef est_mean_score(m0, m1, n, s=0.99):\n    \"\"\"\n    m0: original score\n    m1: score with predictions inverted in some species\n    n: number of species for which predictions were inverted\n    s: expected value of the sum of the original score and the predicted inverted score\n    \"\"\"\n    return 21 / (2 * n) * (m0 - m1) + s / 2\n```",
    "1774394": "It is surprising that the species in the top 11-15 sample size had the largest contribution to the score, even though they were only given less than 20 samples for training.\n\n<a href=\"https://ibb.co/JqjhpP9\"><img src=\"https://i.ibb.co/1MGSXxk/Screen-Shot-2022-05-02-at-15-20-00.png\" alt=\"Screen-Shot-2022-05-02-at-15-20-00\" border=\"0\"></a>\n\n(source: https://www.kaggle.com/code/tatamikenn/birdclef22-eda-on-the-scored-species )",
    "1775366": "> This result indicates that a large frequency of occurrence in the training data does not necessarily mean a large contribution to the public LB score.\n\nThis may be due to the difference in the distribution of labels between the training dataset and the test soundscape.\nEven if the frequency of occurrence of a label in the training dataset is high, it does not necessarily mean that the frequency in the test soundscape is high. What may be happening here is that the top five most frequent labels in the training dataset have a much lower frequency of occurrence in the test soundscape.\nFor the labels that appear more frequently in training, the model is likely to output more positive predictions than for the other species. In contrast, in the test soundscape, the model generates more FPs when the frequency of occurrence is lower than expected. This would result in a smaller TNR and thus a smaller contribution to the score for these species.",
    "1775368": "As evidence to support the above reasoning, I cite the following comment. According to this, 10 of the 21 species were initially targeted by the host, and the other 11 were accidentally recorded on the collection of target calls. Therefore, it is possible that samples of non-target calls in the test soundscape are less frequent than those of the target calls.\n\n> If I recall correctly, the other 11 were recorded incidentally during attempts to get audio for the 10 that are of primary interest.\nhttps://www.kaggle.com/competitions/birdclef-2022/discussion/307752#1730837\n\nAlso, to begin with, most of the sound sources of non-Hawaiian endemic species downloaded from xeno-canto were recorded outside of Hawaii, and it is natural to imagine that the frequency of labels in the training data set is considerably higher than the frequency of labels actually recorded in Hawaii .",
    "1775373": "The implications derived from the above are as follows. The frequency of species labels in the training data could be adjusted to be closer to that in the test soundscape, thereby increasing the score contribution of species with low score contribution at this point in time.",
    "1776347": "You've been publishing very useful information, thank you! 😄\n\nDo you think the most likely scenario is that the most frequent species regain their contribution to the scores in the private set?",
    "1776362": "flrotm While we cannot say for sure about the distribution of the private test set, the conclusion of this thread is that we need to be aware of the difference in species distribution between the training data set and the public (private) data set.",
    "1778251": "# Theoretical error in scores\n\nIn general, let U be a subset of the species, and the average score s_U of U can be calculated as follows.\n\n$$\ns_U = \\frac{21}{2\\|U\\|}(m_0 - m_1) + 0.99 \\tag{7}\n$$\n\nIn this case, since m_0 and m_1 have significant digits to the third decimal place, the theoretical upper bound of error for s_U can be calculated as follows.\n\n$$\n\\text{error}(s_U) = \\frac{21}{2\\|U\\|}0.01 \\tag{8}\n$$\n\nSince|S|=5 for top5, mid_top5 and mid_low_5, the error of the score is as follows.\n\n$$\n\\text{error}(s_{\\text{top5}}) = 0.021 \\tag{9}\n$$\n\nAlso, since|S|=6 for low6, we obtain the following result.\n\n$$\n\\text{error}(s_{\\text{low6}}) = 0.0175 \\tag{10}\n$$\n\nFurthermore, for |S|=1, we get the following result.\n\n$$\n\\text{error}(s_i) = 0.105 \\tag{11}\n$$"
  },
  "source": "meta"
}