{
  "id": 321883,
  "title": "Summary: difficulties when discussing CV/LB correlation",
  "url": "/competitions/birdclef-2022/discussion/321883",
  "author_name": "Bilzard",
  "post_date": "2022-04-29T06:42:24.606000",
  "votes": 17,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Let me summarize the difficulties when discussing the correlation between CV and LB in this competition.</p>\n<p>I think we should be quite cautious when discussing the correlation between CV and LB from several perspectives.</p>\n<h2>1. The evaluation metrics is vague</h2>\n<p>First of all, it should be noted that the evaluation metrics are vague for this competition. There is little point in discussing the correlation between CV and LB when the evaluation metrics are not settled.</p>\n<p>There have been several threads[1-2] in the past two months about evaluation metrics, but we have not received clear answers from the hosts. At this point, the only two pieces of information we have from the host are these[3]:</p>\n<ol>\n<li>the effects of negative and positive cases are adjusted to be equal</li>\n<li>the scores for each bird species are adjusted to be equal</li>\n</ol>\n<p>We also know from a brief submission experiment[4] that</p>\n<ol>\n<li>the scores for predicting all positive examples, predicting all negative examples, and predicting random negative and positive examples are each in the neighborhood of 0.5.</li>\n</ol>\n<p>Based on the above, I came up with is <a href=\"https://scikit-learn.org/stable/modules/model_evaluation.html#balanced-accuracy-score\" target=\"_blank\">balanced accuracy score</a>[5].</p>\n<h2>2. There are no labeled train soundscapes</h2>\n<p>Second, it should also be noted that this competition does not give a labeled soundscape for local validation. We are nearly clueless about the domain of soundscapes for evaluation (with the exception of one downloadable sample).</p>\n<h2>3. Public LB samples are few</h2>\n<p>Third, the data from the public leaderboards is small (16% of the total). This fact makes the use of public leaderboard scores as a substitute for local validation also risky.</p>\n<h2>4. Idea for local validation data</h2>\n<p>In light of the above, the ideas I have for the local validation at this point are as follows:</p>\n<ol>\n<li>hypothesize a CV evaluation score that fits the requirements</li>\n<li>prepare sufficiently reliable evaluation data (e.g., diverting evaluation data from the 2021 BirdCLEF data or other external data).</li>\n<li>confirm by submitting that the hypothesized CVs generally correlate with LBs (LB probe if necessary)</li>\n<li>if the hypotheses in the CV data differ substantially from LB, reestablish the hypotheses.</li>\n<li>repeat 1-4 until CV data are sufficiently reliable</li>\n</ol>\n<h2>Reference</h2>\n<ul>\n<li>[1] <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/311493\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/311493</a></li>\n<li>[2] <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/314999\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/314999</a></li>\n<li>[3] <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/311493#1716290\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/311493#1716290</a></li>\n<li>[4] <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/314999#1735156\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/314999#1735156</a></li>\n<li>[5] <a href=\"https://scikit-learn.org/stable/modules/model_evaluation.html#balanced-accuracy-score\" target=\"_blank\">https://scikit-learn.org/stable/modules/model_evaluation.html#balanced-accuracy-score</a></li>\n</ul>",
  "messages": [
    {
      "id": 1771367,
      "postDate": "2022-04-29T06:42:24.607Z",
      "content": "<p>Let me summarize the difficulties when discussing the correlation between CV and LB in this competition.</p>\n<p>I think we should be quite cautious when discussing the correlation between CV and LB from several perspectives.</p>\n<h2>1. The evaluation metrics is vague</h2>\n<p>First of all, it should be noted that the evaluation metrics are vague for this competition. There is little point in discussing the correlation between CV and LB when the evaluation metrics are not settled.</p>\n<p>There have been several threads[1-2] in the past two months about evaluation metrics, but we have not received clear answers from the hosts. At this point, the only two pieces of information we have from the host are these[3]:</p>\n<ol>\n<li>the effects of negative and positive cases are adjusted to be equal</li>\n<li>the scores for each bird species are adjusted to be equal</li>\n</ol>\n<p>We also know from a brief submission experiment[4] that</p>\n<ol>\n<li>the scores for predicting all positive examples, predicting all negative examples, and predicting random negative and positive examples are each in the neighborhood of 0.5.</li>\n</ol>\n<p>Based on the above, I came up with is <a href=\"https://scikit-learn.org/stable/modules/model_evaluation.html#balanced-accuracy-score\" target=\"_blank\">balanced accuracy score</a>[5].</p>\n<h2>2. There are no labeled train soundscapes</h2>\n<p>Second, it should also be noted that this competition does not give a labeled soundscape for local validation. We are nearly clueless about the domain of soundscapes for evaluation (with the exception of one downloadable sample).</p>\n<h2>3. Public LB samples are few</h2>\n<p>Third, the data from the public leaderboards is small (16% of the total). This fact makes the use of public leaderboard scores as a substitute for local validation also risky.</p>\n<h2>4. Idea for local validation data</h2>\n<p>In light of the above, the ideas I have for the local validation at this point are as follows:</p>\n<ol>\n<li>hypothesize a CV evaluation score that fits the requirements</li>\n<li>prepare sufficiently reliable evaluation data (e.g., diverting evaluation data from the 2021 BirdCLEF data or other external data).</li>\n<li>confirm by submitting that the hypothesized CVs generally correlate with LBs (LB probe if necessary)</li>\n<li>if the hypotheses in the CV data differ substantially from LB, reestablish the hypotheses.</li>\n<li>repeat 1-4 until CV data are sufficiently reliable</li>\n</ol>\n<h2>Reference</h2>\n<ul>\n<li>[1] <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/311493\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/311493</a></li>\n<li>[2] <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/314999\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/314999</a></li>\n<li>[3] <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/311493#1716290\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/311493#1716290</a></li>\n<li>[4] <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/314999#1735156\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/314999#1735156</a></li>\n<li>[5] <a href=\"https://scikit-learn.org/stable/modules/model_evaluation.html#balanced-accuracy-score\" target=\"_blank\">https://scikit-learn.org/stable/modules/model_evaluation.html#balanced-accuracy-score</a></li>\n</ul>",
      "rawMarkdown": "Let me summarize the difficulties when discussing the correlation between CV and LB in this competition.\n\nI think we should be quite cautious when discussing the correlation between CV and LB from several perspectives.\n\n## 1. The evaluation metrics is vague\n\nFirst of all, it should be noted that the evaluation metrics are vague for this competition. There is little point in discussing the correlation between CV and LB when the evaluation metrics are not settled.\n\nThere have been several threads[1-2] in the past two months about evaluation metrics, but we have not received clear answers from the hosts. At this point, the only two pieces of information we have from the host are these[3]:\n\n1. the effects of negative and positive cases are adjusted to be equal\n2. the scores for each bird species are adjusted to be equal\n\nWe also know from a brief submission experiment[4] that\n\n3. the scores for predicting all positive examples, predicting all negative examples, and predicting random negative and positive examples are each in the neighborhood of 0.5.\n\nBased on the above, I came up with is [balanced accuracy score][5].\n\n## 2. There are no labeled train soundscapes\n\nSecond, it should also be noted that this competition does not give a labeled soundscape for local validation. We are nearly clueless about the domain of soundscapes for evaluation (with the exception of one downloadable sample).\n\n## 3. Public LB samples are few\n\nThird, the data from the public leaderboards is small (16% of the total). This fact makes the use of public leaderboard scores as a substitute for local validation also risky.\n\n## 4. Idea for local validation data\n\nIn light of the above, the ideas I have for the local validation at this point are as follows:\n\n1. hypothesize a CV evaluation score that fits the requirements\n2. prepare sufficiently reliable evaluation data (e.g., diverting evaluation data from the 2021 BirdCLEF data or other external data).\n3. confirm by submitting that the hypothesized CVs generally correlate with LBs (LB probe if necessary)\n4. if the hypotheses in the CV data differ substantially from LB, reestablish the hypotheses.\n5. repeat 1-4 until CV data are sufficiently reliable\n\n## Reference\n\n- [1] https://www.kaggle.com/competitions/birdclef-2022/discussion/311493\n- [2] https://www.kaggle.com/competitions/birdclef-2022/discussion/314999\n- [3] https://www.kaggle.com/competitions/birdclef-2022/discussion/311493#1716290\n- [4] https://www.kaggle.com/competitions/birdclef-2022/discussion/314999#1735156\n- [5] https://scikit-learn.org/stable/modules/model_evaluation.html#balanced-accuracy-score\n\n[balanced accuracy score]: https://scikit-learn.org/stable/modules/model_evaluation.html#balanced-accuracy-score\n",
      "votes": 17
    },
    {
      "id": 1771378,
      "postDate": "2022-04-29T06:53:16.417Z",
      "content": "<p><a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a> <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>\n<p>Just to be sure, let me confirm. It seems to me that this competition avoids naming the details of the evaluation metrics. Would it be inconvenient for us to know the algorithmic details of the evaluation metrics?</p>",
      "rawMarkdown": "@tomdenton @stefankahl @sohier \n\nJust to be sure, let me confirm. It seems to me that this competition avoids naming the details of the evaluation metrics. Would it be inconvenient for us to know the algorithmic details of the evaluation metrics?",
      "votes": 3,
      "replies": [
        {
          "id": 1771614,
          "postDate": "2022-04-29T11:59:58.797Z",
          "content": "<p>+1 <br>\nI agree, the details of evaluation metric should be clarified from the hosts for transparency. I really don't get why to keep it secret  </p>",
          "rawMarkdown": "+1 \nI agree, the details of evaluation metric should be clarified from the hosts for transparency. I really don't get why to keep it secret  ",
          "votes": 1
        },
        {
          "id": 1773612,
          "postDate": "2022-05-01T10:23:21.507Z",
          "content": "<p><a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a> <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a></p>\n<p>Let me ask one more question.<br>\nThe following is a quote from the competition evaluation metrics description, but am I correct in understanding that \"removed unscored rows\" means that every test soundscape contains at least one scored species (i.e., soundscapes that contain only species other than 21 species or \"nocall\" are excluded)?</p>\n<blockquote>\n  <p><strong>After dropping all of the un-scored rows</strong> we technically run a weighted classification accuracy with the weights set such that all of the species are assigned the same total weight and the true negatives and true positives for each species have the same weight.</p>\n</blockquote>",
          "rawMarkdown": "@tomdenton @stefankahl @sohier\n\nLet me ask one more question.\nThe following is a quote from the competition evaluation metrics description, but am I correct in understanding that \"removed unscored rows\" means that every test soundscape contains at least one scored species (i.e., soundscapes that contain only species other than 21 species or \"nocall\" are excluded)?\n\n> **After dropping all of the un-scored rows** we technically run a weighted classification accuracy with the weights set such that all of the species are assigned the same total weight and the true negatives and true positives for each species have the same weight."
        }
      ]
    },
    {
      "id": 1774417,
      "postDate": "2022-05-02T06:53:01.633Z",
      "content": "<p>Tips: About setting the threshold value</p>\n<p>Assuming that the evaluation metric is balanced accuracy, the threshold for the model's prediction should be chosen so that the balanced accuracy, or TPR+TNR, is maximized. If the model's correct answer rate is constant, the TPR and TNR will remain constant even if the number of negative or positive examples is changed. Therefore, the score should be constant regardless of the number of positive and negative examples in the test data.</p>",
      "rawMarkdown": "Tips: About setting the threshold value\n\nAssuming that the evaluation metric is balanced accuracy, the threshold for the model's prediction should be chosen so that the balanced accuracy, or TPR+TNR, is maximized. If the model's correct answer rate is constant, the TPR and TNR will remain constant even if the number of negative or positive examples is changed. Therefore, the score should be constant regardless of the number of positive and negative examples in the test data."
    },
    {
      "id": 1772071,
      "postDate": "2022-04-29T19:35:52.237Z",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a>!</p>\n<p>Just in case it helps, your post made me think about the metric and lead me to something interesting, in particular:</p>\n<ul>\n<li><em>the effects of negative and positive cases are adjusted to be equal</em>. It implicitly remarks that F1 is somehow not symmetric with respect to negative or possitive samples. It turns out that <strong>F1 scores better false possitives than false negatives for a fixed accuracy value</strong>.</li>\n</ul>\n<p>Then, if you calculate the same metric but reversing all values (so taking negative as possitive and viceversa) you end up with a metric that scores better false negatives than false possitives (for a fixed accuracy value). Taking the average of F1 and this reversed version seems a good way to avoid dissimetry here.</p>\n<ul>\n<li><em>the scores for predicting all positive examples, predicting all negative examples, and predicting random negative and positive examples are each in the neighborhood of 0.5</em>. After some nummerical testing this new metric seems to behave like that <strong>when there is heavy imbalance between possitive and negative samples</strong>, as it is our case.</li>\n</ul>\n<p>This proposal with some modifications may be a way to use it in training: <a href=\"https://www.kaggle.com/code/rejpalcz/best-loss-function-for-f1-score-metric/notebook\" target=\"_blank\">https://www.kaggle.com/code/rejpalcz/best-loss-function-for-f1-score-metric/notebook</a></p>\n<p>Best!</p>",
      "rawMarkdown": "Thanks @tatamikenn!\n\nJust in case it helps, your post made me think about the metric and lead me to something interesting, in particular:\n- *the effects of negative and positive cases are adjusted to be equal*. It implicitly remarks that F1 is somehow not symmetric with respect to negative or possitive samples. It turns out that **F1 scores better false possitives than false negatives for a fixed accuracy value**.\n\nThen, if you calculate the same metric but reversing all values (so taking negative as possitive and viceversa) you end up with a metric that scores better false negatives than false possitives (for a fixed accuracy value). Taking the average of F1 and this reversed version seems a good way to avoid dissimetry here.\n\n- *the scores for predicting all positive examples, predicting all negative examples, and predicting random negative and positive examples are each in the neighborhood of 0.5*. After some nummerical testing this new metric seems to behave like that **when there is heavy imbalance between possitive and negative samples**, as it is our case.\n\nThis proposal with some modifications may be a way to use it in training: https://www.kaggle.com/code/rejpalcz/best-loss-function-for-f1-score-metric/notebook\n\nBest!",
      "replies": [
        {
          "id": 1772162,
          "postDate": "2022-04-29T23:05:57.893Z",
          "content": "<p>Interesting suggestion. Certainly that method would also eliminate the class imbalance between positive and negative examples.</p>\n<p>However, we have another observation: <em>the public LB score is higher if the threshold is much smaller</em>[1].<br>\nThis indicates that the evaluation metrics are more tolerant of false positives, but the F1 score (and its inverted positive and negative examples) imposes a severe penalty for false positives (this can be confirmed from a simple numerical calculation).<br>\nThus, as I stated in my first post, I think balanced accuracy is the evaluation metric that best fits the observation at this time.</p>\n<p>Thanks for your comment.</p>\n<h2>Reference</h2>\n<ul>\n<li>[1] <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/318999\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/318999</a></li>\n</ul>",
          "rawMarkdown": "Interesting suggestion. Certainly that method would also eliminate the class imbalance between positive and negative examples.\n\nHowever, we have another observation: *the public LB score is higher if the threshold is much smaller*[1].\nThis indicates that the evaluation metrics are more tolerant of false positives, but the F1 score (and its inverted positive and negative examples) imposes a severe penalty for false positives (this can be confirmed from a simple numerical calculation).\nThus, as I stated in my first post, I think balanced accuracy is the evaluation metric that best fits the observation at this time.\n\nThanks for your comment.\n\n## Reference\n\n- [1] https://www.kaggle.com/competitions/birdclef-2022/discussion/318999",
          "votes": 3
        },
        {
          "id": 1772176,
          "postDate": "2022-04-29T23:46:10.300Z",
          "content": "<p>For reference, I shared the results of the numerical simulation here:<br>\n<a href=\"https://www.kaggle.com/code/tatamikenn/birdclef22-simulation-of-evaluation-metrics/notebook?scriptVersionId=94370692\" target=\"_blank\">https://www.kaggle.com/code/tatamikenn/birdclef22-simulation-of-evaluation-metrics/notebook?scriptVersionId=94370692</a></p>",
          "rawMarkdown": "For reference, I shared the results of the numerical simulation here:\nhttps://www.kaggle.com/code/tatamikenn/birdclef22-simulation-of-evaluation-metrics/notebook?scriptVersionId=94370692",
          "votes": 2
        },
        {
          "id": 1772210,
          "postDate": "2022-04-30T02:01:31.770Z",
          "content": "<p>Note: Another feature of balanced accuracy is that it is difficult to estimate the proportion of positive or negative cases in LB probing. As the simulation results show, changing the percentage of positive cases from 0.1%, 1%, to 10% does not change the mean value of the calculated score (however, the variance increases with a smaller percentage of positive cases). In a sense, this is consistent with the host's intention (if any) to hide information on private test data.</p>",
          "rawMarkdown": "Note: Another feature of balanced accuracy is that it is difficult to estimate the proportion of positive or negative cases in LB probing. As the simulation results show, changing the percentage of positive cases from 0.1%, 1%, to 10% does not change the mean value of the calculated score (however, the variance increases with a smaller percentage of positive cases). In a sense, this is consistent with the host's intention (if any) to hide information on private test data.",
          "votes": 2
        },
        {
          "id": 1772215,
          "postDate": "2022-04-30T02:11:57.647Z",
          "content": "<p>Nice insights <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> . But we can not see the notebook you shared, it is probably private</p>",
          "rawMarkdown": "Nice insights @tatamikenn . But we can not see the notebook you shared, it is probably private",
          "votes": 1
        },
        {
          "id": 1772221,
          "postDate": "2022-04-30T02:34:01.030Z",
          "content": "<p><a href=\"https://www.kaggle.com/hinepo\" target=\"_blank\">@hinepo</a> Thanks for pointing out. Now the notebook is public.</p>",
          "rawMarkdown": "@hinepo Thanks for pointing out. Now the notebook is public.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1773712,
      "postDate": "2022-05-01T11:47:03.787Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1772794,
      "postDate": "2022-04-30T15:24:23.930Z",
      "content": "<p>thanks 😃 </p>",
      "rawMarkdown": "thanks 😃 "
    }
  ],
  "comments": [
    {
      "id": 1771378,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2022-04-29T06:53:16.417000",
      "content": "<p><a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a> <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>\n<p>Just to be sure, let me confirm. It seems to me that this competition avoids naming the details of the evaluation metrics. Would it be inconvenient for us to know the algorithmic details of the evaluation metrics?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1771614,
          "author_name": "Ioannis M",
          "author_url": "",
          "post_date": "2022-04-29T11:59:58.797000",
          "content": "<p>+1 <br>\nI agree, the details of evaluation metric should be clarified from the hosts for transparency. I really don't get why to keep it secret  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1773612,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-05-01T10:23:21.507000",
          "content": "<p><a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a> <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a></p>\n<p>Let me ask one more question.<br>\nThe following is a quote from the competition evaluation metrics description, but am I correct in understanding that \"removed unscored rows\" means that every test soundscape contains at least one scored species (i.e., soundscapes that contain only species other than 21 species or \"nocall\" are excluded)?</p>\n<blockquote>\n  <p><strong>After dropping all of the un-scored rows</strong> we technically run a weighted classification accuracy with the weights set such that all of the species are assigned the same total weight and the true negatives and true positives for each species have the same weight.</p>\n</blockquote>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1774417,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2022-05-02T06:53:01.633000",
      "content": "<p>Tips: About setting the threshold value</p>\n<p>Assuming that the evaluation metric is balanced accuracy, the threshold for the model's prediction should be chosen so that the balanced accuracy, or TPR+TNR, is maximized. If the model's correct answer rate is constant, the TPR and TNR will remain constant even if the number of negative or positive examples is changed. Therefore, the score should be constant regardless of the number of positive and negative examples in the test data.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1772071,
      "author_name": "Enrique Gurdiel",
      "author_url": "",
      "post_date": "2022-04-29T19:35:52.237000",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a>!</p>\n<p>Just in case it helps, your post made me think about the metric and lead me to something interesting, in particular:</p>\n<ul>\n<li><em>the effects of negative and positive cases are adjusted to be equal</em>. It implicitly remarks that F1 is somehow not symmetric with respect to negative or possitive samples. It turns out that <strong>F1 scores better false possitives than false negatives for a fixed accuracy value</strong>.</li>\n</ul>\n<p>Then, if you calculate the same metric but reversing all values (so taking negative as possitive and viceversa) you end up with a metric that scores better false negatives than false possitives (for a fixed accuracy value). Taking the average of F1 and this reversed version seems a good way to avoid dissimetry here.</p>\n<ul>\n<li><em>the scores for predicting all positive examples, predicting all negative examples, and predicting random negative and positive examples are each in the neighborhood of 0.5</em>. After some nummerical testing this new metric seems to behave like that <strong>when there is heavy imbalance between possitive and negative samples</strong>, as it is our case.</li>\n</ul>\n<p>This proposal with some modifications may be a way to use it in training: <a href=\"https://www.kaggle.com/code/rejpalcz/best-loss-function-for-f1-score-metric/notebook\" target=\"_blank\">https://www.kaggle.com/code/rejpalcz/best-loss-function-for-f1-score-metric/notebook</a></p>\n<p>Best!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1772162,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-04-29T23:05:57.893000",
          "content": "<p>Interesting suggestion. Certainly that method would also eliminate the class imbalance between positive and negative examples.</p>\n<p>However, we have another observation: <em>the public LB score is higher if the threshold is much smaller</em>[1].<br>\nThis indicates that the evaluation metrics are more tolerant of false positives, but the F1 score (and its inverted positive and negative examples) imposes a severe penalty for false positives (this can be confirmed from a simple numerical calculation).<br>\nThus, as I stated in my first post, I think balanced accuracy is the evaluation metric that best fits the observation at this time.</p>\n<p>Thanks for your comment.</p>\n<h2>Reference</h2>\n<ul>\n<li>[1] <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/318999\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/318999</a></li>\n</ul>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1772176,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-04-29T23:46:10.300000",
          "content": "<p>For reference, I shared the results of the numerical simulation here:<br>\n<a href=\"https://www.kaggle.com/code/tatamikenn/birdclef22-simulation-of-evaluation-metrics/notebook?scriptVersionId=94370692\" target=\"_blank\">https://www.kaggle.com/code/tatamikenn/birdclef22-simulation-of-evaluation-metrics/notebook?scriptVersionId=94370692</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1772210,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-04-30T02:01:31.770000",
          "content": "<p>Note: Another feature of balanced accuracy is that it is difficult to estimate the proportion of positive or negative cases in LB probing. As the simulation results show, changing the percentage of positive cases from 0.1%, 1%, to 10% does not change the mean value of the calculated score (however, the variance increases with a smaller percentage of positive cases). In a sense, this is consistent with the host's intention (if any) to hide information on private test data.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1772215,
          "author_name": "HinePo",
          "author_url": "",
          "post_date": "2022-04-30T02:11:57.647000",
          "content": "<p>Nice insights <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> . But we can not see the notebook you shared, it is probably private</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1772221,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-04-30T02:34:01.030000",
          "content": "<p><a href=\"https://www.kaggle.com/hinepo\" target=\"_blank\">@hinepo</a> Thanks for pointing out. Now the notebook is public.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1773712,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-01T11:47:03.787000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1772794,
      "author_name": "Alex Parkhomenko",
      "author_url": "",
      "post_date": "2022-04-30T15:24:23.930000",
      "content": "<p>thanks 😃 </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1771367": "Let me summarize the difficulties when discussing the correlation between CV and LB in this competition.\n\nI think we should be quite cautious when discussing the correlation between CV and LB from several perspectives.\n\n## 1. The evaluation metrics is vague\n\nFirst of all, it should be noted that the evaluation metrics are vague for this competition. There is little point in discussing the correlation between CV and LB when the evaluation metrics are not settled.\n\nThere have been several threads[1-2] in the past two months about evaluation metrics, but we have not received clear answers from the hosts. At this point, the only two pieces of information we have from the host are these[3]:\n\n1. the effects of negative and positive cases are adjusted to be equal\n2. the scores for each bird species are adjusted to be equal\n\nWe also know from a brief submission experiment[4] that\n\n3. the scores for predicting all positive examples, predicting all negative examples, and predicting random negative and positive examples are each in the neighborhood of 0.5.\n\nBased on the above, I came up with is [balanced accuracy score][5].\n\n## 2. There are no labeled train soundscapes\n\nSecond, it should also be noted that this competition does not give a labeled soundscape for local validation. We are nearly clueless about the domain of soundscapes for evaluation (with the exception of one downloadable sample).\n\n## 3. Public LB samples are few\n\nThird, the data from the public leaderboards is small (16% of the total). This fact makes the use of public leaderboard scores as a substitute for local validation also risky.\n\n## 4. Idea for local validation data\n\nIn light of the above, the ideas I have for the local validation at this point are as follows:\n\n1. hypothesize a CV evaluation score that fits the requirements\n2. prepare sufficiently reliable evaluation data (e.g., diverting evaluation data from the 2021 BirdCLEF data or other external data).\n3. confirm by submitting that the hypothesized CVs generally correlate with LBs (LB probe if necessary)\n4. if the hypotheses in the CV data differ substantially from LB, reestablish the hypotheses.\n5. repeat 1-4 until CV data are sufficiently reliable\n\n## Reference\n\n- [1] https://www.kaggle.com/competitions/birdclef-2022/discussion/311493\n- [2] https://www.kaggle.com/competitions/birdclef-2022/discussion/314999\n- [3] https://www.kaggle.com/competitions/birdclef-2022/discussion/311493#1716290\n- [4] https://www.kaggle.com/competitions/birdclef-2022/discussion/314999#1735156\n- [5] https://scikit-learn.org/stable/modules/model_evaluation.html#balanced-accuracy-score\n\n[balanced accuracy score]: https://scikit-learn.org/stable/modules/model_evaluation.html#balanced-accuracy-score\n",
    "1771378": "@tomdenton @stefankahl @sohier \n\nJust to be sure, let me confirm. It seems to me that this competition avoids naming the details of the evaluation metrics. Would it be inconvenient for us to know the algorithmic details of the evaluation metrics?",
    "1774417": "Tips: About setting the threshold value\n\nAssuming that the evaluation metric is balanced accuracy, the threshold for the model's prediction should be chosen so that the balanced accuracy, or TPR+TNR, is maximized. If the model's correct answer rate is constant, the TPR and TNR will remain constant even if the number of negative or positive examples is changed. Therefore, the score should be constant regardless of the number of positive and negative examples in the test data.",
    "1772071": "Thanks @tatamikenn!\n\nJust in case it helps, your post made me think about the metric and lead me to something interesting, in particular:\n- *the effects of negative and positive cases are adjusted to be equal*. It implicitly remarks that F1 is somehow not symmetric with respect to negative or possitive samples. It turns out that **F1 scores better false possitives than false negatives for a fixed accuracy value**.\n\nThen, if you calculate the same metric but reversing all values (so taking negative as possitive and viceversa) you end up with a metric that scores better false negatives than false possitives (for a fixed accuracy value). Taking the average of F1 and this reversed version seems a good way to avoid dissimetry here.\n\n- *the scores for predicting all positive examples, predicting all negative examples, and predicting random negative and positive examples are each in the neighborhood of 0.5*. After some nummerical testing this new metric seems to behave like that **when there is heavy imbalance between possitive and negative samples**, as it is our case.\n\nThis proposal with some modifications may be a way to use it in training: https://www.kaggle.com/code/rejpalcz/best-loss-function-for-f1-score-metric/notebook\n\nBest!",
    "1773712": "",
    "1772794": "thanks 😃 "
  }
}