{
  "id": 322413,
  "title": "Q: Why does lowering the threshold not improve the Public Score.",
  "url": "/competitions/birdclef-2022/discussion/322413",
  "author_name": "",
  "post_date": "2022-05-02T03:04:29.776151200Z",
  "votes": 1,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I do not really understand why lowering the threshold lowers the PUBLIC SCORE.(In the first place, I don't understand Evaluation Metric very well.)</p>\n<p>In my opinion it is better to lower the threshold and detect more birds.</p>\n<p>When I applied this idea to my model, the results were as follows.</p>\n<table>\n<thead>\n<tr>\n<th>threshold</th>\n<th>PB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.1</td>\n<td>0.60</td>\n</tr>\n<tr>\n<td>0.05</td>\n<td>0.55</td>\n</tr>\n</tbody>\n</table>\n<p>I am thinking that this is caused by too many <code>fp</code> in <code>tp/(tp+fp)</code> in <strong>macro F1 score</strong>, but is there any other possible reason?</p>\n<p>Thank you very much.</p>\n<p>(This text is translated at DeepL and may be difficult to read.)</p>",
  "messages": [
    {
      "id": "1774266",
      "postDate": "05/02/2022 03:04:29",
      "content": "<p>I do not really understand why lowering the threshold lowers the PUBLIC SCORE.(In the first place, I don't understand Evaluation Metric very well.)</p>\n<p>In my opinion it is better to lower the threshold and detect more birds.</p>\n<p>When I applied this idea to my model, the results were as follows.</p>\n<table>\n<thead>\n<tr>\n<th>threshold</th>\n<th>PB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.1</td>\n<td>0.60</td>\n</tr>\n<tr>\n<td>0.05</td>\n<td>0.55</td>\n</tr>\n</tbody>\n</table>\n<p>I am thinking that this is caused by too many <code>fp</code> in <code>tp/(tp+fp)</code> in <strong>macro F1 score</strong>, but is there any other possible reason?</p>\n<p>Thank you very much.</p>\n<p>(This text is translated at DeepL and may be difficult to read.)</p>",
      "rawMarkdown": "I do not really understand why lowering the threshold lowers the PUBLIC SCORE.(In the first place, I don't understand Evaluation Metric very well.)\n\nIn my opinion it is better to lower the threshold and detect more birds.\n\nWhen I applied this idea to my model, the results were as follows.\n\n| threshold | PB |\n| --- | --- |\n| 0.1 |0.60 |\n| 0.05 |0.55 |\n\nI am thinking that this is caused by too many `fp` in `tp/(tp+fp)` in **macro F1 score**, but is there any other possible reason?\n\nThank you very much.\n\n(This text is translated at DeepL and may be difficult to read.)",
      "votes": null
    },
    {
      "id": "1774468",
      "postDate": "05/02/2022 07:34:38",
      "content": "<p>I think you are right.</p>\n<p>The host said that they are weighting positive and negative examples.<br>\nFalse positives are tolerated to some extent if one expects that the test data will have a higher percentage of negative examples.<br>\nExceeding tolerance will hurt your leaderboard.</p>",
      "rawMarkdown": "I think you are right.\n\nThe host said that they are weighting positive and negative examples.\nFalse positives are tolerated to some extent if one expects that the test data will have a higher percentage of negative examples.\nExceeding tolerance will hurt your leaderboard.",
      "votes": null
    },
    {
      "id": "1774540",
      "postDate": "05/02/2022 09:01:05",
      "content": "<p>In a normal problem setting, it's not the case that \"the lower the prediction threshold, the better\".<br>\nIn the extreme case, a model that predicts everything as a positive example is not making any meaningful predictions.<br>\nWe are not told exactly what evaluation metrics the host employs, but in any case, it is reasonable to assume that models that predict everything as a positive example are set to have a lower score.<br>\nThus, as you say, we are naturally expected to devise ways to reduce FP (or increase TN).</p>",
      "rawMarkdown": "In a normal problem setting, it's not the case that \"the lower the prediction threshold, the better\".\nIn the extreme case, a model that predicts everything as a positive example is not making any meaningful predictions.\nWe are not told exactly what evaluation metrics the host employs, but in any case, it is reasonable to assume that models that predict everything as a positive example are set to have a lower score.\nThus, as you say, we are naturally expected to devise ways to reduce FP (or increase TN).",
      "votes": null
    },
    {
      "id": "1774549",
      "postDate": "05/02/2022 09:10:40",
      "content": "<p>By the way, you can also refer to the discussion below in which I stated what is the most probable evaluation metrics.<br>\n<a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/321883\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/321883</a></p>",
      "rawMarkdown": "By the way, you can also refer to the discussion below in which I stated what is the most probable evaluation metrics.\nhttps://www.kaggle.com/competitions/birdclef-2022/discussion/321883",
      "votes": null
    },
    {
      "id": "1774790",
      "postDate": "05/02/2022 12:47:35",
      "content": "<p>Thanks for the useful information!</p>\n<p>I have one point of concern.</p>\n<p><em>False positives are tolerated to some extent if one expects that the test data will have a higher percentage of negative examples.</em></p>\n<p>Does this sentence mean that there is a set number of false positives allowed for each label?</p>\n<p>For example, \"maupar\" is rarely included in the training data, so up to 1000 false positives will not damage the leaderboard, while \"skylar\" is present in large numbers, so in order to reduce the number of false positives, up to 10 false positives will not damage the leaderboard? Is that correct?</p>",
      "rawMarkdown": "Thanks for the useful information!\n\nI have one point of concern.\n\n*False positives are tolerated to some extent if one expects that the test data will have a higher percentage of negative examples.*\n\nDoes this sentence mean that there is a set number of false positives allowed for each label?\n\nFor example, \"maupar\" is rarely included in the training data, so up to 1000 false positives will not damage the leaderboard, while \"skylar\" is present in large numbers, so in order to reduce the number of false positives, up to 10 false positives will not damage the leaderboard? Is that correct?",
      "votes": null
    },
    {
      "id": "1774791",
      "postDate": "05/02/2022 12:50:19",
      "content": "<p>Thanks for the useful information!</p>",
      "rawMarkdown": "Thanks for the useful information!",
      "votes": null
    },
    {
      "id": "1774823",
      "postDate": "05/02/2022 13:24:02",
      "content": "<p>My opinion is that TP and TN do not have the same value depending on the weighting.<br>\nPerhaps TP is more valuable(This may vary with different species).<br>\nI believe that sacrificing FP for TP is acceptable within the range of weights.</p>",
      "rawMarkdown": "My opinion is that TP and TN do not have the same value depending on the weighting.\nPerhaps TP is more valuable(This may vary with different species).\nI believe that sacrificing FP for TP is acceptable within the range of weights.",
      "votes": null
    },
    {
      "id": "1774829",
      "postDate": "05/02/2022 13:32:05",
      "content": "<p><a href=\"https://www.kaggle.com/tatei1000\" target=\"_blank\">@tatei1000</a> </p>\n<p>I may be wrong, but it seemed to me that you does not understand the basics, so I thought it would be more efficient to first get the basics down with the following resources and then ask questions.</p>\n<p><a href=\"https://en.wikipedia.org/wiki/F-score\" target=\"_blank\">https://en.wikipedia.org/wiki/F-score</a></p>\n<p>I'll be frank, since we seem to be the same Japanese, but the question below also seems to be asking something quite off the mark.</p>",
      "rawMarkdown": "tatei1000 \n\nI may be wrong, but it seemed to me that you does not understand the basics, so I thought it would be more efficient to first get the basics down with the following resources and then ask questions.\n\nhttps://en.wikipedia.org/wiki/F-score\n\nI'll be frank, since we seem to be the same Japanese, but the question below also seems to be asking something quite off the mark.",
      "votes": null
    },
    {
      "id": "1776879",
      "postDate": "05/04/2022 09:03:46",
      "content": "<p>Hmmm…</p>\n<p>You are right, my lack of knowledge may have made the question worse.<br>\nI will look into it some more. Thank you very much.</p>",
      "rawMarkdown": "Hmmm...\n\nYou are right, my lack of knowledge may have made the question worse.\nI will look into it some more. Thank you very much.",
      "votes": null
    },
    {
      "id": "1778631",
      "postDate": "05/05/2022 14:25:53",
      "content": "<p>Don't get me wrong, but I am not accusing you of lack of knowledge, I am just making a suggestion about the order in which we understand things. But this is just my impression, and I could be wrong.</p>",
      "rawMarkdown": "Don't get me wrong, but I am not accusing you of lack of knowledge, I am just making a suggestion about the order in which we understand things. But this is just my impression, and I could be wrong.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1774468,
      "author_name": "atsunorifujita",
      "author_url": "",
      "post_date": "05/02/2022 07:34:38",
      "content": "<p>I think you are right.</p>\n<p>The host said that they are weighting positive and negative examples.<br>\nFalse positives are tolerated to some extent if one expects that the test data will have a higher percentage of negative examples.<br>\nExceeding tolerance will hurt your leaderboard.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1774790,
          "author_name": "tatei1000",
          "author_url": "",
          "post_date": "05/02/2022 12:47:35",
          "content": "<p>Thanks for the useful information!</p>\n<p>I have one point of concern.</p>\n<p><em>False positives are tolerated to some extent if one expects that the test data will have a higher percentage of negative examples.</em></p>\n<p>Does this sentence mean that there is a set number of false positives allowed for each label?</p>\n<p>For example, \"maupar\" is rarely included in the training data, so up to 1000 false positives will not damage the leaderboard, while \"skylar\" is present in large numbers, so in order to reduce the number of false positives, up to 10 false positives will not damage the leaderboard? Is that correct?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1774823,
          "author_name": "atsunorifujita",
          "author_url": "",
          "post_date": "05/02/2022 13:24:02",
          "content": "<p>My opinion is that TP and TN do not have the same value depending on the weighting.<br>\nPerhaps TP is more valuable(This may vary with different species).<br>\nI believe that sacrificing FP for TP is acceptable within the range of weights.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1774540,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "05/02/2022 09:01:05",
      "content": "<p>In a normal problem setting, it's not the case that \"the lower the prediction threshold, the better\".<br>\nIn the extreme case, a model that predicts everything as a positive example is not making any meaningful predictions.<br>\nWe are not told exactly what evaluation metrics the host employs, but in any case, it is reasonable to assume that models that predict everything as a positive example are set to have a lower score.<br>\nThus, as you say, we are naturally expected to devise ways to reduce FP (or increase TN).</p>",
      "votes": null,
      "replies": [
        {
          "id": 1774549,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "05/02/2022 09:10:40",
          "content": "<p>By the way, you can also refer to the discussion below in which I stated what is the most probable evaluation metrics.<br>\n<a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/321883\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/321883</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1774791,
          "author_name": "tatei1000",
          "author_url": "",
          "post_date": "05/02/2022 12:50:19",
          "content": "<p>Thanks for the useful information!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1774829,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "05/02/2022 13:32:05",
      "content": "<p><a href=\"https://www.kaggle.com/tatei1000\" target=\"_blank\">@tatei1000</a> </p>\n<p>I may be wrong, but it seemed to me that you does not understand the basics, so I thought it would be more efficient to first get the basics down with the following resources and then ask questions.</p>\n<p><a href=\"https://en.wikipedia.org/wiki/F-score\" target=\"_blank\">https://en.wikipedia.org/wiki/F-score</a></p>\n<p>I'll be frank, since we seem to be the same Japanese, but the question below also seems to be asking something quite off the mark.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1776879,
          "author_name": "tatei1000",
          "author_url": "",
          "post_date": "05/04/2022 09:03:46",
          "content": "<p>Hmmm…</p>\n<p>You are right, my lack of knowledge may have made the question worse.<br>\nI will look into it some more. Thank you very much.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1778631,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "05/05/2022 14:25:53",
          "content": "<p>Don't get me wrong, but I am not accusing you of lack of knowledge, I am just making a suggestion about the order in which we understand things. But this is just my impression, and I could be wrong.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1774266": "I do not really understand why lowering the threshold lowers the PUBLIC SCORE.(In the first place, I don't understand Evaluation Metric very well.)\n\nIn my opinion it is better to lower the threshold and detect more birds.\n\nWhen I applied this idea to my model, the results were as follows.\n\n| threshold | PB |\n| --- | --- |\n| 0.1 |0.60 |\n| 0.05 |0.55 |\n\nI am thinking that this is caused by too many `fp` in `tp/(tp+fp)` in **macro F1 score**, but is there any other possible reason?\n\nThank you very much.\n\n(This text is translated at DeepL and may be difficult to read.)",
    "1774468": "I think you are right.\n\nThe host said that they are weighting positive and negative examples.\nFalse positives are tolerated to some extent if one expects that the test data will have a higher percentage of negative examples.\nExceeding tolerance will hurt your leaderboard.",
    "1774540": "In a normal problem setting, it's not the case that \"the lower the prediction threshold, the better\".\nIn the extreme case, a model that predicts everything as a positive example is not making any meaningful predictions.\nWe are not told exactly what evaluation metrics the host employs, but in any case, it is reasonable to assume that models that predict everything as a positive example are set to have a lower score.\nThus, as you say, we are naturally expected to devise ways to reduce FP (or increase TN).",
    "1774549": "By the way, you can also refer to the discussion below in which I stated what is the most probable evaluation metrics.\nhttps://www.kaggle.com/competitions/birdclef-2022/discussion/321883",
    "1774790": "Thanks for the useful information!\n\nI have one point of concern.\n\n*False positives are tolerated to some extent if one expects that the test data will have a higher percentage of negative examples.*\n\nDoes this sentence mean that there is a set number of false positives allowed for each label?\n\nFor example, \"maupar\" is rarely included in the training data, so up to 1000 false positives will not damage the leaderboard, while \"skylar\" is present in large numbers, so in order to reduce the number of false positives, up to 10 false positives will not damage the leaderboard? Is that correct?",
    "1774791": "Thanks for the useful information!",
    "1774823": "My opinion is that TP and TN do not have the same value depending on the weighting.\nPerhaps TP is more valuable(This may vary with different species).\nI believe that sacrificing FP for TP is acceptable within the range of weights.",
    "1774829": "tatei1000 \n\nI may be wrong, but it seemed to me that you does not understand the basics, so I thought it would be more efficient to first get the basics down with the following resources and then ask questions.\n\nhttps://en.wikipedia.org/wiki/F-score\n\nI'll be frank, since we seem to be the same Japanese, but the question below also seems to be asking something quite off the mark.",
    "1776879": "Hmmm...\n\nYou are right, my lack of knowledge may have made the question worse.\nI will look into it some more. Thank you very much.",
    "1778631": "Don't get me wrong, but I am not accusing you of lack of knowledge, I am just making a suggestion about the order in which we understand things. But this is just my impression, and I could be wrong."
  },
  "source": "meta"
}