{
  "id": 177992,
  "title": "The Importance of Threshold 0.525 - 0.552",
  "url": "/competitions/birdsong-recognition/discussion/177992",
  "author_name": "",
  "post_date": "2020-08-28T06:46:08.459423800Z",
  "votes": 5,
  "comment_count": 16,
  "views": 0,
  "content": "<p>My submission process (like many) involves a threshold where all classes that score above that threshold are considered \"present\" and all below \"absent\". </p>\n<p>I submitted a model with a threshold of 0.31 (based on the threshold that worked best locally) and achieved 525 on the lb. I then resubmitted exactly the same model with a threshold of 0.71  and got 552. </p>\n<p>This suggests that thresholds are crazy important for this comp, and I'm wondering, since my initial estimation mechanism performed so poorly, how the heck we're supposed to estimate good thresholds? </p>\n<p>I picked 0.71 simply because that threshold produced a <code>submission.csv</code> (in test mode) that looked pretty similar to the <code>submission.csv</code> that one of the smart guys got <a href=\"https://www.kaggle.com/ttahara/training-birdsong-baseline-resnest50-fast\" target=\"_blank\">here</a>, but other than copying Tawara, I'm at a complete loss for choosing a good threshold.</p>\n<p>Each model you train probably has its own optimal threshold which makes things even more annoying. Any ideas?</p>",
  "messages": [
    {
      "id": "988612",
      "postDate": "08/28/2020 06:46:08",
      "content": "<p>My submission process (like many) involves a threshold where all classes that score above that threshold are considered \"present\" and all below \"absent\". </p>\n<p>I submitted a model with a threshold of 0.31 (based on the threshold that worked best locally) and achieved 525 on the lb. I then resubmitted exactly the same model with a threshold of 0.71  and got 552. </p>\n<p>This suggests that thresholds are crazy important for this comp, and I'm wondering, since my initial estimation mechanism performed so poorly, how the heck we're supposed to estimate good thresholds? </p>\n<p>I picked 0.71 simply because that threshold produced a <code>submission.csv</code> (in test mode) that looked pretty similar to the <code>submission.csv</code> that one of the smart guys got <a href=\"https://www.kaggle.com/ttahara/training-birdsong-baseline-resnest50-fast\" target=\"_blank\">here</a>, but other than copying Tawara, I'm at a complete loss for choosing a good threshold.</p>\n<p>Each model you train probably has its own optimal threshold which makes things even more annoying. Any ideas?</p>",
      "rawMarkdown": "My submission process (like many) involves a threshold where all classes that score above that threshold are considered \"present\" and all below \"absent\". \n\nI submitted a model with a threshold of 0.31 (based on the threshold that worked best locally) and achieved 525 on the lb. I then resubmitted exactly the same model with a threshold of 0.71  and got 552. \n\nThis suggests that thresholds are crazy important for this comp, and I'm wondering, since my initial estimation mechanism performed so poorly, how the heck we're supposed to estimate good thresholds? \n\nI picked 0.71 simply because that threshold produced a `submission.csv` (in test mode) that looked pretty similar to the `submission.csv` that one of the smart guys got [here](https://www.kaggle.com/ttahara/training-birdsong-baseline-resnest50-fast), but other than copying Tawara, I'm at a complete loss for choosing a good threshold.\n\nEach model you train probably has its own optimal threshold which makes things even more annoying. Any ideas?",
      "votes": null
    },
    {
      "id": "988616",
      "postDate": "08/28/2020 06:49:51",
      "content": "<p>Have anyone tried to use threshold optimizer in this competition?</p>",
      "rawMarkdown": "Have anyone tried to use threshold optimizer in this competition?",
      "votes": null
    },
    {
      "id": "988627",
      "postDate": "08/28/2020 07:01:11",
      "content": "<p>What's a threshold optimizer? That sounds like it could be extremely useful!</p>",
      "rawMarkdown": "What's a threshold optimizer? That sounds like it could be extremely useful!",
      "votes": null
    },
    {
      "id": "988645",
      "postDate": "08/28/2020 07:20:26",
      "content": "<blockquote>\n  <p>This suggests that thresholds are crazy important for this comp</p>\n</blockquote>\n<p>It's often the case for competitions with evaluation metric F1</p>",
      "rawMarkdown": "> This suggests that thresholds are crazy important for this comp\n\nIt's often the case for competitions with evaluation metric F1",
      "votes": null
    },
    {
      "id": "988682",
      "postDate": "08/28/2020 07:42:04",
      "content": "<p>You can optimize threshold using pycaret. It might will meet your requirement. </p>",
      "rawMarkdown": "You can optimize threshold using pycaret. It might will meet your requirement.",
      "votes": null
    },
    {
      "id": "988703",
      "postDate": "08/28/2020 08:09:37",
      "content": "<p><a href=\"https://www.kaggle.com/stecasasso\" target=\"_blank\">@stecasasso</a> were there any good examples of past comps where this was the case?</p>",
      "rawMarkdown": "stecasasso were there any good examples of past comps where this was the case?",
      "votes": null
    },
    {
      "id": "988836",
      "postDate": "08/28/2020 10:15:39",
      "content": "<p>I would advise using this function to pick your threshold :</p>\n<pre><code>def optimize_threshold(predictions):\n     return 0.5\n</code></pre>",
      "rawMarkdown": "I would advise using this function to pick your threshold :\n\n``` \ndef optimize_threshold(predictions):\n     return 0.5\n```",
      "votes": null
    },
    {
      "id": "988837",
      "postDate": "08/28/2020 10:17:21",
      "content": "<p>More seriously, heavily tweaking your threshold, either on validation or on public LB, is not usually a good thing to do </p>",
      "rawMarkdown": "More seriously, heavily tweaking your threshold, either on validation or on public LB, is not usually a good thing to do",
      "votes": null
    },
    {
      "id": "988914",
      "postDate": "08/28/2020 12:04:20",
      "content": "<p>Yes tweaking the probability threshold makes no sense, the nature of the value, but if the model in average hits the target slightly to the left/right based on validation, maybe tweaking it can help with specific labels, know the behavior of the error.<br>\nCrazy idea, can one optimize the optimizing, instead of optimize the whole tree pick out say top 10 worst thresh/prob/label errors, use them in the inference, and like with opt label classification works better with less better model?….</p>",
      "rawMarkdown": "Yes tweaking the probability threshold makes no sense, the nature of the value, but if the model in average hits the target slightly to the left/right based on validation, maybe tweaking it can help with specific labels, know the behavior of the error.\nCrazy idea, can one optimize the optimizing, instead of optimize the whole tree pick out say top 10 worst thresh/prob/label errors, use them in the inference, and like with opt label classification works better with less better model?....",
      "votes": null
    },
    {
      "id": "988926",
      "postDate": "08/28/2020 12:15:30",
      "content": "<p>Recently I remember Human Atlas and Tensorflow Q&amp;A, but I am sure there are others (there does not seem a way to select competitions based on metric).<br>\nThat being said, it depends on the comp. I remember in Human Atlas we heavily optimised the threshold and penalized a bit by that.</p>",
      "rawMarkdown": "Recently I remember Human Atlas and Tensorflow Q&A, but I am sure there are others (there does not seem a way to select competitions based on metric).\nThat being said, it depends on the comp. I remember in Human Atlas we heavily optimised the threshold and penalized a bit by that.",
      "votes": null
    },
    {
      "id": "989312",
      "postDate": "08/28/2020 17:47:53",
      "content": "<p>I have an idea to do it. The one possible way just to optimize threshold on some set and than apply it to final prediction. Or you can use a model and return unique treshold for each sample base on some features (model output, summary of the hidden states, what have you). </p>",
      "rawMarkdown": "I have an idea to do it. The one possible way just to optimize threshold on some set and than apply it to final prediction. Or you can use a model and return unique treshold for each sample base on some features (model output, summary of the hidden states, what have you).",
      "votes": null
    },
    {
      "id": "989487",
      "postDate": "08/28/2020 20:42:10",
      "content": "<p>So I agree, optimizing thresholds will probably lead to over-fitting. However, clearly the threshold can make an immense difference. It would suck to make an excellent model and come last due to a poor threshold. There must be something better than just picking at random.</p>\n<p><a href=\"https://www.kaggle.com/kirderf\" target=\"_blank\">@kirderf</a> I couldn't quite follow what you were suggesting but it sounds interesting. Can you expand?</p>",
      "rawMarkdown": "So I agree, optimizing thresholds will probably lead to over-fitting. However, clearly the threshold can make an immense difference. It would suck to make an excellent model and come last due to a poor threshold. There must be something better than just picking at random.\n\n@kirderf I couldn't quite follow what you were suggesting but it sounds interesting. Can you expand?",
      "votes": null
    },
    {
      "id": "989493",
      "postDate": "08/28/2020 20:50:41",
      "content": "<p>I tried your first idea, that's what gave me my initial threshold, but apparently my threshold optimization set was very far away from the public test set.</p>",
      "rawMarkdown": "I tried your first idea, that's what gave me my initial threshold, but apparently my threshold optimization set was very far away from the public test set.",
      "votes": null
    },
    {
      "id": "989527",
      "postDate": "08/28/2020 21:55:41",
      "content": "<p>The test dataset have so many nocall. So I advise using </p>\n<pre><code>threshold &gt; 0.5\n</code></pre>",
      "rawMarkdown": "The test dataset have so many nocall. So I advise using \n```\nthreshold > 0.5\n```",
      "votes": null
    },
    {
      "id": "989614",
      "postDate": "08/29/2020 00:42:20",
      "content": "<p>why not 0.55 or 0.45?</p>",
      "rawMarkdown": "why not 0.55 or 0.45?",
      "votes": null
    },
    {
      "id": "989972",
      "postDate": "08/29/2020 08:42:25",
      "content": "<p>because The test dataset have so many nocall.</p>",
      "rawMarkdown": "because The test dataset have so many nocall.",
      "votes": null
    },
    {
      "id": "990067",
      "postDate": "08/29/2020 10:02:15",
      "content": "<p>It's just from the observation of <strong>public</strong> LB. <br>\nWorking on finding optimal threshold is ok, but it's more or less a lottery. There are a lot many things to do before you cast the dice.</p>",
      "rawMarkdown": "It's just from the observation of **public** LB. \nWorking on finding optimal threshold is ok, but it's more or less a lottery. There are a lot many things to do before you cast the dice.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 988616,
      "author_name": "kirderf",
      "author_url": "",
      "post_date": "08/28/2020 06:49:51",
      "content": "<p>Have anyone tried to use threshold optimizer in this competition?</p>",
      "votes": null,
      "replies": [
        {
          "id": 988627,
          "author_name": "lewington",
          "author_url": "",
          "post_date": "08/28/2020 07:01:11",
          "content": "<p>What's a threshold optimizer? That sounds like it could be extremely useful!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 989312,
          "author_name": "koza4ukdmitrij",
          "author_url": "",
          "post_date": "08/28/2020 17:47:53",
          "content": "<p>I have an idea to do it. The one possible way just to optimize threshold on some set and than apply it to final prediction. Or you can use a model and return unique treshold for each sample base on some features (model output, summary of the hidden states, what have you). </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 989493,
          "author_name": "lewington",
          "author_url": "",
          "post_date": "08/28/2020 20:50:41",
          "content": "<p>I tried your first idea, that's what gave me my initial threshold, but apparently my threshold optimization set was very far away from the public test set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 988645,
      "author_name": "stecasasso",
      "author_url": "",
      "post_date": "08/28/2020 07:20:26",
      "content": "<blockquote>\n  <p>This suggests that thresholds are crazy important for this comp</p>\n</blockquote>\n<p>It's often the case for competitions with evaluation metric F1</p>",
      "votes": null,
      "replies": [
        {
          "id": 988703,
          "author_name": "lewington",
          "author_url": "",
          "post_date": "08/28/2020 08:09:37",
          "content": "<p><a href=\"https://www.kaggle.com/stecasasso\" target=\"_blank\">@stecasasso</a> were there any good examples of past comps where this was the case?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 988926,
          "author_name": "stecasasso",
          "author_url": "",
          "post_date": "08/28/2020 12:15:30",
          "content": "<p>Recently I remember Human Atlas and Tensorflow Q&amp;A, but I am sure there are others (there does not seem a way to select competitions based on metric).<br>\nThat being said, it depends on the comp. I remember in Human Atlas we heavily optimised the threshold and penalized a bit by that.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 988682,
      "author_name": "acanacar",
      "author_url": "",
      "post_date": "08/28/2020 07:42:04",
      "content": "<p>You can optimize threshold using pycaret. It might will meet your requirement. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 988836,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "08/28/2020 10:15:39",
      "content": "<p>I would advise using this function to pick your threshold :</p>\n<pre><code>def optimize_threshold(predictions):\n     return 0.5\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 988837,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "08/28/2020 10:17:21",
          "content": "<p>More seriously, heavily tweaking your threshold, either on validation or on public LB, is not usually a good thing to do </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 988914,
          "author_name": "kirderf",
          "author_url": "",
          "post_date": "08/28/2020 12:04:20",
          "content": "<p>Yes tweaking the probability threshold makes no sense, the nature of the value, but if the model in average hits the target slightly to the left/right based on validation, maybe tweaking it can help with specific labels, know the behavior of the error.<br>\nCrazy idea, can one optimize the optimizing, instead of optimize the whole tree pick out say top 10 worst thresh/prob/label errors, use them in the inference, and like with opt label classification works better with less better model?….</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 989487,
          "author_name": "lewington",
          "author_url": "",
          "post_date": "08/28/2020 20:42:10",
          "content": "<p>So I agree, optimizing thresholds will probably lead to over-fitting. However, clearly the threshold can make an immense difference. It would suck to make an excellent model and come last due to a poor threshold. There must be something better than just picking at random.</p>\n<p><a href=\"https://www.kaggle.com/kirderf\" target=\"_blank\">@kirderf</a> I couldn't quite follow what you were suggesting but it sounds interesting. Can you expand?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 989527,
      "author_name": "doanquanvietnamca",
      "author_url": "",
      "post_date": "08/28/2020 21:55:41",
      "content": "<p>The test dataset have so many nocall. So I advise using </p>\n<pre><code>threshold &gt; 0.5\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 989614,
          "author_name": "lewington",
          "author_url": "",
          "post_date": "08/29/2020 00:42:20",
          "content": "<p>why not 0.55 or 0.45?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 989972,
          "author_name": "doanquanvietnamca",
          "author_url": "",
          "post_date": "08/29/2020 08:42:25",
          "content": "<p>because The test dataset have so many nocall.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 990067,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "08/29/2020 10:02:15",
          "content": "<p>It's just from the observation of <strong>public</strong> LB. <br>\nWorking on finding optimal threshold is ok, but it's more or less a lottery. There are a lot many things to do before you cast the dice.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "988612": "My submission process (like many) involves a threshold where all classes that score above that threshold are considered \"present\" and all below \"absent\". \n\nI submitted a model with a threshold of 0.31 (based on the threshold that worked best locally) and achieved 525 on the lb. I then resubmitted exactly the same model with a threshold of 0.71  and got 552. \n\nThis suggests that thresholds are crazy important for this comp, and I'm wondering, since my initial estimation mechanism performed so poorly, how the heck we're supposed to estimate good thresholds? \n\nI picked 0.71 simply because that threshold produced a `submission.csv` (in test mode) that looked pretty similar to the `submission.csv` that one of the smart guys got [here](https://www.kaggle.com/ttahara/training-birdsong-baseline-resnest50-fast), but other than copying Tawara, I'm at a complete loss for choosing a good threshold.\n\nEach model you train probably has its own optimal threshold which makes things even more annoying. Any ideas?",
    "988616": "Have anyone tried to use threshold optimizer in this competition?",
    "988627": "What's a threshold optimizer? That sounds like it could be extremely useful!",
    "988645": "> This suggests that thresholds are crazy important for this comp\n\nIt's often the case for competitions with evaluation metric F1",
    "988682": "You can optimize threshold using pycaret. It might will meet your requirement.",
    "988703": "stecasasso were there any good examples of past comps where this was the case?",
    "988836": "I would advise using this function to pick your threshold :\n\n``` \ndef optimize_threshold(predictions):\n     return 0.5\n```",
    "988837": "More seriously, heavily tweaking your threshold, either on validation or on public LB, is not usually a good thing to do",
    "988914": "Yes tweaking the probability threshold makes no sense, the nature of the value, but if the model in average hits the target slightly to the left/right based on validation, maybe tweaking it can help with specific labels, know the behavior of the error.\nCrazy idea, can one optimize the optimizing, instead of optimize the whole tree pick out say top 10 worst thresh/prob/label errors, use them in the inference, and like with opt label classification works better with less better model?....",
    "988926": "Recently I remember Human Atlas and Tensorflow Q&A, but I am sure there are others (there does not seem a way to select competitions based on metric).\nThat being said, it depends on the comp. I remember in Human Atlas we heavily optimised the threshold and penalized a bit by that.",
    "989312": "I have an idea to do it. The one possible way just to optimize threshold on some set and than apply it to final prediction. Or you can use a model and return unique treshold for each sample base on some features (model output, summary of the hidden states, what have you).",
    "989487": "So I agree, optimizing thresholds will probably lead to over-fitting. However, clearly the threshold can make an immense difference. It would suck to make an excellent model and come last due to a poor threshold. There must be something better than just picking at random.\n\n@kirderf I couldn't quite follow what you were suggesting but it sounds interesting. Can you expand?",
    "989493": "I tried your first idea, that's what gave me my initial threshold, but apparently my threshold optimization set was very far away from the public test set.",
    "989527": "The test dataset have so many nocall. So I advise using \n```\nthreshold > 0.5\n```",
    "989614": "why not 0.55 or 0.45?",
    "989972": "because The test dataset have so many nocall.",
    "990067": "It's just from the observation of **public** LB. \nWorking on finding optimal threshold is ok, but it's more or less a lottery. There are a lot many things to do before you cast the dice."
  },
  "source": "meta"
}