{
  "id": 236859,
  "title": "How to choose threshold?",
  "url": "/competitions/birdclef-2021/discussion/236859",
  "author_name": "",
  "post_date": "2021-05-06T06:44:20.593452700Z",
  "votes": 26,
  "comment_count": 18,
  "views": 0,
  "content": "<p>As the topic says how are you guys choosing thresholds? </p>\n<p>I assume everyone who is in the Top-10 uses an ensemble for their submission…</p>\n<ol>\n<li>How do you guys decide on which threshold is the best to use?</li>\n<li>How are you confident that it might score well on the Hidden test set?</li>\n</ol>\n<p>Any resources or guidance will be really helpful..</p>\n<p>Thank you</p>\n<p>EDIT: <br>\nFound 2 different techniques </p>\n<ol>\n<li>Model1-&gt; Validation F1 for thresholds from 0.1-0.9 -&gt; pick best threshold<br>\nRepeat the same process for all models and use the respective thresholds for each models and do a voting ensemble</li>\n<li>Average the predictions for all models -&gt; Validation F1 for thresholds 0.1-0.9 -&gt; pick best threshold and use it</li>\n</ol>\n<p>EDIT2:<br>\nYesterday while going through a forum I found a good comment which said:</p>\n<blockquote>\n  <p>the best threshold lies somewhere near <code>F1score/2</code> where the F1 score is calculated on your validation set</p>\n</blockquote>",
  "messages": [
    {
      "id": "1295034",
      "postDate": "05/06/2021 06:44:20",
      "content": "<p>As the topic says how are you guys choosing thresholds? </p>\n<p>I assume everyone who is in the Top-10 uses an ensemble for their submission…</p>\n<ol>\n<li>How do you guys decide on which threshold is the best to use?</li>\n<li>How are you confident that it might score well on the Hidden test set?</li>\n</ol>\n<p>Any resources or guidance will be really helpful..</p>\n<p>Thank you</p>\n<p>EDIT: <br>\nFound 2 different techniques </p>\n<ol>\n<li>Model1-&gt; Validation F1 for thresholds from 0.1-0.9 -&gt; pick best threshold<br>\nRepeat the same process for all models and use the respective thresholds for each models and do a voting ensemble</li>\n<li>Average the predictions for all models -&gt; Validation F1 for thresholds 0.1-0.9 -&gt; pick best threshold and use it</li>\n</ol>\n<p>EDIT2:<br>\nYesterday while going through a forum I found a good comment which said:</p>\n<blockquote>\n  <p>the best threshold lies somewhere near <code>F1score/2</code> where the F1 score is calculated on your validation set</p>\n</blockquote>",
      "rawMarkdown": "As the topic says how are you guys choosing thresholds? \n\nI assume everyone who is in the Top-10 uses an ensemble for their submission...\n1. How do you guys decide on which threshold is the best to use?\n2. How are you confident that it might score well on the Hidden test set?\n\nAny resources or guidance will be really helpful..\n\nThank you\n\nEDIT: \nFound 2 different techniques \n1. Model1-> Validation F1 for thresholds from 0.1-0.9 -> pick best threshold\nRepeat the same process for all models and use the respective thresholds for each models and do a voting ensemble\n2. Average the predictions for all models -> Validation F1 for thresholds 0.1-0.9 -> pick best threshold and use it\n\nEDIT2:\nYesterday while going through a forum I found a good comment which said:\n> the best threshold lies somewhere near `F1score/2` where the F1 score is calculated on your validation set",
      "votes": null
    },
    {
      "id": "1295140",
      "postDate": "05/06/2021 08:30:09",
      "content": "<blockquote>\n  <p>How are you confident that it might score well on the Hidden test set?</p>\n</blockquote>\n<p>I have no clue on public vs private split hence everything is possible.</p>",
      "rawMarkdown": "> How are you confident that it might score well on the Hidden test set?\n\nI have no clue on public vs private split hence everything is possible.",
      "votes": null
    },
    {
      "id": "1295177",
      "postDate": "05/06/2021 09:06:37",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> <br>\nHow are you deciding which threshold to use?<br>\nI think it has an important role in determining how our models score on Public and Private LB</p>",
      "rawMarkdown": "cpmpml \nHow are you deciding which threshold to use?\nI think it has an important role in determining how our models score on Public and Private LB",
      "votes": null
    },
    {
      "id": "1295189",
      "postDate": "05/06/2021 09:15:00",
      "content": "<blockquote>\n  <p>How are you deciding which threshold to use?</p>\n</blockquote>\n<p>I'll disclose after competition end.  I doubt anyone in LB top 10 will disclose specifics like this during the competition.</p>\n<p>I can only encourage you to experiment and see what works best.</p>",
      "rawMarkdown": "> How are you deciding which threshold to use?\n\nI'll disclose after competition end.  I doubt anyone in LB top 10 will disclose specifics like this during the competition.\n\nI can only encourage you to experiment and see what works best.",
      "votes": null
    },
    {
      "id": "1295214",
      "postDate": "05/06/2021 09:36:20",
      "content": "<p>Cool by asking which threshold to use I did not mean specifics I just wanted to know if there are any resources which you use as a rules to better understand which thresholds work. <br>\nAnyways looking forward for your solution :) </p>",
      "rawMarkdown": "Cool by asking which threshold to use I did not mean specifics I just wanted to know if there are any resources which you use as a rules to better understand which thresholds work. \nAnyways looking forward for your solution :)",
      "votes": null
    },
    {
      "id": "1295234",
      "postDate": "05/06/2021 10:03:34",
      "content": "<p>The only rule I use is to see if I get better metric value or not.</p>",
      "rawMarkdown": "The only rule I use is to see if I get better metric value or not.",
      "votes": null
    },
    {
      "id": "1295434",
      "postDate": "05/06/2021 13:02:46",
      "content": "<p>One way I may think is to use train soundscapes data to validate thresholds comparing obs labels with pred labels from a model trained only with train short audio data.<br>\nFor each species, searching in a grid, select the threshold values which gave highest f1val overall</p>",
      "rawMarkdown": "One way I may think is to use train soundscapes data to validate thresholds comparing obs labels with pred labels from a model trained only with train short audio data.\nFor each species, searching in a grid, select the threshold values which gave highest f1val overall",
      "votes": null
    },
    {
      "id": "1295474",
      "postDate": "05/06/2021 13:39:32",
      "content": "<p>That's what I am trying currently <br>\nBut with the public and private test set we may have a domain shift… </p>",
      "rawMarkdown": "That's what I am trying currently \nBut with the public and private test set we may have a domain shift...",
      "votes": null
    },
    {
      "id": "1295483",
      "postDate": "05/06/2021 13:49:40",
      "content": "<p>If we had the ability to guess that shift we would not be participating here</p>",
      "rawMarkdown": "If we had the ability to guess that shift we would not be participating here",
      "votes": null
    },
    {
      "id": "1295516",
      "postDate": "05/06/2021 14:10:01",
      "content": "<p>That's True :P </p>",
      "rawMarkdown": "That's True :P",
      "votes": null
    },
    {
      "id": "1301735",
      "postDate": "05/11/2021 08:09:29",
      "content": "<p>This is a good topic.Keep following….</p>",
      "rawMarkdown": "This is a good topic.Keep following....",
      "votes": null
    },
    {
      "id": "1306544",
      "postDate": "05/13/2021 21:28:36",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> Sounds like you might not use a threshold at all.</p>",
      "rawMarkdown": "cpmpml Sounds like you might not use a threshold at all.",
      "votes": null
    },
    {
      "id": "1306547",
      "postDate": "05/13/2021 21:37:34",
      "content": "<p>That's the most obvious way we can think - however the CV on train soundscapes does not correlate at all with public LB scores (at least in my early experiments)</p>",
      "rawMarkdown": "That's the most obvious way we can think - however the CV on train soundscapes does not correlate at all with public LB scores (at least in my early experiments)",
      "votes": null
    },
    {
      "id": "1310932",
      "postDate": "05/17/2021 04:23:19",
      "content": "<p>I use <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a>  dataset, train/pre notebook. It is really helpful. (@kneroma, Thank you !!!!)<br>\nif use low threshold the result is good.<br>\nThe main contribution to the score is nocall.<br>\nSo I will do some aug to the data set to improve model output propbility, and chose the best threshold after without evaluating nocall data</p>",
      "rawMarkdown": "I use @kneroma  dataset, train/pre notebook. It is really helpful. (@kneroma, Thank you !!!!)\nif use low threshold the result is good.\nThe main contribution to the score is nocall.\nSo I will do some aug to the data set to improve model output propbility, and chose the best threshold after without evaluating nocall data",
      "votes": null
    },
    {
      "id": "1312673",
      "postDate": "05/18/2021 07:21:15",
      "content": "<p>Do you augment test dataset (TTA)?</p>",
      "rawMarkdown": "Do you augment test dataset (TTA)?",
      "votes": null
    },
    {
      "id": "1312708",
      "postDate": "05/18/2021 07:43:03",
      "content": "<p>have not done yet. I focus nocall data these days. </p>",
      "rawMarkdown": "have not done yet. I focus nocall data these days.",
      "votes": null
    },
    {
      "id": "1313810",
      "postDate": "05/18/2021 18:19:44",
      "content": "<p><a href=\"https://www.kaggle.com/atamazian\" target=\"_blank\">@atamazian</a> I tried TTA it worked for me</p>",
      "rawMarkdown": "atamazian I tried TTA it worked for me",
      "votes": null
    },
    {
      "id": "1314076",
      "postDate": "05/19/2021 00:01:25",
      "content": "<p>thank you for sharing. I will try it later.</p>",
      "rawMarkdown": "thank you for sharing. I will try it later.",
      "votes": null
    },
    {
      "id": "1315593",
      "postDate": "05/20/2021 01:11:17",
      "content": "<p>Hello,how to use TTA? Is it same as computer vision?</p>",
      "rawMarkdown": "Hello,how to use TTA? Is it same as computer vision?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1295140,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/06/2021 08:30:09",
      "content": "<blockquote>\n  <p>How are you confident that it might score well on the Hidden test set?</p>\n</blockquote>\n<p>I have no clue on public vs private split hence everything is possible.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1295177,
          "author_name": "nitindatta",
          "author_url": "",
          "post_date": "05/06/2021 09:06:37",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> <br>\nHow are you deciding which threshold to use?<br>\nI think it has an important role in determining how our models score on Public and Private LB</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1295189,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/06/2021 09:15:00",
          "content": "<blockquote>\n  <p>How are you deciding which threshold to use?</p>\n</blockquote>\n<p>I'll disclose after competition end.  I doubt anyone in LB top 10 will disclose specifics like this during the competition.</p>\n<p>I can only encourage you to experiment and see what works best.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1295214,
          "author_name": "nitindatta",
          "author_url": "",
          "post_date": "05/06/2021 09:36:20",
          "content": "<p>Cool by asking which threshold to use I did not mean specifics I just wanted to know if there are any resources which you use as a rules to better understand which thresholds work. <br>\nAnyways looking forward for your solution :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1295234,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/06/2021 10:03:34",
          "content": "<p>The only rule I use is to see if I get better metric value or not.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1306544,
          "author_name": "fffrrt",
          "author_url": "",
          "post_date": "05/13/2021 21:28:36",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> Sounds like you might not use a threshold at all.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1295434,
      "author_name": "kittlein",
      "author_url": "",
      "post_date": "05/06/2021 13:02:46",
      "content": "<p>One way I may think is to use train soundscapes data to validate thresholds comparing obs labels with pred labels from a model trained only with train short audio data.<br>\nFor each species, searching in a grid, select the threshold values which gave highest f1val overall</p>",
      "votes": null,
      "replies": [
        {
          "id": 1295474,
          "author_name": "nitindatta",
          "author_url": "",
          "post_date": "05/06/2021 13:39:32",
          "content": "<p>That's what I am trying currently <br>\nBut with the public and private test set we may have a domain shift… </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1295483,
          "author_name": "kittlein",
          "author_url": "",
          "post_date": "05/06/2021 13:49:40",
          "content": "<p>If we had the ability to guess that shift we would not be participating here</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1295516,
          "author_name": "nitindatta",
          "author_url": "",
          "post_date": "05/06/2021 14:10:01",
          "content": "<p>That's True :P </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1306547,
          "author_name": "imeintanis",
          "author_url": "",
          "post_date": "05/13/2021 21:37:34",
          "content": "<p>That's the most obvious way we can think - however the CV on train soundscapes does not correlate at all with public LB scores (at least in my early experiments)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1301735,
      "author_name": "majunfu",
      "author_url": "",
      "post_date": "05/11/2021 08:09:29",
      "content": "<p>This is a good topic.Keep following….</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1310932,
      "author_name": "shigengtian",
      "author_url": "",
      "post_date": "05/17/2021 04:23:19",
      "content": "<p>I use <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a>  dataset, train/pre notebook. It is really helpful. (@kneroma, Thank you !!!!)<br>\nif use low threshold the result is good.<br>\nThe main contribution to the score is nocall.<br>\nSo I will do some aug to the data set to improve model output propbility, and chose the best threshold after without evaluating nocall data</p>",
      "votes": null,
      "replies": [
        {
          "id": 1312673,
          "author_name": "atamazian",
          "author_url": "",
          "post_date": "05/18/2021 07:21:15",
          "content": "<p>Do you augment test dataset (TTA)?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1312708,
          "author_name": "shigengtian",
          "author_url": "",
          "post_date": "05/18/2021 07:43:03",
          "content": "<p>have not done yet. I focus nocall data these days. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1313810,
          "author_name": "arunodhayan",
          "author_url": "",
          "post_date": "05/18/2021 18:19:44",
          "content": "<p><a href=\"https://www.kaggle.com/atamazian\" target=\"_blank\">@atamazian</a> I tried TTA it worked for me</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1314076,
          "author_name": "shigengtian",
          "author_url": "",
          "post_date": "05/19/2021 00:01:25",
          "content": "<p>thank you for sharing. I will try it later.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1315593,
          "author_name": "zekunn",
          "author_url": "",
          "post_date": "05/20/2021 01:11:17",
          "content": "<p>Hello,how to use TTA? Is it same as computer vision?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1295034": "As the topic says how are you guys choosing thresholds? \n\nI assume everyone who is in the Top-10 uses an ensemble for their submission...\n1. How do you guys decide on which threshold is the best to use?\n2. How are you confident that it might score well on the Hidden test set?\n\nAny resources or guidance will be really helpful..\n\nThank you\n\nEDIT: \nFound 2 different techniques \n1. Model1-> Validation F1 for thresholds from 0.1-0.9 -> pick best threshold\nRepeat the same process for all models and use the respective thresholds for each models and do a voting ensemble\n2. Average the predictions for all models -> Validation F1 for thresholds 0.1-0.9 -> pick best threshold and use it\n\nEDIT2:\nYesterday while going through a forum I found a good comment which said:\n> the best threshold lies somewhere near `F1score/2` where the F1 score is calculated on your validation set",
    "1295140": "> How are you confident that it might score well on the Hidden test set?\n\nI have no clue on public vs private split hence everything is possible.",
    "1295177": "cpmpml \nHow are you deciding which threshold to use?\nI think it has an important role in determining how our models score on Public and Private LB",
    "1295189": "> How are you deciding which threshold to use?\n\nI'll disclose after competition end.  I doubt anyone in LB top 10 will disclose specifics like this during the competition.\n\nI can only encourage you to experiment and see what works best.",
    "1295214": "Cool by asking which threshold to use I did not mean specifics I just wanted to know if there are any resources which you use as a rules to better understand which thresholds work. \nAnyways looking forward for your solution :)",
    "1295234": "The only rule I use is to see if I get better metric value or not.",
    "1295434": "One way I may think is to use train soundscapes data to validate thresholds comparing obs labels with pred labels from a model trained only with train short audio data.\nFor each species, searching in a grid, select the threshold values which gave highest f1val overall",
    "1295474": "That's what I am trying currently \nBut with the public and private test set we may have a domain shift...",
    "1295483": "If we had the ability to guess that shift we would not be participating here",
    "1295516": "That's True :P",
    "1301735": "This is a good topic.Keep following....",
    "1306544": "cpmpml Sounds like you might not use a threshold at all.",
    "1306547": "That's the most obvious way we can think - however the CV on train soundscapes does not correlate at all with public LB scores (at least in my early experiments)",
    "1310932": "I use @kneroma  dataset, train/pre notebook. It is really helpful. (@kneroma, Thank you !!!!)\nif use low threshold the result is good.\nThe main contribution to the score is nocall.\nSo I will do some aug to the data set to improve model output propbility, and chose the best threshold after without evaluating nocall data",
    "1312673": "Do you augment test dataset (TTA)?",
    "1312708": "have not done yet. I focus nocall data these days.",
    "1313810": "atamazian I tried TTA it worked for me",
    "1314076": "thank you for sharing. I will try it later.",
    "1315593": "Hello,how to use TTA? Is it same as computer vision?"
  },
  "source": "meta"
}