{
  "id": 180028,
  "title": "Is the devil in the threshold !?",
  "url": "/competitions/birdsong-recognition/discussion/180028",
  "author_name": "",
  "post_date": "2020-09-03T16:21:48.552493300Z",
  "votes": 23,
  "comment_count": 50,
  "views": 0,
  "content": "<p>Needless to recall how hard is it to build a robust cross-validation  scheme for this competition. And when it comes to the choice of the right <strong>threshold</strong> for the <strong>nocall</strong> event detection, things get more harder 😈😔.</p>\n<p>During my experiments, I found that right values for the <strong>thresold</strong> hyperparameter depends not only on the model but also on the dataset. So, even a correctly cross-validated threshold for the training set could be useless for the test one as both datasets don't come from  same sources.</p>\n<p>For example, when using Resnest-like models, the right threshold is around <strong>0.6</strong>, but, if I change the training precedure, the model becomes too much confident and I need to raise the threshold to <strong>0.8</strong> in order to get any interesting score. </p>\n<p>This whole competition could just be a matter of choosing the right threshold ! We're trying experiments to get ride of it and obtain a threshold-agnostic model but till now, we get no consistant  results …</p>\n<p>And you, how do you  choose your <strong>threshold</strong>, if you're using any !?</p>",
  "messages": [
    {
      "id": "996900",
      "postDate": "09/03/2020 16:21:48",
      "content": "<p>Needless to recall how hard is it to build a robust cross-validation  scheme for this competition. And when it comes to the choice of the right <strong>threshold</strong> for the <strong>nocall</strong> event detection, things get more harder 😈😔.</p>\n<p>During my experiments, I found that right values for the <strong>thresold</strong> hyperparameter depends not only on the model but also on the dataset. So, even a correctly cross-validated threshold for the training set could be useless for the test one as both datasets don't come from  same sources.</p>\n<p>For example, when using Resnest-like models, the right threshold is around <strong>0.6</strong>, but, if I change the training precedure, the model becomes too much confident and I need to raise the threshold to <strong>0.8</strong> in order to get any interesting score. </p>\n<p>This whole competition could just be a matter of choosing the right threshold ! We're trying experiments to get ride of it and obtain a threshold-agnostic model but till now, we get no consistant  results …</p>\n<p>And you, how do you  choose your <strong>threshold</strong>, if you're using any !?</p>",
      "rawMarkdown": "Needless to recall how hard is it to build a robust cross-validation  scheme for this competition. And when it comes to the choice of the right **threshold** for the **nocall** event detection, things get more harder 😈😔.\n\nDuring my experiments, I found that right values for the **thresold** hyperparameter depends not only on the model but also on the dataset. So, even a correctly cross-validated threshold for the training set could be useless for the test one as both datasets don't come from  same sources.\n\nFor example, when using Resnest-like models, the right threshold is around **0.6**, but, if I change the training precedure, the model becomes too much confident and I need to raise the threshold to **0.8** in order to get any interesting score. \n\nThis whole competition could just be a matter of choosing the right threshold ! We're trying experiments to get ride of it and obtain a threshold-agnostic model but till now, we get no consistant  results ...\n\nAnd you, how do you  choose your **threshold**, if you're using any !?",
      "votes": null
    },
    {
      "id": "997031",
      "postDate": "09/03/2020 17:33:39",
      "content": "<p>Clearly, predicting nocall is key.  But I have not found how to do it yet ;)</p>",
      "rawMarkdown": "Clearly, predicting nocall is key.  But I have not found how to do it yet ;)",
      "votes": null
    },
    {
      "id": "997090",
      "postDate": "09/03/2020 18:27:31",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1690820%2F66e0855cc68f73c3c2968fcf76929d74%2F4dqozp.jpg?generation=1599157649395881&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1690820%2F66e0855cc68f73c3c2968fcf76929d74%2F4dqozp.jpg?generation=1599157649395881&alt=media)",
      "votes": null
    },
    {
      "id": "997141",
      "postDate": "09/03/2020 19:18:28",
      "content": "<p>may I know your current LB score is using what threshold value? thanks</p>",
      "rawMarkdown": "may I know your current LB score is using what threshold value? thanks",
      "votes": null
    },
    {
      "id": "997171",
      "postDate": "09/03/2020 19:50:03",
      "content": "<blockquote>\n  <p>This whole competition could just be a matter of choosing the right threshold ! We're trying experiments to get ride of it and obtain a threshold-agnostic model but till now, we get no consistant results …</p>\n</blockquote>\n<p>I think you are hitting your head on a wall instead of walking around it.</p>\n<blockquote>\n  <p>And you, how do you choose your threshold, if you're using any !?</p>\n</blockquote>\n<p>Write down all model answers on validation fold, create two 100-element arrays (each element for each % of confidence).<br>\nFor each model answer add +1 to \"total answers\" array at position of confidence of the answer (total_answers[35] = total_answers[35] + 1 if model output was \"whateverbird: 0.3572), add +1 to \"corrects\" array at the same position if model answer was correct.</p>\n<p>At the end you would get two arrays shaped like this:</p>\n<pre><code>corrects = [0,0,0,2,1,3,2,4,6,5,4,7,10,16,8,8,13,13,21,19,18,14,15,20,19,20,24,26,25,17,37,27,22,22,36,20,21,27,21,25,27,16,34,21,23,31,23,28,38,33,28,23,23,23,29,13,25,27,19,18,28,14,21,18,23,17,16,15,26,23,23,15,20,18,22,17,15,15,21,10,18,21,10,16,17,16,25,11,18,23,14,12,11,20,17,6,22,16,12,15]\ntotal = [1,16,44,62,72,81,60,66,77,68,71,81,86,86,76,69,74,70,81,81,73,71,65,67,61,67,79,62,61,63,73,59,56,48,75,51,59,55,52,52,50,31,55,55,48,59,35,42,45,44,49,38,39,39,37,32,34,37,30,24,35,22,28,27,29,20,20,20,32,26,23,17,22,23,26,21,19,17,24,12,19,23,11,17,18,16,25,12,19,23,14,13,12,20,17,7,23,16,14,16]\n</code></pre>\n<p>Plot this and get this nice graph: ![<a href=\"https://i.imgur.com/Nb3Br1Z.png\" target=\"_blank\">https://i.imgur.com/Nb3Br1Z.png</a>]</p>\n<p>Then use your trusted eyeballs and pick an accuracy number roughly at score you want to achieve (lets say 0.6). In that case you would use confidence around 0.5, because this is where accuracy reached 0.6. For that model i tried 0.6 and 0.5, and 0.5 got me 0.002 better score.</p>\n<p>Another example - model from this graph was tried on 0.35 and 0.3, and it scored the same in both cases, so this picking strategy seems to work: ![<a href=\"https://i.imgur.com/LS9hlH8.png\" target=\"_blank\">https://i.imgur.com/LS9hlH8.png</a>]</p>\n<p>Keep in mind that threshold optimization does not heavily influence your score, as shown in <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/167263\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition/discussion/167263</a> - with most models you have 0.2 threshold range where your score is in ~0.002 result from best possible score.</p>",
      "rawMarkdown": "> This whole competition could just be a matter of choosing the right threshold ! We're trying experiments to get ride of it and obtain a threshold-agnostic model but till now, we get no consistant results …\n\nI think you are hitting your head on a wall instead of walking around it.\n\n> And you, how do you choose your threshold, if you're using any !?\n\nWrite down all model answers on validation fold, create two 100-element arrays (each element for each % of confidence).\nFor each model answer add +1 to \"total answers\" array at position of confidence of the answer (total_answers[35] = total_answers[35] + 1 if model output was \"whateverbird: 0.3572), add +1 to \"corrects\" array at the same position if model answer was correct.\n\nAt the end you would get two arrays shaped like this:\n```\ncorrects = [0,0,0,2,1,3,2,4,6,5,4,7,10,16,8,8,13,13,21,19,18,14,15,20,19,20,24,26,25,17,37,27,22,22,36,20,21,27,21,25,27,16,34,21,23,31,23,28,38,33,28,23,23,23,29,13,25,27,19,18,28,14,21,18,23,17,16,15,26,23,23,15,20,18,22,17,15,15,21,10,18,21,10,16,17,16,25,11,18,23,14,12,11,20,17,6,22,16,12,15]\ntotal = [1,16,44,62,72,81,60,66,77,68,71,81,86,86,76,69,74,70,81,81,73,71,65,67,61,67,79,62,61,63,73,59,56,48,75,51,59,55,52,52,50,31,55,55,48,59,35,42,45,44,49,38,39,39,37,32,34,37,30,24,35,22,28,27,29,20,20,20,32,26,23,17,22,23,26,21,19,17,24,12,19,23,11,17,18,16,25,12,19,23,14,13,12,20,17,7,23,16,14,16]\n```\nPlot this and get this nice graph: ![https://i.imgur.com/Nb3Br1Z.png]\n\nThen use your trusted eyeballs and pick an accuracy number roughly at score you want to achieve (lets say 0.6). In that case you would use confidence around 0.5, because this is where accuracy reached 0.6. For that model i tried 0.6 and 0.5, and 0.5 got me 0.002 better score.\n\nAnother example - model from this graph was tried on 0.35 and 0.3, and it scored the same in both cases, so this picking strategy seems to work: ![https://i.imgur.com/LS9hlH8.png]\n\nKeep in mind that threshold optimization does not heavily influence your score, as shown in https://www.kaggle.com/c/birdsong-recognition/discussion/167263 - with most models you have 0.2 threshold range where your score is in ~0.002 result from best possible score.",
      "votes": null
    },
    {
      "id": "997187",
      "postDate": "09/03/2020 19:59:10",
      "content": "<p>0.5 goes brrrrrrrr</p>",
      "rawMarkdown": "0.5 goes brrrrrrrr",
      "votes": null
    },
    {
      "id": "997189",
      "postDate": "09/03/2020 19:59:47",
      "content": "<p>0.5, we didn't try other values for this specific model</p>",
      "rawMarkdown": "0.5, we didn't try other values for this specific model",
      "votes": null
    },
    {
      "id": "997203",
      "postDate": "09/03/2020 20:13:24",
      "content": "<p>0.5 for best model here.</p>",
      "rawMarkdown": "0.5 for best model here.",
      "votes": null
    },
    {
      "id": "997220",
      "postDate": "09/03/2020 20:35:32",
      "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> is right since we are so afraid of the <strong>Threshold</strong> devil 😈😈</p>",
      "rawMarkdown": "theoviel is right since we are so afraid of the **Threshold** devil 😈😈",
      "votes": null
    },
    {
      "id": "997226",
      "postDate": "09/03/2020 20:38:52",
      "content": "<p>Don't mind sharing whith me if you ever find how to do it 😄😄😄</p>",
      "rawMarkdown": "Don't mind sharing whith me if you ever find how to do it 😄😄😄",
      "votes": null
    },
    {
      "id": "997328",
      "postDate": "09/04/2020 00:19:43",
      "content": "<p>For me, I use 0.5 for models trained with label smoothing, and 0.6 for models trained without label smoothing.</p>",
      "rawMarkdown": "For me, I use 0.5 for models trained with label smoothing, and 0.6 for models trained without label smoothing.",
      "votes": null
    },
    {
      "id": "997410",
      "postDate": "09/04/2020 03:12:16",
      "content": "<p>Train simple model for call or nocall </p>",
      "rawMarkdown": "Train simple model for call or nocall",
      "votes": null
    },
    {
      "id": "997458",
      "postDate": "09/04/2020 04:16:26",
      "content": "<p><a href=\"https://www.kaggle.com/gopidurgaprasad\" target=\"_blank\">@gopidurgaprasad</a> , theoretically that is sound but it is challenging distinguishing call or nocall with no strong labels. Coming up with nocall labels from scratch is daunting with no potentially proven results, as the only validation you have of your nocall set is through submitting to LB</p>",
      "rawMarkdown": "gopidurgaprasad , theoretically that is sound but it is challenging distinguishing call or nocall with no strong labels. Coming up with nocall labels from scratch is daunting with no potentially proven results, as the only validation you have of your nocall set is through submitting to LB",
      "votes": null
    },
    {
      "id": "997499",
      "postDate": "09/04/2020 04:52:50",
      "content": "<p><a href=\"https://www.kaggle.com/alanchn31\" target=\"_blank\">@alanchn31</a>  i trained nocall/call model it gives me 0.97+ AUC</p>",
      "rawMarkdown": "alanchn31  i trained nocall/call model it gives me 0.97+ AUC",
      "votes": null
    },
    {
      "id": "997500",
      "postDate": "09/04/2020 04:52:51",
      "content": "<p><a href=\"https://www.kaggle.com/alanchn31\" target=\"_blank\">@alanchn31</a>  i trained nocall/call model it gives me 0.97+ AUC</p>",
      "rawMarkdown": "alanchn31  i trained nocall/call model it gives me 0.97+ AUC",
      "votes": null
    },
    {
      "id": "997548",
      "postDate": "09/04/2020 05:29:16",
      "content": "<p>is label smoothing helping you?  I haven't tried yet.</p>",
      "rawMarkdown": "is label smoothing helping you?  I haven't tried yet.",
      "votes": null
    },
    {
      "id": "997558",
      "postDate": "09/04/2020 05:31:23",
      "content": "<p><a href=\"https://www.kaggle.com/gopidurgaprasad\" target=\"_blank\">@gopidurgaprasad</a> Does it help you for test data?  Also, how do you train and validate a nocall model given train data has zero nocall clips.  Sure, you can extract nocall parts, but how do you know you do it right?  The issue is to get reliable nocall labels.</p>",
      "rawMarkdown": "gopidurgaprasad Does it help you for test data?  Also, how do you train and validate a nocall model given train data has zero nocall clips.  Sure, you can extract nocall parts, but how do you know you do it right?  The issue is to get reliable nocall labels.",
      "votes": null
    },
    {
      "id": "997577",
      "postDate": "09/04/2020 05:48:24",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>  hai,</p>\n<p>i trained that nocall model 5 folds it gives 0.976+ AUC, just finesed 5th fold.<br>\nI am not yet submitted, i don't have GPU for this week, i will going to submit it tomorrow.</p>",
      "rawMarkdown": "cpmpml  hai,\n\ni trained that nocall model 5 folds it gives 0.976+ AUC, just finesed 5th fold.\nI am not yet submitted, i don't have GPU for this week, i will going to submit it tomorrow.",
      "votes": null
    },
    {
      "id": "997586",
      "postDate": "09/04/2020 05:52:59",
      "content": "<p>call/nocall idea is like we alredy have 5 fold model trained on randam 5sec right.<br>\nstep 1 : create &lt;= 1sec clips<br>\nstep 2: predict 5 fold model on 1sec clips<br>\nstep 3: find threshold that based on your model on 1 sec clips<br>\nstep 4: below threshold 1sec clips are probably noise</p>",
      "rawMarkdown": "call/nocall idea is like we alredy have 5 fold model trained on randam 5sec right.\nstep 1 : create <= 1sec clips\nstep 2: predict 5 fold model on 1sec clips\nstep 3: find threshold that based on your model on 1 sec clips\nstep 4: below threshold 1sec clips are probably noise",
      "votes": null
    },
    {
      "id": "997603",
      "postDate": "09/04/2020 06:07:15",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> </p>\n<p>the same idea a lot of people doing it inference time, but not at training time.</p>\n<p>my idea for nocall model most of the same, but not exactly same.</p>",
      "rawMarkdown": "cpmpml \n\nthe same idea a lot of people doing it inference time, but not at training time.\n\nmy idea for nocall model most of the same, but not exactly same.",
      "votes": null
    },
    {
      "id": "997620",
      "postDate": "09/04/2020 06:22:14",
      "content": "<blockquote>\n  <p>find threshold that based on your model on 1 sec clips</p>\n</blockquote>\n<p>find threshold for what?  How do you decide that the threshold is the right threshold?</p>",
      "rawMarkdown": "> find threshold that based on your model on 1 sec clips\n\nfind threshold for what?  How do you decide that the threshold is the right threshold?",
      "votes": null
    },
    {
      "id": "997628",
      "postDate": "09/04/2020 06:26:07",
      "content": "<p>the idea is like how we do at inference time, the same idea applies for one-second clip</p>\n<p>for example : </p>\n<p>one-sec clip audio predictions are like - [0.1, 0.2, 0.0, …..]</p>\n<p>all class predictions are very low my threshold &lt;0.3. if all class predictions &lt;0.3 that most probably noise/nocall  </p>",
      "rawMarkdown": "the idea is like how we do at inference time, the same idea applies for one-second clip\n\nfor example : \n\none-sec clip audio predictions are like - [0.1, 0.2, 0.0, .....]\n\nall class predictions are very low my threshold <0.3. if all class predictions <0.3 that most probably noise/nocall",
      "votes": null
    },
    {
      "id": "997631",
      "postDate": "09/04/2020 06:28:02",
      "content": "<p>For single model, it only helps my CV score. However, it helps a lot when I use model ensembling.</p>",
      "rawMarkdown": "For single model, it only helps my CV score. However, it helps a lot when I use model ensembling.",
      "votes": null
    },
    {
      "id": "997633",
      "postDate": "09/04/2020 06:28:42",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> </p>\n<p>again those are noise labels, using a student-teacher training method, incrementally we remove noise labels</p>\n<p>one more idea is like do clustering for noise labels</p>",
      "rawMarkdown": "cpmpml \n\nagain those are noise labels, using a student-teacher training method, incrementally we remove noise labels\n\none more idea is like do clustering for noise labels",
      "votes": null
    },
    {
      "id": "997644",
      "postDate": "09/04/2020 06:35:45",
      "content": "<p>You don't answer my question: how do you decide that the 0.3 threshold is the right threshold?</p>\n<p>I am not criticizing, I am trying to understand what you do ;)</p>",
      "rawMarkdown": "You don't answer my question: how do you decide that the 0.3 threshold is the right threshold?\n\nI am not criticizing, I am trying to understand what you do ;)",
      "votes": null
    },
    {
      "id": "997654",
      "postDate": "09/04/2020 06:43:09",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> </p>\n<p>sorry about that.</p>\n<p>I just assuming 0.3 is my threshold, as like others assuming 0.5 at inference time.</p>",
      "rawMarkdown": "cpmpml \n\nsorry about that.\n\nI just assuming 0.3 is my threshold, as like others assuming 0.5 at inference time.",
      "votes": null
    },
    {
      "id": "997656",
      "postDate": "09/04/2020 06:44:27",
      "content": "<p><code>again those are noise labels, using a student-teacher training method, incrementally we remove noise labels</code></p>\n<p>after this process, I got 0.97+ AUC.</p>\n<p>for me it's working.</p>",
      "rawMarkdown": "`again those are noise labels, using a student-teacher training method, incrementally we remove noise labels`\n\nafter this process, I got 0.97+ AUC.\n\nfor me it's working.",
      "votes": null
    },
    {
      "id": "997671",
      "postDate": "09/04/2020 06:58:26",
      "content": "<p><a href=\"https://www.kaggle.com/gopidurgaprasad\" target=\"_blank\">@gopidurgaprasad</a> That's the point I am trying to make, can current results be trusted? Granted you got good validation results, but in the first place, we don't know the distribution of test set and what nocalls sound like in the test set. The only way to know is to submit. Either way, please keep me posted, will love to see if your nocall set was effective.</p>",
      "rawMarkdown": "gopidurgaprasad That's the point I am trying to make, can current results be trusted? Granted you got good validation results, but in the first place, we don't know the distribution of test set and what nocalls sound like in the test set. The only way to know is to submit. Either way, please keep me posted, will love to see if your nocall set was effective.",
      "votes": null
    },
    {
      "id": "997675",
      "postDate": "09/04/2020 07:02:08",
      "content": "<p><a href=\"https://www.kaggle.com/alanchn31\" target=\"_blank\">@alanchn31</a> yes sure I will update tomorrow </p>",
      "rawMarkdown": "alanchn31 yes sure I will update tomorrow",
      "votes": null
    },
    {
      "id": "997698",
      "postDate": "09/04/2020 07:25:05",
      "content": "<p>I'll try a threshold optimizer, can it help, I'm skeptical…</p>",
      "rawMarkdown": "I'll try a threshold optimizer, can it help, I'm skeptical...",
      "votes": null
    },
    {
      "id": "997706",
      "postDate": "09/04/2020 07:28:40",
      "content": "<p><a href=\"https://www.kaggle.com/gopidurgaprasad\" target=\"_blank\">@gopidurgaprasad</a>, if you have not already done so, next time, save and work with your notebooks in cpu mode, and then submit them in gpu mode from cpu mode, it helps the last days of gpu quota. Just open a notebook in GPU mode, burns many minutes. </p>",
      "rawMarkdown": "gopidurgaprasad, if you have not already done so, next time, save and work with your notebooks in cpu mode, and then submit them in gpu mode from cpu mode, it helps the last days of gpu quota. Just open a notebook in GPU mode, burns many minutes.",
      "votes": null
    },
    {
      "id": "997962",
      "postDate": "09/04/2020 11:19:56",
      "content": "<p>OK, makes sense because you add diversity to models.</p>",
      "rawMarkdown": "OK, makes sense because you add diversity to models.",
      "votes": null
    },
    {
      "id": "998337",
      "postDate": "09/04/2020 17:07:04",
      "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a>  thanks for reply. Looking forward to your team's solution after the competition ended 😄</p>",
      "rawMarkdown": "theoviel @kneroma  thanks for reply. Looking forward to your team's solution after the competition ended 😄",
      "votes": null
    },
    {
      "id": "998586",
      "postDate": "09/04/2020 20:59:15",
      "content": "<p>It feels like it would go towards overfitting to LB. I feel like 0.5 is ideal and I have seen a lot of high scoring people using 0.5 too.</p>",
      "rawMarkdown": "It feels like it would go towards overfitting to LB. I feel like 0.5 is ideal and I have seen a lot of high scoring people using 0.5 too.",
      "votes": null
    },
    {
      "id": "998587",
      "postDate": "09/04/2020 21:00:52",
      "content": "<p>Am I the only one here who had to struggle to find the threshold between 0.99, 0.999, 0.999… 😳</p>",
      "rawMarkdown": "Am I the only one here who had to struggle to find the threshold between 0.99, 0.999, 0.999... 😳",
      "votes": null
    },
    {
      "id": "998590",
      "postDate": "09/04/2020 21:03:38",
      "content": "<p>You should stick to one threshold first try to improve your model in general and when you are done with everything then you could tweak it a little to see if it changes majorly. But I think this whole threshold tweaking will just lead to overfitted public LB and significantly lower private LB.</p>",
      "rawMarkdown": "You should stick to one threshold first try to improve your model in general and when you are done with everything then you could tweak it a little to see if it changes majorly. But I think this whole threshold tweaking will just lead to overfitted public LB and significantly lower private LB.",
      "votes": null
    },
    {
      "id": "998752",
      "postDate": "09/05/2020 03:22:43",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> <a href=\"https://www.kaggle.com/kirderf\" target=\"_blank\">@kirderf</a> </p>\n<p>Here is the update for nocall model</p>\n<p><code>\"nocall\" if pred_proba &gt; 0.5 else \"none\"</code>    -&gt; LB : 0.427</p>\n<p>only with nocall LB max is 0.55 my model predicts LB: 0.427 with a simple threshold 0.5 </p>",
      "rawMarkdown": "cpmpml @kirderf \n\nHere is the update for nocall model\n\n` \"nocall\" if pred_proba > 0.5 else \"none\" `    -> LB : 0.427\n\nonly with nocall LB max is 0.55 my model predicts LB: 0.427 with a simple threshold 0.5",
      "votes": null
    },
    {
      "id": "998923",
      "postDate": "09/05/2020 07:44:34",
      "content": "<p>understood…</p>\n<p>but every model I trained for this competition (3-4 now) converge crazily and the predictions (sigmoid values) I get on sample tests are 1 or very close to 1 (like 0.9,0.99,0.999). </p>\n<p>and everyone in discussion forums is talking about 0.5,0.6,0.7.</p>\n<p>I use Keras and as far as my models are concerned I don't find any other significant difference from public starter notebooks other than that I use LeakyRelu instead of Relu</p>",
      "rawMarkdown": "understood...\n\nbut every model I trained for this competition (3-4 now) converge crazily and the predictions (sigmoid values) I get on sample tests are 1 or very close to 1 (like 0.9,0.99,0.999). \n\nand everyone in discussion forums is talking about 0.5,0.6,0.7.\n\nI use Keras and as far as my models are concerned I don't find any other significant difference from public starter notebooks other than that I use LeakyRelu instead of Relu",
      "votes": null
    },
    {
      "id": "998930",
      "postDate": "09/05/2020 07:51:14",
      "content": "<p><a href=\"https://www.kaggle.com/gopidurgaprasad\" target=\"_blank\">@gopidurgaprasad</a> Thanks for reporting.</p>\n<p>I guess we all want to see how this helps models that predict ebird_code as well.  You are probably going to work on it now, I'd be interested in whatever result you get.</p>",
      "rawMarkdown": "gopidurgaprasad Thanks for reporting.\n\nI guess we all want to see how this helps models that predict ebird_code as well.  You are probably going to work on it now, I'd be interested in whatever result you get.",
      "votes": null
    },
    {
      "id": "998931",
      "postDate": "09/05/2020 07:54:43",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a></p>\n<p>Yes, today I am going to combine with ebird_code model. </p>\n<p>Once I get that results I will update that one. </p>\n<p>How much your models predicting nocall on LB? </p>",
      "rawMarkdown": "cpmpml\n\nYes, today I am going to combine with ebird_code model. \n\nOnce I get that results I will update that one. \n\nHow much your models predicting nocall on LB?",
      "votes": null
    },
    {
      "id": "998951",
      "postDate": "09/05/2020 08:16:30",
      "content": "<p><a href=\"https://www.kaggle.com/vsvrp1995\" target=\"_blank\">@vsvrp1995</a> threshold should not be between 0.99, 0.999, 0.999.<br>\nIm a layman and from my understanding if your prob prediction of each class after sigmoid is so close to each other (like yr case 0.9,0.99,0.999),  your model actually failed to do the classification job and just blindly predict everything. The LB score improvement probably come from no call/aldfly only. </p>\n<p>I encounter similar situation before and this was due improper training and model architecture.</p>",
      "rawMarkdown": "vsvrp1995 threshold should not be between 0.99, 0.999, 0.999.\nIm a layman and from my understanding if your prob prediction of each class after sigmoid is so close to each other (like yr case 0.9,0.99,0.999),  your model actually failed to do the classification job and just blindly predict everything. The LB score improvement probably come from no call/aldfly only. \n \nI encounter similar situation before and this was due improper training and model architecture.",
      "votes": null
    },
    {
      "id": "998956",
      "postDate": "09/05/2020 08:21:38",
      "content": "<p>I have been experimenting with the example test audio since it is the only soundscape example we have for test.  Still looking for other soundscape audios we are allowed to use.  Think it is hard to consider what works with a validation set from train if the audio quality is better more focused.  But who knows what the hidden test set is truly like.  Or what the private LB will score!   What if most of the nocalls are in the public LB?  That would be harsh.</p>\n<p>Have tried different threshold strategies for each site since in theory hidden test should be 3 different North American locations, possibly different dates/seasons. Also tried a 2 step threshold to look at most likely species to be there then use 0.5 for them and raise the threshold for the rest.  That seems promising sort of.  Still working on it…</p>",
      "rawMarkdown": "I have been experimenting with the example test audio since it is the only soundscape example we have for test.  Still looking for other soundscape audios we are allowed to use.  Think it is hard to consider what works with a validation set from train if the audio quality is better more focused.  But who knows what the hidden test set is truly like.  Or what the private LB will score!   What if most of the nocalls are in the public LB?  That would be harsh.\n\nHave tried different threshold strategies for each site since in theory hidden test should be 3 different North American locations, possibly different dates/seasons. Also tried a 2 step threshold to look at most likely species to be there then use 0.5 for them and raise the threshold for the rest.  That seems promising sort of.  Still working on it...",
      "votes": null
    },
    {
      "id": "998969",
      "postDate": "09/05/2020 08:47:30",
      "content": "<blockquote>\n  <p>How much your models predicting nocall on LB? </p>\n</blockquote>\n<p>I did not submit a nocall model, I don't have one.</p>",
      "rawMarkdown": "> How much your models predicting nocall on LB? \n\nI did not submit a nocall model, I don't have one.",
      "votes": null
    },
    {
      "id": "998982",
      "postDate": "09/05/2020 08:54:32",
      "content": "<p><a href=\"https://www.kaggle.com/fiyeroleung\" target=\"_blank\">@fiyeroleung</a> </p>\n<p>I have this dataset for testing my code it has like 30 different species… these are 60 longest audio files of our training dataset.</p>\n<p>link: <a href=\"https://www.kaggle.com/vsvrp1995/faketest\" target=\"_blank\">https://www.kaggle.com/vsvrp1995/faketest</a><br>\n(created this to test memory usage, timeout, and other errors)</p>\n<p>it does a decent job of predicting these nearly 30 different species…<br>\nmany times it perfectly predicts existing positives but again with many false positives forcing me to further increase my threshold.</p>",
      "rawMarkdown": "fiyeroleung \n\nI have this dataset for testing my code it has like 30 different species... these are 60 longest audio files of our training dataset.\n\nlink: https://www.kaggle.com/vsvrp1995/faketest\n(created this to test memory usage, timeout, and other errors)\n\nit does a decent job of predicting these nearly 30 different species...\nmany times it perfectly predicts existing positives but again with many false positives forcing me to further increase my threshold.",
      "votes": null
    },
    {
      "id": "998991",
      "postDate": "09/05/2020 08:57:26",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> </p>\n<p>From the ebird_code model you can do that.<br>\nThen you know how much your ebird_code model able to predict nocalls, give 'none' for all ebird_codes then you only have 'nocall'</p>\n<p>We already know that 'nocall' max on LB</p>",
      "rawMarkdown": "cpmpml \n\nFrom the ebird_code model you can do that.\nThen you know how much your ebird_code model able to predict nocalls, give 'none' for all ebird_codes then you only have 'nocall'\n\nWe already know that 'nocall' max on LB",
      "votes": null
    },
    {
      "id": "999127",
      "postDate": "09/05/2020 11:19:04",
      "content": "<p>I will not waste a sub for that ;)  If we had more subs a day I would do it probably.</p>",
      "rawMarkdown": "I will not waste a sub for that ;)  If we had more subs a day I would do it probably.",
      "votes": null
    },
    {
      "id": "1000147",
      "postDate": "09/06/2020 10:31:55",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a><br>\nI was quite demoralized when nothing seemed to work for the last 2 days and didn't make submissions 2 days in a row.</p>\n<p>Last night, I saw your post about preserving submissions and decided to train and make a submission based on an idea …AND it worked :)</p>\n<p>Strange inspiration, but Thank you :)<br>\nAlso, I see you Jump many hundred ranks too. Good luck to you too</p>",
      "rawMarkdown": "cpmpml\nI was quite demoralized when nothing seemed to work for the last 2 days and didn't make submissions 2 days in a row.\n\nLast night, I saw your post about preserving submissions and decided to train and make a submission based on an idea ...AND it worked :)\n\nStrange inspiration, but Thank you :)\nAlso, I see you Jump many hundred ranks too. Good luck to you too",
      "votes": null
    },
    {
      "id": "1000158",
      "postDate": "09/06/2020 10:51:25",
      "content": "<p>I am glad my comment helped you :)</p>\n<p>A competition is only lost or won at the end.  It is normal to be demoralized form time to time, but there is always hope till the end.  In a recent competition I moved from 49 to 15 place at the very last minute of the competition…</p>\n<blockquote>\n  <p>I see you Jump many hundred ranks too</p>\n</blockquote>\n<p>We jump a lot once we past all the people who just submitted the best public notebook.  This notebook distorts the LB quite  a bit.  </p>\n<p>Good luck to you as well!</p>",
      "rawMarkdown": "I am glad my comment helped you :)\n\nA competition is only lost or won at the end.  It is normal to be demoralized form time to time, but there is always hope till the end.  In a recent competition I moved from 49 to 15 place at the very last minute of the competition...\n\n>  I see you Jump many hundred ranks too\n\nWe jump a lot once we past all the people who just submitted the best public notebook.  This notebook distorts the LB quite  a bit.  \n\nGood luck to you as well!",
      "votes": null
    },
    {
      "id": "1000185",
      "postDate": "09/06/2020 11:13:20",
      "content": "<blockquote>\n  <p>A competition is only lost or won at the end. </p>\n</blockquote>\n<p>Well said - thanks for your kind words. </p>",
      "rawMarkdown": "> A competition is only lost or won at the end. \n\nWell said - thanks for your kind words.",
      "votes": null
    },
    {
      "id": "1000617",
      "postDate": "09/06/2020 16:56:36",
      "content": "<p>The two  steps approach sounds great ! Don't mind sharing more about it whenever you feel it !</p>",
      "rawMarkdown": "The two  steps approach sounds great ! Don't mind sharing more about it whenever you feel it !",
      "votes": null
    },
    {
      "id": "1001378",
      "postDate": "09/07/2020 09:45:42",
      "content": "<p>I guess it will depend on your model and thresholds working for it, but the idea with the example test audio I tried was to start with a high threshold to see what ebird codes are being predicted that are not nocall. So for me that was ['bewwre', 'stejay', 'westan', 'olsfly'].  <br>\nThen look at the counts for them so only 1 is maybe not a good candidate for reduced threshold.  But where counts as % of non nocalls is X  or ebird code appears in sequential 5 secs or alternate 5 secs, these could be good candidates.  <br>\nSo for me, that was  'stejay' and 'westan'.  So dropping the threshold for these 2 to 0.5 or even 0.3 and predicting again picks up a few more of these and improves the F1 score, (best was 0.538).  Sometimes there are several things calling at the same time or using the example metadata start/end to check, the call may be a small fragment in the 5 sec so maybe a lower threshold finds these. Hard to say if it can work for hidden test also.  Probably needs to be per site and audio clip - the data page says 150 recordings roughly 10 mins long.  The example audio clips are around 7 and 5 mins so hoping they are comparable for this checking.  Have you tried predictions for example test audio? </p>\n<p>Feel free also to share about getting to 3rd!!   Good luck!</p>",
      "rawMarkdown": "I guess it will depend on your model and thresholds working for it, but the idea with the example test audio I tried was to start with a high threshold to see what ebird codes are being predicted that are not nocall. So for me that was ['bewwre', 'stejay', 'westan', 'olsfly'].  \nThen look at the counts for them so only 1 is maybe not a good candidate for reduced threshold.  But where counts as % of non nocalls is X  or ebird code appears in sequential 5 secs or alternate 5 secs, these could be good candidates.  \nSo for me, that was  'stejay' and 'westan'.  So dropping the threshold for these 2 to 0.5 or even 0.3 and predicting again picks up a few more of these and improves the F1 score, (best was 0.538).  Sometimes there are several things calling at the same time or using the example metadata start/end to check, the call may be a small fragment in the 5 sec so maybe a lower threshold finds these. Hard to say if it can work for hidden test also.  Probably needs to be per site and audio clip - the data page says 150 recordings roughly 10 mins long.  The example audio clips are around 7 and 5 mins so hoping they are comparable for this checking.  Have you tried predictions for example test audio? \n \nFeel free also to share about getting to 3rd!!   Good luck!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 997031,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "09/03/2020 17:33:39",
      "content": "<p>Clearly, predicting nocall is key.  But I have not found how to do it yet ;)</p>",
      "votes": null,
      "replies": [
        {
          "id": 997226,
          "author_name": "kneroma",
          "author_url": "",
          "post_date": "09/03/2020 20:38:52",
          "content": "<p>Don't mind sharing whith me if you ever find how to do it 😄😄😄</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 997410,
          "author_name": "gopidurgaprasad",
          "author_url": "",
          "post_date": "09/04/2020 03:12:16",
          "content": "<p>Train simple model for call or nocall </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 997458,
          "author_name": "alanchn31",
          "author_url": "",
          "post_date": "09/04/2020 04:16:26",
          "content": "<p><a href=\"https://www.kaggle.com/gopidurgaprasad\" target=\"_blank\">@gopidurgaprasad</a> , theoretically that is sound but it is challenging distinguishing call or nocall with no strong labels. Coming up with nocall labels from scratch is daunting with no potentially proven results, as the only validation you have of your nocall set is through submitting to LB</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 997499,
          "author_name": "gopidurgaprasad",
          "author_url": "",
          "post_date": "09/04/2020 04:52:50",
          "content": "<p><a href=\"https://www.kaggle.com/alanchn31\" target=\"_blank\">@alanchn31</a>  i trained nocall/call model it gives me 0.97+ AUC</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 997500,
          "author_name": "gopidurgaprasad",
          "author_url": "",
          "post_date": "09/04/2020 04:52:51",
          "content": "<p><a href=\"https://www.kaggle.com/alanchn31\" target=\"_blank\">@alanchn31</a>  i trained nocall/call model it gives me 0.97+ AUC</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 997558,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/04/2020 05:31:23",
          "content": "<p><a href=\"https://www.kaggle.com/gopidurgaprasad\" target=\"_blank\">@gopidurgaprasad</a> Does it help you for test data?  Also, how do you train and validate a nocall model given train data has zero nocall clips.  Sure, you can extract nocall parts, but how do you know you do it right?  The issue is to get reliable nocall labels.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 997577,
          "author_name": "gopidurgaprasad",
          "author_url": "",
          "post_date": "09/04/2020 05:48:24",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>  hai,</p>\n<p>i trained that nocall model 5 folds it gives 0.976+ AUC, just finesed 5th fold.<br>\nI am not yet submitted, i don't have GPU for this week, i will going to submit it tomorrow.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 997586,
          "author_name": "gopidurgaprasad",
          "author_url": "",
          "post_date": "09/04/2020 05:52:59",
          "content": "<p>call/nocall idea is like we alredy have 5 fold model trained on randam 5sec right.<br>\nstep 1 : create &lt;= 1sec clips<br>\nstep 2: predict 5 fold model on 1sec clips<br>\nstep 3: find threshold that based on your model on 1 sec clips<br>\nstep 4: below threshold 1sec clips are probably noise</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 997603,
          "author_name": "gopidurgaprasad",
          "author_url": "",
          "post_date": "09/04/2020 06:07:15",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> </p>\n<p>the same idea a lot of people doing it inference time, but not at training time.</p>\n<p>my idea for nocall model most of the same, but not exactly same.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 997620,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/04/2020 06:22:14",
          "content": "<blockquote>\n  <p>find threshold that based on your model on 1 sec clips</p>\n</blockquote>\n<p>find threshold for what?  How do you decide that the threshold is the right threshold?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 997628,
          "author_name": "gopidurgaprasad",
          "author_url": "",
          "post_date": "09/04/2020 06:26:07",
          "content": "<p>the idea is like how we do at inference time, the same idea applies for one-second clip</p>\n<p>for example : </p>\n<p>one-sec clip audio predictions are like - [0.1, 0.2, 0.0, …..]</p>\n<p>all class predictions are very low my threshold &lt;0.3. if all class predictions &lt;0.3 that most probably noise/nocall  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 997633,
          "author_name": "gopidurgaprasad",
          "author_url": "",
          "post_date": "09/04/2020 06:28:42",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> </p>\n<p>again those are noise labels, using a student-teacher training method, incrementally we remove noise labels</p>\n<p>one more idea is like do clustering for noise labels</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 997644,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/04/2020 06:35:45",
          "content": "<p>You don't answer my question: how do you decide that the 0.3 threshold is the right threshold?</p>\n<p>I am not criticizing, I am trying to understand what you do ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 997654,
          "author_name": "gopidurgaprasad",
          "author_url": "",
          "post_date": "09/04/2020 06:43:09",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> </p>\n<p>sorry about that.</p>\n<p>I just assuming 0.3 is my threshold, as like others assuming 0.5 at inference time.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 997656,
          "author_name": "gopidurgaprasad",
          "author_url": "",
          "post_date": "09/04/2020 06:44:27",
          "content": "<p><code>again those are noise labels, using a student-teacher training method, incrementally we remove noise labels</code></p>\n<p>after this process, I got 0.97+ AUC.</p>\n<p>for me it's working.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 997671,
          "author_name": "alanchn31",
          "author_url": "",
          "post_date": "09/04/2020 06:58:26",
          "content": "<p><a href=\"https://www.kaggle.com/gopidurgaprasad\" target=\"_blank\">@gopidurgaprasad</a> That's the point I am trying to make, can current results be trusted? Granted you got good validation results, but in the first place, we don't know the distribution of test set and what nocalls sound like in the test set. The only way to know is to submit. Either way, please keep me posted, will love to see if your nocall set was effective.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 997675,
          "author_name": "gopidurgaprasad",
          "author_url": "",
          "post_date": "09/04/2020 07:02:08",
          "content": "<p><a href=\"https://www.kaggle.com/alanchn31\" target=\"_blank\">@alanchn31</a> yes sure I will update tomorrow </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 997698,
          "author_name": "kirderf",
          "author_url": "",
          "post_date": "09/04/2020 07:25:05",
          "content": "<p>I'll try a threshold optimizer, can it help, I'm skeptical…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 997706,
          "author_name": "kirderf",
          "author_url": "",
          "post_date": "09/04/2020 07:28:40",
          "content": "<p><a href=\"https://www.kaggle.com/gopidurgaprasad\" target=\"_blank\">@gopidurgaprasad</a>, if you have not already done so, next time, save and work with your notebooks in cpu mode, and then submit them in gpu mode from cpu mode, it helps the last days of gpu quota. Just open a notebook in GPU mode, burns many minutes. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 998586,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "09/04/2020 20:59:15",
          "content": "<p>It feels like it would go towards overfitting to LB. I feel like 0.5 is ideal and I have seen a lot of high scoring people using 0.5 too.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 998752,
          "author_name": "gopidurgaprasad",
          "author_url": "",
          "post_date": "09/05/2020 03:22:43",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> <a href=\"https://www.kaggle.com/kirderf\" target=\"_blank\">@kirderf</a> </p>\n<p>Here is the update for nocall model</p>\n<p><code>\"nocall\" if pred_proba &gt; 0.5 else \"none\"</code>    -&gt; LB : 0.427</p>\n<p>only with nocall LB max is 0.55 my model predicts LB: 0.427 with a simple threshold 0.5 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 998930,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/05/2020 07:51:14",
          "content": "<p><a href=\"https://www.kaggle.com/gopidurgaprasad\" target=\"_blank\">@gopidurgaprasad</a> Thanks for reporting.</p>\n<p>I guess we all want to see how this helps models that predict ebird_code as well.  You are probably going to work on it now, I'd be interested in whatever result you get.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 998931,
          "author_name": "gopidurgaprasad",
          "author_url": "",
          "post_date": "09/05/2020 07:54:43",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a></p>\n<p>Yes, today I am going to combine with ebird_code model. </p>\n<p>Once I get that results I will update that one. </p>\n<p>How much your models predicting nocall on LB? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 998969,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/05/2020 08:47:30",
          "content": "<blockquote>\n  <p>How much your models predicting nocall on LB? </p>\n</blockquote>\n<p>I did not submit a nocall model, I don't have one.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 998991,
          "author_name": "gopidurgaprasad",
          "author_url": "",
          "post_date": "09/05/2020 08:57:26",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> </p>\n<p>From the ebird_code model you can do that.<br>\nThen you know how much your ebird_code model able to predict nocalls, give 'none' for all ebird_codes then you only have 'nocall'</p>\n<p>We already know that 'nocall' max on LB</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 999127,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/05/2020 11:19:04",
          "content": "<p>I will not waste a sub for that ;)  If we had more subs a day I would do it probably.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1000147,
          "author_name": "watzisname",
          "author_url": "",
          "post_date": "09/06/2020 10:31:55",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a><br>\nI was quite demoralized when nothing seemed to work for the last 2 days and didn't make submissions 2 days in a row.</p>\n<p>Last night, I saw your post about preserving submissions and decided to train and make a submission based on an idea …AND it worked :)</p>\n<p>Strange inspiration, but Thank you :)<br>\nAlso, I see you Jump many hundred ranks too. Good luck to you too</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1000158,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/06/2020 10:51:25",
          "content": "<p>I am glad my comment helped you :)</p>\n<p>A competition is only lost or won at the end.  It is normal to be demoralized form time to time, but there is always hope till the end.  In a recent competition I moved from 49 to 15 place at the very last minute of the competition…</p>\n<blockquote>\n  <p>I see you Jump many hundred ranks too</p>\n</blockquote>\n<p>We jump a lot once we past all the people who just submitted the best public notebook.  This notebook distorts the LB quite  a bit.  </p>\n<p>Good luck to you as well!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1000185,
          "author_name": "watzisname",
          "author_url": "",
          "post_date": "09/06/2020 11:13:20",
          "content": "<blockquote>\n  <p>A competition is only lost or won at the end. </p>\n</blockquote>\n<p>Well said - thanks for your kind words. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 997090,
      "author_name": "vladimirsydor",
      "author_url": "",
      "post_date": "09/03/2020 18:27:31",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1690820%2F66e0855cc68f73c3c2968fcf76929d74%2F4dqozp.jpg?generation=1599157649395881&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 997187,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "09/03/2020 19:59:10",
          "content": "<p>0.5 goes brrrrrrrr</p>",
          "votes": null,
          "replies": [
            {
              "id": 997203,
              "author_name": "mpware",
              "author_url": "",
              "post_date": "09/03/2020 20:13:24",
              "content": "<p>0.5 for best model here.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 997141,
      "author_name": "fiyeroleung",
      "author_url": "",
      "post_date": "09/03/2020 19:18:28",
      "content": "<p>may I know your current LB score is using what threshold value? thanks</p>",
      "votes": null,
      "replies": [
        {
          "id": 997189,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "09/03/2020 19:59:47",
          "content": "<p>0.5, we didn't try other values for this specific model</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 997220,
          "author_name": "kneroma",
          "author_url": "",
          "post_date": "09/03/2020 20:35:32",
          "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> is right since we are so afraid of the <strong>Threshold</strong> devil 😈😈</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 998337,
          "author_name": "fiyeroleung",
          "author_url": "",
          "post_date": "09/04/2020 17:07:04",
          "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> <a href=\"https://www.kaggle.com/kneroma\" target=\"_blank\">@kneroma</a>  thanks for reply. Looking forward to your team's solution after the competition ended 😄</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 997171,
      "author_name": "fffrrt",
      "author_url": "",
      "post_date": "09/03/2020 19:50:03",
      "content": "<blockquote>\n  <p>This whole competition could just be a matter of choosing the right threshold ! We're trying experiments to get ride of it and obtain a threshold-agnostic model but till now, we get no consistant results …</p>\n</blockquote>\n<p>I think you are hitting your head on a wall instead of walking around it.</p>\n<blockquote>\n  <p>And you, how do you choose your threshold, if you're using any !?</p>\n</blockquote>\n<p>Write down all model answers on validation fold, create two 100-element arrays (each element for each % of confidence).<br>\nFor each model answer add +1 to \"total answers\" array at position of confidence of the answer (total_answers[35] = total_answers[35] + 1 if model output was \"whateverbird: 0.3572), add +1 to \"corrects\" array at the same position if model answer was correct.</p>\n<p>At the end you would get two arrays shaped like this:</p>\n<pre><code>corrects = [0,0,0,2,1,3,2,4,6,5,4,7,10,16,8,8,13,13,21,19,18,14,15,20,19,20,24,26,25,17,37,27,22,22,36,20,21,27,21,25,27,16,34,21,23,31,23,28,38,33,28,23,23,23,29,13,25,27,19,18,28,14,21,18,23,17,16,15,26,23,23,15,20,18,22,17,15,15,21,10,18,21,10,16,17,16,25,11,18,23,14,12,11,20,17,6,22,16,12,15]\ntotal = [1,16,44,62,72,81,60,66,77,68,71,81,86,86,76,69,74,70,81,81,73,71,65,67,61,67,79,62,61,63,73,59,56,48,75,51,59,55,52,52,50,31,55,55,48,59,35,42,45,44,49,38,39,39,37,32,34,37,30,24,35,22,28,27,29,20,20,20,32,26,23,17,22,23,26,21,19,17,24,12,19,23,11,17,18,16,25,12,19,23,14,13,12,20,17,7,23,16,14,16]\n</code></pre>\n<p>Plot this and get this nice graph: ![<a href=\"https://i.imgur.com/Nb3Br1Z.png\" target=\"_blank\">https://i.imgur.com/Nb3Br1Z.png</a>]</p>\n<p>Then use your trusted eyeballs and pick an accuracy number roughly at score you want to achieve (lets say 0.6). In that case you would use confidence around 0.5, because this is where accuracy reached 0.6. For that model i tried 0.6 and 0.5, and 0.5 got me 0.002 better score.</p>\n<p>Another example - model from this graph was tried on 0.35 and 0.3, and it scored the same in both cases, so this picking strategy seems to work: ![<a href=\"https://i.imgur.com/LS9hlH8.png\" target=\"_blank\">https://i.imgur.com/LS9hlH8.png</a>]</p>\n<p>Keep in mind that threshold optimization does not heavily influence your score, as shown in <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/167263\" target=\"_blank\">https://www.kaggle.com/c/birdsong-recognition/discussion/167263</a> - with most models you have 0.2 threshold range where your score is in ~0.002 result from best possible score.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 997328,
      "author_name": "nvnnghia",
      "author_url": "",
      "post_date": "09/04/2020 00:19:43",
      "content": "<p>For me, I use 0.5 for models trained with label smoothing, and 0.6 for models trained without label smoothing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 997548,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/04/2020 05:29:16",
          "content": "<p>is label smoothing helping you?  I haven't tried yet.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 997631,
          "author_name": "nvnnghia",
          "author_url": "",
          "post_date": "09/04/2020 06:28:02",
          "content": "<p>For single model, it only helps my CV score. However, it helps a lot when I use model ensembling.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 997962,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "09/04/2020 11:19:56",
          "content": "<p>OK, makes sense because you add diversity to models.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 998587,
      "author_name": "vsvrp1995",
      "author_url": "",
      "post_date": "09/04/2020 21:00:52",
      "content": "<p>Am I the only one here who had to struggle to find the threshold between 0.99, 0.999, 0.999… 😳</p>",
      "votes": null,
      "replies": [
        {
          "id": 998590,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "09/04/2020 21:03:38",
          "content": "<p>You should stick to one threshold first try to improve your model in general and when you are done with everything then you could tweak it a little to see if it changes majorly. But I think this whole threshold tweaking will just lead to overfitted public LB and significantly lower private LB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 998923,
          "author_name": "vsvrp1995",
          "author_url": "",
          "post_date": "09/05/2020 07:44:34",
          "content": "<p>understood…</p>\n<p>but every model I trained for this competition (3-4 now) converge crazily and the predictions (sigmoid values) I get on sample tests are 1 or very close to 1 (like 0.9,0.99,0.999). </p>\n<p>and everyone in discussion forums is talking about 0.5,0.6,0.7.</p>\n<p>I use Keras and as far as my models are concerned I don't find any other significant difference from public starter notebooks other than that I use LeakyRelu instead of Relu</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 998951,
          "author_name": "fiyeroleung",
          "author_url": "",
          "post_date": "09/05/2020 08:16:30",
          "content": "<p><a href=\"https://www.kaggle.com/vsvrp1995\" target=\"_blank\">@vsvrp1995</a> threshold should not be between 0.99, 0.999, 0.999.<br>\nIm a layman and from my understanding if your prob prediction of each class after sigmoid is so close to each other (like yr case 0.9,0.99,0.999),  your model actually failed to do the classification job and just blindly predict everything. The LB score improvement probably come from no call/aldfly only. </p>\n<p>I encounter similar situation before and this was due improper training and model architecture.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 998982,
          "author_name": "vsvrp1995",
          "author_url": "",
          "post_date": "09/05/2020 08:54:32",
          "content": "<p><a href=\"https://www.kaggle.com/fiyeroleung\" target=\"_blank\">@fiyeroleung</a> </p>\n<p>I have this dataset for testing my code it has like 30 different species… these are 60 longest audio files of our training dataset.</p>\n<p>link: <a href=\"https://www.kaggle.com/vsvrp1995/faketest\" target=\"_blank\">https://www.kaggle.com/vsvrp1995/faketest</a><br>\n(created this to test memory usage, timeout, and other errors)</p>\n<p>it does a decent job of predicting these nearly 30 different species…<br>\nmany times it perfectly predicts existing positives but again with many false positives forcing me to further increase my threshold.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 998956,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "09/05/2020 08:21:38",
      "content": "<p>I have been experimenting with the example test audio since it is the only soundscape example we have for test.  Still looking for other soundscape audios we are allowed to use.  Think it is hard to consider what works with a validation set from train if the audio quality is better more focused.  But who knows what the hidden test set is truly like.  Or what the private LB will score!   What if most of the nocalls are in the public LB?  That would be harsh.</p>\n<p>Have tried different threshold strategies for each site since in theory hidden test should be 3 different North American locations, possibly different dates/seasons. Also tried a 2 step threshold to look at most likely species to be there then use 0.5 for them and raise the threshold for the rest.  That seems promising sort of.  Still working on it…</p>",
      "votes": null,
      "replies": [
        {
          "id": 1000617,
          "author_name": "kneroma",
          "author_url": "",
          "post_date": "09/06/2020 16:56:36",
          "content": "<p>The two  steps approach sounds great ! Don't mind sharing more about it whenever you feel it !</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1001378,
          "author_name": "something4kag",
          "author_url": "",
          "post_date": "09/07/2020 09:45:42",
          "content": "<p>I guess it will depend on your model and thresholds working for it, but the idea with the example test audio I tried was to start with a high threshold to see what ebird codes are being predicted that are not nocall. So for me that was ['bewwre', 'stejay', 'westan', 'olsfly'].  <br>\nThen look at the counts for them so only 1 is maybe not a good candidate for reduced threshold.  But where counts as % of non nocalls is X  or ebird code appears in sequential 5 secs or alternate 5 secs, these could be good candidates.  <br>\nSo for me, that was  'stejay' and 'westan'.  So dropping the threshold for these 2 to 0.5 or even 0.3 and predicting again picks up a few more of these and improves the F1 score, (best was 0.538).  Sometimes there are several things calling at the same time or using the example metadata start/end to check, the call may be a small fragment in the 5 sec so maybe a lower threshold finds these. Hard to say if it can work for hidden test also.  Probably needs to be per site and audio clip - the data page says 150 recordings roughly 10 mins long.  The example audio clips are around 7 and 5 mins so hoping they are comparable for this checking.  Have you tried predictions for example test audio? </p>\n<p>Feel free also to share about getting to 3rd!!   Good luck!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "996900": "Needless to recall how hard is it to build a robust cross-validation  scheme for this competition. And when it comes to the choice of the right **threshold** for the **nocall** event detection, things get more harder 😈😔.\n\nDuring my experiments, I found that right values for the **thresold** hyperparameter depends not only on the model but also on the dataset. So, even a correctly cross-validated threshold for the training set could be useless for the test one as both datasets don't come from  same sources.\n\nFor example, when using Resnest-like models, the right threshold is around **0.6**, but, if I change the training precedure, the model becomes too much confident and I need to raise the threshold to **0.8** in order to get any interesting score. \n\nThis whole competition could just be a matter of choosing the right threshold ! We're trying experiments to get ride of it and obtain a threshold-agnostic model but till now, we get no consistant  results ...\n\nAnd you, how do you  choose your **threshold**, if you're using any !?",
    "997031": "Clearly, predicting nocall is key.  But I have not found how to do it yet ;)",
    "997090": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1690820%2F66e0855cc68f73c3c2968fcf76929d74%2F4dqozp.jpg?generation=1599157649395881&alt=media)",
    "997141": "may I know your current LB score is using what threshold value? thanks",
    "997171": "> This whole competition could just be a matter of choosing the right threshold ! We're trying experiments to get ride of it and obtain a threshold-agnostic model but till now, we get no consistant results …\n\nI think you are hitting your head on a wall instead of walking around it.\n\n> And you, how do you choose your threshold, if you're using any !?\n\nWrite down all model answers on validation fold, create two 100-element arrays (each element for each % of confidence).\nFor each model answer add +1 to \"total answers\" array at position of confidence of the answer (total_answers[35] = total_answers[35] + 1 if model output was \"whateverbird: 0.3572), add +1 to \"corrects\" array at the same position if model answer was correct.\n\nAt the end you would get two arrays shaped like this:\n```\ncorrects = [0,0,0,2,1,3,2,4,6,5,4,7,10,16,8,8,13,13,21,19,18,14,15,20,19,20,24,26,25,17,37,27,22,22,36,20,21,27,21,25,27,16,34,21,23,31,23,28,38,33,28,23,23,23,29,13,25,27,19,18,28,14,21,18,23,17,16,15,26,23,23,15,20,18,22,17,15,15,21,10,18,21,10,16,17,16,25,11,18,23,14,12,11,20,17,6,22,16,12,15]\ntotal = [1,16,44,62,72,81,60,66,77,68,71,81,86,86,76,69,74,70,81,81,73,71,65,67,61,67,79,62,61,63,73,59,56,48,75,51,59,55,52,52,50,31,55,55,48,59,35,42,45,44,49,38,39,39,37,32,34,37,30,24,35,22,28,27,29,20,20,20,32,26,23,17,22,23,26,21,19,17,24,12,19,23,11,17,18,16,25,12,19,23,14,13,12,20,17,7,23,16,14,16]\n```\nPlot this and get this nice graph: ![https://i.imgur.com/Nb3Br1Z.png]\n\nThen use your trusted eyeballs and pick an accuracy number roughly at score you want to achieve (lets say 0.6). In that case you would use confidence around 0.5, because this is where accuracy reached 0.6. For that model i tried 0.6 and 0.5, and 0.5 got me 0.002 better score.\n\nAnother example - model from this graph was tried on 0.35 and 0.3, and it scored the same in both cases, so this picking strategy seems to work: ![https://i.imgur.com/LS9hlH8.png]\n\nKeep in mind that threshold optimization does not heavily influence your score, as shown in https://www.kaggle.com/c/birdsong-recognition/discussion/167263 - with most models you have 0.2 threshold range where your score is in ~0.002 result from best possible score.",
    "997187": "0.5 goes brrrrrrrr",
    "997189": "0.5, we didn't try other values for this specific model",
    "997203": "0.5 for best model here.",
    "997220": "theoviel is right since we are so afraid of the **Threshold** devil 😈😈",
    "997226": "Don't mind sharing whith me if you ever find how to do it 😄😄😄",
    "997328": "For me, I use 0.5 for models trained with label smoothing, and 0.6 for models trained without label smoothing.",
    "997410": "Train simple model for call or nocall",
    "997458": "gopidurgaprasad , theoretically that is sound but it is challenging distinguishing call or nocall with no strong labels. Coming up with nocall labels from scratch is daunting with no potentially proven results, as the only validation you have of your nocall set is through submitting to LB",
    "997499": "alanchn31  i trained nocall/call model it gives me 0.97+ AUC",
    "997500": "alanchn31  i trained nocall/call model it gives me 0.97+ AUC",
    "997548": "is label smoothing helping you?  I haven't tried yet.",
    "997558": "gopidurgaprasad Does it help you for test data?  Also, how do you train and validate a nocall model given train data has zero nocall clips.  Sure, you can extract nocall parts, but how do you know you do it right?  The issue is to get reliable nocall labels.",
    "997577": "cpmpml  hai,\n\ni trained that nocall model 5 folds it gives 0.976+ AUC, just finesed 5th fold.\nI am not yet submitted, i don't have GPU for this week, i will going to submit it tomorrow.",
    "997586": "call/nocall idea is like we alredy have 5 fold model trained on randam 5sec right.\nstep 1 : create <= 1sec clips\nstep 2: predict 5 fold model on 1sec clips\nstep 3: find threshold that based on your model on 1 sec clips\nstep 4: below threshold 1sec clips are probably noise",
    "997603": "cpmpml \n\nthe same idea a lot of people doing it inference time, but not at training time.\n\nmy idea for nocall model most of the same, but not exactly same.",
    "997620": "> find threshold that based on your model on 1 sec clips\n\nfind threshold for what?  How do you decide that the threshold is the right threshold?",
    "997628": "the idea is like how we do at inference time, the same idea applies for one-second clip\n\nfor example : \n\none-sec clip audio predictions are like - [0.1, 0.2, 0.0, .....]\n\nall class predictions are very low my threshold <0.3. if all class predictions <0.3 that most probably noise/nocall",
    "997631": "For single model, it only helps my CV score. However, it helps a lot when I use model ensembling.",
    "997633": "cpmpml \n\nagain those are noise labels, using a student-teacher training method, incrementally we remove noise labels\n\none more idea is like do clustering for noise labels",
    "997644": "You don't answer my question: how do you decide that the 0.3 threshold is the right threshold?\n\nI am not criticizing, I am trying to understand what you do ;)",
    "997654": "cpmpml \n\nsorry about that.\n\nI just assuming 0.3 is my threshold, as like others assuming 0.5 at inference time.",
    "997656": "`again those are noise labels, using a student-teacher training method, incrementally we remove noise labels`\n\nafter this process, I got 0.97+ AUC.\n\nfor me it's working.",
    "997671": "gopidurgaprasad That's the point I am trying to make, can current results be trusted? Granted you got good validation results, but in the first place, we don't know the distribution of test set and what nocalls sound like in the test set. The only way to know is to submit. Either way, please keep me posted, will love to see if your nocall set was effective.",
    "997675": "alanchn31 yes sure I will update tomorrow",
    "997698": "I'll try a threshold optimizer, can it help, I'm skeptical...",
    "997706": "gopidurgaprasad, if you have not already done so, next time, save and work with your notebooks in cpu mode, and then submit them in gpu mode from cpu mode, it helps the last days of gpu quota. Just open a notebook in GPU mode, burns many minutes.",
    "997962": "OK, makes sense because you add diversity to models.",
    "998337": "theoviel @kneroma  thanks for reply. Looking forward to your team's solution after the competition ended 😄",
    "998586": "It feels like it would go towards overfitting to LB. I feel like 0.5 is ideal and I have seen a lot of high scoring people using 0.5 too.",
    "998587": "Am I the only one here who had to struggle to find the threshold between 0.99, 0.999, 0.999... 😳",
    "998590": "You should stick to one threshold first try to improve your model in general and when you are done with everything then you could tweak it a little to see if it changes majorly. But I think this whole threshold tweaking will just lead to overfitted public LB and significantly lower private LB.",
    "998752": "cpmpml @kirderf \n\nHere is the update for nocall model\n\n` \"nocall\" if pred_proba > 0.5 else \"none\" `    -> LB : 0.427\n\nonly with nocall LB max is 0.55 my model predicts LB: 0.427 with a simple threshold 0.5",
    "998923": "understood...\n\nbut every model I trained for this competition (3-4 now) converge crazily and the predictions (sigmoid values) I get on sample tests are 1 or very close to 1 (like 0.9,0.99,0.999). \n\nand everyone in discussion forums is talking about 0.5,0.6,0.7.\n\nI use Keras and as far as my models are concerned I don't find any other significant difference from public starter notebooks other than that I use LeakyRelu instead of Relu",
    "998930": "gopidurgaprasad Thanks for reporting.\n\nI guess we all want to see how this helps models that predict ebird_code as well.  You are probably going to work on it now, I'd be interested in whatever result you get.",
    "998931": "cpmpml\n\nYes, today I am going to combine with ebird_code model. \n\nOnce I get that results I will update that one. \n\nHow much your models predicting nocall on LB?",
    "998951": "vsvrp1995 threshold should not be between 0.99, 0.999, 0.999.\nIm a layman and from my understanding if your prob prediction of each class after sigmoid is so close to each other (like yr case 0.9,0.99,0.999),  your model actually failed to do the classification job and just blindly predict everything. The LB score improvement probably come from no call/aldfly only. \n \nI encounter similar situation before and this was due improper training and model architecture.",
    "998956": "I have been experimenting with the example test audio since it is the only soundscape example we have for test.  Still looking for other soundscape audios we are allowed to use.  Think it is hard to consider what works with a validation set from train if the audio quality is better more focused.  But who knows what the hidden test set is truly like.  Or what the private LB will score!   What if most of the nocalls are in the public LB?  That would be harsh.\n\nHave tried different threshold strategies for each site since in theory hidden test should be 3 different North American locations, possibly different dates/seasons. Also tried a 2 step threshold to look at most likely species to be there then use 0.5 for them and raise the threshold for the rest.  That seems promising sort of.  Still working on it...",
    "998969": "> How much your models predicting nocall on LB? \n\nI did not submit a nocall model, I don't have one.",
    "998982": "fiyeroleung \n\nI have this dataset for testing my code it has like 30 different species... these are 60 longest audio files of our training dataset.\n\nlink: https://www.kaggle.com/vsvrp1995/faketest\n(created this to test memory usage, timeout, and other errors)\n\nit does a decent job of predicting these nearly 30 different species...\nmany times it perfectly predicts existing positives but again with many false positives forcing me to further increase my threshold.",
    "998991": "cpmpml \n\nFrom the ebird_code model you can do that.\nThen you know how much your ebird_code model able to predict nocalls, give 'none' for all ebird_codes then you only have 'nocall'\n\nWe already know that 'nocall' max on LB",
    "999127": "I will not waste a sub for that ;)  If we had more subs a day I would do it probably.",
    "1000147": "cpmpml\nI was quite demoralized when nothing seemed to work for the last 2 days and didn't make submissions 2 days in a row.\n\nLast night, I saw your post about preserving submissions and decided to train and make a submission based on an idea ...AND it worked :)\n\nStrange inspiration, but Thank you :)\nAlso, I see you Jump many hundred ranks too. Good luck to you too",
    "1000158": "I am glad my comment helped you :)\n\nA competition is only lost or won at the end.  It is normal to be demoralized form time to time, but there is always hope till the end.  In a recent competition I moved from 49 to 15 place at the very last minute of the competition...\n\n>  I see you Jump many hundred ranks too\n\nWe jump a lot once we past all the people who just submitted the best public notebook.  This notebook distorts the LB quite  a bit.  \n\nGood luck to you as well!",
    "1000185": "> A competition is only lost or won at the end. \n\nWell said - thanks for your kind words.",
    "1000617": "The two  steps approach sounds great ! Don't mind sharing more about it whenever you feel it !",
    "1001378": "I guess it will depend on your model and thresholds working for it, but the idea with the example test audio I tried was to start with a high threshold to see what ebird codes are being predicted that are not nocall. So for me that was ['bewwre', 'stejay', 'westan', 'olsfly'].  \nThen look at the counts for them so only 1 is maybe not a good candidate for reduced threshold.  But where counts as % of non nocalls is X  or ebird code appears in sequential 5 secs or alternate 5 secs, these could be good candidates.  \nSo for me, that was  'stejay' and 'westan'.  So dropping the threshold for these 2 to 0.5 or even 0.3 and predicting again picks up a few more of these and improves the F1 score, (best was 0.538).  Sometimes there are several things calling at the same time or using the example metadata start/end to check, the call may be a small fragment in the 5 sec so maybe a lower threshold finds these. Hard to say if it can work for hidden test also.  Probably needs to be per site and audio clip - the data page says 150 recordings roughly 10 mins long.  The example audio clips are around 7 and 5 mins so hoping they are comparable for this checking.  Have you tried predictions for example test audio? \n \nFeel free also to share about getting to 3rd!!   Good luck!"
  },
  "source": "meta"
}