{
  "id": 175097,
  "title": "Model Output Observations",
  "url": "/competitions/birdsong-recognition/discussion/175097",
  "author_name": "",
  "post_date": "2020-08-17T06:06:39.107153400Z",
  "votes": 6,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I have some observations with output of <a href=\"https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast\" target=\"_blank\">one of the best current public solution</a> and want to share it with you. So here is the deal.</p>\n<p><strong>Max Output Probability</strong> (max output probability for each sample -&gt; sorting)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2Fdc7fdb2ce0efe96929622d00686769f0%2Fmax_proba.png?generation=1597597236007920&amp;alt=media\" alt=\"\"></p>\n<p>We can see that there are lots of cases when max sample probability is lower than for instance popular threshold 0.5. It means, that for this cases we guarantee a zero f1 score (inside train set there is no \"nocall\" target). Maybe, that's problem of the model, but there are chances, that it happens cause of data and samples with small probability have some problem (large noise level, weak bird's voces, what have you). </p>\n<p>I <a href=\"https://www.kaggle.com/koza4ukdmitrij/small-output-probabilities-cbi?select=proba_less_than_0.5.csv\" target=\"_blank\">save</a> samples with probabilities less than threshold for some thresholds. If this is problem of the data, you can check whether your model has the same challenge. </p>\n<p>Conclusions:</p>\n<ul>\n<li>a big part of errors related to nocall prediction instead of other bird prediction.</li>\n<li>maybe, samples with small probabilities have problem in input data.</li>\n</ul>\n<p><strong>Best Threshold</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F76e5d75fe74f8da83776e6a0c1b93e56%2Fbest_thresh.png?generation=1597597248131764&amp;alt=media\" alt=\"\"></p>\n<p>Best threshhold by a dev set is about 0.2, but setting this threshold for test submission decreased LB score so much (about 0.568 -&gt; 0.543). So I think there are two conclusions there:</p>\n<ul>\n<li>f1 score highly depends on the chosen threshold (cause curve above has a not null derivative about a popular 0.5 area).</li>\n<li>best threshold highly depends on the set for what we evaluate it. So choosing best threshold for hidden test set is an open problem.</li>\n</ul>\n<p>What do you think about it? Do you have the same problem with thresholds?</p>",
  "messages": [
    {
      "id": "973137",
      "postDate": "08/17/2020 06:06:39",
      "content": "<p>I have some observations with output of <a href=\"https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast\" target=\"_blank\">one of the best current public solution</a> and want to share it with you. So here is the deal.</p>\n<p><strong>Max Output Probability</strong> (max output probability for each sample -&gt; sorting)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2Fdc7fdb2ce0efe96929622d00686769f0%2Fmax_proba.png?generation=1597597236007920&amp;alt=media\" alt=\"\"></p>\n<p>We can see that there are lots of cases when max sample probability is lower than for instance popular threshold 0.5. It means, that for this cases we guarantee a zero f1 score (inside train set there is no \"nocall\" target). Maybe, that's problem of the model, but there are chances, that it happens cause of data and samples with small probability have some problem (large noise level, weak bird's voces, what have you). </p>\n<p>I <a href=\"https://www.kaggle.com/koza4ukdmitrij/small-output-probabilities-cbi?select=proba_less_than_0.5.csv\" target=\"_blank\">save</a> samples with probabilities less than threshold for some thresholds. If this is problem of the data, you can check whether your model has the same challenge. </p>\n<p>Conclusions:</p>\n<ul>\n<li>a big part of errors related to nocall prediction instead of other bird prediction.</li>\n<li>maybe, samples with small probabilities have problem in input data.</li>\n</ul>\n<p><strong>Best Threshold</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F76e5d75fe74f8da83776e6a0c1b93e56%2Fbest_thresh.png?generation=1597597248131764&amp;alt=media\" alt=\"\"></p>\n<p>Best threshhold by a dev set is about 0.2, but setting this threshold for test submission decreased LB score so much (about 0.568 -&gt; 0.543). So I think there are two conclusions there:</p>\n<ul>\n<li>f1 score highly depends on the chosen threshold (cause curve above has a not null derivative about a popular 0.5 area).</li>\n<li>best threshold highly depends on the set for what we evaluate it. So choosing best threshold for hidden test set is an open problem.</li>\n</ul>\n<p>What do you think about it? Do you have the same problem with thresholds?</p>",
      "rawMarkdown": "I have some observations with output of [one of the best current public solution](https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast) and want to share it with you. So here is the deal.\n\n**Max Output Probability** (max output probability for each sample -> sorting)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2Fdc7fdb2ce0efe96929622d00686769f0%2Fmax_proba.png?generation=1597597236007920&alt=media)\n\nWe can see that there are lots of cases when max sample probability is lower than for instance popular threshold 0.5. It means, that for this cases we guarantee a zero f1 score (inside train set there is no \"nocall\" target). Maybe, that's problem of the model, but there are chances, that it happens cause of data and samples with small probability have some problem (large noise level, weak bird's voces, what have you). \n\nI [save](https://www.kaggle.com/koza4ukdmitrij/small-output-probabilities-cbi?select=proba_less_than_0.5.csv) samples with probabilities less than threshold for some thresholds. If this is problem of the data, you can check whether your model has the same challenge. \n\nConclusions:\n- a big part of errors related to nocall prediction instead of other bird prediction.\n- maybe, samples with small probabilities have problem in input data.\n\n**Best Threshold**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F76e5d75fe74f8da83776e6a0c1b93e56%2Fbest_thresh.png?generation=1597597248131764&alt=media)\n\nBest threshhold by a dev set is about 0.2, but setting this threshold for test submission decreased LB score so much (about 0.568 -> 0.543). So I think there are two conclusions there:\n- f1 score highly depends on the chosen threshold (cause curve above has a not null derivative about a popular 0.5 area).\n- best threshold highly depends on the set for what we evaluate it. So choosing best threshold for hidden test set is an open problem.\n\nWhat do you think about it? Do you have the same problem with thresholds?",
      "votes": null
    },
    {
      "id": "974328",
      "postDate": "08/17/2020 22:21:37",
      "content": "<p>I have same problem with thresholds. Very inconsistent, even for the same model just using different epochs, the % of nocall predicted can be very different for the same thresholds. That is one of the main challenges of this competition, and I dont think you will find any who want to reveal their secrets.</p>",
      "rawMarkdown": "I have same problem with thresholds. Very inconsistent, even for the same model just using different epochs, the % of nocall predicted can be very different for the same thresholds. That is one of the main challenges of this competition, and I dont think you will find any who want to reveal their secrets.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 974328,
      "author_name": "returnofsputnik",
      "author_url": "",
      "post_date": "08/17/2020 22:21:37",
      "content": "<p>I have same problem with thresholds. Very inconsistent, even for the same model just using different epochs, the % of nocall predicted can be very different for the same thresholds. That is one of the main challenges of this competition, and I dont think you will find any who want to reveal their secrets.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "973137": "I have some observations with output of [one of the best current public solution](https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast) and want to share it with you. So here is the deal.\n\n**Max Output Probability** (max output probability for each sample -> sorting)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2Fdc7fdb2ce0efe96929622d00686769f0%2Fmax_proba.png?generation=1597597236007920&alt=media)\n\nWe can see that there are lots of cases when max sample probability is lower than for instance popular threshold 0.5. It means, that for this cases we guarantee a zero f1 score (inside train set there is no \"nocall\" target). Maybe, that's problem of the model, but there are chances, that it happens cause of data and samples with small probability have some problem (large noise level, weak bird's voces, what have you). \n\nI [save](https://www.kaggle.com/koza4ukdmitrij/small-output-probabilities-cbi?select=proba_less_than_0.5.csv) samples with probabilities less than threshold for some thresholds. If this is problem of the data, you can check whether your model has the same challenge. \n\nConclusions:\n- a big part of errors related to nocall prediction instead of other bird prediction.\n- maybe, samples with small probabilities have problem in input data.\n\n**Best Threshold**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1659719%2F76e5d75fe74f8da83776e6a0c1b93e56%2Fbest_thresh.png?generation=1597597248131764&alt=media)\n\nBest threshhold by a dev set is about 0.2, but setting this threshold for test submission decreased LB score so much (about 0.568 -> 0.543). So I think there are two conclusions there:\n- f1 score highly depends on the chosen threshold (cause curve above has a not null derivative about a popular 0.5 area).\n- best threshold highly depends on the set for what we evaluate it. So choosing best threshold for hidden test set is an open problem.\n\nWhat do you think about it? Do you have the same problem with thresholds?",
    "974328": "I have same problem with thresholds. Very inconsistent, even for the same model just using different epochs, the % of nocall predicted can be very different for the same thresholds. That is one of the main challenges of this competition, and I dont think you will find any who want to reveal their secrets."
  },
  "source": "meta"
}