{
  "id": 167263,
  "title": "How should we make validation and decide threshold?",
  "url": "/competitions/birdsong-recognition/discussion/167263",
  "author_name": "",
  "post_date": "2020-07-15T20:18:47.316004700Z",
  "votes": 25,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi.\nI'd like to share public scores of <a href=\"https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast\">my baseline notebook</a> with thresholds.</p>\n\n<p>| Threshold | Public Score |\n|:----------:|:-------------:|\n| 0.1             |  0.249 |\n| 0.3            |  0.538 |\n| 0.4            | 0.560 |\n| 0.5            | 0.566 |\n| <strong><em>0.6</em></strong>            | <strong><em>0.568</em></strong> |\n| 0.7            | 0.566 |\n| 0.8            | 0.564 |\n| 0.9            | 0.562 |</p>\n\n<p><br>\nIt seems that 0.6 is good for <strong>this model</strong>.\nBut, you know, best threshold is depends on  each model (architecture, training method, ...) and <strong><em>public LB is only 27% of the test data</em></strong>. \nIn addition to those, there are still difficulties (for example, see <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/167202\">this topic</a>).</p>\n\n<p>I think how to make validation is one of the keys and interesting points in this competition :)\n<br>\nAny ideas?</p>",
  "messages": [
    {
      "id": "930926",
      "postDate": "07/15/2020 20:18:47",
      "content": "<p>Hi.\nI'd like to share public scores of <a href=\"https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast\">my baseline notebook</a> with thresholds.</p>\n\n<p>| Threshold | Public Score |\n|:----------:|:-------------:|\n| 0.1             |  0.249 |\n| 0.3            |  0.538 |\n| 0.4            | 0.560 |\n| 0.5            | 0.566 |\n| <strong><em>0.6</em></strong>            | <strong><em>0.568</em></strong> |\n| 0.7            | 0.566 |\n| 0.8            | 0.564 |\n| 0.9            | 0.562 |</p>\n\n<p><br>\nIt seems that 0.6 is good for <strong>this model</strong>.\nBut, you know, best threshold is depends on  each model (architecture, training method, ...) and <strong><em>public LB is only 27% of the test data</em></strong>. \nIn addition to those, there are still difficulties (for example, see <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/167202\">this topic</a>).</p>\n\n<p>I think how to make validation is one of the keys and interesting points in this competition :)\n<br>\nAny ideas?</p>",
      "rawMarkdown": "Hi.\nI'd like to share public scores of [my baseline notebook](https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast) with thresholds.\n\n| Threshold | Public Score |\n|:----------:|:-------------:|\n| 0.1             |  0.249 |\n| 0.3            |  0.538 |\n| 0.4            | 0.560 |\n| 0.5            | 0.566 |\n| **_0.6_**            | **_0.568_** |\n| 0.7            | 0.566 |\n| 0.8            | 0.564 |\n| 0.9            | 0.562 |\n\n<br>\nIt seems that 0.6 is good for **this model**.\nBut, you know, best threshold is depends on  each model (architecture, training method, ...) and **_public LB is only 27% of the test data_**. \nIn addition to those, there are still difficulties (for example, see [this topic](https://www.kaggle.com/c/birdsong-recognition/discussion/167202)).\n\nI think how to make validation is one of the keys and interesting points in this competition :)\n<br>\nAny ideas?",
      "votes": null
    },
    {
      "id": "931609",
      "postDate": "07/16/2020 09:54:12",
      "content": "<p>The other issue is it is  possible to have multiple threshold depending of the class. I feel it is a bit tricky to have a proper validation process, as we do not have the same kind of data between train/test. Maybe we should pseudo label the validation and working on it for the threshold, but it will probably be a bit bias</p>",
      "rawMarkdown": "The other issue is it is  possible to have multiple threshold depending of the class. I feel it is a bit tricky to have a proper validation process, as we do not have the same kind of data between train/test. Maybe we should pseudo label the validation and working on it for the threshold, but it will probably be a bit bias",
      "votes": null
    },
    {
      "id": "932300",
      "postDate": "07/16/2020 23:24:18",
      "content": "<p>Thanks. </p>\n\n<p>Yes, pseudo labeling sounds good.\nBut for it, we need a model which can predict precisely. Hmm ...</p>",
      "rawMarkdown": "Thanks. \n\nYes, pseudo labeling sounds good.\nBut for it, we need a model which can predict precisely. Hmm ...",
      "votes": null
    },
    {
      "id": "933619",
      "postDate": "07/17/2020 20:30:30",
      "content": "<p>It seems like any attempt I make to pseudolabel/clean the data makes the classifier worse at handling the noisy test data.</p>\n\n<p>It's definitely the right way to go, but we'll need other models in the pipeline to make the whole system more robust.</p>",
      "rawMarkdown": "It seems like any attempt I make to pseudolabel/clean the data makes the classifier worse at handling the noisy test data.\n\nIt's definitely the right way to go, but we'll need other models in the pipeline to make the whole system more robust.",
      "votes": null
    },
    {
      "id": "933700",
      "postDate": "07/17/2020 23:26:15",
      "content": "<p>I agree. The issue is I feel it will be difficult to experiment well because we are very limited about the cpu speed as we need to load each time the audio. Excepted if we have a very good computer at home or using GCP, otherwise it seems tricky. </p>",
      "rawMarkdown": "I agree. The issue is I feel it will be difficult to experiment well because we are very limited about the cpu speed as we need to load each time the audio. Excepted if we have a very good computer at home or using GCP, otherwise it seems tricky.",
      "votes": null
    },
    {
      "id": "969067",
      "postDate": "08/13/2020 12:45:18",
      "content": "<p>Have you tried working on threshold optimization separately? For example, assuming your model will output probabilities for each label, by predicting a validation set you can apply ROC opt to get a threshold for each label. I have tried it myself but didn't have great success. Not sure if this is a good approach!</p>",
      "rawMarkdown": "Have you tried working on threshold optimization separately? For example, assuming your model will output probabilities for each label, by predicting a validation set you can apply ROC opt to get a threshold for each label. I have tried it myself but didn't have great success. Not sure if this is a good approach!",
      "votes": null
    },
    {
      "id": "976107",
      "postDate": "08/18/2020 16:33:12",
      "content": "<p>I think the problem is that we don't have a reliable validation set with labels, so tuning by submitting seems to be the next-best thing</p>",
      "rawMarkdown": "I think the problem is that we don't have a reliable validation set with labels, so tuning by submitting seems to be the next-best thing",
      "votes": null
    },
    {
      "id": "976439",
      "postDate": "08/18/2020 21:09:45",
      "content": "<p><a href=\"https://www.kaggle.com/bsmit1659\" target=\"_blank\">@bsmit1659</a> I have also tried to pseudo labelling and like you got worse results. (prediction for site1/site2)</p>",
      "rawMarkdown": "bsmit1659 I have also tried to pseudo labelling and like you got worse results. (prediction for site1/site2)",
      "votes": null
    },
    {
      "id": "976450",
      "postDate": "08/18/2020 21:27:50",
      "content": "<p><a href=\"https://www.kaggle.com/ludovick\" target=\"_blank\">@ludovick</a> I recently found that adding some pseudolabels can help, but only in cases where we know that the classifier has high precision. It's definitely helping to untangle some of the noisier labels, but I don't know how many iterations I can make it through before the end.</p>",
      "rawMarkdown": "ludovick I recently found that adding some pseudolabels can help, but only in cases where we know that the classifier has high precision. It's definitely helping to untangle some of the noisier labels, but I don't know how many iterations I can make it through before the end.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 969067,
      "author_name": "smerllo",
      "author_url": "",
      "post_date": "08/13/2020 12:45:18",
      "content": "<p>Have you tried working on threshold optimization separately? For example, assuming your model will output probabilities for each label, by predicting a validation set you can apply ROC opt to get a threshold for each label. I have tried it myself but didn't have great success. Not sure if this is a good approach!</p>",
      "votes": null,
      "replies": [
        {
          "id": 976107,
          "author_name": "marcogorelli",
          "author_url": "",
          "post_date": "08/18/2020 16:33:12",
          "content": "<p>I think the problem is that we don't have a reliable validation set with labels, so tuning by submitting seems to be the next-best thing</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 931609,
      "author_name": "ludovick",
      "author_url": "",
      "post_date": "07/16/2020 09:54:12",
      "content": "<p>The other issue is it is  possible to have multiple threshold depending of the class. I feel it is a bit tricky to have a proper validation process, as we do not have the same kind of data between train/test. Maybe we should pseudo label the validation and working on it for the threshold, but it will probably be a bit bias</p>",
      "votes": null,
      "replies": [
        {
          "id": 932300,
          "author_name": "ttahara",
          "author_url": "",
          "post_date": "07/16/2020 23:24:18",
          "content": "<p>Thanks. </p>\n\n<p>Yes, pseudo labeling sounds good.\nBut for it, we need a model which can predict precisely. Hmm ...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 933619,
          "author_name": "bsmit1659",
          "author_url": "",
          "post_date": "07/17/2020 20:30:30",
          "content": "<p>It seems like any attempt I make to pseudolabel/clean the data makes the classifier worse at handling the noisy test data.</p>\n\n<p>It's definitely the right way to go, but we'll need other models in the pipeline to make the whole system more robust.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 933700,
          "author_name": "ludovick",
          "author_url": "",
          "post_date": "07/17/2020 23:26:15",
          "content": "<p>I agree. The issue is I feel it will be difficult to experiment well because we are very limited about the cpu speed as we need to load each time the audio. Excepted if we have a very good computer at home or using GCP, otherwise it seems tricky. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 976439,
          "author_name": "ludovick",
          "author_url": "",
          "post_date": "08/18/2020 21:09:45",
          "content": "<p><a href=\"https://www.kaggle.com/bsmit1659\" target=\"_blank\">@bsmit1659</a> I have also tried to pseudo labelling and like you got worse results. (prediction for site1/site2)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 976450,
          "author_name": "bsmit1659",
          "author_url": "",
          "post_date": "08/18/2020 21:27:50",
          "content": "<p><a href=\"https://www.kaggle.com/ludovick\" target=\"_blank\">@ludovick</a> I recently found that adding some pseudolabels can help, but only in cases where we know that the classifier has high precision. It's definitely helping to untangle some of the noisier labels, but I don't know how many iterations I can make it through before the end.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "930926": "Hi.\nI'd like to share public scores of [my baseline notebook](https://www.kaggle.com/ttahara/inference-birdsong-baseline-resnest50-fast) with thresholds.\n\n| Threshold | Public Score |\n|:----------:|:-------------:|\n| 0.1             |  0.249 |\n| 0.3            |  0.538 |\n| 0.4            | 0.560 |\n| 0.5            | 0.566 |\n| **_0.6_**            | **_0.568_** |\n| 0.7            | 0.566 |\n| 0.8            | 0.564 |\n| 0.9            | 0.562 |\n\n<br>\nIt seems that 0.6 is good for **this model**.\nBut, you know, best threshold is depends on  each model (architecture, training method, ...) and **_public LB is only 27% of the test data_**. \nIn addition to those, there are still difficulties (for example, see [this topic](https://www.kaggle.com/c/birdsong-recognition/discussion/167202)).\n\nI think how to make validation is one of the keys and interesting points in this competition :)\n<br>\nAny ideas?",
    "931609": "The other issue is it is  possible to have multiple threshold depending of the class. I feel it is a bit tricky to have a proper validation process, as we do not have the same kind of data between train/test. Maybe we should pseudo label the validation and working on it for the threshold, but it will probably be a bit bias",
    "932300": "Thanks. \n\nYes, pseudo labeling sounds good.\nBut for it, we need a model which can predict precisely. Hmm ...",
    "933619": "It seems like any attempt I make to pseudolabel/clean the data makes the classifier worse at handling the noisy test data.\n\nIt's definitely the right way to go, but we'll need other models in the pipeline to make the whole system more robust.",
    "933700": "I agree. The issue is I feel it will be difficult to experiment well because we are very limited about the cpu speed as we need to load each time the audio. Excepted if we have a very good computer at home or using GCP, otherwise it seems tricky.",
    "969067": "Have you tried working on threshold optimization separately? For example, assuming your model will output probabilities for each label, by predicting a validation set you can apply ROC opt to get a threshold for each label. I have tried it myself but didn't have great success. Not sure if this is a good approach!",
    "976107": "I think the problem is that we don't have a reliable validation set with labels, so tuning by submitting seems to be the next-best thing",
    "976439": "bsmit1659 I have also tried to pseudo labelling and like you got worse results. (prediction for site1/site2)",
    "976450": "ludovick I recently found that adding some pseudolabels can help, but only in cases where we know that the classifier has high precision. It's definitely helping to untangle some of the noisier labels, but I don't know how many iterations I can make it through before the end."
  },
  "source": "meta"
}