{
  "id": 70376,
  "title": "Threshold selection",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/70376",
  "author_name": "",
  "post_date": "2018-11-02T16:33:39.495727400Z",
  "votes": 11,
  "comment_count": 5,
  "views": 0,
  "content": "<p>This is very important for this is challenge, where the data of some class is very small, making the precision recall curve unsmooth. This is what you can do;</p>\n\n<ol>\n<li><p>Plot the precision  vs recall curve or threshold vs f1 curve. Smooth the curve by hand. Select the threshold from the graph ( and not from argmax from code)</p></li>\n<li><p>Make perturbations of your validation set until you have enough samples to construct a smooth curve.</p></li>\n</ol>\n\n<p>Generally, the curve should be smooth and the f1 value should be fairly constant over a wide range of threshold if you have enough validation samples. This ensures good generalisation </p>",
  "messages": [
    {
      "id": "414388",
      "postDate": "11/02/2018 16:33:39",
      "content": "<p>This is very important for this is challenge, where the data of some class is very small, making the precision recall curve unsmooth. This is what you can do;</p>\n\n<ol>\n<li><p>Plot the precision  vs recall curve or threshold vs f1 curve. Smooth the curve by hand. Select the threshold from the graph ( and not from argmax from code)</p></li>\n<li><p>Make perturbations of your validation set until you have enough samples to construct a smooth curve.</p></li>\n</ol>\n\n<p>Generally, the curve should be smooth and the f1 value should be fairly constant over a wide range of threshold if you have enough validation samples. This ensures good generalisation </p>",
      "rawMarkdown": "This is very important for this is challenge, where the data of some class is very small, making the precision recall curve unsmooth. This is what you can do;\n\n1. Plot the precision  vs recall curve or threshold vs f1 curve. Smooth the curve by hand. Select the threshold from the graph ( and not from argmax from code)\n\n2. Make perturbations of your validation set until you have enough samples to construct a smooth curve.\n\nGenerally, the curve should be smooth and the f1 value should be fairly constant over a wide range of threshold if you have enough validation samples. This ensures good generalisation",
      "votes": null
    },
    {
      "id": "414460",
      "postDate": "11/02/2018 20:04:45",
      "content": "<p>Here is an iteresting study:\n<a href=\"https://www.csie.ntu.edu.tw/~cjlin/papers/threshold.pdf\">Threshold Selection for Multi-label Classification</a></p>",
      "rawMarkdown": "Here is an iteresting study:\n[Threshold Selection for Multi-label Classification](https://www.csie.ntu.edu.tw/~cjlin/papers/threshold.pdf)",
      "votes": null
    },
    {
      "id": "415994",
      "postDate": "11/06/2018 02:38:27",
      "content": "<p>A proper threshold definitely helps on this one. I've found that the threshold will plateau for each one. Doing a threshold search from highest to lowest and from lowest to highest ends up with either the low side of the plateau or high side. The lower threshold did better.</p>",
      "rawMarkdown": "A proper threshold definitely helps on this one. I've found that the threshold will plateau for each one. Doing a threshold search from highest to lowest and from lowest to highest ends up with either the low side of the plateau or high side. The lower threshold did better.",
      "votes": null
    },
    {
      "id": "422130",
      "postDate": "11/15/2018 20:11:27",
      "content": "<p>I assume you run threshold search on your test set. How about LB probability? ie: match the LB prob with local prob</p>",
      "rawMarkdown": "I assume you run threshold search on your test set. How about LB probability? ie: match the LB prob with local prob",
      "votes": null
    },
    {
      "id": "422169",
      "postDate": "11/15/2018 21:39:24",
      "content": "<p>I don't use the LB probability, I think the public LB doesn't quite represent the entire LB so trying to fit to that specifically may give a worse result in the end. Using sklearn's f1 macro score, the best public score has always been with lower thresholds than the local best.</p>",
      "rawMarkdown": "I don't use the LB probability, I think the public LB doesn't quite represent the entire LB so trying to fit to that specifically may give a worse result in the end. Using sklearn's f1 macro score, the best public score has always been with lower thresholds than the local best.",
      "votes": null
    },
    {
      "id": "422235",
      "postDate": "11/16/2018 00:05:51",
      "content": "<p>Choosing threshold from the validation data just did not work for me. I ended up using a simple power law to choose thresholds based on the label count - the lower counts are favoured - consider it a back-door \"los weighting\". Its convenient in that searching the parameters is done after training.</p>",
      "rawMarkdown": "Choosing threshold from the validation data just did not work for me. I ended up using a simple power law to choose thresholds based on the label count - the lower counts are favoured - consider it a back-door \"los weighting\". Its convenient in that searching the parameters is done after training.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 414460,
      "author_name": "pestipeti",
      "author_url": "",
      "post_date": "11/02/2018 20:04:45",
      "content": "<p>Here is an iteresting study:\n<a href=\"https://www.csie.ntu.edu.tw/~cjlin/papers/threshold.pdf\">Threshold Selection for Multi-label Classification</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 415994,
      "author_name": "ldm314",
      "author_url": "",
      "post_date": "11/06/2018 02:38:27",
      "content": "<p>A proper threshold definitely helps on this one. I've found that the threshold will plateau for each one. Doing a threshold search from highest to lowest and from lowest to highest ends up with either the low side of the plateau or high side. The lower threshold did better.</p>",
      "votes": null,
      "replies": [
        {
          "id": 422130,
          "author_name": "kokecacao",
          "author_url": "",
          "post_date": "11/15/2018 20:11:27",
          "content": "<p>I assume you run threshold search on your test set. How about LB probability? ie: match the LB prob with local prob</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 422169,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "11/15/2018 21:39:24",
          "content": "<p>I don't use the LB probability, I think the public LB doesn't quite represent the entire LB so trying to fit to that specifically may give a worse result in the end. Using sklearn's f1 macro score, the best public score has always been with lower thresholds than the local best.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 422235,
      "author_name": "petewills",
      "author_url": "",
      "post_date": "11/16/2018 00:05:51",
      "content": "<p>Choosing threshold from the validation data just did not work for me. I ended up using a simple power law to choose thresholds based on the label count - the lower counts are favoured - consider it a back-door \"los weighting\". Its convenient in that searching the parameters is done after training.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "414388": "This is very important for this is challenge, where the data of some class is very small, making the precision recall curve unsmooth. This is what you can do;\n\n1. Plot the precision  vs recall curve or threshold vs f1 curve. Smooth the curve by hand. Select the threshold from the graph ( and not from argmax from code)\n\n2. Make perturbations of your validation set until you have enough samples to construct a smooth curve.\n\nGenerally, the curve should be smooth and the f1 value should be fairly constant over a wide range of threshold if you have enough validation samples. This ensures good generalisation",
    "414460": "Here is an iteresting study:\n[Threshold Selection for Multi-label Classification](https://www.csie.ntu.edu.tw/~cjlin/papers/threshold.pdf)",
    "415994": "A proper threshold definitely helps on this one. I've found that the threshold will plateau for each one. Doing a threshold search from highest to lowest and from lowest to highest ends up with either the low side of the plateau or high side. The lower threshold did better.",
    "422130": "I assume you run threshold search on your test set. How about LB probability? ie: match the LB prob with local prob",
    "422169": "I don't use the LB probability, I think the public LB doesn't quite represent the entire LB so trying to fit to that specifically may give a worse result in the end. Using sklearn's f1 macro score, the best public score has always been with lower thresholds than the local best.",
    "422235": "Choosing threshold from the validation data just did not work for me. I ended up using a simple power law to choose thresholds based on the label count - the lower counts are favoured - consider it a back-door \"los weighting\". Its convenient in that searching the parameters is done after training."
  },
  "source": "meta"
}