{
  "id": 75214,
  "title": "Min threshold?",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/75214",
  "author_name": "",
  "post_date": "2018-12-19T14:09:47.731452400Z",
  "votes": null,
  "comment_count": 9,
  "views": 0,
  "content": "<p>For people who are optimizing the thresholds on a per-class basis, have you tried to \"correct\" the threshold afterward by setting a minimum value? I ended up with <code>0.038</code> as a threshold for one class (class 9), and that seems like it's overfit to the validation distribution.</p>",
  "messages": [
    {
      "id": "442126",
      "postDate": "12/19/2018 14:09:47",
      "content": "<p>For people who are optimizing the thresholds on a per-class basis, have you tried to \"correct\" the threshold afterward by setting a minimum value? I ended up with <code>0.038</code> as a threshold for one class (class 9), and that seems like it's overfit to the validation distribution.</p>",
      "rawMarkdown": "For people who are optimizing the thresholds on a per-class basis, have you tried to \"correct\" the threshold afterward by setting a minimum value? I ended up with `0.038` as a threshold for one class (class 9), and that seems like it's overfit to the validation distribution.",
      "votes": null
    },
    {
      "id": "442130",
      "postDate": "12/19/2018 14:13:01",
      "content": "<p>Is this for a single validation set or over a k-fold run? Either way, 0.038 seems fairly low...have you plotted the distribution of model output for that class?</p>",
      "rawMarkdown": "Is this for a single validation set or over a k-fold run? Either way, 0.038 seems fairly low...have you plotted the distribution of model output for that class?",
      "votes": null
    },
    {
      "id": "442136",
      "postDate": "12/19/2018 14:18:50",
      "content": "<p>This is with a single validation set</p>",
      "rawMarkdown": "This is with a single validation set",
      "votes": null
    },
    {
      "id": "442206",
      "postDate": "12/19/2018 15:59:01",
      "content": "<p>When you posted few days ago about suddenly over-fitting, I forgot that you started using class weighting. This has the effect of changing output probabilities. If you were to use \"true\" class weights, that would actually bring the threshold to 0.5 for each class. I think you do not use true class weights, so you'll have to find new thresholds that will likely be higher compared to before-weighting.</p>",
      "rawMarkdown": "When you posted few days ago about suddenly over-fitting, I forgot that you started using class weighting. This has the effect of changing output probabilities. If you were to use \"true\" class weights, that would actually bring the threshold to 0.5 for each class. I think you do not use true class weights, so you'll have to find new thresholds that will likely be higher compared to before-weighting.",
      "votes": null
    },
    {
      "id": "442462",
      "postDate": "12/20/2018 01:48:56",
      "content": "<p>How would I find the new thresholds for each individual class after using the class weights in my loss function?</p>",
      "rawMarkdown": "How would I find the new thresholds for each individual class after using the class weights in my loss function?",
      "votes": null
    },
    {
      "id": "442508",
      "postDate": "12/20/2018 03:28:06",
      "content": "<p>I found threshold searching is harmful to the performance on LB... Even I search with unseen data...</p>",
      "rawMarkdown": "I found threshold searching is harmful to the performance on LB... Even I search with unseen data...",
      "votes": null
    },
    {
      "id": "443995",
      "postDate": "12/22/2018 21:53:04",
      "content": "<p>I also found the same. It makes me wonder if the public LB class distribution has been tweaked to penalize those over-fitting to the LB. As others have noted elsewhere, there is quite the gap in local validation vs the public LB. Just a random thought...</p>",
      "rawMarkdown": "I also found the same. It makes me wonder if the public LB class distribution has been tweaked to penalize those over-fitting to the LB. As others have noted elsewhere, there is quite the gap in local validation vs the public LB. Just a random thought...",
      "votes": null
    },
    {
      "id": "444001",
      "postDate": "12/22/2018 22:12:28",
      "content": "<p>What do you use if you’re not searching? Some fixed value?</p>",
      "rawMarkdown": "What do you use if you’re not searching? Some fixed value?",
      "votes": null
    },
    {
      "id": "444003",
      "postDate": "12/22/2018 22:20:04",
      "content": "<p>In general, a constant threshold of 0.16 - 0.2 has worked pretty well for me on the LB. Finding an optimal constant threshold or per class thresholds has looked promising on unseen local data however it has always given worse results on the public LB.</p>",
      "rawMarkdown": "In general, a constant threshold of 0.16 - 0.2 has worked pretty well for me on the LB. Finding an optimal constant threshold or per class thresholds has looked promising on unseen local data however it has always given worse results on the public LB.",
      "votes": null
    },
    {
      "id": "444013",
      "postDate": "12/22/2018 23:16:58",
      "content": "<p>I’m using weighted random sampling so I imagine my best thresholds will be closer to 0.5. I’ll give constant thresholds a shot and see if I get any improvements</p>",
      "rawMarkdown": "I’m using weighted random sampling so I imagine my best thresholds will be closer to 0.5. I’ll give constant thresholds a shot and see if I get any improvements",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 442130,
      "author_name": "maw501",
      "author_url": "",
      "post_date": "12/19/2018 14:13:01",
      "content": "<p>Is this for a single validation set or over a k-fold run? Either way, 0.038 seems fairly low...have you plotted the distribution of model output for that class?</p>",
      "votes": null,
      "replies": [
        {
          "id": 442136,
          "author_name": "hortonhearsafoo",
          "author_url": "",
          "post_date": "12/19/2018 14:18:50",
          "content": "<p>This is with a single validation set</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 442206,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "12/19/2018 15:59:01",
      "content": "<p>When you posted few days ago about suddenly over-fitting, I forgot that you started using class weighting. This has the effect of changing output probabilities. If you were to use \"true\" class weights, that would actually bring the threshold to 0.5 for each class. I think you do not use true class weights, so you'll have to find new thresholds that will likely be higher compared to before-weighting.</p>",
      "votes": null,
      "replies": [
        {
          "id": 442462,
          "author_name": "criminal",
          "author_url": "",
          "post_date": "12/20/2018 01:48:56",
          "content": "<p>How would I find the new thresholds for each individual class after using the class weights in my loss function?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 442508,
      "author_name": "manyfoldcv",
      "author_url": "",
      "post_date": "12/20/2018 03:28:06",
      "content": "<p>I found threshold searching is harmful to the performance on LB... Even I search with unseen data...</p>",
      "votes": null,
      "replies": [
        {
          "id": 443995,
          "author_name": "pjbutcher",
          "author_url": "",
          "post_date": "12/22/2018 21:53:04",
          "content": "<p>I also found the same. It makes me wonder if the public LB class distribution has been tweaked to penalize those over-fitting to the LB. As others have noted elsewhere, there is quite the gap in local validation vs the public LB. Just a random thought...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 444001,
          "author_name": "hortonhearsafoo",
          "author_url": "",
          "post_date": "12/22/2018 22:12:28",
          "content": "<p>What do you use if you’re not searching? Some fixed value?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 444003,
          "author_name": "pjbutcher",
          "author_url": "",
          "post_date": "12/22/2018 22:20:04",
          "content": "<p>In general, a constant threshold of 0.16 - 0.2 has worked pretty well for me on the LB. Finding an optimal constant threshold or per class thresholds has looked promising on unseen local data however it has always given worse results on the public LB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 444013,
          "author_name": "hortonhearsafoo",
          "author_url": "",
          "post_date": "12/22/2018 23:16:58",
          "content": "<p>I’m using weighted random sampling so I imagine my best thresholds will be closer to 0.5. I’ll give constant thresholds a shot and see if I get any improvements</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "442126": "For people who are optimizing the thresholds on a per-class basis, have you tried to \"correct\" the threshold afterward by setting a minimum value? I ended up with `0.038` as a threshold for one class (class 9), and that seems like it's overfit to the validation distribution.",
    "442130": "Is this for a single validation set or over a k-fold run? Either way, 0.038 seems fairly low...have you plotted the distribution of model output for that class?",
    "442136": "This is with a single validation set",
    "442206": "When you posted few days ago about suddenly over-fitting, I forgot that you started using class weighting. This has the effect of changing output probabilities. If you were to use \"true\" class weights, that would actually bring the threshold to 0.5 for each class. I think you do not use true class weights, so you'll have to find new thresholds that will likely be higher compared to before-weighting.",
    "442462": "How would I find the new thresholds for each individual class after using the class weights in my loss function?",
    "442508": "I found threshold searching is harmful to the performance on LB... Even I search with unseen data...",
    "443995": "I also found the same. It makes me wonder if the public LB class distribution has been tweaked to penalize those over-fitting to the LB. As others have noted elsewhere, there is quite the gap in local validation vs the public LB. Just a random thought...",
    "444001": "What do you use if you’re not searching? Some fixed value?",
    "444003": "In general, a constant threshold of 0.16 - 0.2 has worked pretty well for me on the LB. Finding an optimal constant threshold or per class thresholds has looked promising on unseen local data however it has always given worse results on the public LB.",
    "444013": "I’m using weighted random sampling so I imagine my best thresholds will be closer to 0.5. I’ll give constant thresholds a shot and see if I get any improvements"
  },
  "source": "meta"
}