{
  "id": 465659,
  "title": "How to determine thresholds？",
  "url": "/competitions/blood-vessel-segmentation/discussion/465659",
  "author_name": "tbt",
  "post_date": "2024-01-05T05:52:54.926000",
  "votes": 10,
  "comment_count": 7,
  "views": 0,
  "content": "<p>The recently published notebook (<a href=\"https://www.kaggle.com/code/misakimatsutomo/inference-1024-should-have-a-percentile-of-0-00149\" target=\"_blank\">https://www.kaggle.com/code/misakimatsutomo/inference-1024-should-have-a-percentile-of-0-00149</a>) shows a high public score. However, the threshold is determined by the percentile method, which may result in overfitting the public LB.</p>\n<p>What are your thoughts on how to determine an appropriate threshold? I believe the most basic way to determine the threshold is by looking at the CV of the training data, but I am struggling with that method as the sample is small for this competition.</p>\n<p>For example, the percentage of positive examples for each training data set is as follows and varies from sample to sample, even allowing for poor segmentation.</p>\n<table>\n<thead>\n<tr>\n<th>sample</th>\n<th>proportion of positive value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>kidney_1_dense</td>\n<td>0.0062</td>\n</tr>\n<tr>\n<td>kidney_1_voi</td>\n<td>0.012</td>\n</tr>\n<tr>\n<td>kidney_2</td>\n<td>0.0041</td>\n</tr>\n<tr>\n<td>kidney_3_dense</td>\n<td>0.0022</td>\n</tr>\n<tr>\n<td>kidney_3_sparse</td>\n<td>0.0027</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": 2587923,
      "postDate": "2024-01-05T05:52:54.927Z",
      "content": "<p>The recently published notebook (<a href=\"https://www.kaggle.com/code/misakimatsutomo/inference-1024-should-have-a-percentile-of-0-00149\" target=\"_blank\">https://www.kaggle.com/code/misakimatsutomo/inference-1024-should-have-a-percentile-of-0-00149</a>) shows a high public score. However, the threshold is determined by the percentile method, which may result in overfitting the public LB.</p>\n<p>What are your thoughts on how to determine an appropriate threshold? I believe the most basic way to determine the threshold is by looking at the CV of the training data, but I am struggling with that method as the sample is small for this competition.</p>\n<p>For example, the percentage of positive examples for each training data set is as follows and varies from sample to sample, even allowing for poor segmentation.</p>\n<table>\n<thead>\n<tr>\n<th>sample</th>\n<th>proportion of positive value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>kidney_1_dense</td>\n<td>0.0062</td>\n</tr>\n<tr>\n<td>kidney_1_voi</td>\n<td>0.012</td>\n</tr>\n<tr>\n<td>kidney_2</td>\n<td>0.0041</td>\n</tr>\n<tr>\n<td>kidney_3_dense</td>\n<td>0.0022</td>\n</tr>\n<tr>\n<td>kidney_3_sparse</td>\n<td>0.0027</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "The recently published notebook ([https://www.kaggle.com/code/misakimatsutomo/inference-1024-should-have-a-percentile-of-0-00149](https://www.kaggle.com/code/misakimatsutomo/inference-1024-should-have-a-percentile-of-0-00149)) shows a high public score. However, the threshold is determined by the percentile method, which may result in overfitting the public LB.\n\nWhat are your thoughts on how to determine an appropriate threshold? I believe the most basic way to determine the threshold is by looking at the CV of the training data, but I am struggling with that method as the sample is small for this competition.\n\nFor example, the percentage of positive examples for each training data set is as follows and varies from sample to sample, even allowing for poor segmentation.\n\n| sample | proportion of positive value|\n| --- | --- |\n| kidney_1_dense | 0.0062 |\n| kidney_1_voi | 0.012 |\n| kidney_2 | 0.0041 |\n| kidney_3_dense | 0.0022 |\n| kidney_3_sparse | 0.0027 |\n",
      "votes": 10
    },
    {
      "id": 2588316,
      "postDate": "2024-01-05T12:28:43.933Z",
      "content": "<p>can you show the actual probability, surface dice, hit rate, fp rate values for each percentile indicated in your table above?</p>\n<pre><code>percentile=\n\n = np(probability,percentile)\n\n\npredict = probability&gt;\n)\n\nhit_rate = (predict*truth)()/truth()\n\n....\n</code></pre>",
      "rawMarkdown": "can you show the actual probability, surface dice, hit rate, fp rate values for each percentile indicated in your table above?\n\n```\npercentile=0.0185\n\nth = np.percentile(probability,percentile)\nprint('actual threshold is', th)\n\npredict = probability>th\nprint('surface dice', compute_surface_dice(predict, truth))\n\nhit_rate = (predict*truth).sum()/truth.sum()\nprint('hit_rate', hit_rate)\n....\n\n\n```",
      "votes": 1,
      "replies": [
        {
          "id": 2588353,
          "postDate": "2024-01-05T12:50:21.720Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2588481,
          "postDate": "2024-01-05T14:22:10.803Z",
          "content": "<p>Thank you for comment. In the meantime, I did the calculations for kidney_2, kidney_3_dense.<br>\nAlthough the value of surface dice is low, it is confirmed that by increasing the threshold value, the value can be increased to 0.74,0.83.</p>\n<table>\n<thead>\n<tr>\n<th>sample</th>\n<th>th</th>\n<th>surface dice</th>\n<th>hit rate</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>kidney_2</td>\n<td>2.65e-8</td>\n<td>0.162</td>\n<td>0.903</td>\n</tr>\n<tr>\n<td>kidney_3_dense</td>\n<td>1.21e-6</td>\n<td>0.527</td>\n<td>0.882</td>\n</tr>\n</tbody>\n</table>",
          "rawMarkdown": "Thank you for comment. In the meantime, I did the calculations for kidney_2, kidney_3_dense.\nAlthough the value of surface dice is low, it is confirmed that by increasing the threshold value, the value can be increased to 0.74,0.83.\n\n| sample | th | surface dice | hit rate |\n| --- | --- | --- | --- | \n| kidney_2 | 2.65e-8 | 0.162 | 0.903 |\n| kidney_3_dense | 1.21e-6 | 0.527 | 0.882 |",
          "replies": [
            {
              "id": 2588732,
              "postDate": "2024-01-05T17:25:03.607Z",
              "content": "<p>it should be somthing like:</p>\n<pre><code> = kidney_1_dense  (sparsity=%)\n = kidney_3_dense (sparsity=%)\n\n        th      dice (CV)       hit rate      fp rate       dice (lb)    \n.             .          .           .      .          xxx\n.            .          .             .         .        yyy\n.g. only ....\n. ...\n</code></pre>\n<p>build tables for train/ validation and for different fold etc ….</p>\n<p>you should note that lower fp usually gives better results.</p>\n<p>when we says low fp, it means low mean and low std of fp values over a range of theshold (e.g. from 0.25 to 0.35). we want performance metrics not sensitive the threshold</p>",
              "rawMarkdown": "it should be somthing like:\n\n```\ntrain = kidney_1_dense  (sparsity=100%)\nvalid = kidney_3_dense (sparsity=100%)\n\npercentile        th      dice (CV)       hit rate      fp rate       dice (lb)    \n0.0001             0.1          0.81           0.91      0.111          xxx\n0.0002            0.2          0.82             0.92         0.222        yyy\ne.g. only ....\n0.010 ...\n\n```\n\n\nbuild tables for train/ validation and for different fold etc ....\n\nyou should note that lower fp usually gives better results.\n\nwhen we says low fp, it means low mean and low std of fp values over a range of theshold (e.g. from 0.25 to 0.35). we want performance metrics not sensitive the threshold",
              "votes": 2
            },
            {
              "id": 2588737,
              "postDate": "2024-01-05T17:27:45.523Z",
              "content": "<p>percentile is not important. it is the \"threshold value at that percentile\"  is important</p>",
              "rawMarkdown": "percentile is not important. it is the \"threshold value at that percentile\"  is important",
              "votes": 1
            },
            {
              "id": 2589453,
              "postDate": "2024-01-06T11:50:15.110Z",
              "content": "<p>Thanks again for the advice. The results are as follows.<br>\nAs you point out, it seems more reasonable to refer to the threshold determined based on the percentile, rather than the percentile itself.<br>\nI will continue to investigate.</p>\n<p>train = kidney_1_dense, valid = kidney_3_dense</p>\n<table>\n<thead>\n<tr>\n<th>percentile</th>\n<th>th</th>\n<th>dice(CV)</th>\n<th>hit rate</th>\n<th>fp</th>\n<th>dice(LB)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.0010</td>\n<td>0.93</td>\n<td>0.54</td>\n<td>0.44</td>\n<td>0.003</td>\n<td>?</td>\n</tr>\n<tr>\n<td>0.0020</td>\n<td>0.55</td>\n<td>0.83</td>\n<td>0.87</td>\n<td>0.016</td>\n<td>0.852</td>\n</tr>\n<tr>\n<td>0.0025</td>\n<td>0.11</td>\n<td>0.82</td>\n<td>0.93</td>\n<td>0.154</td>\n<td>?</td>\n</tr>\n<tr>\n<td>0.0030</td>\n<td>0.02</td>\n<td>0.75</td>\n<td>0.96</td>\n<td>0.280</td>\n<td>?</td>\n</tr>\n</tbody>\n</table>",
              "rawMarkdown": "Thanks again for the advice. The results are as follows.\nAs you point out, it seems more reasonable to refer to the threshold determined based on the percentile, rather than the percentile itself.\nI will continue to investigate.\n\ntrain = kidney_1_dense, valid = kidney_3_dense\n\n| percentile | th | dice(CV) | hit rate | fp | dice(LB) |\n| --- | --- | --- | --- | --- | --- |\n| 0.0010 | 0.93 | 0.54 | 0.44 | 0.003  | ? | \n| 0.0020 | 0.55 | 0.83 | 0.87 | 0.016  | 0.852 | \n| 0.0025 | 0.11 | 0.82 | 0.93 | 0.154  | ? | \n| 0.0030 | 0.02 | 0.75 | 0.96 | 0.280  | ? | \n",
              "votes": 2
            },
            {
              "id": 2589499,
              "postDate": "2024-01-06T12:41:15.930Z",
              "content": "<p>check this :<br>\n<a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/456118#2589498\" target=\"_blank\">https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/456118#2589498</a></p>",
              "rawMarkdown": "check this :\nhttps://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/456118#2589498",
              "votes": 1
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2588316,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-01-05T12:28:43.933000",
      "content": "<p>can you show the actual probability, surface dice, hit rate, fp rate values for each percentile indicated in your table above?</p>\n<pre><code>percentile=\n\n = np(probability,percentile)\n\n\npredict = probability&gt;\n)\n\nhit_rate = (predict*truth)()/truth()\n\n....\n</code></pre>",
      "votes": 1,
      "replies": [
        {
          "id": 2588353,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-01-05T12:50:21.720000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2588481,
          "author_name": "tbt",
          "author_url": "",
          "post_date": "2024-01-05T14:22:10.803000",
          "content": "<p>Thank you for comment. In the meantime, I did the calculations for kidney_2, kidney_3_dense.<br>\nAlthough the value of surface dice is low, it is confirmed that by increasing the threshold value, the value can be increased to 0.74,0.83.</p>\n<table>\n<thead>\n<tr>\n<th>sample</th>\n<th>th</th>\n<th>surface dice</th>\n<th>hit rate</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>kidney_2</td>\n<td>2.65e-8</td>\n<td>0.162</td>\n<td>0.903</td>\n</tr>\n<tr>\n<td>kidney_3_dense</td>\n<td>1.21e-6</td>\n<td>0.527</td>\n<td>0.882</td>\n</tr>\n</tbody>\n</table>",
          "votes": 0,
          "replies": [
            {
              "id": 2588732,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-01-05T17:25:03.607000",
              "content": "<p>it should be somthing like:</p>\n<pre><code> = kidney_1_dense  (sparsity=%)\n = kidney_3_dense (sparsity=%)\n\n        th      dice (CV)       hit rate      fp rate       dice (lb)    \n.             .          .           .      .          xxx\n.            .          .             .         .        yyy\n.g. only ....\n. ...\n</code></pre>\n<p>build tables for train/ validation and for different fold etc ….</p>\n<p>you should note that lower fp usually gives better results.</p>\n<p>when we says low fp, it means low mean and low std of fp values over a range of theshold (e.g. from 0.25 to 0.35). we want performance metrics not sensitive the threshold</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2588737,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-01-05T17:27:45.523000",
              "content": "<p>percentile is not important. it is the \"threshold value at that percentile\"  is important</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2589453,
              "author_name": "tbt",
              "author_url": "",
              "post_date": "2024-01-06T11:50:15.110000",
              "content": "<p>Thanks again for the advice. The results are as follows.<br>\nAs you point out, it seems more reasonable to refer to the threshold determined based on the percentile, rather than the percentile itself.<br>\nI will continue to investigate.</p>\n<p>train = kidney_1_dense, valid = kidney_3_dense</p>\n<table>\n<thead>\n<tr>\n<th>percentile</th>\n<th>th</th>\n<th>dice(CV)</th>\n<th>hit rate</th>\n<th>fp</th>\n<th>dice(LB)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.0010</td>\n<td>0.93</td>\n<td>0.54</td>\n<td>0.44</td>\n<td>0.003</td>\n<td>?</td>\n</tr>\n<tr>\n<td>0.0020</td>\n<td>0.55</td>\n<td>0.83</td>\n<td>0.87</td>\n<td>0.016</td>\n<td>0.852</td>\n</tr>\n<tr>\n<td>0.0025</td>\n<td>0.11</td>\n<td>0.82</td>\n<td>0.93</td>\n<td>0.154</td>\n<td>?</td>\n</tr>\n<tr>\n<td>0.0030</td>\n<td>0.02</td>\n<td>0.75</td>\n<td>0.96</td>\n<td>0.280</td>\n<td>?</td>\n</tr>\n</tbody>\n</table>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2589499,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-01-06T12:41:15.930000",
              "content": "<p>check this :<br>\n<a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/456118#2589498\" target=\"_blank\">https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/456118#2589498</a></p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2587923": "The recently published notebook ([https://www.kaggle.com/code/misakimatsutomo/inference-1024-should-have-a-percentile-of-0-00149](https://www.kaggle.com/code/misakimatsutomo/inference-1024-should-have-a-percentile-of-0-00149)) shows a high public score. However, the threshold is determined by the percentile method, which may result in overfitting the public LB.\n\nWhat are your thoughts on how to determine an appropriate threshold? I believe the most basic way to determine the threshold is by looking at the CV of the training data, but I am struggling with that method as the sample is small for this competition.\n\nFor example, the percentage of positive examples for each training data set is as follows and varies from sample to sample, even allowing for poor segmentation.\n\n| sample | proportion of positive value|\n| --- | --- |\n| kidney_1_dense | 0.0062 |\n| kidney_1_voi | 0.012 |\n| kidney_2 | 0.0041 |\n| kidney_3_dense | 0.0022 |\n| kidney_3_sparse | 0.0027 |\n",
    "2588316": "can you show the actual probability, surface dice, hit rate, fp rate values for each percentile indicated in your table above?\n\n```\npercentile=0.0185\n\nth = np.percentile(probability,percentile)\nprint('actual threshold is', th)\n\npredict = probability>th\nprint('surface dice', compute_surface_dice(predict, truth))\n\nhit_rate = (predict*truth).sum()/truth.sum()\nprint('hit_rate', hit_rate)\n....\n\n\n```"
  }
}