{
  "id": 469046,
  "title": "Some questions about the codes in public notebooks",
  "url": "/competitions/blood-vessel-segmentation/discussion/469046",
  "author_name": "",
  "post_date": "2024-01-18T21:36:50.209091400Z",
  "votes": 7,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Its my first segmentation competition and I have some hard time understanding the following code blocks + the purpose of using them. <br>\n1-</p>\n<pre><code> ():\n    dim=((,x.ndim))\n    mean=x.mean(dim=dim,keepdim=)\n    std=x.std(dim=dim,keepdim=)\n    x=(x-mean)/(std+smooth)\n    x[x&gt;]=(x[x&gt;]-)* +\n    x[x&lt;-]=(x[x&lt;-]+)*-\n     x\n</code></pre>\n<p>Why not just normalizing? <br>\nAlso, its recalled for each dataset separately. So, aren't we getting different scaling for each dataset? What is the intuition behind this? Based on my knowledge I don't think this is generalizable. Or maybe people just used it because it gave high public lb (overfitting)?</p>\n<p>2-</p>\n<pre><code>\nTH=x.reshape(-).numpy()\nindex = -((TH) * CFG.chopping_percentile)\nTH: = np.partition(TH, index)[index]\nx[x&gt;TH]=(TH)\n\nTH=x.reshape(-).numpy()\nindex = -((TH) * CFG.chopping_percentile)\nTH: = np.partition(TH, -index)[-index]\nx[x&lt;TH]=(TH)\n\nx=(min_max_normalization(x.to(tc.float16)[])[]*).to(tc.uint8)\n</code></pre>\n<p>What is happening here exactly? and how do we determine the chopping_percentile? hyperparam?</p>\n<p>and the same code goes here in inference but with a totally different percentile value:</p>\n<pre><code>TH=[x.flatten().numpy()  x  output]\nTH=np.concatenate(TH)\nindex = -((TH) * CFG.th_percentile)\nTH: = np.partition(TH, index)[index]\n(TH)\n</code></pre>\n<p>Also, why the first code is used for the training \"images\", while the second one is used for the predicted \"masks\" (labels)?</p>\n<p>3-<br>\nAnd the main question is: <br>\ndo i really need all these things :)?<br>\nThanks!</p>",
  "messages": [
    {
      "id": "2608526",
      "postDate": "01/18/2024 21:36:50",
      "content": "<p>Its my first segmentation competition and I have some hard time understanding the following code blocks + the purpose of using them. <br>\n1-</p>\n<pre><code> ():\n    dim=((,x.ndim))\n    mean=x.mean(dim=dim,keepdim=)\n    std=x.std(dim=dim,keepdim=)\n    x=(x-mean)/(std+smooth)\n    x[x&gt;]=(x[x&gt;]-)* +\n    x[x&lt;-]=(x[x&lt;-]+)*-\n     x\n</code></pre>\n<p>Why not just normalizing? <br>\nAlso, its recalled for each dataset separately. So, aren't we getting different scaling for each dataset? What is the intuition behind this? Based on my knowledge I don't think this is generalizable. Or maybe people just used it because it gave high public lb (overfitting)?</p>\n<p>2-</p>\n<pre><code>\nTH=x.reshape(-).numpy()\nindex = -((TH) * CFG.chopping_percentile)\nTH: = np.partition(TH, index)[index]\nx[x&gt;TH]=(TH)\n\nTH=x.reshape(-).numpy()\nindex = -((TH) * CFG.chopping_percentile)\nTH: = np.partition(TH, -index)[-index]\nx[x&lt;TH]=(TH)\n\nx=(min_max_normalization(x.to(tc.float16)[])[]*).to(tc.uint8)\n</code></pre>\n<p>What is happening here exactly? and how do we determine the chopping_percentile? hyperparam?</p>\n<p>and the same code goes here in inference but with a totally different percentile value:</p>\n<pre><code>TH=[x.flatten().numpy()  x  output]\nTH=np.concatenate(TH)\nindex = -((TH) * CFG.th_percentile)\nTH: = np.partition(TH, index)[index]\n(TH)\n</code></pre>\n<p>Also, why the first code is used for the training \"images\", while the second one is used for the predicted \"masks\" (labels)?</p>\n<p>3-<br>\nAnd the main question is: <br>\ndo i really need all these things :)?<br>\nThanks!</p>",
      "rawMarkdown": "Its my first segmentation competition and I have some hard time understanding the following code blocks + the purpose of using them. \n1-\n```python\ndef norm_with_clip(x:tc.Tensor,smooth=1e-5):\n    dim=list(range(1,x.ndim))\n    mean=x.mean(dim=dim,keepdim=True)\n    std=x.std(dim=dim,keepdim=True)\n    x=(x-mean)/(std+smooth)\n    x[x>5]=(x[x>5]-5)*1e-3 +5\n    x[x<-3]=(x[x<-3]+3)*1e-3-3\n    return x\n```\nWhy not just normalizing? \nAlso, its recalled for each dataset separately. So, aren't we getting different scaling for each dataset? What is the intuition behind this? Based on my knowledge I don't think this is generalizable. Or maybe people just used it because it gave high public lb (overfitting)?\n\n2-\n```python\n########################################################################\nTH=x.reshape(-1).numpy()\nindex = -int(len(TH) * CFG.chopping_percentile)\nTH:int = np.partition(TH, index)[index]\nx[x>TH]=int(TH)\n########################################################################\nTH=x.reshape(-1).numpy()\nindex = -int(len(TH) * CFG.chopping_percentile)\nTH:int = np.partition(TH, -index)[-index]\nx[x<TH]=int(TH)\n########################################################################\nx=(min_max_normalization(x.to(tc.float16)[None])[0]*255).to(tc.uint8)\n```\nWhat is happening here exactly? and how do we determine the chopping_percentile? hyperparam?\n\nand the same code goes here in inference but with a totally different percentile value:\n\n```python\nTH=[x.flatten().numpy() for x in output]\nTH=np.concatenate(TH)\nindex = -int(len(TH) * CFG.th_percentile)\nTH:int = np.partition(TH, index)[index]\nprint(TH)\n```\n\nAlso, why the first code is used for the training \"images\", while the second one is used for the predicted \"masks\" (labels)?\n\n3-\nAnd the main question is: \ndo i really need all these things :)?\nThanks!",
      "votes": null
    },
    {
      "id": "2608534",
      "postDate": "01/18/2024 22:06:15",
      "content": "<p>My approach is slighly different. But as I understand, first training TH is for clip outliers while second mask TH is for artificially binarize the final output (allowing more or less positives).<br>\nBut yes, general opinion is that score is mask TH overfitted.</p>",
      "rawMarkdown": "My approach is slighly different. But as I understand, first training TH is for clip outliers while second mask TH is for artificially binarize the final output (allowing more or less positives).\nBut yes, general opinion is that score is mask TH overfitted.",
      "votes": null
    },
    {
      "id": "2608573",
      "postDate": "01/18/2024 23:20:38",
      "content": "<p>I would say this is slightly related to this discussion <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/469022\" target=\"_blank\">https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/469022</a> </p>\n<p>I think you don’t need them but it could help if done well. I do think that the structure in this particular case is unnecessary. Scaling each dataset separately seems strange to me as well. I would have thought it to be better to scale together based off train and apply that scale to test. But that may also be different with the different resolutions of the train and test.</p>",
      "rawMarkdown": "I would say this is slightly related to this discussion https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/469022 \n\nI think you don’t need them but it could help if done well. I do think that the structure in this particular case is unnecessary. Scaling each dataset separately seems strange to me as well. I would have thought it to be better to scale together based off train and apply that scale to test. But that may also be different with the different resolutions of the train and test.",
      "votes": null
    },
    {
      "id": "2611100",
      "postDate": "01/20/2024 14:59:40",
      "content": "<p>It looks like shamanism, but it works quite effectively. I’ll show you using the first era as an example:</p>\n<p>With normalization:<br>\nepoch:0,loss:0.9154,score:0.2735,lr2.4942e-05: 100%|██████████| 4603/4603 [16:50&lt;00:00,  4.56it/s]  \nval--&gt;loss:0.7888,score:0.7206: 100%|██████████| 250/250 [00:29&lt;00:00,  8.56it/s]</p>\n<p>Without normalization:<br>\nepoch:0,loss:0.9674,score:0.0396,lr5.0430e-06: 100%|██████████| 1151/1151 [04:37&lt;00:00,  4.15it/s]  \nval--&gt;loss:0.9779,score:0.0471: 100%|██████████| 125/125 [00:08&lt;00:00, 14.91it/s]</p>\n<p>The result is obvious: score:0.7206 vs score:0.0471.</p>\n<p>An example to understand the work:</p>\n<p>To find the top k or bottom k elements without having to first sort the entire array.<br>\nimport numpy as np<br>\nimport torch as tc</p>\n<p>x = np.random.randint(low=0, high=255, size=(10,10))<br>\nx = tc.tensor(x, dtype=tc.int32)<br>\nchopping_percentile = 1e-1<br>\ny = x.reshape(-1).numpy()<br>\nprint(f\"y.shape={y.shape}\")<br>\nk = -int(len(y) * chopping_percentile)<br>\nprint(f\"k={k}\")</p>\n<p>// To find the top k:<br>\ntop_k = np.partition(y, -k)[-k:]  // Results are not sorted<br>\nprint(f\"top_k={top_k}\")<br>\nSorted from largest to smallest<br>\n// np.sort(top_k)[::-1]  </p>\n<p>// To find the lower k:<br>\nbottom_k = np.partition(y, k)[:k]  // Results are not sorted<br>\nprint(f\"bottom_k={bottom_k}\")<br>\n// np.sort(bottom_k)  // Sorted from smallest to largest</p>\n<p>TH = x.reshape(-1).numpy()<br>\nindex = -int(len(TH) * chopping_percentile)<br>\nTH:int = np.partition(TH, index)[index]<br>\nprint(f\"TH (High Limit)={TH}\")<br>\nx[x &gt; TH] = int(TH)  // Replace values exceeding threshold with threshold value</p>\n<p>TH = x.reshape(-1).numpy()<br>\nindex = -int(len(TH) * chopping_percentile)<br>\nTH:int = np.partition(TH, -index)[-index]<br>\nprint(f\"TH (Low Limit)={TH}\")<br>\nx[x &lt; TH] = int(TH)  // Replace value less than threshold with threshold value</p>\n<p>print(f\"x={x}\")</p>",
      "rawMarkdown": "It looks like shamanism, but it works quite effectively. I’ll show you using the first era as an example:\n\nWith normalization:\nepoch:0,loss:0.9154,score:0.2735,lr2.4942e-05: 100%|██████████| 4603/4603 [16:50<00:00,  4.56it/s]  \nval-->loss:0.7888,score:0.7206: 100%|██████████| 250/250 [00:29<00:00,  8.56it/s]\n\nWithout normalization:\nepoch:0,loss:0.9674,score:0.0396,lr5.0430e-06: 100%|██████████| 1151/1151 [04:37<00:00,  4.15it/s]  \nval-->loss:0.9779,score:0.0471: 100%|██████████| 125/125 [00:08<00:00, 14.91it/s]\n\nThe result is obvious: score:0.7206 vs score:0.0471.\n\nAn example to understand the work:\n\nTo find the top k or bottom k elements without having to first sort the entire array.\nimport numpy as np\nimport torch as tc\n\nx = np.random.randint(low=0, high=255, size=(10,10))\nx = tc.tensor(x, dtype=tc.int32)\nchopping_percentile = 1e-1\ny = x.reshape(-1).numpy()\nprint(f\"y.shape={y.shape}\")\nk = -int(len(y) * chopping_percentile)\nprint(f\"k={k}\")\n\n// To find the top k:\ntop_k = np.partition(y, -k)[-k:]  // Results are not sorted\nprint(f\"top_k={top_k}\")\nSorted from largest to smallest\n// np.sort(top_k)[::-1]  \n\n// To find the lower k:\nbottom_k = np.partition(y, k)[:k]  // Results are not sorted\nprint(f\"bottom_k={bottom_k}\")\n// np.sort(bottom_k)  // Sorted from smallest to largest\n\nTH = x.reshape(-1).numpy()\nindex = -int(len(TH) * chopping_percentile)\nTH:int = np.partition(TH, index)[index]\nprint(f\"TH (High Limit)={TH}\")\nx[x > TH] = int(TH)  // Replace values exceeding threshold with threshold value\n\nTH = x.reshape(-1).numpy()\nindex = -int(len(TH) * chopping_percentile)\nTH:int = np.partition(TH, -index)[-index]\nprint(f\"TH (Low Limit)={TH}\")\nx[x < TH] = int(TH)  // Replace value less than threshold with threshold value\n\nprint(f\"x={x}\")",
      "votes": null
    },
    {
      "id": "2635495",
      "postDate": "02/04/2024 12:46:55",
      "content": "<p>Similarly, in the public training notebook, there's an application of a threshold to the model input with the expression y = data['mask'] &gt;= 127. Is it better to apply a threshold of 0.5 to the training input as well?</p>",
      "rawMarkdown": "Similarly, in the public training notebook, there's an application of a threshold to the model input with the expression y = data['mask'] >= 127. Is it better to apply a threshold of 0.5 to the training input as well?",
      "votes": null
    },
    {
      "id": "2635520",
      "postDate": "02/04/2024 12:53:54",
      "content": "<p>If is applied to input, input labels are 0 or 255. It doesn't matter as far is &gt; 0</p>\n<p>EDIT: Unless they are interpolating the input masks, then yes, 127 seems reasonable. I haven't good experience interpolating masks.</p>",
      "rawMarkdown": "If is applied to input, input labels are 0 or 255. It doesn't matter as far is > 0\n\nEDIT: Unless they are interpolating the input masks, then yes, 127 seems reasonable. I haven't good experience interpolating masks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2608534,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "01/18/2024 22:06:15",
      "content": "<p>My approach is slighly different. But as I understand, first training TH is for clip outliers while second mask TH is for artificially binarize the final output (allowing more or less positives).<br>\nBut yes, general opinion is that score is mask TH overfitted.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2608573,
      "author_name": "cody11null",
      "author_url": "",
      "post_date": "01/18/2024 23:20:38",
      "content": "<p>I would say this is slightly related to this discussion <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/469022\" target=\"_blank\">https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/469022</a> </p>\n<p>I think you don’t need them but it could help if done well. I do think that the structure in this particular case is unnecessary. Scaling each dataset separately seems strange to me as well. I would have thought it to be better to scale together based off train and apply that scale to test. But that may also be different with the different resolutions of the train and test.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2611100,
      "author_name": "konstantinboyko",
      "author_url": "",
      "post_date": "01/20/2024 14:59:40",
      "content": "<p>It looks like shamanism, but it works quite effectively. I’ll show you using the first era as an example:</p>\n<p>With normalization:<br>\nepoch:0,loss:0.9154,score:0.2735,lr2.4942e-05: 100%|██████████| 4603/4603 [16:50&lt;00:00,  4.56it/s]  \nval--&gt;loss:0.7888,score:0.7206: 100%|██████████| 250/250 [00:29&lt;00:00,  8.56it/s]</p>\n<p>Without normalization:<br>\nepoch:0,loss:0.9674,score:0.0396,lr5.0430e-06: 100%|██████████| 1151/1151 [04:37&lt;00:00,  4.15it/s]  \nval--&gt;loss:0.9779,score:0.0471: 100%|██████████| 125/125 [00:08&lt;00:00, 14.91it/s]</p>\n<p>The result is obvious: score:0.7206 vs score:0.0471.</p>\n<p>An example to understand the work:</p>\n<p>To find the top k or bottom k elements without having to first sort the entire array.<br>\nimport numpy as np<br>\nimport torch as tc</p>\n<p>x = np.random.randint(low=0, high=255, size=(10,10))<br>\nx = tc.tensor(x, dtype=tc.int32)<br>\nchopping_percentile = 1e-1<br>\ny = x.reshape(-1).numpy()<br>\nprint(f\"y.shape={y.shape}\")<br>\nk = -int(len(y) * chopping_percentile)<br>\nprint(f\"k={k}\")</p>\n<p>// To find the top k:<br>\ntop_k = np.partition(y, -k)[-k:]  // Results are not sorted<br>\nprint(f\"top_k={top_k}\")<br>\nSorted from largest to smallest<br>\n// np.sort(top_k)[::-1]  </p>\n<p>// To find the lower k:<br>\nbottom_k = np.partition(y, k)[:k]  // Results are not sorted<br>\nprint(f\"bottom_k={bottom_k}\")<br>\n// np.sort(bottom_k)  // Sorted from smallest to largest</p>\n<p>TH = x.reshape(-1).numpy()<br>\nindex = -int(len(TH) * chopping_percentile)<br>\nTH:int = np.partition(TH, index)[index]<br>\nprint(f\"TH (High Limit)={TH}\")<br>\nx[x &gt; TH] = int(TH)  // Replace values exceeding threshold with threshold value</p>\n<p>TH = x.reshape(-1).numpy()<br>\nindex = -int(len(TH) * chopping_percentile)<br>\nTH:int = np.partition(TH, -index)[-index]<br>\nprint(f\"TH (Low Limit)={TH}\")<br>\nx[x &lt; TH] = int(TH)  // Replace value less than threshold with threshold value</p>\n<p>print(f\"x={x}\")</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2635495,
      "author_name": "aravind36",
      "author_url": "",
      "post_date": "02/04/2024 12:46:55",
      "content": "<p>Similarly, in the public training notebook, there's an application of a threshold to the model input with the expression y = data['mask'] &gt;= 127. Is it better to apply a threshold of 0.5 to the training input as well?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2635520,
          "author_name": "sacuscreed",
          "author_url": "",
          "post_date": "02/04/2024 12:53:54",
          "content": "<p>If is applied to input, input labels are 0 or 255. It doesn't matter as far is &gt; 0</p>\n<p>EDIT: Unless they are interpolating the input masks, then yes, 127 seems reasonable. I haven't good experience interpolating masks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2608526": "Its my first segmentation competition and I have some hard time understanding the following code blocks + the purpose of using them. \n1-\n```python\ndef norm_with_clip(x:tc.Tensor,smooth=1e-5):\n    dim=list(range(1,x.ndim))\n    mean=x.mean(dim=dim,keepdim=True)\n    std=x.std(dim=dim,keepdim=True)\n    x=(x-mean)/(std+smooth)\n    x[x>5]=(x[x>5]-5)*1e-3 +5\n    x[x<-3]=(x[x<-3]+3)*1e-3-3\n    return x\n```\nWhy not just normalizing? \nAlso, its recalled for each dataset separately. So, aren't we getting different scaling for each dataset? What is the intuition behind this? Based on my knowledge I don't think this is generalizable. Or maybe people just used it because it gave high public lb (overfitting)?\n\n2-\n```python\n########################################################################\nTH=x.reshape(-1).numpy()\nindex = -int(len(TH) * CFG.chopping_percentile)\nTH:int = np.partition(TH, index)[index]\nx[x>TH]=int(TH)\n########################################################################\nTH=x.reshape(-1).numpy()\nindex = -int(len(TH) * CFG.chopping_percentile)\nTH:int = np.partition(TH, -index)[-index]\nx[x<TH]=int(TH)\n########################################################################\nx=(min_max_normalization(x.to(tc.float16)[None])[0]*255).to(tc.uint8)\n```\nWhat is happening here exactly? and how do we determine the chopping_percentile? hyperparam?\n\nand the same code goes here in inference but with a totally different percentile value:\n\n```python\nTH=[x.flatten().numpy() for x in output]\nTH=np.concatenate(TH)\nindex = -int(len(TH) * CFG.th_percentile)\nTH:int = np.partition(TH, index)[index]\nprint(TH)\n```\n\nAlso, why the first code is used for the training \"images\", while the second one is used for the predicted \"masks\" (labels)?\n\n3-\nAnd the main question is: \ndo i really need all these things :)?\nThanks!",
    "2608534": "My approach is slighly different. But as I understand, first training TH is for clip outliers while second mask TH is for artificially binarize the final output (allowing more or less positives).\nBut yes, general opinion is that score is mask TH overfitted.",
    "2608573": "I would say this is slightly related to this discussion https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/469022 \n\nI think you don’t need them but it could help if done well. I do think that the structure in this particular case is unnecessary. Scaling each dataset separately seems strange to me as well. I would have thought it to be better to scale together based off train and apply that scale to test. But that may also be different with the different resolutions of the train and test.",
    "2611100": "It looks like shamanism, but it works quite effectively. I’ll show you using the first era as an example:\n\nWith normalization:\nepoch:0,loss:0.9154,score:0.2735,lr2.4942e-05: 100%|██████████| 4603/4603 [16:50<00:00,  4.56it/s]  \nval-->loss:0.7888,score:0.7206: 100%|██████████| 250/250 [00:29<00:00,  8.56it/s]\n\nWithout normalization:\nepoch:0,loss:0.9674,score:0.0396,lr5.0430e-06: 100%|██████████| 1151/1151 [04:37<00:00,  4.15it/s]  \nval-->loss:0.9779,score:0.0471: 100%|██████████| 125/125 [00:08<00:00, 14.91it/s]\n\nThe result is obvious: score:0.7206 vs score:0.0471.\n\nAn example to understand the work:\n\nTo find the top k or bottom k elements without having to first sort the entire array.\nimport numpy as np\nimport torch as tc\n\nx = np.random.randint(low=0, high=255, size=(10,10))\nx = tc.tensor(x, dtype=tc.int32)\nchopping_percentile = 1e-1\ny = x.reshape(-1).numpy()\nprint(f\"y.shape={y.shape}\")\nk = -int(len(y) * chopping_percentile)\nprint(f\"k={k}\")\n\n// To find the top k:\ntop_k = np.partition(y, -k)[-k:]  // Results are not sorted\nprint(f\"top_k={top_k}\")\nSorted from largest to smallest\n// np.sort(top_k)[::-1]  \n\n// To find the lower k:\nbottom_k = np.partition(y, k)[:k]  // Results are not sorted\nprint(f\"bottom_k={bottom_k}\")\n// np.sort(bottom_k)  // Sorted from smallest to largest\n\nTH = x.reshape(-1).numpy()\nindex = -int(len(TH) * chopping_percentile)\nTH:int = np.partition(TH, index)[index]\nprint(f\"TH (High Limit)={TH}\")\nx[x > TH] = int(TH)  // Replace values exceeding threshold with threshold value\n\nTH = x.reshape(-1).numpy()\nindex = -int(len(TH) * chopping_percentile)\nTH:int = np.partition(TH, -index)[-index]\nprint(f\"TH (Low Limit)={TH}\")\nx[x < TH] = int(TH)  // Replace value less than threshold with threshold value\n\nprint(f\"x={x}\")",
    "2635495": "Similarly, in the public training notebook, there's an application of a threshold to the model input with the expression y = data['mask'] >= 127. Is it better to apply a threshold of 0.5 to the training input as well?",
    "2635520": "If is applied to input, input labels are 0 or 255. It doesn't matter as far is > 0\n\nEDIT: Unless they are interpolating the input masks, then yes, 127 seems reasonable. I haven't good experience interpolating masks."
  },
  "source": "meta"
}