{
  "id": 468319,
  "title": "Model performing well on validation (`kidney_3_dense`) while failing miserably on the public test set",
  "url": "/competitions/blood-vessel-segmentation/discussion/468319",
  "author_name": "",
  "post_date": "2024-01-16T08:09:59.559927300Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I am training a se_resnext-50_32x4d architecture on the \"kidney_1_dense\" while validating on the \"kidney_3_dense\" one.</p>\n<p>To normalize the images, I first scale the pixel values in the 0, 1 range, then I apply normalization over each image. Then I apply some random data augmentation on the train data using the <code>monai</code> and <code>albumentation</code> libraries, while keeping the validation data unaltered.</p>\n<pre><code>deterministic_transforms = Compose(\n    [\n        LoadImaged(keys=[, ], =PILReader, =),\n        ScaleIntensityRanged(\n            keys=[],\n            =0,\n            =2**16,\n            =0.0,\n            =1.0,\n            =,\n        ),\n        ScaleIntensityRanged(\n            keys=[],\n            =0,\n            =255,\n            =0.0,\n            =1.0,\n            =,\n        ),\n        NormalizeIntensityd(\n            keys=[],\n        ),\n        EnsureTyped(keys=[, ], =device, =),\n    ]\n)\n</code></pre>\n<p>These are the only transforms I apply to the validation data. While I also resize the input training image to a target size of 1024, the validation ones are kept as they are and inference is done in a sliding window manner using the <code>sliding_window_inference</code> function from <code>monai.inferers</code>.</p>\n<p>I trained this model and it seems to be converging. Once I finish training, I load the best metric model and do a final validation pass on the \"kidney_3_dense\" data, I get the following results:</p>\n<pre><code>\n metric = .\n rate = .\n positive rate = .\n\n metric = .\n rate = .\n positive rate = .\n</code></pre>\n<p>which to me appear to be pretty decent validation results with a pretty narrow generalization gap. Also the train and validation losses/metrics curves over the iterations seems okay to me. </p>\n<p>I then happily submitted this model (by applying the same <code>deterministic_transforms</code> to the test images) and got an astonishing score of 0.009!</p>\n<p>I think the problem lies in the way I am normalizing the images, but I saw that the majority of the public notebooks are doing per image normalization with clipping in a similar way than me, by using the following function:</p>\n<pre><code> norm_with_clip(x:tc.Tensor,smooth=e-):\n    =list(range(,x.ndim))\n    =x.mean(dim=dim,keepdim=True)\n    =x.std(dim=dim,keepdim=True)\n    =(x-mean)/(std+smooth)\n    [x&gt;]=(x[x&gt;]-)*e- +\n    [x&lt;-]=(x[x&lt;-]+)*e--\n     x\n</code></pre>\n<p>While (if I understood correctly) still performing training on <code>kidney_1_dense</code> and validation on <code>kidney_3_dense</code>. Also I obviously made sure that no gradient is leaked from the validation set with <code>torch.no_grad()</code> context manager (but this could be a potential reason behind the high validation scores). </p>\n<p>To compute the dice loss I used yet again the <code>monai</code> implementation, while I compute the hit rate and false positive rates as follows.</p>\n<pre><code>        hit_rate = (predictions*ground_truth).()/ground_truth.()\n        false_positive_rate = (predictions* * ( - ground_truth)).() / ground_truth.()\n</code></pre>\n<p>I am missing something big here and I need to reflect more on it, but I wanted to ask here anyway to see if someone could already have some suggestions about this problem. Maybe this only means that using only these two datasets is not enough for achieving a decent result, but the weird thing is that I did my first submission with a basic <code>Unet</code> architecture with no precompiled backbone and using the same datasets as train and validation (but only pixel range scaling without normalization) and I got 0.606 which is still low but nothing compared to this new 0.009. I also tried to remove the per image normalization on the se_resnext architecture, but it overfits straight away as expected…</p>",
  "messages": [
    {
      "id": "2604038",
      "postDate": "01/16/2024 08:09:59",
      "content": "<p>I am training a se_resnext-50_32x4d architecture on the \"kidney_1_dense\" while validating on the \"kidney_3_dense\" one.</p>\n<p>To normalize the images, I first scale the pixel values in the 0, 1 range, then I apply normalization over each image. Then I apply some random data augmentation on the train data using the <code>monai</code> and <code>albumentation</code> libraries, while keeping the validation data unaltered.</p>\n<pre><code>deterministic_transforms = Compose(\n    [\n        LoadImaged(keys=[, ], =PILReader, =),\n        ScaleIntensityRanged(\n            keys=[],\n            =0,\n            =2**16,\n            =0.0,\n            =1.0,\n            =,\n        ),\n        ScaleIntensityRanged(\n            keys=[],\n            =0,\n            =255,\n            =0.0,\n            =1.0,\n            =,\n        ),\n        NormalizeIntensityd(\n            keys=[],\n        ),\n        EnsureTyped(keys=[, ], =device, =),\n    ]\n)\n</code></pre>\n<p>These are the only transforms I apply to the validation data. While I also resize the input training image to a target size of 1024, the validation ones are kept as they are and inference is done in a sliding window manner using the <code>sliding_window_inference</code> function from <code>monai.inferers</code>.</p>\n<p>I trained this model and it seems to be converging. Once I finish training, I load the best metric model and do a final validation pass on the \"kidney_3_dense\" data, I get the following results:</p>\n<pre><code>\n metric = .\n rate = .\n positive rate = .\n\n metric = .\n rate = .\n positive rate = .\n</code></pre>\n<p>which to me appear to be pretty decent validation results with a pretty narrow generalization gap. Also the train and validation losses/metrics curves over the iterations seems okay to me. </p>\n<p>I then happily submitted this model (by applying the same <code>deterministic_transforms</code> to the test images) and got an astonishing score of 0.009!</p>\n<p>I think the problem lies in the way I am normalizing the images, but I saw that the majority of the public notebooks are doing per image normalization with clipping in a similar way than me, by using the following function:</p>\n<pre><code> norm_with_clip(x:tc.Tensor,smooth=e-):\n    =list(range(,x.ndim))\n    =x.mean(dim=dim,keepdim=True)\n    =x.std(dim=dim,keepdim=True)\n    =(x-mean)/(std+smooth)\n    [x&gt;]=(x[x&gt;]-)*e- +\n    [x&lt;-]=(x[x&lt;-]+)*e--\n     x\n</code></pre>\n<p>While (if I understood correctly) still performing training on <code>kidney_1_dense</code> and validation on <code>kidney_3_dense</code>. Also I obviously made sure that no gradient is leaked from the validation set with <code>torch.no_grad()</code> context manager (but this could be a potential reason behind the high validation scores). </p>\n<p>To compute the dice loss I used yet again the <code>monai</code> implementation, while I compute the hit rate and false positive rates as follows.</p>\n<pre><code>        hit_rate = (predictions*ground_truth).()/ground_truth.()\n        false_positive_rate = (predictions* * ( - ground_truth)).() / ground_truth.()\n</code></pre>\n<p>I am missing something big here and I need to reflect more on it, but I wanted to ask here anyway to see if someone could already have some suggestions about this problem. Maybe this only means that using only these two datasets is not enough for achieving a decent result, but the weird thing is that I did my first submission with a basic <code>Unet</code> architecture with no precompiled backbone and using the same datasets as train and validation (but only pixel range scaling without normalization) and I got 0.606 which is still low but nothing compared to this new 0.009. I also tried to remove the per image normalization on the se_resnext architecture, but it overfits straight away as expected…</p>",
      "rawMarkdown": "I am training a se_resnext-50_32x4d architecture on the \"kidney_1_dense\" while validating on the \"kidney_3_dense\" one.\n\nTo normalize the images, I first scale the pixel values in the 0, 1 range, then I apply normalization over each image. Then I apply some random data augmentation on the train data using the `monai` and `albumentation` libraries, while keeping the validation data unaltered.\n\n```\ndeterministic_transforms = Compose(\n    [\n        LoadImaged(keys=[\"image\", \"label\"], reader=PILReader, ensure_channel_first=True),\n        ScaleIntensityRanged(\n            keys=[\"image\"],\n            a_min=0,\n            a_max=2**16,\n            b_min=0.0,\n            b_max=1.0,\n            clip=True,\n        ),\n        ScaleIntensityRanged(\n            keys=[\"label\"],\n            a_min=0,\n            a_max=255,\n            b_min=0.0,\n            b_max=1.0,\n            clip=True,\n        ),\n        NormalizeIntensityd(\n            keys=[\"image\"],\n        ),\n        EnsureTyped(keys=[\"image\", \"label\"], device=device, track_meta=False),\n    ]\n)\n```\nThese are the only transforms I apply to the validation data. While I also resize the input training image to a target size of 1024, the validation ones are kept as they are and inference is done in a sliding window manner using the `sliding_window_inference` function from `monai.inferers`.\n\nI trained this model and it seems to be converging. Once I finish training, I load the best metric model and do a final validation pass on the \"kidney_3_dense\" data, I get the following results:\n\n```\n# results on kidney_3_dense (validation)\ndice metric = 0.8816532492637634\nhit rate = 0.6700526073811546\nfalse positive rate = 0.025235866517981605\n# results on kidney_1_dense (training)\ndice metric = 0.9021822810173035\nhit rate = 0.8241333725159629\nfalse positive rate = 0.055753962012628715\n```\n\nwhich to me appear to be pretty decent validation results with a pretty narrow generalization gap. Also the train and validation losses/metrics curves over the iterations seems okay to me. \n\nI then happily submitted this model (by applying the same `deterministic_transforms` to the test images) and got an astonishing score of 0.009!\n\nI think the problem lies in the way I am normalizing the images, but I saw that the majority of the public notebooks are doing per image normalization with clipping in a similar way than me, by using the following function:\n\n```\ndef norm_with_clip(x:tc.Tensor,smooth=1e-5):\n    dim=list(range(1,x.ndim))\n    mean=x.mean(dim=dim,keepdim=True)\n    std=x.std(dim=dim,keepdim=True)\n    x=(x-mean)/(std+smooth)\n    x[x>5]=(x[x>5]-5)*1e-3 +5\n    x[x<-3]=(x[x<-3]+3)*1e-3-3\n    return x\n```\n\nWhile (if I understood correctly) still performing training on `kidney_1_dense` and validation on `kidney_3_dense`. Also I obviously made sure that no gradient is leaked from the validation set with `torch.no_grad()` context manager (but this could be a potential reason behind the high validation scores). \n\nTo compute the dice loss I used yet again the `monai` implementation, while I compute the hit rate and false positive rates as follows.\n\n```\n\n        hit_rate = (predictions*ground_truth).sum()/ground_truth.sum()\n        false_positive_rate = (predictions* * (1 - ground_truth)).sum() / ground_truth.sum()\n```\n\nI am missing something big here and I need to reflect more on it, but I wanted to ask here anyway to see if someone could already have some suggestions about this problem. Maybe this only means that using only these two datasets is not enough for achieving a decent result, but the weird thing is that I did my first submission with a basic `Unet` architecture with no precompiled backbone and using the same datasets as train and validation (but only pixel range scaling without normalization) and I got 0.606 which is still low but nothing compared to this new 0.009. I also tried to remove the per image normalization on the se_resnext architecture, but it overfits straight away as expected...",
      "votes": null
    },
    {
      "id": "2604566",
      "postDate": "01/16/2024 14:07:53",
      "content": "<p>did you visualize your normalised image?<br>\ne.g. save normalsied image as grayscale (0,255) png and take a look</p>",
      "rawMarkdown": "did you visualize your normalised image?\ne.g. save normalsied image as grayscale (0,255) png and take a look",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2604566,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "01/16/2024 14:07:53",
      "content": "<p>did you visualize your normalised image?<br>\ne.g. save normalsied image as grayscale (0,255) png and take a look</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2604038": "I am training a se_resnext-50_32x4d architecture on the \"kidney_1_dense\" while validating on the \"kidney_3_dense\" one.\n\nTo normalize the images, I first scale the pixel values in the 0, 1 range, then I apply normalization over each image. Then I apply some random data augmentation on the train data using the `monai` and `albumentation` libraries, while keeping the validation data unaltered.\n\n```\ndeterministic_transforms = Compose(\n    [\n        LoadImaged(keys=[\"image\", \"label\"], reader=PILReader, ensure_channel_first=True),\n        ScaleIntensityRanged(\n            keys=[\"image\"],\n            a_min=0,\n            a_max=2**16,\n            b_min=0.0,\n            b_max=1.0,\n            clip=True,\n        ),\n        ScaleIntensityRanged(\n            keys=[\"label\"],\n            a_min=0,\n            a_max=255,\n            b_min=0.0,\n            b_max=1.0,\n            clip=True,\n        ),\n        NormalizeIntensityd(\n            keys=[\"image\"],\n        ),\n        EnsureTyped(keys=[\"image\", \"label\"], device=device, track_meta=False),\n    ]\n)\n```\nThese are the only transforms I apply to the validation data. While I also resize the input training image to a target size of 1024, the validation ones are kept as they are and inference is done in a sliding window manner using the `sliding_window_inference` function from `monai.inferers`.\n\nI trained this model and it seems to be converging. Once I finish training, I load the best metric model and do a final validation pass on the \"kidney_3_dense\" data, I get the following results:\n\n```\n# results on kidney_3_dense (validation)\ndice metric = 0.8816532492637634\nhit rate = 0.6700526073811546\nfalse positive rate = 0.025235866517981605\n# results on kidney_1_dense (training)\ndice metric = 0.9021822810173035\nhit rate = 0.8241333725159629\nfalse positive rate = 0.055753962012628715\n```\n\nwhich to me appear to be pretty decent validation results with a pretty narrow generalization gap. Also the train and validation losses/metrics curves over the iterations seems okay to me. \n\nI then happily submitted this model (by applying the same `deterministic_transforms` to the test images) and got an astonishing score of 0.009!\n\nI think the problem lies in the way I am normalizing the images, but I saw that the majority of the public notebooks are doing per image normalization with clipping in a similar way than me, by using the following function:\n\n```\ndef norm_with_clip(x:tc.Tensor,smooth=1e-5):\n    dim=list(range(1,x.ndim))\n    mean=x.mean(dim=dim,keepdim=True)\n    std=x.std(dim=dim,keepdim=True)\n    x=(x-mean)/(std+smooth)\n    x[x>5]=(x[x>5]-5)*1e-3 +5\n    x[x<-3]=(x[x<-3]+3)*1e-3-3\n    return x\n```\n\nWhile (if I understood correctly) still performing training on `kidney_1_dense` and validation on `kidney_3_dense`. Also I obviously made sure that no gradient is leaked from the validation set with `torch.no_grad()` context manager (but this could be a potential reason behind the high validation scores). \n\nTo compute the dice loss I used yet again the `monai` implementation, while I compute the hit rate and false positive rates as follows.\n\n```\n\n        hit_rate = (predictions*ground_truth).sum()/ground_truth.sum()\n        false_positive_rate = (predictions* * (1 - ground_truth)).sum() / ground_truth.sum()\n```\n\nI am missing something big here and I need to reflect more on it, but I wanted to ask here anyway to see if someone could already have some suggestions about this problem. Maybe this only means that using only these two datasets is not enough for achieving a decent result, but the weird thing is that I did my first submission with a basic `Unet` architecture with no precompiled backbone and using the same datasets as train and validation (but only pixel range scaling without normalization) and I got 0.606 which is still low but nothing compared to this new 0.009. I also tried to remove the per image normalization on the se_resnext architecture, but it overfits straight away as expected...",
    "2604566": "did you visualize your normalised image?\ne.g. save normalsied image as grayscale (0,255) png and take a look"
  },
  "source": "meta"
}