{
  "id": 469022,
  "title": "Normalization, please stop it.",
  "url": "/competitions/blood-vessel-segmentation/discussion/469022",
  "author_name": "SSS",
  "post_date": "2024-01-18T18:55:22.335000",
  "votes": 17,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Hi,</p>\n<h2>Intro</h2>\n<p>I was going thru the notebooks and again and again I saw the same pattern.<br>\nPeople trained and normalized the data by using stats (min, max, std or mean, percentile depening on your choice) of the individual images. Or normalize each cuboid/volume by its own stats. And I even saw some folks normalize by percentile and on top of that use min-max, that's trully A Novel Approach in ML. </p>\n<p>This is what happens if you take middle slice along z axis and look into the distibutions:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F96912c5b5c17c77d2cb801d477f5fb6f%2FFigure_1.png?generation=1705593658292533&amp;alt=media\"></p>\n<p>1) The first subplot shows the distibution if we were to take all volumes and normalize it by its own stats, e.g.:</p>\n<pre><code> ():\n     (img - img_min) / (img_max - img_min)\n\nkid1_min, kid1_max = fit_cub_img[].(), fit_cub_img[].()\nkid2_min, kid2_max = fit_cub_img[].(), fit_cub_img[].()\nkid3_min, kid3_max = fit_cub_img[].(), fit_cub_img[].()\n\nscaled_kid1 = min_max_scaler(fit_cub_img[], kid1_min, kid1_max)\nscaled_kid2 = min_max_scaler(fit_cub_img[], kid2_min, kid2_max)\nscaled_kid3 = min_max_scaler(fit_cub_img[], kid3_min, kid3_max)\n</code></pre>\n<p>2) The second subplot shows if we were to train on kidney1 and apply its stats to kid2 or kid3 during the validation.</p>\n<pre><code>scaled_kid1 = min_max_scaler(fit_cub_img[], kid1_min, kid1_max)\nscaled_kid2 = min_max_scaler(fit_cub_img[], kid2_min, kid2_max)\nscaled_kid3 = min_max_scaler(fit_cub_img[], kid3_min, kid3_max)\n</code></pre>\n<p>Yes, we brought distribution of kid2 and kid3 closer to kid1.</p>\n<p>Adding the whole volume normalized histograms:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F6f26b0d58142e8f2180862ef59ec0645%2Ffig_volumes_hist.png?generation=1705633638305910&amp;alt=media\"></p>\n<p>Now the kidney3 closer to kidney1 with min_max by kidney1.</p>\n<p>I thought it was always <strong>a rule of thumb</strong> to fit and transform the train data and only transform the validation set.<br>\nIf you are lucky enough, the first scenario will yield a good score for you, but just because test set has close or the same stats as the one which you trained on. </p>\n<h2>Outro</h2>\n<p>It seems folks just normalize because they heard models perform better if you push everything to gaussian from 0 to 1. But I just want to remind you the reason we do <code>fit_transform</code> on train and <code>transform</code> on the test. In this competition the cv score dances all over the places once you slightly change a percentile threshold for the normalization and I strongly believe it will be a key to success. If you know what you are doing, fine, but if you don't here is a <a href=\"https://sebastianraschka.com/faq/docs/scale-training-test.html\" target=\"_blank\">futher reading for you</a>. </p>\n<p>Good luck, there are still 3 weeks!</p>",
  "messages": [
    {
      "id": 2608410,
      "postDate": "2024-01-18T18:55:22.337Z",
      "content": "<p>Hi,</p>\n<h2>Intro</h2>\n<p>I was going thru the notebooks and again and again I saw the same pattern.<br>\nPeople trained and normalized the data by using stats (min, max, std or mean, percentile depening on your choice) of the individual images. Or normalize each cuboid/volume by its own stats. And I even saw some folks normalize by percentile and on top of that use min-max, that's trully A Novel Approach in ML. </p>\n<p>This is what happens if you take middle slice along z axis and look into the distibutions:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F96912c5b5c17c77d2cb801d477f5fb6f%2FFigure_1.png?generation=1705593658292533&amp;alt=media\"></p>\n<p>1) The first subplot shows the distibution if we were to take all volumes and normalize it by its own stats, e.g.:</p>\n<pre><code> ():\n     (img - img_min) / (img_max - img_min)\n\nkid1_min, kid1_max = fit_cub_img[].(), fit_cub_img[].()\nkid2_min, kid2_max = fit_cub_img[].(), fit_cub_img[].()\nkid3_min, kid3_max = fit_cub_img[].(), fit_cub_img[].()\n\nscaled_kid1 = min_max_scaler(fit_cub_img[], kid1_min, kid1_max)\nscaled_kid2 = min_max_scaler(fit_cub_img[], kid2_min, kid2_max)\nscaled_kid3 = min_max_scaler(fit_cub_img[], kid3_min, kid3_max)\n</code></pre>\n<p>2) The second subplot shows if we were to train on kidney1 and apply its stats to kid2 or kid3 during the validation.</p>\n<pre><code>scaled_kid1 = min_max_scaler(fit_cub_img[], kid1_min, kid1_max)\nscaled_kid2 = min_max_scaler(fit_cub_img[], kid2_min, kid2_max)\nscaled_kid3 = min_max_scaler(fit_cub_img[], kid3_min, kid3_max)\n</code></pre>\n<p>Yes, we brought distribution of kid2 and kid3 closer to kid1.</p>\n<p>Adding the whole volume normalized histograms:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F6f26b0d58142e8f2180862ef59ec0645%2Ffig_volumes_hist.png?generation=1705633638305910&amp;alt=media\"></p>\n<p>Now the kidney3 closer to kidney1 with min_max by kidney1.</p>\n<p>I thought it was always <strong>a rule of thumb</strong> to fit and transform the train data and only transform the validation set.<br>\nIf you are lucky enough, the first scenario will yield a good score for you, but just because test set has close or the same stats as the one which you trained on. </p>\n<h2>Outro</h2>\n<p>It seems folks just normalize because they heard models perform better if you push everything to gaussian from 0 to 1. But I just want to remind you the reason we do <code>fit_transform</code> on train and <code>transform</code> on the test. In this competition the cv score dances all over the places once you slightly change a percentile threshold for the normalization and I strongly believe it will be a key to success. If you know what you are doing, fine, but if you don't here is a <a href=\"https://sebastianraschka.com/faq/docs/scale-training-test.html\" target=\"_blank\">futher reading for you</a>. </p>\n<p>Good luck, there are still 3 weeks!</p>",
      "rawMarkdown": "Hi,\n\n##Intro\nI was going thru the notebooks and again and again I saw the same pattern.\nPeople trained and normalized the data by using stats (min, max, std or mean, percentile depening on your choice) of the individual images. Or normalize each cuboid/volume by its own stats. And I even saw some folks normalize by percentile and on top of that use min-max, that's trully A Novel Approach in ML. \n\nThis is what happens if you take middle slice along z axis and look into the distibutions:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F96912c5b5c17c77d2cb801d477f5fb6f%2FFigure_1.png?generation=1705593658292533&alt=media)\n\n 1) The first subplot shows the distibution if we were to take all volumes and normalize it by its own stats, e.g.:\n```python\ndef min_max_scaler(img, img_min, img_max):\n    return (img - img_min) / (img_max - img_min)\n\nkid1_min, kid1_max = fit_cub_img[0].min(), fit_cub_img[0].max()\nkid2_min, kid2_max = fit_cub_img[1].min(), fit_cub_img[1].max()\nkid3_min, kid3_max = fit_cub_img[2].min(), fit_cub_img[2].max()\n\nscaled_kid1 = min_max_scaler(fit_cub_img[0], kid1_min, kid1_max)\nscaled_kid2 = min_max_scaler(fit_cub_img[1], kid2_min, kid2_max)\nscaled_kid3 = min_max_scaler(fit_cub_img[2], kid3_min, kid3_max)\n```\n 2) The second subplot shows if we were to train on kidney1 and apply its stats to kid2 or kid3 during the validation.\n\n```python\nscaled_kid1 = min_max_scaler(fit_cub_img[0], kid1_min, kid1_max)\nscaled_kid2 = min_max_scaler(fit_cub_img[1], kid2_min, kid2_max)\nscaled_kid3 = min_max_scaler(fit_cub_img[2], kid3_min, kid3_max)\n```\nYes, we brought distribution of kid2 and kid3 closer to kid1.\n\nAdding the whole volume normalized histograms:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F6f26b0d58142e8f2180862ef59ec0645%2Ffig_volumes_hist.png?generation=1705633638305910&alt=media)\n\nNow the kidney3 closer to kidney1 with min_max by kidney1.\n\nI thought it was always **a rule of thumb** to fit and transform the train data and only transform the validation set.\nIf you are lucky enough, the first scenario will yield a good score for you, but just because test set has close or the same stats as the one which you trained on. \n\n##Outro\nIt seems folks just normalize because they heard models perform better if you push everything to gaussian from 0 to 1. But I just want to remind you the reason we do `fit_transform` on train and `transform` on the test. In this competition the cv score dances all over the places once you slightly change a percentile threshold for the normalization and I strongly believe it will be a key to success. If you know what you are doing, fine, but if you don't here is a [futher reading for you](https://sebastianraschka.com/faq/docs/scale-training-test.html). \n\nGood luck, there are still 3 weeks!\n",
      "votes": 17
    },
    {
      "id": 2608667,
      "postDate": "2024-01-19T01:59:55.247Z",
      "content": "<p>the correct approach:</p>\n<ol>\n<li>make intesnity histogram for kidney1,2,3 (and maybe one more kidney from external data)</li>\n<li>normalise. but there is still inconsistency</li>\n<li>create intensity augmentation for the inconsistency (i.e. this is \"input space\" )</li>\n<li>train model</li>\n<li>measure model performance on normalise and inconsistency<br>\n(ensure model is robust within the \"input space\") </li>\n<li>probe public and private and make sure are are also within \"input space\" </li>\n</ol>\n<p>lastly, visualise,visualise,visualise, your results (both input and prediction)!!!<br>\nhow normalisation affects input and output, visualy and in metric ….?</p>\n<hr>\n<p>note:</p>\n<ul>\n<li>features in image data is either shape or intensity</li>\n<li>normalisation enhance either or both of the above.</li>\n<li>shape or intensity … which is more important?</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe89caa5cc37c0d167069498f36d323c9%2FSelection_999(4618).png?generation=1705629801880200&amp;alt=media\"></p>\n<p>this is from the paper. take note of step4</p>",
      "rawMarkdown": "the correct approach:\n1. make intesnity histogram for kidney1,2,3 (and maybe one more kidney from external data)\n2. normalise. but there is still inconsistency\n3. create intensity augmentation for the inconsistency (i.e. this is \"input space\" )\n4. train model\n5. measure model performance on normalise and inconsistency\n(ensure model is robust within the \"input space\") \n6. probe public and private and make sure are are also within \"input space\" \n\nlastly, visualise,visualise,visualise, your results (both input and prediction)!!!\nhow normalisation affects input and output, visualy and in metric ....?\n\n---\n\nnote:\n- features in image data is either shape or intensity\n- normalisation enhance either or both of the above.\n-  shape or intensity ... which is more important?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe89caa5cc37c0d167069498f36d323c9%2FSelection_999(4618).png?generation=1705629801880200&alt=media)\n\nthis is from the paper. take note of step4\n",
      "votes": 4,
      "replies": [
        {
          "id": 2608694,
          "postDate": "2024-01-19T03:02:30.277Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> I actually did stack kid5 and kid6 and those 3 examples each are closer to kidney1 than anything else.</p>\n<p>Here is a 100_000 bin hist for min max on the whole volumes (had to dance with custom hist1d calculation since we cannot just input 1000x1000x1000+ volume into matplotlib hist). If we normalize individually kidney2 distribution is closer to kidney1, otherwise kidney3 is closer to kidney1. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F783c979489b4eabaf563498ae95a66c0%2Ffig_volumes_hist.png?generation=1705633367916191&amp;alt=media\"></p>",
          "rawMarkdown": "@hengck23 I actually did stack kid5 and kid6 and those 3 examples each are closer to kidney1 than anything else.\n\nHere is a 100_000 bin hist for min max on the whole volumes (had to dance with custom hist1d calculation since we cannot just input 1000x1000x1000+ volume into matplotlib hist). If we normalize individually kidney2 distribution is closer to kidney1, otherwise kidney3 is closer to kidney1. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F783c979489b4eabaf563498ae95a66c0%2Ffig_volumes_hist.png?generation=1705633367916191&alt=media)\n\n",
          "replies": [
            {
              "id": 2608696,
              "postDate": "2024-01-19T03:08:40.850Z",
              "content": "<p>\"we cannot just input 1000x1000x1000+ volume into matplotlib\"</p>\n<pre><code>volume = ()   \n #(, , )\nh = np(volume(-))\nplt(h)\nplt()\n</code></pre>",
              "rawMarkdown": "\"we cannot just input 1000x1000x1000+ volume into matplotlib\"\n\n```\nvolume = read_volume()  #e.g 'kidney_2'\nprint(volume.shape) #(2217, 1041, 1511)\nh = np.bincount(volume.reshape(-1))\nplt.plot(h)\nplt.show()\n\n\n```"
            },
            {
              "id": 2608697,
              "postDate": "2024-01-19T03:11:04.607Z",
              "content": "<p>what we want afer normalisation is probably like this:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa147278c8c9c89978e2d3fe113330287%2FSelection_999(4630).png?generation=1705633861970876&amp;alt=media\"></p>",
              "rawMarkdown": "what we want afer normalisation is probably like this:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa147278c8c9c89978e2d3fe113330287%2FSelection_999(4630).png?generation=1705633861970876&alt=media)",
              "votes": 1
            },
            {
              "id": 2608698,
              "postDate": "2024-01-19T03:13:28.783Z",
              "content": "<p>for kidney2, it is unfornatlet like this<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F33d1c3010b775ac1cc3d78edd7484345%2FSelection_999(4631).png?generation=1705634006870849&amp;alt=media\"></p>",
              "rawMarkdown": "for kidney2, it is unfornatlet like this\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F33d1c3010b775ac1cc3d78edd7484345%2FSelection_999(4631).png?generation=1705634006870849&alt=media)"
            },
            {
              "id": 2608706,
              "postDate": "2024-01-19T03:17:52.647Z",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <code>np.bincount()</code> is neat. It works for int input only so you are to do back transform by multipling by max_value after normalization and then cast to int. </p>\n<p>This one took 15 seconds for me.</p>\n<pre><code> fast_histogram  histogram1d\n\nk1 = histogram1d(scaled_kid1.flatten(), =[scaled_kid1.(), scaled_kid1.()], bins=)\nk2 = histogram1d(scaled_kid2.flatten(), =[scaled_kid2.(), scaled_kid2.()], bins=)\nk3 = histogram1d(scaled_kid3.flatten(), =[scaled_kid3.(), scaled_kid3.()], bins=)\n\n\n()\nfig, axes = plt.subplots(, , figsize=(, ), dpi=)\nax = axes.flatten()\nax[].fill_between(x=np.linspace(scaled_kid1.(), scaled_kid1.(), ), y1=k1, alpha=, label=, color=)\nax[].fill_between(x=np.linspace(scaled_kid2.(), scaled_kid2.(), ), y1=k2, alpha=, label=, color=)\nax[].fill_between(x=np.linspace(scaled_kid3.(), scaled_kid3.(), ), y1=k3, alpha=, label=, color=)\nax[].set_title()\nax[].set_ylim(, )\nax[].legend()\n</code></pre>",
              "rawMarkdown": "@hengck23 `np.bincount()` is neat. It works for int input only so you are to do back transform by multipling by max_value after normalization and then cast to int. \n\nThis one took 15 seconds for me.\n```python\nfrom fast_histogram import histogram1d\n\nk1 = histogram1d(scaled_kid1.flatten(), range=[scaled_kid1.min(), scaled_kid1.max()], bins=100_000)\nk2 = histogram1d(scaled_kid2.flatten(), range=[scaled_kid2.min(), scaled_kid2.max()], bins=100_000)\nk3 = histogram1d(scaled_kid3.flatten(), range=[scaled_kid3.min(), scaled_kid3.max()], bins=100_000)\n\n\nprint('visualizing ...')\nfig, axes = plt.subplots(2, 1, figsize=(12, 8), dpi=80)\nax = axes.flatten()\nax[0].fill_between(x=np.linspace(scaled_kid1.min(), scaled_kid1.max(), 100_000), y1=k1, alpha=0.9, label='kid1', color='red')\nax[0].fill_between(x=np.linspace(scaled_kid2.min(), scaled_kid2.max(), 100_000), y1=k2, alpha=0.7, label='kid2', color='green')\nax[0].fill_between(x=np.linspace(scaled_kid3.min(), scaled_kid3.max(), 100_000), y1=k3, alpha=0.5, label='kid3', color='blue')\nax[0].set_title('normalized by individual volume min max values')\nax[0].set_ylim(0, 0.8e6)\nax[0].legend()\n```\n\n",
              "votes": 1
            }
          ]
        },
        {
          "id": 2608806,
          "postDate": "2024-01-19T05:32:17.807Z",
          "content": "<p>The normalization of my 2D slices seems to be inconsistent with the normalization of 3D volumes.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16378016%2Fc0e687c352882f88a13734ec82e4fe19%2F1705642224332.jpg?generation=1705642330726151&amp;alt=media\"></p>",
          "rawMarkdown": "The normalization of my 2D slices seems to be inconsistent with the normalization of 3D volumes.![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16378016%2Fc0e687c352882f88a13734ec82e4fe19%2F1705642224332.jpg?generation=1705642330726151&alt=media)",
          "votes": 2,
          "replies": [
            {
              "id": 2610224,
              "postDate": "2024-01-20T02:46:59.883Z",
              "content": "<p>How did you create your 2D slices? And how to do the distribution?</p>",
              "rawMarkdown": "How did you create your 2D slices? And how to do the distribution?"
            }
          ]
        },
        {
          "id": 2612995,
          "postDate": "2024-01-21T18:45:14.540Z",
          "content": "<p>I find that my models fit best (both validation and public leaderboard) when I standardize per image (2D model). I also tried zero centering the slices by the entire volume's mean, but that didn't work well. I'm still looking at other normalization and standardization strategies, but have you looked at any standardizations instead of normalizations? I'm curious how they've worked out for others.</p>",
          "rawMarkdown": "I find that my models fit best (both validation and public leaderboard) when I standardize per image (2D model). I also tried zero centering the slices by the entire volume's mean, but that didn't work well. I'm still looking at other normalization and standardization strategies, but have you looked at any standardizations instead of normalizations? I'm curious how they've worked out for others."
        }
      ]
    },
    {
      "id": 2608438,
      "postDate": "2024-01-18T19:21:43.163Z",
      "content": "<p>I guess this is why Hamlet said: 'normalise or not,, that is the question!'<br>\n,, didn't he?</p>",
      "rawMarkdown": "I guess this is why Hamlet said: 'normalise or not,, that is the question!'\n,, didn't he?",
      "votes": 1,
      "replies": [
        {
          "id": 2608455,
          "postDate": "2024-01-18T19:36:14.767Z",
          "content": "<p>And \"The rest is silence.\" </p>",
          "rawMarkdown": "And \"The rest is silence.\" ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2608532,
      "postDate": "2024-01-18T21:55:52.713Z",
      "content": "<p>\"This is what happens if you take middle slice along z axis and look into the distibutions\"<br>\nBut will be different structures at each middle slice: more or less vessels, background, etc…<br>\nNormalizing all the volume for its own stats, since they are all kidneys, we can expect more comparable values. I'm wrong?</p>",
      "rawMarkdown": "\"This is what happens if you take middle slice along z axis and look into the distibutions\"\nBut will be different structures at each middle slice: more or less vessels, background, etc...\nNormalizing all the volume for its own stats, since they are all kidneys, we can expect more comparable values. I'm wrong?",
      "votes": 2,
      "replies": [
        {
          "id": 2608705,
          "postDate": "2024-01-19T03:16:18.490Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2608461,
      "postDate": "2024-01-18T19:38:45.427Z",
      "content": "<p>One more thing.<br>\nAbout your normalisation type 1 and 2.<br>\nDid you 'chop' - 'outlier' as done in public notebooks?<br>\nThat process brings the 'own stats min-max normalised version: type 1' (probability densities) closer, which is why it works<br>\n,, probably :-)</p>",
      "rawMarkdown": "One more thing.\nAbout your normalisation type 1 and 2.\nDid you 'chop' - 'outlier' as done in public notebooks?\nThat process brings the 'own stats min-max normalised version: type 1' (probability densities) closer, which is why it works\n,, probably :-)",
      "votes": 2,
      "replies": [
        {
          "id": 2608467,
          "postDate": "2024-01-18T19:50:12.580Z",
          "content": "<p>… and for all this, <a href=\"https://www.kaggle.com/yoyobar\" target=\"_blank\">@yoyobar</a> is to blame sharing that too good baseline,<br>\n.. and all that beating the data to death by chopping - normalising etc.  :-)</p>",
          "rawMarkdown": "... and for all this, @yoyobar is to blame sharing that too good baseline,\n.. and all that beating the data to death by chopping - normalising etc.  :-)",
          "replies": [
            {
              "id": 2608780,
              "postDate": "2024-01-19T05:12:28.180Z",
              "content": "<p>with respect to <a href=\"https://www.kaggle.com/yoyobar\" target=\"_blank\">@yoyobar</a> - that too good baseline is very similar to Vesuvius Ink Detection competition and did quite well there.  cannot comment on the normalisation approaches beween that competition and this one though - layers of ancient papyrus scans may be different, need different approaches?</p>",
              "rawMarkdown": "with respect to @yoyobar - that too good baseline is very similar to Vesuvius Ink Detection competition and did quite well there.  cannot comment on the normalisation approaches beween that competition and this one though - layers of ancient papyrus scans may be different, need different approaches?",
              "votes": 1
            }
          ]
        },
        {
          "id": 2608472,
          "postDate": "2024-01-18T19:53:00.813Z",
          "content": "<p>Probably here is a key word :), as you see values are between 0.2 and 0.5 already scaled, the best scoring public notebook does standard scaling with</p>\n<pre><code>x=(x-mean)/(std+smooth)\n x[x&gt;]=(x[x&gt;]-)* +\n x[x&lt;-]=(x[x&lt;-]+)*-\n</code></pre>\n<p>and that thing comes from the <a href=\"https://www.kaggle.com/code/yoyobar/2-5d-cutting-model-baseline-training\" target=\"_blank\">training baseline</a> here.</p>\n<p>The author used 1st case scenario and applied train stats on train and valid stats on valid and more over it was done per batch, not using global mean and std.</p>\n<pre><code> i,(x,y)  (train_dataset):\n        x=x.cuda().to(tc.float32)\n        y=y.cuda().to(tc.float32)\n        x=norm_with_clip(x.reshape(-,*x.shape[:])).reshape(x.shape)\n</code></pre>\n<p>It looks like a pure Kaggle voodo magic.</p>",
          "rawMarkdown": "Probably here is a key word :), as you see values are between 0.2 and 0.5 already scaled, the best scoring public notebook does standard scaling with\n\n ```python\nx=(x-mean)/(std+smooth)\n x[x>5]=(x[x>5]-5)*1e-3 +5\n x[x<-3]=(x[x<-3]+3)*1e-3-3\n```\nand that thing comes from the [training baseline](https://www.kaggle.com/code/yoyobar/2-5d-cutting-model-baseline-training) here.\n\nThe author used 1st case scenario and applied train stats on train and valid stats on valid and more over it was done per batch, not using global mean and std.\n```python\nfor i,(x,y) in enumerate(train_dataset):\n        x=x.cuda().to(tc.float32)\n        y=y.cuda().to(tc.float32)\n        x=norm_with_clip(x.reshape(-1,*x.shape[2:])).reshape(x.shape)\n```\n\nIt looks like a pure Kaggle voodo magic.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2608667,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-01-19T01:59:55.247000",
      "content": "<p>the correct approach:</p>\n<ol>\n<li>make intesnity histogram for kidney1,2,3 (and maybe one more kidney from external data)</li>\n<li>normalise. but there is still inconsistency</li>\n<li>create intensity augmentation for the inconsistency (i.e. this is \"input space\" )</li>\n<li>train model</li>\n<li>measure model performance on normalise and inconsistency<br>\n(ensure model is robust within the \"input space\") </li>\n<li>probe public and private and make sure are are also within \"input space\" </li>\n</ol>\n<p>lastly, visualise,visualise,visualise, your results (both input and prediction)!!!<br>\nhow normalisation affects input and output, visualy and in metric ….?</p>\n<hr>\n<p>note:</p>\n<ul>\n<li>features in image data is either shape or intensity</li>\n<li>normalisation enhance either or both of the above.</li>\n<li>shape or intensity … which is more important?</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe89caa5cc37c0d167069498f36d323c9%2FSelection_999(4618).png?generation=1705629801880200&amp;alt=media\"></p>\n<p>this is from the paper. take note of step4</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2608694,
          "author_name": "SSS",
          "author_url": "",
          "post_date": "2024-01-19T03:02:30.277000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> I actually did stack kid5 and kid6 and those 3 examples each are closer to kidney1 than anything else.</p>\n<p>Here is a 100_000 bin hist for min max on the whole volumes (had to dance with custom hist1d calculation since we cannot just input 1000x1000x1000+ volume into matplotlib hist). If we normalize individually kidney2 distribution is closer to kidney1, otherwise kidney3 is closer to kidney1. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F783c979489b4eabaf563498ae95a66c0%2Ffig_volumes_hist.png?generation=1705633367916191&amp;alt=media\"></p>",
          "votes": 0,
          "replies": [
            {
              "id": 2608696,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-01-19T03:08:40.850000",
              "content": "<p>\"we cannot just input 1000x1000x1000+ volume into matplotlib\"</p>\n<pre><code>volume = ()   \n #(, , )\nh = np(volume(-))\nplt(h)\nplt()\n</code></pre>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2608697,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-01-19T03:11:04.607000",
              "content": "<p>what we want afer normalisation is probably like this:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fa147278c8c9c89978e2d3fe113330287%2FSelection_999(4630).png?generation=1705633861970876&amp;alt=media\"></p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2608698,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-01-19T03:13:28.783000",
              "content": "<p>for kidney2, it is unfornatlet like this<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F33d1c3010b775ac1cc3d78edd7484345%2FSelection_999(4631).png?generation=1705634006870849&amp;alt=media\"></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2608706,
              "author_name": "SSS",
              "author_url": "",
              "post_date": "2024-01-19T03:17:52.647000",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <code>np.bincount()</code> is neat. It works for int input only so you are to do back transform by multipling by max_value after normalization and then cast to int. </p>\n<p>This one took 15 seconds for me.</p>\n<pre><code> fast_histogram  histogram1d\n\nk1 = histogram1d(scaled_kid1.flatten(), =[scaled_kid1.(), scaled_kid1.()], bins=)\nk2 = histogram1d(scaled_kid2.flatten(), =[scaled_kid2.(), scaled_kid2.()], bins=)\nk3 = histogram1d(scaled_kid3.flatten(), =[scaled_kid3.(), scaled_kid3.()], bins=)\n\n\n()\nfig, axes = plt.subplots(, , figsize=(, ), dpi=)\nax = axes.flatten()\nax[].fill_between(x=np.linspace(scaled_kid1.(), scaled_kid1.(), ), y1=k1, alpha=, label=, color=)\nax[].fill_between(x=np.linspace(scaled_kid2.(), scaled_kid2.(), ), y1=k2, alpha=, label=, color=)\nax[].fill_between(x=np.linspace(scaled_kid3.(), scaled_kid3.(), ), y1=k3, alpha=, label=, color=)\nax[].set_title()\nax[].set_ylim(, )\nax[].legend()\n</code></pre>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2608806,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-01-19T05:32:17.807000",
          "content": "<p>The normalization of my 2D slices seems to be inconsistent with the normalization of 3D volumes.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16378016%2Fc0e687c352882f88a13734ec82e4fe19%2F1705642224332.jpg?generation=1705642330726151&amp;alt=media\"></p>",
          "votes": 2,
          "replies": [
            {
              "id": 2610224,
              "author_name": "豆柴金鯱",
              "author_url": "",
              "post_date": "2024-01-20T02:46:59.883000",
              "content": "<p>How did you create your 2D slices? And how to do the distribution?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2612995,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "2024-01-21T18:45:14.540000",
          "content": "<p>I find that my models fit best (both validation and public leaderboard) when I standardize per image (2D model). I also tried zero centering the slices by the entire volume's mean, but that didn't work well. I'm still looking at other normalization and standardization strategies, but have you looked at any standardizations instead of normalizations? I'm curious how they've worked out for others.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2608438,
      "author_name": "GUNER",
      "author_url": "",
      "post_date": "2024-01-18T19:21:43.163000",
      "content": "<p>I guess this is why Hamlet said: 'normalise or not,, that is the question!'<br>\n,, didn't he?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2608455,
          "author_name": "SSS",
          "author_url": "",
          "post_date": "2024-01-18T19:36:14.767000",
          "content": "<p>And \"The rest is silence.\" </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2608532,
      "author_name": "Ángel Jacinto Sánchez Ruiz",
      "author_url": "",
      "post_date": "2024-01-18T21:55:52.713000",
      "content": "<p>\"This is what happens if you take middle slice along z axis and look into the distibutions\"<br>\nBut will be different structures at each middle slice: more or less vessels, background, etc…<br>\nNormalizing all the volume for its own stats, since they are all kidneys, we can expect more comparable values. I'm wrong?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2608705,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-01-19T03:16:18.490000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2608461,
      "author_name": "GUNER",
      "author_url": "",
      "post_date": "2024-01-18T19:38:45.427000",
      "content": "<p>One more thing.<br>\nAbout your normalisation type 1 and 2.<br>\nDid you 'chop' - 'outlier' as done in public notebooks?<br>\nThat process brings the 'own stats min-max normalised version: type 1' (probability densities) closer, which is why it works<br>\n,, probably :-)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2608467,
          "author_name": "GUNER",
          "author_url": "",
          "post_date": "2024-01-18T19:50:12.580000",
          "content": "<p>… and for all this, <a href=\"https://www.kaggle.com/yoyobar\" target=\"_blank\">@yoyobar</a> is to blame sharing that too good baseline,<br>\n.. and all that beating the data to death by chopping - normalising etc.  :-)</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2608780,
              "author_name": "something4kag",
              "author_url": "",
              "post_date": "2024-01-19T05:12:28.180000",
              "content": "<p>with respect to <a href=\"https://www.kaggle.com/yoyobar\" target=\"_blank\">@yoyobar</a> - that too good baseline is very similar to Vesuvius Ink Detection competition and did quite well there.  cannot comment on the normalisation approaches beween that competition and this one though - layers of ancient papyrus scans may be different, need different approaches?</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2608472,
          "author_name": "SSS",
          "author_url": "",
          "post_date": "2024-01-18T19:53:00.813000",
          "content": "<p>Probably here is a key word :), as you see values are between 0.2 and 0.5 already scaled, the best scoring public notebook does standard scaling with</p>\n<pre><code>x=(x-mean)/(std+smooth)\n x[x&gt;]=(x[x&gt;]-)* +\n x[x&lt;-]=(x[x&lt;-]+)*-\n</code></pre>\n<p>and that thing comes from the <a href=\"https://www.kaggle.com/code/yoyobar/2-5d-cutting-model-baseline-training\" target=\"_blank\">training baseline</a> here.</p>\n<p>The author used 1st case scenario and applied train stats on train and valid stats on valid and more over it was done per batch, not using global mean and std.</p>\n<pre><code> i,(x,y)  (train_dataset):\n        x=x.cuda().to(tc.float32)\n        y=y.cuda().to(tc.float32)\n        x=norm_with_clip(x.reshape(-,*x.shape[:])).reshape(x.shape)\n</code></pre>\n<p>It looks like a pure Kaggle voodo magic.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2608410": "Hi,\n\n##Intro\nI was going thru the notebooks and again and again I saw the same pattern.\nPeople trained and normalized the data by using stats (min, max, std or mean, percentile depening on your choice) of the individual images. Or normalize each cuboid/volume by its own stats. And I even saw some folks normalize by percentile and on top of that use min-max, that's trully A Novel Approach in ML. \n\nThis is what happens if you take middle slice along z axis and look into the distibutions:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F96912c5b5c17c77d2cb801d477f5fb6f%2FFigure_1.png?generation=1705593658292533&alt=media)\n\n 1) The first subplot shows the distibution if we were to take all volumes and normalize it by its own stats, e.g.:\n```python\ndef min_max_scaler(img, img_min, img_max):\n    return (img - img_min) / (img_max - img_min)\n\nkid1_min, kid1_max = fit_cub_img[0].min(), fit_cub_img[0].max()\nkid2_min, kid2_max = fit_cub_img[1].min(), fit_cub_img[1].max()\nkid3_min, kid3_max = fit_cub_img[2].min(), fit_cub_img[2].max()\n\nscaled_kid1 = min_max_scaler(fit_cub_img[0], kid1_min, kid1_max)\nscaled_kid2 = min_max_scaler(fit_cub_img[1], kid2_min, kid2_max)\nscaled_kid3 = min_max_scaler(fit_cub_img[2], kid3_min, kid3_max)\n```\n 2) The second subplot shows if we were to train on kidney1 and apply its stats to kid2 or kid3 during the validation.\n\n```python\nscaled_kid1 = min_max_scaler(fit_cub_img[0], kid1_min, kid1_max)\nscaled_kid2 = min_max_scaler(fit_cub_img[1], kid2_min, kid2_max)\nscaled_kid3 = min_max_scaler(fit_cub_img[2], kid3_min, kid3_max)\n```\nYes, we brought distribution of kid2 and kid3 closer to kid1.\n\nAdding the whole volume normalized histograms:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6259210%2F6f26b0d58142e8f2180862ef59ec0645%2Ffig_volumes_hist.png?generation=1705633638305910&alt=media)\n\nNow the kidney3 closer to kidney1 with min_max by kidney1.\n\nI thought it was always **a rule of thumb** to fit and transform the train data and only transform the validation set.\nIf you are lucky enough, the first scenario will yield a good score for you, but just because test set has close or the same stats as the one which you trained on. \n\n##Outro\nIt seems folks just normalize because they heard models perform better if you push everything to gaussian from 0 to 1. But I just want to remind you the reason we do `fit_transform` on train and `transform` on the test. In this competition the cv score dances all over the places once you slightly change a percentile threshold for the normalization and I strongly believe it will be a key to success. If you know what you are doing, fine, but if you don't here is a [futher reading for you](https://sebastianraschka.com/faq/docs/scale-training-test.html). \n\nGood luck, there are still 3 weeks!\n",
    "2608667": "the correct approach:\n1. make intesnity histogram for kidney1,2,3 (and maybe one more kidney from external data)\n2. normalise. but there is still inconsistency\n3. create intensity augmentation for the inconsistency (i.e. this is \"input space\" )\n4. train model\n5. measure model performance on normalise and inconsistency\n(ensure model is robust within the \"input space\") \n6. probe public and private and make sure are are also within \"input space\" \n\nlastly, visualise,visualise,visualise, your results (both input and prediction)!!!\nhow normalisation affects input and output, visualy and in metric ....?\n\n---\n\nnote:\n- features in image data is either shape or intensity\n- normalisation enhance either or both of the above.\n-  shape or intensity ... which is more important?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fe89caa5cc37c0d167069498f36d323c9%2FSelection_999(4618).png?generation=1705629801880200&alt=media)\n\nthis is from the paper. take note of step4\n",
    "2608438": "I guess this is why Hamlet said: 'normalise or not,, that is the question!'\n,, didn't he?",
    "2608532": "\"This is what happens if you take middle slice along z axis and look into the distibutions\"\nBut will be different structures at each middle slice: more or less vessels, background, etc...\nNormalizing all the volume for its own stats, since they are all kidneys, we can expect more comparable values. I'm wrong?",
    "2608461": "One more thing.\nAbout your normalisation type 1 and 2.\nDid you 'chop' - 'outlier' as done in public notebooks?\nThat process brings the 'own stats min-max normalised version: type 1' (probability densities) closer, which is why it works\n,, probably :-)"
  }
}