{
  "id": 470374,
  "title": "Are there any risks in using the 2.5d method when testing in private public test?",
  "url": "/competitions/blood-vessel-segmentation/discussion/470374",
  "author_name": "tanxxx",
  "post_date": "2024-01-24T02:20:11.629000",
  "votes": 0,
  "comment_count": 16,
  "views": 0,
  "content": "<p>There are about 1000 samples in the public test, and about 500 samples in the private. Has the private been down-sampled along the z-axis? If so, will the connection between slices in 2.5d be affected?</p>",
  "messages": [
    {
      "id": 2617390,
      "postDate": "2024-01-24T07:59:59.933Z",
      "content": "<p><a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/460833\" target=\"_blank\">https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/460833</a><br>\n\"Private Test:<br>\nContinuous 3D part of a whole human kidney imaged with HiP-CT - Originally scanned at 15.77um/voxel binned to 63.08um/voxel (bin x4) before segmentation.\"<br>\nShould be not downsampled, just a section.</p>",
      "rawMarkdown": "https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/460833\n\"Private Test:\nContinuous 3D part of a whole human kidney imaged with HiP-CT - Originally scanned at 15.77um/voxel binned to 63.08um/voxel (bin x4) before segmentation.\"\nShould be not downsampled, just a section.",
      "votes": 1,
      "replies": [
        {
          "id": 2617429,
          "postDate": "2024-01-24T08:36:50.137Z",
          "content": "<p>I'seem have some trouble grasping your point. I feel that the resolution of public test should be similar to the training set, but the resolution of private test should be smaller than the training set, this might result in a visually blurry appearance.</p>",
          "rawMarkdown": "I'seem have some trouble grasping your point. I feel that the resolution of public test should be similar to the training set, but the resolution of private test should be smaller than the training set, this might result in a visually blurry appearance."
        }
      ]
    },
    {
      "id": 2617788,
      "postDate": "2024-01-24T13:05:08.823Z",
      "content": "<p>just to add a note.</p>\n<p>some kagglers may be wondering why is there a fuss about different resolution?</p>\n<ol>\n<li><p>do note that we are talking about surface dice. this is very sensitive. we need to predict the correct mask at the correct pixel location, not near the correct pixel location.</p></li>\n<li><p>hence even choice of resize function like bilinear, bicubic, anti-alisas methods will affects results greatly.</p></li>\n<li><p>more importantly, we have no training annotation for private resolution at all! train and public annotation are different from private one, at least in the \"surface dice sense\"</p></li>\n</ol>",
      "rawMarkdown": "just to add a note.\n\nsome kagglers may be wondering why is there a fuss about different resolution?\n\n1. do note that we are talking about surface dice. this is very sensitive. we need to predict the correct mask at the correct pixel location, not near the correct pixel location.\n\n2. hence even choice of resize function like bilinear, bicubic, anti-alisas methods will affects results greatly.\n\n3. more importantly, we have no training annotation for private resolution at all! train and public annotation are different from private one, at least in the \"surface dice sense\"",
      "votes": 2,
      "replies": [
        {
          "id": 2618927,
          "postDate": "2024-01-25T05:24:32.163Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> and <a href=\"https://www.kaggle.com/tanxxx\" target=\"_blank\">@tanxxx</a> - <br>\nhave you considered using kidney_1_voi - the high-resolution subset of kidney_1, at 5.2um resolution and the code you posted below to try to get different resolution similar to Private Test?  At least it has labels and although kidney_1 may be used for training you had mentioned elsewhere the labels were different. Or for models just with kidney 2 and 3 perhaps?</p>\n<p>Had looked into this earlier from the paper referenced but as code was using matlab (and not open source) was not sure allowed.  But perhaps to do your own testing to see how models perform would be OK? </p>\n<p>Not sure if this helps just thought to mention it here.  Good Luck to you both!</p>",
          "rawMarkdown": "@hengck23 and @tanxxx - \nhave you considered using kidney_1_voi - the high-resolution subset of kidney_1, at 5.2um resolution and the code you posted below to try to get different resolution similar to Private Test?  At least it has labels and although kidney_1 may be used for training you had mentioned elsewhere the labels were different. Or for models just with kidney 2 and 3 perhaps?\n\nHad looked into this earlier from the paper referenced but as code was using matlab (and not open source) was not sure allowed.  But perhaps to do your own testing to see how models perform would be OK? \n\nNot sure if this helps just thought to mention it here.  Good Luck to you both!",
          "votes": 1,
          "replies": [
            {
              "id": 2619102,
              "postDate": "2024-01-25T07:58:18.087Z",
              "content": "<p>This is a really unique perspective. I believe the method you mentioned might be able to solve this problem, so I'm definitely going to give it a try.</p>",
              "rawMarkdown": "This is a really unique perspective. I believe the method you mentioned might be able to solve this problem, so I'm definitely going to give it a try.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2616984,
      "postDate": "2024-01-24T02:20:11.630Z",
      "content": "<p>There are about 1000 samples in the public test, and about 500 samples in the private. Has the private been down-sampled along the z-axis? If so, will the connection between slices in 2.5d be affected?</p>",
      "rawMarkdown": "There are about 1000 samples in the public test, and about 500 samples in the private. Has the private been down-sampled along the z-axis? If so, will the connection between slices in 2.5d be affected?"
    },
    {
      "id": 2619225,
      "postDate": "2024-01-25T10:24:18.397Z",
      "content": "<p>alternatively, you can distill 2.5d, 3d , etc model to single slice 2d model</p>",
      "rawMarkdown": "alternatively, you can distill 2.5d, 3d , etc model to single slice 2d model",
      "replies": [
        {
          "id": 2619263,
          "postDate": "2024-01-25T11:01:51.623Z",
          "content": "<p>thanks, this indeed a solution</p>",
          "rawMarkdown": "thanks, this indeed a solution"
        }
      ]
    },
    {
      "id": 2617104,
      "postDate": "2024-01-24T04:29:33.203Z",
      "content": "<p>\" private been down-sampled along the z-axis? \"</p>\n<p>i think yes. (if you read the paper and the github CT reconstruction code. There are also discussions in the forum).<br>\nit also  depends </p>\n<ul>\n<li>on how many slice in 2.5d, e.g. 2,8,16,32 …</li>\n<li>your strategy for different resolution, e.g.<br>\nare you are going 3d resize private data to 50um voxel/pixel?</li>\n</ul>",
      "rawMarkdown": "\" private been down-sampled along the z-axis? \"\n\ni think yes. (if you read the paper and the github CT reconstruction code. There are also discussions in the forum).\nit also  depends \n- on how many slice in 2.5d, e.g. 2,8,16,32 ...\n- your strategy for different resolution, e.g.\nare you are going 3d resize private data to 50um voxel/pixel?",
      "replies": [
        {
          "id": 2617181,
          "postDate": "2024-01-24T05:23:09.513Z",
          "content": "<p>One of my model uses 3 consecutive slices as input and can get similar results to 2d method in LB ( worse in CV ). But I am worried about whether private test will cause domain differences due to slice continuity. </p>\n<p>\"are you are going 3d resize private data to 50um voxel/pixel?\" </p>\n<p>I did not resize because I can't judge from CV or LB whether resize will have a negative impact. If I want to resize in slice-wise, using torch.nn.functional.interpolate(input,size,scale_factor), what should scale_factor be set to?</p>",
          "rawMarkdown": "One of my model uses 3 consecutive slices as input and can get similar results to 2d method in LB ( worse in CV ). But I am worried about whether private test will cause domain differences due to slice continuity. \n\n\"are you are going 3d resize private data to 50um voxel/pixel?\" \n\nI did not resize because I can't judge from CV or LB whether resize will have a negative impact. If I want to resize in slice-wise, using torch.nn.functional.interpolate(input,size,scale_factor), what should scale_factor be set to?\n",
          "replies": [
            {
              "id": 2617313,
              "postDate": "2024-01-24T07:03:57.123Z",
              "content": "<p>this is also a problem that we are facing.</p>\n<p>the issue is that we have no annotation ground truth in private 63um/voxel.<br>\ni think private and public annotation are different.</p>\n<blockquote>\n  <blockquote>\n    <p>I did not resize because I can't judge from CV or LB whether resize will have a negative impact. If I want to resize in &gt;&gt;slice-wise, using torch.nn.functional.interpolate(input,size,scale_factor), what should scale_factor be set to?</p>\n  </blockquote>\n</blockquote>\n<p>you will have to check the code</p>\n<p><a href=\"https://github.com/HiPCTProject/Tomo_Recon\" target=\"_blank\">https://github.com/HiPCTProject/Tomo_Recon</a></p>\n<p>/tif_bin_3d_master_OAR.m<br>\n/tif_bin_3d_slave_OAR.m<br>\n% that function aims to calculate 3D binning from tif stack, resulting<br>\n% in a new 8 tif stack, parameters are the binning factor, and the cluster<br>\n% name (OAR or condor)</p>\n<pre><code>\n\n\n    =sprintf(fname{i});\n    =single(imread(init_slice_name));\n\n     =1:binning_factor-1\n         i+j&lt;number_of_files-1\n            =sprintf(fname{i+j});\n            =single(imread(slice_name));\n            =init_slice+slice;\n        end\n    end\n\n    =init_slice/binning_factor;\n</code></pre>",
              "rawMarkdown": "this is also a problem that we are facing.\n\nthe issue is that we have no annotation ground truth in private 63um/voxel.\ni think private and public annotation are different.\n\n\n>>I did not resize because I can't judge from CV or LB whether resize will have a negative impact. If I want to resize in >>slice-wise, using torch.nn.functional.interpolate(input,size,scale_factor), what should scale_factor be set to?\n\nyou will have to check the code\n\nhttps://github.com/HiPCTProject/Tomo_Recon\n\n/tif_bin_3d_master_OAR.m\n/tif_bin_3d_slave_OAR.m\n% that function aims to calculate 3D binning from tif stack, resulting\n% in a new 8 tif stack, parameters are the binning factor, and the cluster\n% name (OAR or condor)\n\n```\n### code snippet for averaging(binning) in the z-direction\n\n\n    init_slice_name=sprintf(fname{i});\n    init_slice=single(imread(init_slice_name));\n    \n    for j=1:binning_factor-1\n        if i+j<number_of_files-1\n            slice_name=sprintf(fname{i+j});\n            slice=single(imread(slice_name));\n            init_slice=init_slice+slice;\n        end\n    end\n    \n    final_slice=init_slice/binning_factor;\n\n\n```"
            },
            {
              "id": 2617381,
              "postDate": "2024-01-24T07:50:47.020Z",
              "content": "<p>Thanks a lot, hope both of us will get a good result😀</p>",
              "rawMarkdown": "Thanks a lot, hope both of us will get a good result😀"
            },
            {
              "id": 2617796,
              "postDate": "2024-01-24T13:09:31.087Z",
              "content": "<p>from competition point of view, you need NOT to be very correct. you just need to be more correct then your opponents. </p>\n<p>let's us consider the small vessel. i don't think anyone can confidently do better than the rest. because lb score can be very sensitive to annotation truth.</p>\n<p>how about big vessels? I think they are less sensitive to resolution. So if we have more correct big vessels, maybe we can have better score?</p>\n<p>and of course, lower your fp would help.</p>",
              "rawMarkdown": "from competition point of view, you need NOT to be very correct. you just need to be more correct then your opponents. \n\nlet's us consider the small vessel. i don't think anyone can confidently do better than the rest. because lb score can be very sensitive to annotation truth.\n\nhow about big vessels? I think they are less sensitive to resolution. So if we have more correct big vessels, maybe we can have better score?\n\nand of course, lower your fp would help."
            },
            {
              "id": 2618006,
              "postDate": "2024-01-24T14:32:29.617Z",
              "content": "<p>thank, I will take a try,  and a higher threshold may be helpful for our private test.</p>",
              "rawMarkdown": "thank, I will take a try,  and a higher threshold may be helpful for our private test.\n"
            },
            {
              "id": 2619092,
              "postDate": "2024-01-25T07:45:19.470Z",
              "content": "<p>Why would you say that lower FP would help? None of your charts suggest that. \"Lower FP would help\" would mean correlation between lower FP and higher LB. That doesn't happen </p>",
              "rawMarkdown": "Why would you say that lower FP would help? None of your charts suggest that. \"Lower FP would help\" would mean correlation between lower FP and higher LB. That doesn't happen "
            },
            {
              "id": 2623117,
              "postDate": "2024-01-28T01:46:09.723Z",
              "content": "<p>my experiments suggest that change of scale causes large change of fp.</p>",
              "rawMarkdown": "my experiments suggest that change of scale causes large change of fp."
            }
          ]
        }
      ]
    },
    {
      "id": 2618050,
      "postDate": "2024-01-24T14:53:22.920Z",
      "rawMarkdown": "",
      "votes": -4,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2617390,
      "author_name": "Ángel Jacinto Sánchez Ruiz",
      "author_url": "",
      "post_date": "2024-01-24T07:59:59.933000",
      "content": "<p><a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/460833\" target=\"_blank\">https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/460833</a><br>\n\"Private Test:<br>\nContinuous 3D part of a whole human kidney imaged with HiP-CT - Originally scanned at 15.77um/voxel binned to 63.08um/voxel (bin x4) before segmentation.\"<br>\nShould be not downsampled, just a section.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2617429,
          "author_name": "tanxxx",
          "author_url": "",
          "post_date": "2024-01-24T08:36:50.137000",
          "content": "<p>I'seem have some trouble grasping your point. I feel that the resolution of public test should be similar to the training set, but the resolution of private test should be smaller than the training set, this might result in a visually blurry appearance.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2617788,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-01-24T13:05:08.823000",
      "content": "<p>just to add a note.</p>\n<p>some kagglers may be wondering why is there a fuss about different resolution?</p>\n<ol>\n<li><p>do note that we are talking about surface dice. this is very sensitive. we need to predict the correct mask at the correct pixel location, not near the correct pixel location.</p></li>\n<li><p>hence even choice of resize function like bilinear, bicubic, anti-alisas methods will affects results greatly.</p></li>\n<li><p>more importantly, we have no training annotation for private resolution at all! train and public annotation are different from private one, at least in the \"surface dice sense\"</p></li>\n</ol>",
      "votes": 2,
      "replies": [
        {
          "id": 2618927,
          "author_name": "something4kag",
          "author_url": "",
          "post_date": "2024-01-25T05:24:32.163000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> and <a href=\"https://www.kaggle.com/tanxxx\" target=\"_blank\">@tanxxx</a> - <br>\nhave you considered using kidney_1_voi - the high-resolution subset of kidney_1, at 5.2um resolution and the code you posted below to try to get different resolution similar to Private Test?  At least it has labels and although kidney_1 may be used for training you had mentioned elsewhere the labels were different. Or for models just with kidney 2 and 3 perhaps?</p>\n<p>Had looked into this earlier from the paper referenced but as code was using matlab (and not open source) was not sure allowed.  But perhaps to do your own testing to see how models perform would be OK? </p>\n<p>Not sure if this helps just thought to mention it here.  Good Luck to you both!</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2619102,
              "author_name": "tanxxx",
              "author_url": "",
              "post_date": "2024-01-25T07:58:18.087000",
              "content": "<p>This is a really unique perspective. I believe the method you mentioned might be able to solve this problem, so I'm definitely going to give it a try.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2619225,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-01-25T10:24:18.397000",
      "content": "<p>alternatively, you can distill 2.5d, 3d , etc model to single slice 2d model</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2619263,
          "author_name": "tanxxx",
          "author_url": "",
          "post_date": "2024-01-25T11:01:51.623000",
          "content": "<p>thanks, this indeed a solution</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2617104,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-01-24T04:29:33.203000",
      "content": "<p>\" private been down-sampled along the z-axis? \"</p>\n<p>i think yes. (if you read the paper and the github CT reconstruction code. There are also discussions in the forum).<br>\nit also  depends </p>\n<ul>\n<li>on how many slice in 2.5d, e.g. 2,8,16,32 …</li>\n<li>your strategy for different resolution, e.g.<br>\nare you are going 3d resize private data to 50um voxel/pixel?</li>\n</ul>",
      "votes": 0,
      "replies": [
        {
          "id": 2617181,
          "author_name": "tanxxx",
          "author_url": "",
          "post_date": "2024-01-24T05:23:09.513000",
          "content": "<p>One of my model uses 3 consecutive slices as input and can get similar results to 2d method in LB ( worse in CV ). But I am worried about whether private test will cause domain differences due to slice continuity. </p>\n<p>\"are you are going 3d resize private data to 50um voxel/pixel?\" </p>\n<p>I did not resize because I can't judge from CV or LB whether resize will have a negative impact. If I want to resize in slice-wise, using torch.nn.functional.interpolate(input,size,scale_factor), what should scale_factor be set to?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2617313,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-01-24T07:03:57.123000",
              "content": "<p>this is also a problem that we are facing.</p>\n<p>the issue is that we have no annotation ground truth in private 63um/voxel.<br>\ni think private and public annotation are different.</p>\n<blockquote>\n  <blockquote>\n    <p>I did not resize because I can't judge from CV or LB whether resize will have a negative impact. If I want to resize in &gt;&gt;slice-wise, using torch.nn.functional.interpolate(input,size,scale_factor), what should scale_factor be set to?</p>\n  </blockquote>\n</blockquote>\n<p>you will have to check the code</p>\n<p><a href=\"https://github.com/HiPCTProject/Tomo_Recon\" target=\"_blank\">https://github.com/HiPCTProject/Tomo_Recon</a></p>\n<p>/tif_bin_3d_master_OAR.m<br>\n/tif_bin_3d_slave_OAR.m<br>\n% that function aims to calculate 3D binning from tif stack, resulting<br>\n% in a new 8 tif stack, parameters are the binning factor, and the cluster<br>\n% name (OAR or condor)</p>\n<pre><code>\n\n\n    =sprintf(fname{i});\n    =single(imread(init_slice_name));\n\n     =1:binning_factor-1\n         i+j&lt;number_of_files-1\n            =sprintf(fname{i+j});\n            =single(imread(slice_name));\n            =init_slice+slice;\n        end\n    end\n\n    =init_slice/binning_factor;\n</code></pre>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2617381,
              "author_name": "tanxxx",
              "author_url": "",
              "post_date": "2024-01-24T07:50:47.020000",
              "content": "<p>Thanks a lot, hope both of us will get a good result😀</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2617796,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-01-24T13:09:31.087000",
              "content": "<p>from competition point of view, you need NOT to be very correct. you just need to be more correct then your opponents. </p>\n<p>let's us consider the small vessel. i don't think anyone can confidently do better than the rest. because lb score can be very sensitive to annotation truth.</p>\n<p>how about big vessels? I think they are less sensitive to resolution. So if we have more correct big vessels, maybe we can have better score?</p>\n<p>and of course, lower your fp would help.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2618006,
              "author_name": "tanxxx",
              "author_url": "",
              "post_date": "2024-01-24T14:32:29.617000",
              "content": "<p>thank, I will take a try,  and a higher threshold may be helpful for our private test.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2619092,
              "author_name": "Ivan Panshin",
              "author_url": "",
              "post_date": "2024-01-25T07:45:19.470000",
              "content": "<p>Why would you say that lower FP would help? None of your charts suggest that. \"Lower FP would help\" would mean correlation between lower FP and higher LB. That doesn't happen </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2623117,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-01-28T01:46:09.723000",
              "content": "<p>my experiments suggest that change of scale causes large change of fp.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2618050,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-24T14:53:22.920000",
      "content": "",
      "votes": -4,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2617390": "https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/460833\n\"Private Test:\nContinuous 3D part of a whole human kidney imaged with HiP-CT - Originally scanned at 15.77um/voxel binned to 63.08um/voxel (bin x4) before segmentation.\"\nShould be not downsampled, just a section.",
    "2617788": "just to add a note.\n\nsome kagglers may be wondering why is there a fuss about different resolution?\n\n1. do note that we are talking about surface dice. this is very sensitive. we need to predict the correct mask at the correct pixel location, not near the correct pixel location.\n\n2. hence even choice of resize function like bilinear, bicubic, anti-alisas methods will affects results greatly.\n\n3. more importantly, we have no training annotation for private resolution at all! train and public annotation are different from private one, at least in the \"surface dice sense\"",
    "2616984": "There are about 1000 samples in the public test, and about 500 samples in the private. Has the private been down-sampled along the z-axis? If so, will the connection between slices in 2.5d be affected?",
    "2619225": "alternatively, you can distill 2.5d, 3d , etc model to single slice 2d model",
    "2617104": "\" private been down-sampled along the z-axis? \"\n\ni think yes. (if you read the paper and the github CT reconstruction code. There are also discussions in the forum).\nit also  depends \n- on how many slice in 2.5d, e.g. 2,8,16,32 ...\n- your strategy for different resolution, e.g.\nare you are going 3d resize private data to 50um voxel/pixel?",
    "2618050": ""
  }
}