{
  "id": 472762,
  "title": "[0.891 !!] resnet50 is all you need in public lb, but what will happen in private?",
  "url": "/competitions/blood-vessel-segmentation/discussion/472762",
  "author_name": "lhwcv",
  "post_date": "2024-02-02T02:08:18.173000",
  "votes": 18,
  "comment_count": 31,
  "views": 0,
  "content": "<p>I'm currently very hesitant about how to handle these ResNet50. They perform poorly in cross-validation, but they can achieve e.g 0.891, 0.885, 0.881   on the public leaderboard. (Below comments are my results of one)<br>\nHow can we safely ensemble them to prevent shake-up? Should we trust the public leaderboard?</p>",
  "messages": [
    {
      "id": 2631856,
      "postDate": "2024-02-02T02:08:18.173Z",
      "content": "<p>I'm currently very hesitant about how to handle these ResNet50. They perform poorly in cross-validation, but they can achieve e.g 0.891, 0.885, 0.881   on the public leaderboard. (Below comments are my results of one)<br>\nHow can we safely ensemble them to prevent shake-up? Should we trust the public leaderboard?</p>",
      "rawMarkdown": "I'm currently very hesitant about how to handle these ResNet50. They perform poorly in cross-validation, but they can achieve e.g 0.891, 0.885, 0.881   on the public leaderboard. (Below comments are my results of one)\nHow can we safely ensemble them to prevent shake-up? Should we trust the public leaderboard?",
      "votes": 18
    },
    {
      "id": 2632404,
      "postDate": "2024-02-02T10:43:21.447Z",
      "content": "<p>One of the challenges here is not enough data to have a holdout set for validation to use for all, unless outside data was found. If a holdout set was used for all 3 kidneys then at least it might give a good idea of cross validation and reliability. </p>\n<p>Also the original image sizes vary amongst kidney 1, 2,3 - (1303, 912), (1041, 1511), (1706, 1510) respectively, whereas kidney 1 voi is (1928, 1928) but hard to find a way to use.  It is possible kidney 1 performs well because manipulating image size may lose less information.  Also there may be artifacts introduced that cause false positives when trying to get square images.  </p>\n<p>Then there are issues/challenges with the metric as discussed <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/472209\" target=\"_blank\">here</a></p>\n<p>The image sizes for Public and Private could be different to any of Train as well as each other.  Maybe doing inference testing on different image sizes and seeing the results on the validation sets would give an idea of how models perform.  Sometimes inference can be done on 2-3x image size but have not tried here.  </p>",
      "rawMarkdown": "One of the challenges here is not enough data to have a holdout set for validation to use for all, unless outside data was found. If a holdout set was used for all 3 kidneys then at least it might give a good idea of cross validation and reliability. \n\nAlso the original image sizes vary amongst kidney 1, 2,3 - (1303, 912), (1041, 1511), (1706, 1510) respectively, whereas kidney 1 voi is (1928, 1928) but hard to find a way to use.  It is possible kidney 1 performs well because manipulating image size may lose less information.  Also there may be artifacts introduced that cause false positives when trying to get square images.  \n\nThen there are issues/challenges with the metric as discussed [here](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/472209)\n\nThe image sizes for Public and Private could be different to any of Train as well as each other.  Maybe doing inference testing on different image sizes and seeing the results on the validation sets would give an idea of how models perform.  Sometimes inference can be done on 2-3x image size but have not tried here.  ",
      "votes": 3,
      "replies": [
        {
          "id": 2634960,
          "postDate": "2024-02-04T05:50:11.423Z",
          "content": "<p>thanks for sharing! have you tried infer on 2-3x image size now?</p>",
          "rawMarkdown": "thanks for sharing! have you tried infer on 2-3x image size now?",
          "replies": [
            {
              "id": 2636441,
              "postDate": "2024-02-05T05:16:21.373Z",
              "content": "<p>Basic test on inference on larger image size (not even 2x) unfortunately will get timeout. But could be different depending on models, etc. used.   Maybe try on the validation sets you use, see if it helps? </p>",
              "rawMarkdown": "Basic test on inference on larger image size (not even 2x) unfortunately will get timeout. But could be different depending on models, etc. used.   Maybe try on the validation sets you use, see if it helps? "
            }
          ]
        }
      ]
    },
    {
      "id": 2631858,
      "postDate": "2024-02-02T02:10:02.573Z",
      "content": "<p>kidney1 and kidney2 results:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F0491e5b17eaa1864ec477111d58555f3%2F5751706839303_.pic.jpg?generation=1706839782485189&amp;alt=media\"></p>",
      "rawMarkdown": "kidney1 and kidney2 results:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F0491e5b17eaa1864ec477111d58555f3%2F5751706839303_.pic.jpg?generation=1706839782485189&alt=media)",
      "votes": 3,
      "replies": [
        {
          "id": 2631877,
          "postDate": "2024-02-02T02:51:16.343Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2631924,
          "postDate": "2024-02-02T03:28:06.443Z",
          "content": "<p>I have a slight concern regarding the performance of my resnet50 model. The model was trained using data from Kidney 1 and Kidney 3. When I evaluate it on Kidney 2 (instances from 900 onwards), the local dice coefficient score of 0.95. However, the score on the lb is lower, specifically 0.842. I'm curious if the scoring function used by you differs from the one am using or if there might be another reason for this gap. Below is the scoring function I'm using:</p>\n<p>Image size: 1024</p>\n<pre><code>def dice_coef(y_pred: tc.Tensor, y_true: tc.Tensor, thr=, =(-, -), =):\n    y_pred = y_pred.sigmoid()\n    y_true = y_true.to(tc.float32)\n    y_pred = (y_pred &gt; thr).to(tc.float32)\n    inter = (y_true * y_pred).(=)\n    den = y_true.(=) + y_pred.(=)\n    dice = (( * inter + ) / (den + )).mean()\n     dice\n</code></pre>",
          "rawMarkdown": "I have a slight concern regarding the performance of my resnet50 model. The model was trained using data from Kidney 1 and Kidney 3. When I evaluate it on Kidney 2 (instances from 900 onwards), the local dice coefficient score of 0.95. However, the score on the lb is lower, specifically 0.842. I'm curious if the scoring function used by you differs from the one am using or if there might be another reason for this gap. Below is the scoring function I'm using:\n\nImage size: 1024\n\n```\ndef dice_coef(y_pred: tc.Tensor, y_true: tc.Tensor, thr=0.5, dim=(-1, -2), epsilon=0.001):\n    y_pred = y_pred.sigmoid()\n    y_true = y_true.to(tc.float32)\n    y_pred = (y_pred > thr).to(tc.float32)\n    inter = (y_true * y_pred).sum(dim=dim)\n    den = y_true.sum(dim=dim) + y_pred.sum(dim=dim)\n    dice = ((2 * inter + epsilon) / (den + epsilon)).mean()\n    return dice\n```",
          "replies": [
            {
              "id": 2631931,
              "postDate": "2024-02-02T03:34:22.017Z",
              "content": "<p>I'm using from here: <a href=\"https://www.kaggle.com/code/junkoda/fast-surface-dice-computation\" target=\"_blank\">https://www.kaggle.com/code/junkoda/fast-surface-dice-computation</a></p>",
              "rawMarkdown": "I'm using from here: https://www.kaggle.com/code/junkoda/fast-surface-dice-computation",
              "votes": 4
            }
          ]
        }
      ]
    },
    {
      "id": 2631859,
      "postDate": "2024-02-02T02:10:50.697Z",
      "content": "<p>kidney3 results:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F481ea1b09ce61429dc3a7e76840e9bc8%2F5761706839365_.pic.jpg?generation=1706839835920031&amp;alt=media\"></p>",
      "rawMarkdown": "kidney3 results:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F481ea1b09ce61429dc3a7e76840e9bc8%2F5761706839365_.pic.jpg?generation=1706839835920031&alt=media)",
      "votes": 4,
      "replies": [
        {
          "id": 2632338,
          "postDate": "2024-02-02T09:34:57Z",
          "content": "<p>Are you training on k1-k3 and validating on k3? </p>",
          "rawMarkdown": "Are you training on k1-k3 and validating on k3? ",
          "votes": 4,
          "replies": [
            {
              "id": 2634711,
              "postDate": "2024-02-04T00:33:10.853Z",
              "content": "<p>yes, I trained on kidney1 and kidney3, val on kidney3 with early stoping </p>",
              "rawMarkdown": "yes, I trained on kidney1 and kidney3, val on kidney3 with early stoping "
            }
          ]
        }
      ]
    },
    {
      "id": 2634079,
      "postDate": "2024-02-03T13:37:14.693Z",
      "content": "<p>May I ask how to build a CV in this competition? It seems unreasonable to split a kidney volume into K parts.</p>",
      "rawMarkdown": "May I ask how to build a CV in this competition? It seems unreasonable to split a kidney volume into K parts.",
      "votes": 1,
      "replies": [
        {
          "id": 2634092,
          "postDate": "2024-02-03T13:47:54.840Z",
          "content": "<p>Why do you think splitting the kidney volume into k parts is unreasonable? It seems reasonable to me to partition the kidney into contiguous chunks and to use those chunks for CV.</p>",
          "rawMarkdown": "Why do you think splitting the kidney volume into k parts is unreasonable? It seems reasonable to me to partition the kidney into contiguous chunks and to use those chunks for CV.",
          "replies": [
            {
              "id": 2635175,
              "postDate": "2024-02-04T07:41:34.313Z",
              "content": "<p>because the metric, surface dice, needs to be evaluated on a 3D volume. A chunk may not be accurate enough. </p>",
              "rawMarkdown": "because the metric, surface dice, needs to be evaluated on a 3D volume. A chunk may not be accurate enough. "
            }
          ]
        },
        {
          "id": 2634713,
          "postDate": "2024-02-04T00:34:05.177Z",
          "content": "<p>I trained on kidney1 and kidney3, val on kidney3 or kidney2 for different models</p>",
          "rawMarkdown": "I trained on kidney1 and kidney3, val on kidney3 or kidney2 for different models",
          "votes": 4,
          "replies": [
            {
              "id": 2635177,
              "postDate": "2024-02-04T07:42:17.967Z",
              "content": "<p>Thank you for this. I trained on kidney1 and validated on kidney3. </p>",
              "rawMarkdown": "Thank you for this. I trained on kidney1 and validated on kidney3. "
            }
          ]
        }
      ]
    },
    {
      "id": 2631985,
      "postDate": "2024-02-02T04:45:02.207Z",
      "content": "<p>If you use early stopping during training it may give you \"optimistic\" cv score which is not generalizable. <br>\nSometimes this can be noticed by getting worse score on lb for that model. So, I'd say the best option is to choose what improves both cv and lb at the same time.</p>",
      "rawMarkdown": "If you use early stopping during training it may give you \"optimistic\" cv score which is not generalizable. \nSometimes this can be noticed by getting worse score on lb for that model. So, I'd say the best option is to choose what improves both cv and lb at the same time.",
      "votes": 2,
      "replies": [
        {
          "id": 2632083,
          "postDate": "2024-02-02T06:07:13.250Z",
          "content": "<p>I think your methods is useful, now I decide to choose the best of  avg(cv + lb) </p>",
          "rawMarkdown": "I think your methods is useful, now I decide to choose the best of  avg(cv + lb) ",
          "votes": 1,
          "replies": [
            {
              "id": 2632430,
              "postDate": "2024-02-02T11:03:46.560Z",
              "content": "<p>\"So, I'd say the best option is to choose what improves both cv and lb at the same time.\"</p>\n<p>unfortunately this is not enough.<br>\nboth cv and lb = only 2 object samples.</p>\n<p>we can validate on both kidney2 and 3, e.g. improving both at the same time and yet there is poor correlations in private lb.</p>\n<hr>\n<p>even if we train on one and validate on the remaining 2, we have at most<br>\ncv and lb = only3 object samples … still too few!</p>",
              "rawMarkdown": "\"So, I'd say the best option is to choose what improves both cv and lb at the same time.\"\n\nunfortunately this is not enough.\nboth cv and lb = only 2 object samples.\n\nwe can validate on both kidney2 and 3, e.g. improving both at the same time and yet there is poor correlations in private lb.\n\n---\n\neven if we train on one and validate on the remaining 2, we have at most\ncv and lb = only3 object samples ... still too few!\n",
              "votes": 5
            },
            {
              "id": 2632483,
              "postDate": "2024-02-02T11:29:22.340Z",
              "content": "<p>What I'm worring about is that all kidneys in training data have resolution 50um. And the public lb has 50um as well. But private has 63um. So we don't have similar training images. Even if we sought after cv or lb (which are both 50um) what could guarantee we generalize to 63um well?<br>\nI mean, if we trained on kidney_1_voi which is ~5um i guess, and tried to validate on the other kidneys (which have less resolution) you get worse results right? So maybe the same could happen for our case?</p>",
              "rawMarkdown": "What I'm worring about is that all kidneys in training data have resolution 50um. And the public lb has 50um as well. But private has 63um. So we don't have similar training images. Even if we sought after cv or lb (which are both 50um) what could guarantee we generalize to 63um well?\nI mean, if we trained on kidney_1_voi which is ~5um i guess, and tried to validate on the other kidneys (which have less resolution) you get worse results right? So maybe the same could happen for our case?",
              "votes": 1
            },
            {
              "id": 2632484,
              "postDate": "2024-02-02T11:31:55.440Z",
              "content": "<p>Maybe the best cv is to validate on kidney_1_voi given its different resolution? Idk tbh.</p>",
              "rawMarkdown": "Maybe the best cv is to validate on kidney_1_voi given its different resolution? Idk tbh."
            },
            {
              "id": 2632504,
              "postDate": "2024-02-02T11:50:30.663Z",
              "content": "<p>why not downsample voi to private res?</p>",
              "rawMarkdown": "why not downsample voi to private res?",
              "votes": 1
            },
            {
              "id": 2632518,
              "postDate": "2024-02-02T12:12:56.457Z",
              "content": "<p>Kidney_1_voi is a subeset, while the others are full images.<br>\nBut why not downsampling kidney 3 to private res and try to validate on that? if we managed to get good score then it looks we are good to go with our 50um.<br>\nOr other thing is, Why not making several versions of kidneys 1_voi, 2 and 3, where each version has different resolution (downsampled ones) then validate on these kidneys?</p>",
              "rawMarkdown": "Kidney_1_voi is a subeset, while the others are full images.\nBut why not downsampling kidney 3 to private res and try to validate on that? if we managed to get good score then it looks we are good to go with our 50um.\nOr other thing is, Why not making several versions of kidneys 1_voi, 2 and 3, where each version has different resolution (downsampled ones) then validate on these kidneys?"
            },
            {
              "id": 2637770,
              "postDate": "2024-02-05T21:11:02.173Z",
              "content": "<p>Ensembling different models trained on varied image resolutions like 1024x1024 and 512x512 may be a viable approach.<br>\nThis could reduce the overfitting and boost the accuracy. </p>",
              "rawMarkdown": "Ensembling different models trained on varied image resolutions like 1024x1024 and 512x512 may be a viable approach.\nThis could reduce the overfitting and boost the accuracy. "
            }
          ]
        }
      ]
    },
    {
      "id": 2634877,
      "postDate": "2024-02-04T04:37:09.657Z",
      "content": "<p>While I'm not aware of the specific models you might have used aside from ResNet, could the larger channel sizes from ResNet(64, 256, 512, 1024, 2048) potentially contribute to its robustness?</p>",
      "rawMarkdown": "While I'm not aware of the specific models you might have used aside from ResNet, could the larger channel sizes from ResNet(64, 256, 512, 1024, 2048) potentially contribute to its robustness?",
      "replies": [
        {
          "id": 2634953,
          "postDate": "2024-02-04T05:48:03.333Z",
          "content": "<p>I use smp models and didn't change channel size, what's your result if using  larger channels?</p>",
          "rawMarkdown": "I use smp models and didn't change channel size, what's your result if using  larger channels?",
          "replies": [
            {
              "id": 2634970,
              "postDate": "2024-02-04T05:57:50.893Z",
              "content": "<p>what I meant was the output channels at each stage of the encoder.<br>\nI haven't experimented with that myself, so I'm not sure! 😃</p>",
              "rawMarkdown": "what I meant was the output channels at each stage of the encoder.\nI haven't experimented with that myself, so I'm not sure! 😃"
            }
          ]
        }
      ]
    },
    {
      "id": 2634057,
      "postDate": "2024-02-03T13:14:33.433Z",
      "content": "<p>May I ask if you are using the \"slide\" scheme or the \"full size\" scheme?</p>",
      "rawMarkdown": "May I ask if you are using the \"slide\" scheme or the \"full size\" scheme?",
      "replies": [
        {
          "id": 2634951,
          "postDate": "2024-02-04T05:46:28.500Z",
          "content": "<p>I'm using full size </p>",
          "rawMarkdown": "I'm using full size "
        }
      ]
    },
    {
      "id": 2637632,
      "postDate": "2024-02-05T19:50:05.713Z",
      "rawMarkdown": "",
      "votes": -2,
      "isDeleted": true
    },
    {
      "id": 2637628,
      "postDate": "2024-02-05T19:49:34.693Z",
      "rawMarkdown": "",
      "votes": -2,
      "isDeleted": true
    },
    {
      "id": 2631908,
      "postDate": "2024-02-02T03:10:22.947Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2632404,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "2024-02-02T10:43:21.447000",
      "content": "<p>One of the challenges here is not enough data to have a holdout set for validation to use for all, unless outside data was found. If a holdout set was used for all 3 kidneys then at least it might give a good idea of cross validation and reliability. </p>\n<p>Also the original image sizes vary amongst kidney 1, 2,3 - (1303, 912), (1041, 1511), (1706, 1510) respectively, whereas kidney 1 voi is (1928, 1928) but hard to find a way to use.  It is possible kidney 1 performs well because manipulating image size may lose less information.  Also there may be artifacts introduced that cause false positives when trying to get square images.  </p>\n<p>Then there are issues/challenges with the metric as discussed <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/472209\" target=\"_blank\">here</a></p>\n<p>The image sizes for Public and Private could be different to any of Train as well as each other.  Maybe doing inference testing on different image sizes and seeing the results on the validation sets would give an idea of how models perform.  Sometimes inference can be done on 2-3x image size but have not tried here.  </p>",
      "votes": 3,
      "replies": [
        {
          "id": 2634960,
          "author_name": "lhwcv",
          "author_url": "",
          "post_date": "2024-02-04T05:50:11.423000",
          "content": "<p>thanks for sharing! have you tried infer on 2-3x image size now?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2636441,
              "author_name": "something4kag",
              "author_url": "",
              "post_date": "2024-02-05T05:16:21.373000",
              "content": "<p>Basic test on inference on larger image size (not even 2x) unfortunately will get timeout. But could be different depending on models, etc. used.   Maybe try on the validation sets you use, see if it helps? </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2631858,
      "author_name": "lhwcv",
      "author_url": "",
      "post_date": "2024-02-02T02:10:02.573000",
      "content": "<p>kidney1 and kidney2 results:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F0491e5b17eaa1864ec477111d58555f3%2F5751706839303_.pic.jpg?generation=1706839782485189&amp;alt=media\"></p>",
      "votes": 3,
      "replies": [
        {
          "id": 2631877,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-02-02T02:51:16.343000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2631924,
          "author_name": "Aravind S",
          "author_url": "",
          "post_date": "2024-02-02T03:28:06.443000",
          "content": "<p>I have a slight concern regarding the performance of my resnet50 model. The model was trained using data from Kidney 1 and Kidney 3. When I evaluate it on Kidney 2 (instances from 900 onwards), the local dice coefficient score of 0.95. However, the score on the lb is lower, specifically 0.842. I'm curious if the scoring function used by you differs from the one am using or if there might be another reason for this gap. Below is the scoring function I'm using:</p>\n<p>Image size: 1024</p>\n<pre><code>def dice_coef(y_pred: tc.Tensor, y_true: tc.Tensor, thr=, =(-, -), =):\n    y_pred = y_pred.sigmoid()\n    y_true = y_true.to(tc.float32)\n    y_pred = (y_pred &gt; thr).to(tc.float32)\n    inter = (y_true * y_pred).(=)\n    den = y_true.(=) + y_pred.(=)\n    dice = (( * inter + ) / (den + )).mean()\n     dice\n</code></pre>",
          "votes": 0,
          "replies": [
            {
              "id": 2631931,
              "author_name": "lhwcv",
              "author_url": "",
              "post_date": "2024-02-02T03:34:22.017000",
              "content": "<p>I'm using from here: <a href=\"https://www.kaggle.com/code/junkoda/fast-surface-dice-computation\" target=\"_blank\">https://www.kaggle.com/code/junkoda/fast-surface-dice-computation</a></p>",
              "votes": 4,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2631859,
      "author_name": "lhwcv",
      "author_url": "",
      "post_date": "2024-02-02T02:10:50.697000",
      "content": "<p>kidney3 results:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F481ea1b09ce61429dc3a7e76840e9bc8%2F5761706839365_.pic.jpg?generation=1706839835920031&amp;alt=media\"></p>",
      "votes": 4,
      "replies": [
        {
          "id": 2632338,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2024-02-02T09:34:57",
          "content": "<p>Are you training on k1-k3 and validating on k3? </p>",
          "votes": 4,
          "replies": [
            {
              "id": 2634711,
              "author_name": "lhwcv",
              "author_url": "",
              "post_date": "2024-02-04T00:33:10.853000",
              "content": "<p>yes, I trained on kidney1 and kidney3, val on kidney3 with early stoping </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2634079,
      "author_name": "豆柴金鯱",
      "author_url": "",
      "post_date": "2024-02-03T13:37:14.693000",
      "content": "<p>May I ask how to build a CV in this competition? It seems unreasonable to split a kidney volume into K parts.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2634092,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "2024-02-03T13:47:54.840000",
          "content": "<p>Why do you think splitting the kidney volume into k parts is unreasonable? It seems reasonable to me to partition the kidney into contiguous chunks and to use those chunks for CV.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2635175,
              "author_name": "豆柴金鯱",
              "author_url": "",
              "post_date": "2024-02-04T07:41:34.313000",
              "content": "<p>because the metric, surface dice, needs to be evaluated on a 3D volume. A chunk may not be accurate enough. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2634713,
          "author_name": "lhwcv",
          "author_url": "",
          "post_date": "2024-02-04T00:34:05.177000",
          "content": "<p>I trained on kidney1 and kidney3, val on kidney3 or kidney2 for different models</p>",
          "votes": 4,
          "replies": [
            {
              "id": 2635177,
              "author_name": "豆柴金鯱",
              "author_url": "",
              "post_date": "2024-02-04T07:42:17.967000",
              "content": "<p>Thank you for this. I trained on kidney1 and validated on kidney3. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2631985,
      "author_name": "Mohamed Eltayeb",
      "author_url": "",
      "post_date": "2024-02-02T04:45:02.207000",
      "content": "<p>If you use early stopping during training it may give you \"optimistic\" cv score which is not generalizable. <br>\nSometimes this can be noticed by getting worse score on lb for that model. So, I'd say the best option is to choose what improves both cv and lb at the same time.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2632083,
          "author_name": "lhwcv",
          "author_url": "",
          "post_date": "2024-02-02T06:07:13.250000",
          "content": "<p>I think your methods is useful, now I decide to choose the best of  avg(cv + lb) </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2632430,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-02-02T11:03:46.560000",
              "content": "<p>\"So, I'd say the best option is to choose what improves both cv and lb at the same time.\"</p>\n<p>unfortunately this is not enough.<br>\nboth cv and lb = only 2 object samples.</p>\n<p>we can validate on both kidney2 and 3, e.g. improving both at the same time and yet there is poor correlations in private lb.</p>\n<hr>\n<p>even if we train on one and validate on the remaining 2, we have at most<br>\ncv and lb = only3 object samples … still too few!</p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2632483,
              "author_name": "Mohamed Eltayeb",
              "author_url": "",
              "post_date": "2024-02-02T11:29:22.340000",
              "content": "<p>What I'm worring about is that all kidneys in training data have resolution 50um. And the public lb has 50um as well. But private has 63um. So we don't have similar training images. Even if we sought after cv or lb (which are both 50um) what could guarantee we generalize to 63um well?<br>\nI mean, if we trained on kidney_1_voi which is ~5um i guess, and tried to validate on the other kidneys (which have less resolution) you get worse results right? So maybe the same could happen for our case?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2632484,
              "author_name": "Mohamed Eltayeb",
              "author_url": "",
              "post_date": "2024-02-02T11:31:55.440000",
              "content": "<p>Maybe the best cv is to validate on kidney_1_voi given its different resolution? Idk tbh.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2632504,
              "author_name": "Ángel Jacinto Sánchez Ruiz",
              "author_url": "",
              "post_date": "2024-02-02T11:50:30.663000",
              "content": "<p>why not downsample voi to private res?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2632518,
              "author_name": "Mohamed Eltayeb",
              "author_url": "",
              "post_date": "2024-02-02T12:12:56.457000",
              "content": "<p>Kidney_1_voi is a subeset, while the others are full images.<br>\nBut why not downsampling kidney 3 to private res and try to validate on that? if we managed to get good score then it looks we are good to go with our 50um.<br>\nOr other thing is, Why not making several versions of kidneys 1_voi, 2 and 3, where each version has different resolution (downsampled ones) then validate on these kidneys?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2637770,
              "author_name": "Min-Hsien Weng",
              "author_url": "",
              "post_date": "2024-02-05T21:11:02.173000",
              "content": "<p>Ensembling different models trained on varied image resolutions like 1024x1024 and 512x512 may be a viable approach.<br>\nThis could reduce the overfitting and boost the accuracy. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2634877,
      "author_name": "siwooyong",
      "author_url": "",
      "post_date": "2024-02-04T04:37:09.657000",
      "content": "<p>While I'm not aware of the specific models you might have used aside from ResNet, could the larger channel sizes from ResNet(64, 256, 512, 1024, 2048) potentially contribute to its robustness?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2634953,
          "author_name": "lhwcv",
          "author_url": "",
          "post_date": "2024-02-04T05:48:03.333000",
          "content": "<p>I use smp models and didn't change channel size, what's your result if using  larger channels?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2634970,
              "author_name": "siwooyong",
              "author_url": "",
              "post_date": "2024-02-04T05:57:50.893000",
              "content": "<p>what I meant was the output channels at each stage of the encoder.<br>\nI haven't experimented with that myself, so I'm not sure! 😃</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2634057,
      "author_name": "pky",
      "author_url": "",
      "post_date": "2024-02-03T13:14:33.433000",
      "content": "<p>May I ask if you are using the \"slide\" scheme or the \"full size\" scheme?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2634951,
          "author_name": "lhwcv",
          "author_url": "",
          "post_date": "2024-02-04T05:46:28.500000",
          "content": "<p>I'm using full size </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2637632,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-05T19:50:05.713000",
      "content": "",
      "votes": -2,
      "replies": []
    },
    {
      "id": 2637628,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-05T19:49:34.693000",
      "content": "",
      "votes": -2,
      "replies": []
    },
    {
      "id": 2631908,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-02T03:10:22.947000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2631856": "I'm currently very hesitant about how to handle these ResNet50. They perform poorly in cross-validation, but they can achieve e.g 0.891, 0.885, 0.881   on the public leaderboard. (Below comments are my results of one)\nHow can we safely ensemble them to prevent shake-up? Should we trust the public leaderboard?",
    "2632404": "One of the challenges here is not enough data to have a holdout set for validation to use for all, unless outside data was found. If a holdout set was used for all 3 kidneys then at least it might give a good idea of cross validation and reliability. \n\nAlso the original image sizes vary amongst kidney 1, 2,3 - (1303, 912), (1041, 1511), (1706, 1510) respectively, whereas kidney 1 voi is (1928, 1928) but hard to find a way to use.  It is possible kidney 1 performs well because manipulating image size may lose less information.  Also there may be artifacts introduced that cause false positives when trying to get square images.  \n\nThen there are issues/challenges with the metric as discussed [here](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/472209)\n\nThe image sizes for Public and Private could be different to any of Train as well as each other.  Maybe doing inference testing on different image sizes and seeing the results on the validation sets would give an idea of how models perform.  Sometimes inference can be done on 2-3x image size but have not tried here.  ",
    "2631858": "kidney1 and kidney2 results:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F0491e5b17eaa1864ec477111d58555f3%2F5751706839303_.pic.jpg?generation=1706839782485189&alt=media)",
    "2631859": "kidney3 results:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F481ea1b09ce61429dc3a7e76840e9bc8%2F5761706839365_.pic.jpg?generation=1706839835920031&alt=media)",
    "2634079": "May I ask how to build a CV in this competition? It seems unreasonable to split a kidney volume into K parts.",
    "2631985": "If you use early stopping during training it may give you \"optimistic\" cv score which is not generalizable. \nSometimes this can be noticed by getting worse score on lb for that model. So, I'd say the best option is to choose what improves both cv and lb at the same time.",
    "2634877": "While I'm not aware of the specific models you might have used aside from ResNet, could the larger channel sizes from ResNet(64, 256, 512, 1024, 2048) potentially contribute to its robustness?",
    "2634057": "May I ask if you are using the \"slide\" scheme or the \"full size\" scheme?",
    "2637632": "",
    "2637628": "",
    "2631908": ""
  }
}