{
  "id": 456714,
  "title": "CV vs public LB",
  "url": "/competitions/blood-vessel-segmentation/discussion/456714",
  "author_name": "YYama",
  "post_date": "2023-11-21T11:12:40.185000",
  "votes": 11,
  "comment_count": 25,
  "views": 0,
  "content": "<p>I couldn't find the usual discussion thread, so I created one.</p>\n<p>I used the <a href=\"https://www.kaggle.com/kashiwaba\" target=\"_blank\">@kashiwaba</a> notebook (<a href=\"https://www.kaggle.com/code/kashiwaba/sennet-hoa-train-unet-simple-baseline\" target=\"_blank\">https://www.kaggle.com/code/kashiwaba/sennet-hoa-train-unet-simple-baseline</a>) with minor changes. <br>\nTherefore, I used 'kidney_3_dense' for validation.</p>\n<p>CV 0.166, LB 0.262.</p>",
  "messages": [
    {
      "id": 2532825,
      "postDate": "2023-11-21T11:12:40.187Z",
      "content": "<p>I couldn't find the usual discussion thread, so I created one.</p>\n<p>I used the <a href=\"https://www.kaggle.com/kashiwaba\" target=\"_blank\">@kashiwaba</a> notebook (<a href=\"https://www.kaggle.com/code/kashiwaba/sennet-hoa-train-unet-simple-baseline\" target=\"_blank\">https://www.kaggle.com/code/kashiwaba/sennet-hoa-train-unet-simple-baseline</a>) with minor changes. <br>\nTherefore, I used 'kidney_3_dense' for validation.</p>\n<p>CV 0.166, LB 0.262.</p>",
      "rawMarkdown": "I couldn't find the usual discussion thread, so I created one.\n\nI used the @kashiwaba notebook (https://www.kaggle.com/code/kashiwaba/sennet-hoa-train-unet-simple-baseline) with minor changes. \nTherefore, I used 'kidney_3_dense' for validation.\n\nCV 0.166, LB 0.262.",
      "votes": 11
    },
    {
      "id": 2533262,
      "postDate": "2023-11-21T17:55:16.780Z",
      "content": "<p>I train the model on only kidney_1_dense, hflip+vflip+randomrotate90 augs, I don't have a steady validation yet</p>\n<p>CV kidney_3_dense: 0.821<br>\nLB: 0.782</p>",
      "rawMarkdown": "I train the model on only kidney_1_dense, hflip+vflip+randomrotate90 augs, I don't have a steady validation yet\n\nCV kidney_3_dense: 0.821\nLB: 0.782",
      "votes": 5,
      "replies": [
        {
          "id": 2533456,
          "postDate": "2023-11-22T00:29:11.870Z",
          "content": "<p>What metrics u use to calculate cv </p>",
          "rawMarkdown": "What metrics u use to calculate cv ",
          "replies": [
            {
              "id": 2533518,
              "postDate": "2023-11-22T02:06:35.293Z",
              "content": "<p>Surface Dice, the competition metric</p>",
              "rawMarkdown": "Surface Dice, the competition metric"
            },
            {
              "id": 2553147,
              "postDate": "2023-12-08T03:38:45.320Z",
              "content": "<p>I am trying to understand why train on one kidney_1_dense? Why not all? <br>\nThanks in adavance.</p>",
              "rawMarkdown": "I am trying to understand why train on one kidney_1_dense? Why not all? \nThanks in adavance.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2562407,
      "postDate": "2023-12-15T10:52:50.473Z",
      "content": "<p>I follow hengck's CV setup [train on kidney_1_dense, validate on kidney_3_dense], 2d, single model</p>\n<p>I use hengck's <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/461213#2561478\" target=\"_blank\">helper.py</a> to compute the score</p>\n<table>\n<thead>\n<tr>\n<th>kidney_3_dense</th>\n<th>kidney_2</th>\n<th>public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.8549</td>\n<td>-</td>\n<td>0.781</td>\n</tr>\n<tr>\n<td>0.8788</td>\n<td>-</td>\n<td>0.793</td>\n</tr>\n<tr>\n<td>0.8920</td>\n<td>-</td>\n<td>0.816</td>\n</tr>\n<tr>\n<td>0.9019</td>\n<td>0.8538</td>\n<td>0.839</td>\n</tr>\n<tr>\n<td>-</td>\n<td>0.8607</td>\n<td>0.789…?</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "I follow hengck's CV setup [train on kidney_1_dense, validate on kidney_3_dense], 2d, single model\n\nI use hengck's [helper.py](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/461213#2561478) to compute the score\n\n| kidney_3_dense | kidney_2 | public LB |\n| ---|---|---|\n| 0.8549 |-| 0.781 |\n| 0.8788 |- | 0.793 |\n| 0.8920 |- | 0.816 |\n| 0.9019 | 0.8538 | 0.839 |  \n| -  | 0.8607 | 0.789...? |",
      "votes": 1,
      "replies": [
        {
          "id": 2562408,
          "postDate": "2023-12-15T10:53:20.190Z",
          "content": "<p>I'm not sure about correlation between CV and LB, I'm trying to tune threshold on kidney_3_dense, <br>\nbut when using TTA the optimal choice for public LB is to use higher threshold</p>\n<p>for my models, 0.10 is optimal threshold for kindey_3_dense and ~0.23 for kidney_2</p>",
          "rawMarkdown": "I'm not sure about correlation between CV and LB, I'm trying to tune threshold on kidney_3_dense, \nbut when using TTA the optimal choice for public LB is to use higher threshold\n\nfor my models, 0.10 is optimal threshold for kindey_3_dense and ~0.23 for kidney_2",
          "replies": [
            {
              "id": 2562420,
              "postDate": "2023-12-15T11:04:30.470Z",
              "content": "<p>Looks too high for me, my best lb model's cv is about 0.85</p>",
              "rawMarkdown": "Looks too high for me, my best lb model's cv is about 0.85",
              "votes": 1
            },
            {
              "id": 2562446,
              "postDate": "2023-12-15T11:57:52.463Z",
              "content": "<p>I found a bug, - I forgot to sort slices in validation [this is not the first time glob does something like this] :D<br>\nI'll fix the post later<br>\nP.S. done</p>",
              "rawMarkdown": "I found a bug, - I forgot to sort slices in validation [this is not the first time glob does something like this] :D\nI'll fix the post later\nP.S. done"
            }
          ]
        },
        {
          "id": 2562410,
          "postDate": "2023-12-15T10:54:47.347Z",
          "content": "<p>For a single model, threshold values are expected to be smooth judging from this plot <br>\n[2d + XY/XY/YZ inference, kidney_3_dense]<br>\nhowever, the threshold is not optimal on public LB, +-0.05 threshold results in +-2.5 surface dice shift, <br>\nhence we should expect a moderate shakeup</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5487737%2Ffbde6d50bb97e27e7b5a8c520c2ccdfa%2Fkidney_3.jpg?generation=1703019570155688&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "For a single model, threshold values are expected to be smooth judging from this plot \n[2d + XY/XY/YZ inference, kidney_3_dense]\nhowever, the threshold is not optimal on public LB, +-0.05 threshold results in +-2.5 surface dice shift, \nhence we should expect a moderate shakeup\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5487737%2Ffbde6d50bb97e27e7b5a8c520c2ccdfa%2Fkidney_3.jpg?generation=1703019570155688&alt=media)",
          "replies": [
            {
              "id": 2562701,
              "postDate": "2023-12-15T16:00:01.950Z",
              "content": "<p>threshold for hidden test and local may be different<br>\n(sign for shakeup)</p>",
              "rawMarkdown": "threshold for hidden test and local may be different\n(sign for shakeup)"
            }
          ]
        }
      ]
    },
    {
      "id": 2533092,
      "postDate": "2023-11-21T15:26:37.837Z",
      "content": "<p>I naively took an 80:20 split for train and validation that worked pretty well for the start, but that was mostly just for marble testing my workflow. I'm switching up my strategy now to properly hold out a full partition which I expect to provide better insight in to public LB score. I posted more information on my experimental configuration in another <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/456787\" target=\"_blank\">post</a>, but here are my current results:</p>\n<table>\n<thead>\n<tr>\n<th>Threshold</th>\n<th>CV (Kidney_3_dense)</th>\n<th>LB</th>\n<th>Notes</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.5</td>\n<td>0.53</td>\n<td>0.65</td>\n<td>B3</td>\n</tr>\n<tr>\n<td>0.05</td>\n<td>0.55</td>\n<td>Running</td>\n<td>B3 + Reduced Score Threshold</td>\n</tr>\n<tr>\n<td>0.05</td>\n<td>0.55</td>\n<td>OOM</td>\n<td>B3 + Size Adjustment + Reduced Score Threshold</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "I naively took an 80:20 split for train and validation that worked pretty well for the start, but that was mostly just for marble testing my workflow. I'm switching up my strategy now to properly hold out a full partition which I expect to provide better insight in to public LB score. I posted more information on my experimental configuration in another [post](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/456787), but here are my current results:\n\n| Threshold | CV (Kidney_3_dense)  | LB | Notes |\n| --- | --- | --- | --- |\n| 0.5 |  0.53 | 0.65 | B3 |\n| 0.05 |  0.55 | Running| B3 + Reduced Score Threshold |\n| 0.05 |  0.55 | OOM | B3 + Size Adjustment + Reduced Score Threshold |",
      "votes": 1
    },
    {
      "id": 2539822,
      "postDate": "2023-11-27T10:38:10.307Z",
      "content": "<p>Train: kidney_1<br>\nvalid: kidney_3<br>\nmodel: resnet50-unet2d<br>\npostprocess: remove small objects (min_area: 10)</p>\n<blockquote>\n  <table>\n  <thead>\n  <tr>\n  <th>Threshold</th>\n  <th>CV (Kidney_3_dense)</th>\n  <th>LB</th>\n  </tr>\n  </thead>\n  <tbody>\n  <tr>\n  <td>0.1</td>\n  <td>0.752</td>\n  <td>0.648</td>\n  </tr>\n  <tr>\n  <td>0.5</td>\n  <td>0.818</td>\n  <td>0.694</td>\n  </tr>\n  </tbody>\n  </table>\n</blockquote>",
      "rawMarkdown": "Train: kidney_1\nvalid: kidney_3\nmodel: resnet50-unet2d\npostprocess: remove small objects (min_area: 10)\n\n> \n> | Threshold | CV (Kidney_3_dense)  | LB |\n> | --- | --- | --- |\n> | 0.1 |  0.752 | 0.648 |\n> | 0.5 |  0.818 | 0.694|",
      "votes": 2,
      "replies": [
        {
          "id": 2539923,
          "postDate": "2023-11-27T11:55:37.567Z",
          "content": "<p>May i ask what size you use for training? also only used one mask?</p>",
          "rawMarkdown": "May i ask what size you use for training? also only used one mask?",
          "votes": 1,
          "replies": [
            {
              "id": 2539952,
              "postDate": "2023-11-27T12:17:28.243Z",
              "content": "<p>full resolution(train &amp; test), and yes only blood-vessel mask.</p>",
              "rawMarkdown": "full resolution(train & test), and yes only blood-vessel mask.",
              "votes": 1
            },
            {
              "id": 2539957,
              "postDate": "2023-11-27T12:23:41.267Z",
              "content": "<p>Thank you for reply! good luck with this competition!!</p>",
              "rawMarkdown": "Thank you for reply! good luck with this competition!!",
              "votes": 2
            },
            {
              "id": 2542020,
              "postDate": "2023-11-29T02:31:40.637Z",
              "content": "<p><a href=\"https://www.kaggle.com/ihebch\" target=\"_blank\">@ihebch</a> <br>\nCould you kindly guide me on utilizing the full size as input? Typically, models require symmetry, like 512x512. Additionally, should the test set be resized to the input size employed during the training of the kidney_1 model? <br>\nThank you very much in advance for your time and consideration.</p>",
              "rawMarkdown": "@ihebch \nCould you kindly guide me on utilizing the full size as input? Typically, models require symmetry, like 512x512. Additionally, should the test set be resized to the input size employed during the training of the kidney_1 model? \nThank you very much in advance for your time and consideration.",
              "votes": 1
            }
          ]
        },
        {
          "id": 2543783,
          "postDate": "2023-11-30T11:09:28.887Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2533054,
      "postDate": "2023-11-21T14:51:08.883Z",
      "content": "<p>I am currently doing validation on kidney_3_sparse (note that the comp score between sparse &amp; dense annotation gives 0.981 so I think I'd rather validate on the largest possible data).</p>\n<p>This OOM issue is killing me, took me 12 submissions to get a LB score!</p>\n<table>\n<thead>\n<tr>\n<th>Threshold</th>\n<th>CV (kidney_3_sparse)</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.5</td>\n<td>0.778</td>\n<td>OOM</td>\n</tr>\n<tr>\n<td>0.85</td>\n<td>0.642</td>\n<td>OOM</td>\n</tr>\n<tr>\n<td>0.95</td>\n<td>0.466</td>\n<td>0.451</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "I am currently doing validation on kidney_3_sparse (note that the comp score between sparse & dense annotation gives 0.981 so I think I'd rather validate on the largest possible data).\n\nThis OOM issue is killing me, took me 12 submissions to get a LB score!\n\n| Threshold |         CV (kidney_3_sparse)         |    LB    |\n|:---------:|:------------------:|:--------:|\n|    0.5    |       0.778        |   OOM    |\n|   0.85    |   0.642    |   OOM    |\n|   0.95    |       0.466        |  0.451   |",
      "votes": 2,
      "replies": [
        {
          "id": 2552194,
          "postDate": "2023-12-07T09:28:31.523Z",
          "content": "<p>After metric's fix:</p>\n<table>\n<thead>\n<tr>\n<th>Threshold</th>\n<th>CV (kidney_3_sparse)</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.5</td>\n<td>0.778</td>\n<td>0.729</td>\n</tr>\n<tr>\n<td>0.85</td>\n<td>0.642</td>\n<td>0.598</td>\n</tr>\n<tr>\n<td>0.95</td>\n<td>0.466</td>\n<td>0.451</td>\n</tr>\n</tbody>\n</table>",
          "rawMarkdown": "After metric's fix:\n| Threshold |         CV (kidney_3_sparse)         |    LB    |\n|:---------:|:------------------:|:--------:|\n|    0.5    |       0.778        |   0.729    |\n|   0.85    |   0.642    |   0.598    |\n|   0.95    |       0.466        |  0.451   |",
          "votes": 2,
          "replies": [
            {
              "id": 2552238,
              "postDate": "2023-12-07T10:18:27.450Z",
              "content": "<p>0.981 is only for part of kidney 3. validate on kidney 3 dense may be better</p>",
              "rawMarkdown": " 0.981 is only for part of kidney 3. validate on kidney 3 dense may be better"
            },
            {
              "id": 2552240,
              "postDate": "2023-12-07T10:19:28.310Z",
              "content": "<p>cv vs lb on kidney 3 should be about 0.04</p>",
              "rawMarkdown": "cv vs lb on kidney 3 should be about 0.04"
            },
            {
              "id": 2554157,
              "postDate": "2023-12-08T20:38:19.960Z",
              "content": "<p><a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> could you tell which metric version (i.e v24) are you using and which image resolution?</p>",
              "rawMarkdown": "@optimo could you tell which metric version (i.e v24) are you using and which image resolution?"
            },
            {
              "id": 2554190,
              "postDate": "2023-12-08T21:52:15.087Z",
              "content": "<p>This is the latest version of the official metric, these are results from a 3D model (crops) at full resolution.</p>",
              "rawMarkdown": "This is the latest version of the official metric, these are results from a 3D model (crops) at full resolution."
            }
          ]
        }
      ]
    },
    {
      "id": 2538287,
      "postDate": "2023-11-26T01:26:35.137Z",
      "content": "<p>does your model overfit</p>",
      "rawMarkdown": "does your model overfit\n"
    },
    {
      "id": 2537244,
      "postDate": "2023-11-25T00:17:18.243Z",
      "content": "<p>efficientnet-b1 + U-Net<br>\nInput image size is 1536 x 1536.<br>\nCV score is 0.544, LB score is 0.493 (the threshold for CV is 0.5 and for LB is 0.1, I forgot to match them).</p>\n<p>Since the surface dice score is the evaluation metric, missing tiny blood vessels worse the score. Therefore, is it important to increase the image size or to use tiling?</p>\n<p>edit: Using both VOI and dense for training resulted in worse CV and LB scores, so I used only dense of kidney1  for training and validated with kidney3_dense.</p>",
      "rawMarkdown": "efficientnet-b1 + U-Net\nInput image size is 1536 x 1536.\nCV score is 0.544, LB score is 0.493 (the threshold for CV is 0.5 and for LB is 0.1, I forgot to match them).\n\nSince the surface dice score is the evaluation metric, missing tiny blood vessels worse the score. Therefore, is it important to increase the image size or to use tiling?\n\nedit: Using both VOI and dense for training resulted in worse CV and LB scores, so I used only dense of kidney1  for training and validated with kidney3_dense."
    }
  ],
  "comments": [
    {
      "id": 2533262,
      "author_name": "Harshit Sheoran",
      "author_url": "",
      "post_date": "2023-11-21T17:55:16.780000",
      "content": "<p>I train the model on only kidney_1_dense, hflip+vflip+randomrotate90 augs, I don't have a steady validation yet</p>\n<p>CV kidney_3_dense: 0.821<br>\nLB: 0.782</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2533456,
          "author_name": "Arunodhayan",
          "author_url": "",
          "post_date": "2023-11-22T00:29:11.870000",
          "content": "<p>What metrics u use to calculate cv </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2533518,
              "author_name": "Harshit Sheoran",
              "author_url": "",
              "post_date": "2023-11-22T02:06:35.293000",
              "content": "<p>Surface Dice, the competition metric</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2553147,
              "author_name": "FdotRK",
              "author_url": "",
              "post_date": "2023-12-08T03:38:45.320000",
              "content": "<p>I am trying to understand why train on one kidney_1_dense? Why not all? <br>\nThanks in adavance.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2562407,
      "author_name": "slime",
      "author_url": "",
      "post_date": "2023-12-15T10:52:50.473000",
      "content": "<p>I follow hengck's CV setup [train on kidney_1_dense, validate on kidney_3_dense], 2d, single model</p>\n<p>I use hengck's <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/461213#2561478\" target=\"_blank\">helper.py</a> to compute the score</p>\n<table>\n<thead>\n<tr>\n<th>kidney_3_dense</th>\n<th>kidney_2</th>\n<th>public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.8549</td>\n<td>-</td>\n<td>0.781</td>\n</tr>\n<tr>\n<td>0.8788</td>\n<td>-</td>\n<td>0.793</td>\n</tr>\n<tr>\n<td>0.8920</td>\n<td>-</td>\n<td>0.816</td>\n</tr>\n<tr>\n<td>0.9019</td>\n<td>0.8538</td>\n<td>0.839</td>\n</tr>\n<tr>\n<td>-</td>\n<td>0.8607</td>\n<td>0.789…?</td>\n</tr>\n</tbody>\n</table>",
      "votes": 1,
      "replies": [
        {
          "id": 2562408,
          "author_name": "slime",
          "author_url": "",
          "post_date": "2023-12-15T10:53:20.190000",
          "content": "<p>I'm not sure about correlation between CV and LB, I'm trying to tune threshold on kidney_3_dense, <br>\nbut when using TTA the optimal choice for public LB is to use higher threshold</p>\n<p>for my models, 0.10 is optimal threshold for kindey_3_dense and ~0.23 for kidney_2</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2562420,
              "author_name": "Snorf",
              "author_url": "",
              "post_date": "2023-12-15T11:04:30.470000",
              "content": "<p>Looks too high for me, my best lb model's cv is about 0.85</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2562446,
              "author_name": "slime",
              "author_url": "",
              "post_date": "2023-12-15T11:57:52.463000",
              "content": "<p>I found a bug, - I forgot to sort slices in validation [this is not the first time glob does something like this] :D<br>\nI'll fix the post later<br>\nP.S. done</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2562410,
          "author_name": "slime",
          "author_url": "",
          "post_date": "2023-12-15T10:54:47.347000",
          "content": "<p>For a single model, threshold values are expected to be smooth judging from this plot <br>\n[2d + XY/XY/YZ inference, kidney_3_dense]<br>\nhowever, the threshold is not optimal on public LB, +-0.05 threshold results in +-2.5 surface dice shift, <br>\nhence we should expect a moderate shakeup</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5487737%2Ffbde6d50bb97e27e7b5a8c520c2ccdfa%2Fkidney_3.jpg?generation=1703019570155688&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": [
            {
              "id": 2562701,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-12-15T16:00:01.950000",
              "content": "<p>threshold for hidden test and local may be different<br>\n(sign for shakeup)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2533092,
      "author_name": "kcetskcaz",
      "author_url": "",
      "post_date": "2023-11-21T15:26:37.837000",
      "content": "<p>I naively took an 80:20 split for train and validation that worked pretty well for the start, but that was mostly just for marble testing my workflow. I'm switching up my strategy now to properly hold out a full partition which I expect to provide better insight in to public LB score. I posted more information on my experimental configuration in another <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/456787\" target=\"_blank\">post</a>, but here are my current results:</p>\n<table>\n<thead>\n<tr>\n<th>Threshold</th>\n<th>CV (Kidney_3_dense)</th>\n<th>LB</th>\n<th>Notes</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.5</td>\n<td>0.53</td>\n<td>0.65</td>\n<td>B3</td>\n</tr>\n<tr>\n<td>0.05</td>\n<td>0.55</td>\n<td>Running</td>\n<td>B3 + Reduced Score Threshold</td>\n</tr>\n<tr>\n<td>0.05</td>\n<td>0.55</td>\n<td>OOM</td>\n<td>B3 + Size Adjustment + Reduced Score Threshold</td>\n</tr>\n</tbody>\n</table>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2539822,
      "author_name": "Reacher",
      "author_url": "",
      "post_date": "2023-11-27T10:38:10.307000",
      "content": "<p>Train: kidney_1<br>\nvalid: kidney_3<br>\nmodel: resnet50-unet2d<br>\npostprocess: remove small objects (min_area: 10)</p>\n<blockquote>\n  <table>\n  <thead>\n  <tr>\n  <th>Threshold</th>\n  <th>CV (Kidney_3_dense)</th>\n  <th>LB</th>\n  </tr>\n  </thead>\n  <tbody>\n  <tr>\n  <td>0.1</td>\n  <td>0.752</td>\n  <td>0.648</td>\n  </tr>\n  <tr>\n  <td>0.5</td>\n  <td>0.818</td>\n  <td>0.694</td>\n  </tr>\n  </tbody>\n  </table>\n</blockquote>",
      "votes": 2,
      "replies": [
        {
          "id": 2539923,
          "author_name": "kaggler",
          "author_url": "",
          "post_date": "2023-11-27T11:55:37.567000",
          "content": "<p>May i ask what size you use for training? also only used one mask?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2539952,
              "author_name": "Reacher",
              "author_url": "",
              "post_date": "2023-11-27T12:17:28.243000",
              "content": "<p>full resolution(train &amp; test), and yes only blood-vessel mask.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2539957,
              "author_name": "kaggler",
              "author_url": "",
              "post_date": "2023-11-27T12:23:41.267000",
              "content": "<p>Thank you for reply! good luck with this competition!!</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2542020,
              "author_name": "JamesLearnsToCode",
              "author_url": "",
              "post_date": "2023-11-29T02:31:40.637000",
              "content": "<p><a href=\"https://www.kaggle.com/ihebch\" target=\"_blank\">@ihebch</a> <br>\nCould you kindly guide me on utilizing the full size as input? Typically, models require symmetry, like 512x512. Additionally, should the test set be resized to the input size employed during the training of the kidney_1 model? <br>\nThank you very much in advance for your time and consideration.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2543783,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-11-30T11:09:28.887000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2533054,
      "author_name": "Optimo",
      "author_url": "",
      "post_date": "2023-11-21T14:51:08.883000",
      "content": "<p>I am currently doing validation on kidney_3_sparse (note that the comp score between sparse &amp; dense annotation gives 0.981 so I think I'd rather validate on the largest possible data).</p>\n<p>This OOM issue is killing me, took me 12 submissions to get a LB score!</p>\n<table>\n<thead>\n<tr>\n<th>Threshold</th>\n<th>CV (kidney_3_sparse)</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.5</td>\n<td>0.778</td>\n<td>OOM</td>\n</tr>\n<tr>\n<td>0.85</td>\n<td>0.642</td>\n<td>OOM</td>\n</tr>\n<tr>\n<td>0.95</td>\n<td>0.466</td>\n<td>0.451</td>\n</tr>\n</tbody>\n</table>",
      "votes": 2,
      "replies": [
        {
          "id": 2552194,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2023-12-07T09:28:31.523000",
          "content": "<p>After metric's fix:</p>\n<table>\n<thead>\n<tr>\n<th>Threshold</th>\n<th>CV (kidney_3_sparse)</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.5</td>\n<td>0.778</td>\n<td>0.729</td>\n</tr>\n<tr>\n<td>0.85</td>\n<td>0.642</td>\n<td>0.598</td>\n</tr>\n<tr>\n<td>0.95</td>\n<td>0.466</td>\n<td>0.451</td>\n</tr>\n</tbody>\n</table>",
          "votes": 2,
          "replies": [
            {
              "id": 2552238,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-12-07T10:18:27.450000",
              "content": "<p>0.981 is only for part of kidney 3. validate on kidney 3 dense may be better</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2552240,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-12-07T10:19:28.310000",
              "content": "<p>cv vs lb on kidney 3 should be about 0.04</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2554157,
              "author_name": "moth",
              "author_url": "",
              "post_date": "2023-12-08T20:38:19.960000",
              "content": "<p><a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> could you tell which metric version (i.e v24) are you using and which image resolution?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2554190,
              "author_name": "Optimo",
              "author_url": "",
              "post_date": "2023-12-08T21:52:15.087000",
              "content": "<p>This is the latest version of the official metric, these are results from a 3D model (crops) at full resolution.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2538287,
      "author_name": "Arunodhayan",
      "author_url": "",
      "post_date": "2023-11-26T01:26:35.137000",
      "content": "<p>does your model overfit</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2537244,
      "author_name": "YYama",
      "author_url": "",
      "post_date": "2023-11-25T00:17:18.243000",
      "content": "<p>efficientnet-b1 + U-Net<br>\nInput image size is 1536 x 1536.<br>\nCV score is 0.544, LB score is 0.493 (the threshold for CV is 0.5 and for LB is 0.1, I forgot to match them).</p>\n<p>Since the surface dice score is the evaluation metric, missing tiny blood vessels worse the score. Therefore, is it important to increase the image size or to use tiling?</p>\n<p>edit: Using both VOI and dense for training resulted in worse CV and LB scores, so I used only dense of kidney1  for training and validated with kidney3_dense.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2532825": "I couldn't find the usual discussion thread, so I created one.\n\nI used the @kashiwaba notebook (https://www.kaggle.com/code/kashiwaba/sennet-hoa-train-unet-simple-baseline) with minor changes. \nTherefore, I used 'kidney_3_dense' for validation.\n\nCV 0.166, LB 0.262.",
    "2533262": "I train the model on only kidney_1_dense, hflip+vflip+randomrotate90 augs, I don't have a steady validation yet\n\nCV kidney_3_dense: 0.821\nLB: 0.782",
    "2562407": "I follow hengck's CV setup [train on kidney_1_dense, validate on kidney_3_dense], 2d, single model\n\nI use hengck's [helper.py](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/461213#2561478) to compute the score\n\n| kidney_3_dense | kidney_2 | public LB |\n| ---|---|---|\n| 0.8549 |-| 0.781 |\n| 0.8788 |- | 0.793 |\n| 0.8920 |- | 0.816 |\n| 0.9019 | 0.8538 | 0.839 |  \n| -  | 0.8607 | 0.789...? |",
    "2533092": "I naively took an 80:20 split for train and validation that worked pretty well for the start, but that was mostly just for marble testing my workflow. I'm switching up my strategy now to properly hold out a full partition which I expect to provide better insight in to public LB score. I posted more information on my experimental configuration in another [post](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/456787), but here are my current results:\n\n| Threshold | CV (Kidney_3_dense)  | LB | Notes |\n| --- | --- | --- | --- |\n| 0.5 |  0.53 | 0.65 | B3 |\n| 0.05 |  0.55 | Running| B3 + Reduced Score Threshold |\n| 0.05 |  0.55 | OOM | B3 + Size Adjustment + Reduced Score Threshold |",
    "2539822": "Train: kidney_1\nvalid: kidney_3\nmodel: resnet50-unet2d\npostprocess: remove small objects (min_area: 10)\n\n> \n> | Threshold | CV (Kidney_3_dense)  | LB |\n> | --- | --- | --- |\n> | 0.1 |  0.752 | 0.648 |\n> | 0.5 |  0.818 | 0.694|",
    "2533054": "I am currently doing validation on kidney_3_sparse (note that the comp score between sparse & dense annotation gives 0.981 so I think I'd rather validate on the largest possible data).\n\nThis OOM issue is killing me, took me 12 submissions to get a LB score!\n\n| Threshold |         CV (kidney_3_sparse)         |    LB    |\n|:---------:|:------------------:|:--------:|\n|    0.5    |       0.778        |   OOM    |\n|   0.85    |   0.642    |   OOM    |\n|   0.95    |       0.466        |  0.451   |",
    "2538287": "does your model overfit\n",
    "2537244": "efficientnet-b1 + U-Net\nInput image size is 1536 x 1536.\nCV score is 0.544, LB score is 0.493 (the threshold for CV is 0.5 and for LB is 0.1, I forgot to match them).\n\nSince the surface dice score is the evaluation metric, missing tiny blood vessels worse the score. Therefore, is it important to increase the image size or to use tiling?\n\nedit: Using both VOI and dense for training resulted in worse CV and LB scores, so I used only dense of kidney1  for training and validated with kidney3_dense."
  }
}