{
  "id": 468525,
  "title": "How to reduce shake-up, has anyone found a correlation between CV and LB? Here are some my results",
  "url": "/competitions/blood-vessel-segmentation/discussion/468525",
  "author_name": "",
  "post_date": "2024-01-17T00:33:06.549802100Z",
  "votes": 22,
  "comment_count": 22,
  "views": 0,
  "content": "<p>I trained on K1 and val on K3,  or  train K1,K3 val on K2,  both not found any correlation between CV and LB.<br>\nHas anyone use K2 sparse, K3 sparse or other mixture of them  for train, what's the result?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F6ffe902ddc7efeaddfbc65dd31c2c3f7%2F5631705451322_.pic.jpg?generation=1705451397035794&amp;alt=media\"></p>",
  "messages": [
    {
      "id": "2605231",
      "postDate": "01/17/2024 00:33:06",
      "content": "<p>I trained on K1 and val on K3,  or  train K1,K3 val on K2,  both not found any correlation between CV and LB.<br>\nHas anyone use K2 sparse, K3 sparse or other mixture of them  for train, what's the result?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F6ffe902ddc7efeaddfbc65dd31c2c3f7%2F5631705451322_.pic.jpg?generation=1705451397035794&amp;alt=media\"></p>",
      "rawMarkdown": "I trained on K1 and val on K3,  or  train K1,K3 val on K2,  both not found any correlation between CV and LB.\nHas anyone use K2 sparse, K3 sparse or other mixture of them  for train, what's the result?\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F6ffe902ddc7efeaddfbc65dd31c2c3f7%2F5631705451322_.pic.jpg?generation=1705451397035794&alt=media)",
      "votes": null
    },
    {
      "id": "2605446",
      "postDate": "01/17/2024 04:38:21",
      "content": "<p>the possible problem is the public/private dataset distribution? if they are same type of data.  <br>\neither models trained on all data would be little bit lower on Public while stable on Private,<br>\nor models with 2 stages(classification and prediction) would be sustainable to shake up.</p>\n<p>However, without local computation power, I am lazy to try out my simple ideas.</p>",
      "rawMarkdown": "the possible problem is the public/private dataset distribution? if they are same type of data.  \neither models trained on all data would be little bit lower on Public while stable on Private,\nor models with 2 stages(classification and prediction) would be sustainable to shake up.\n\nHowever, without local computation power, I am lazy to try out my simple ideas.",
      "votes": null
    },
    {
      "id": "2606272",
      "postDate": "01/17/2024 14:29:25",
      "content": "<p>Can you please clarify what you mean by k2 and k2 sparse? So that I can share comparable results, thanks!</p>",
      "rawMarkdown": "Can you please clarify what you mean by k2 and k2 sparse? So that I can share comparable results, thanks!",
      "votes": null
    },
    {
      "id": "2606278",
      "postDate": "01/17/2024 14:36:11",
      "content": "<p>May be k2 is semipseudolabeled</p>",
      "rawMarkdown": "May be k2 is semipseudolabeled",
      "votes": null
    },
    {
      "id": "2607125",
      "postDate": "01/18/2024 04:07:37",
      "content": "<p>I mean use kidney2 to train or kidney3 sparse to train, what’s the difference between using kidney1 dense in LB</p>",
      "rawMarkdown": "I mean use kidney2 to train or kidney3 sparse to train, what’s the difference between using kidney1 dense in LB",
      "votes": null
    },
    {
      "id": "2607654",
      "postDate": "01/18/2024 11:00:59",
      "content": "<p>is your score for the entire k2 or k2 on slices &gt; 900 as your teammate reports <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/468826\" target=\"_blank\">here</a> ?</p>",
      "rawMarkdown": "is your score for the entire k2 or k2 on slices > 900 as your teammate reports [here](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/468826) ?",
      "votes": null
    },
    {
      "id": "2607884",
      "postDate": "01/18/2024 13:49:22",
      "content": "<p>Out of curiosity, have you tried training using different seeds and checking the LB vs CV correlation (for base models and averaged over seeds)?</p>",
      "rawMarkdown": "Out of curiosity, have you tried training using different seeds and checking the LB vs CV correlation (for base models and averaged over seeds)?",
      "votes": null
    },
    {
      "id": "2607904",
      "postDate": "01/18/2024 13:58:17",
      "content": "<p>Here are the results of one experiment I just did (k2 here is the original kidney with all slices), LB threshold is 0.5:</p>\n<table>\n<thead>\n<tr>\n<th>Train</th>\n<th>Valid</th>\n<th>k1_dense</th>\n<th>k2</th>\n<th>k3_dense</th>\n<th>LB (Th 0.5)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>k1_dense, k3_dense</td>\n<td>k2</td>\n<td></td>\n<td>Thresh 0.10: 0.7446<br>Thresh 0.20: 0.7780<br>Thresh 0.30: 0.7909<br>Thresh 0.35: 0.7941<br>Thresh 0.45: 0.7968<br>Thresh 0.50: 0.7970<br>Thresh 0.55: 0.7963<br>Thresh 0.60: 0.7940<br>Thresh 0.70: 0.7839</td>\n<td></td>\n<td>0.799</td>\n</tr>\n<tr>\n<td>k2</td>\n<td>k1_dense, k3_dense</td>\n<td>Thresh 0.10: 0.6312<br>Thresh 0.20: 0.7462<br>Thresh 0.30: 0.8034<br>Thresh 0.35: 0.8215<br>Thresh 0.45: 0.8428<br>Thresh 0.50: 0.8477<br>Thresh 0.55: 0.8487<br>Thresh 0.60: 0.8459<br>Thresh 0.70: 0.8264</td>\n<td></td>\n<td>Thresh 0.10: 0.8047<br>Thresh 0.20: 0.8712<br>Thresh 0.30: 0.8991<br>Thresh 0.35: 0.9064<br>Thresh 0.45: 0.9122<br>Thresh 0.50: 0.9114<br>Thresh 0.55: 0.9084<br>Thresh 0.60: 0.9030<br>Thresh 0.70: 0.8835</td>\n<td>0.724</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "Here are the results of one experiment I just did (k2 here is the original kidney with all slices), LB threshold is 0.5:\n\n| Train         | Valid         | k1_dense |    k2      | k3_dense | LB (Th 0.5)     |\n|-----------|----------|--------------|----------|-------------|------|\n| k1_dense, k3_dense    |  k2  |          |   Thresh 0.10: 0.7446<br>Thresh 0.20: 0.7780<br>Thresh 0.30: 0.7909<br>Thresh 0.35: 0.7941<br>Thresh 0.45: 0.7968<br>Thresh 0.50: 0.7970<br>Thresh 0.55: 0.7963<br>Thresh 0.60: 0.7940<br>Thresh 0.70: 0.7839 |        |   0.799 |\n| k2    |  k1_dense, k3_dense  |  Thresh 0.10: 0.6312<br>Thresh 0.20: 0.7462<br>Thresh 0.30: 0.8034<br>Thresh 0.35: 0.8215<br>Thresh 0.45: 0.8428<br>Thresh 0.50: 0.8477<br>Thresh 0.55: 0.8487<br>Thresh 0.60: 0.8459<br>Thresh 0.70: 0.8264        |    |  Thresh 0.10: 0.8047<br>Thresh 0.20: 0.8712<br>Thresh 0.30: 0.8991<br>Thresh 0.35: 0.9064<br>Thresh 0.45: 0.9122<br>Thresh 0.50: 0.9114<br>Thresh 0.55: 0.9084<br>Thresh 0.60: 0.9030<br>Thresh 0.70: 0.8835      |   0.724 |",
      "votes": null
    },
    {
      "id": "2608274",
      "postDate": "01/18/2024 16:57:38",
      "content": "<p>Still same experiment, threshold adapted to CV (0.2)</p>\n<table>\n<thead>\n<tr>\n<th>Train</th>\n<th>Valid</th>\n<th>k1_dense</th>\n<th>k2</th>\n<th>k3_dense</th>\n<th>LB (Th 0.2)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>k1_dense</td>\n<td>k2, k3_dense</td>\n<td></td>\n<td>Thresh 0.10: 0.7884<br>Thresh 0.20: 0.8002<br>Thresh 0.30: 0.8008<br>Thresh 0.35: 0.7991<br>Thresh 0.45: 0.7930<br>Thresh 0.50: 0.7890<br>Thresh 0.55: 0.7842<br>Thresh 0.60: 0.7785<br>Thresh 0.70: 0.7629</td>\n<td>Thresh 0.10: 0.9005<br>Thresh 0.20: 0.9117<br>Thresh 0.30: 0.9089<br>Thresh 0.35: 0.9052<br>Thresh 0.45: 0.8945<br>Thresh 0.50: 0.8873<br>Thresh 0.55: 0.8785<br>Thresh 0.60: 0.8682<br>Thresh 0.70: 0.8406</td>\n<td>0.838</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "Still same experiment, threshold adapted to CV (0.2)\n\n| Train         | Valid         | k1_dense |    k2      | k3_dense | LB (Th 0.2)     |\n|-----------|----------|--------------|----------|-------------|------|\n| k1_dense   |  k2, k3_dense   |          |   Thresh 0.10: 0.7884<br>Thresh 0.20: 0.8002<br>Thresh 0.30: 0.8008<br>Thresh 0.35: 0.7991<br>Thresh 0.45: 0.7930<br>Thresh 0.50: 0.7890<br>Thresh 0.55: 0.7842<br>Thresh 0.60: 0.7785<br>Thresh 0.70: 0.7629  | Thresh 0.10: 0.9005<br>Thresh 0.20: 0.9117<br>Thresh 0.30: 0.9089<br>Thresh 0.35: 0.9052<br>Thresh 0.45: 0.8945<br>Thresh 0.50: 0.8873<br>Thresh 0.55: 0.8785<br>Thresh 0.60: 0.8682<br>Thresh 0.70: 0.8406       |   0.838 |",
      "votes": null
    },
    {
      "id": "2608617",
      "postDate": "01/19/2024 00:39:50",
      "content": "<p>Thanks for sharing! in my experiments, k1_dense, k3_dense train together is slightly higher in LB than only use k1 dense.<br>\nOnly use kidney2 train, LB is pretty low also around 0.72, similar with yours</p>",
      "rawMarkdown": "Thanks for sharing! in my experiments, k1_dense, k3_dense train together is slightly higher in LB than only use k1 dense.\nOnly use kidney2 train, LB is pretty low also around 0.72, similar with yours",
      "votes": null
    },
    {
      "id": "2608618",
      "postDate": "01/19/2024 00:40:41",
      "content": "<p>I tried different seeds and also not found better correlation, have you found?</p>",
      "rawMarkdown": "I tried different seeds and also not found better correlation, have you found?",
      "votes": null
    },
    {
      "id": "2608640",
      "postDate": "01/19/2024 01:26:38",
      "content": "<p>Train a simple NN classifier online,  class1:  kidney1, public; class2:  kidney2; class3:  kidney3;   and then we can prob out the private?  but this is maybe classified by intensity, actually we need the metric domain.</p>",
      "rawMarkdown": "Train a simple NN classifier online,  class1:  kidney1, public; class2:  kidney2; class3:  kidney3;   and then we can prob out the private?  but this is maybe classified by intensity, actually we need the metric domain.",
      "votes": null
    },
    {
      "id": "2608646",
      "postDate": "01/19/2024 01:37:34",
      "content": "<p>Useful links:</p>\n<ul>\n<li><a href=\"https://towardsdatascience.com/modality-tests-and-kernel-density-estimations-3f349bb9e595\" target=\"_blank\">https://towardsdatascience.com/modality-tests-and-kernel-density-estimations-3f349bb9e595</a></li>\n<li><a href=\"https://towardsdatascience.com/how-to-find-the-best-theoretical-distribution-for-your-data-a26e5673b4bd\" target=\"_blank\">https://towardsdatascience.com/how-to-find-the-best-theoretical-distribution-for-your-data-a26e5673b4bd</a></li>\n</ul>",
      "rawMarkdown": "Useful links:\n- https://towardsdatascience.com/modality-tests-and-kernel-density-estimations-3f349bb9e595\n- https://towardsdatascience.com/how-to-find-the-best-theoretical-distribution-for-your-data-a26e5673b4bd",
      "votes": null
    },
    {
      "id": "2609098",
      "postDate": "01/19/2024 08:45:44",
      "content": "<p>May I inquire whether the metrics 'hit' and 'fp' are calculated based on a single axis or three axes?</p>",
      "rawMarkdown": "May I inquire whether the metrics 'hit' and 'fp' are calculated based on a single axis or three axes?",
      "votes": null
    },
    {
      "id": "2609230",
      "postDate": "01/19/2024 10:25:23",
      "content": "<p>Not really. I need to do further experiments but I am lacking time. 😬</p>",
      "rawMarkdown": "Not really. I need to do further experiments but I am lacking time. 😬",
      "votes": null
    },
    {
      "id": "2609359",
      "postDate": "01/19/2024 12:19:37",
      "content": "<p>my scores on kidney 2 seem much lower than yours… Do you report scores for the lower part of kindey 2 or the entire kidney 2? (on lower part of kidney 2 I have scores around 0.84-0.85)</p>",
      "rawMarkdown": "my scores on kidney 2 seem much lower than yours... Do you report scores for the lower part of kindey 2 or the entire kidney 2? (on lower part of kidney 2 I have scores around 0.84-0.85)",
      "votes": null
    },
    {
      "id": "2610212",
      "postDate": "01/20/2024 02:02:24",
      "content": "<p>three axes</p>",
      "rawMarkdown": "three axes",
      "votes": null
    },
    {
      "id": "2610213",
      "postDate": "01/20/2024 02:04:15",
      "content": "<p>the result is on lower part of kidney 2, after 900 slice. my score is not stable range from 0.82--0.89</p>",
      "rawMarkdown": "the result is on lower part of kidney 2, after 900 slice. my score is not stable range from 0.82--0.89",
      "votes": null
    },
    {
      "id": "2610228",
      "postDate": "01/20/2024 02:53:45",
      "content": "<p>if you perform TTA, you can easily go beyond 0.9+ for validation on kidney1,2(lower) or 3 !!</p>",
      "rawMarkdown": "if you perform TTA, you can easily go beyond 0.9+ for validation on kidney1,2(lower) or 3 !!",
      "votes": null
    },
    {
      "id": "2612187",
      "postDate": "01/21/2024 08:49:51",
      "content": "<p><a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> - have you considered to use kidney 2 like for fine tuning on one of your kidney 1 or 1 and 3 models?  <br>\nwas considering doing that on the lower part of kidney 2 but not sure what to use for validation unless split part of it off, or use the rest but not sure how good it will be, if it is data leakage. etc.  may have a go, see what happens…</p>",
      "rawMarkdown": "optimo - have you considered to use kidney 2 like for fine tuning on one of your kidney 1 or 1 and 3 models?  \nwas considering doing that on the lower part of kidney 2 but not sure what to use for validation unless split part of it off, or use the rest but not sure how good it will be, if it is data leakage. etc.  may have a go, see what happens...",
      "votes": null
    },
    {
      "id": "2616346",
      "postDate": "01/23/2024 15:32:30",
      "content": "<p>is your score for the entire k2 or k2 on slices &gt;= 900 ?</p>",
      "rawMarkdown": "is your score for the entire k2 or k2 on slices >= 900 ?",
      "votes": null
    },
    {
      "id": "2616820",
      "postDate": "01/23/2024 21:14:21",
      "content": "<p>i think we may made a fundamental mistake.</p>\n<p>When the validation data is scare, you should use validation data + their augmentation for CV.<br>\nyou cannot validate with only \"one dataset\"<br>\nThis will have a more robust CV.</p>\n<p>Note: this require you to understand what is the different image characteristics of your test set (e.g. size, resolution, intensity ….)</p>",
      "rawMarkdown": "i think we may made a fundamental mistake.\n\nWhen the validation data is scare, you should use validation data + their augmentation for CV.\nyou cannot validate with only \"one dataset\"\nThis will have a more robust CV.\n\nNote: this require you to understand what is the different image characteristics of your test set (e.g. size, resolution, intensity ....)",
      "votes": null
    },
    {
      "id": "2629724",
      "postDate": "02/01/2024 01:47:15",
      "content": "<p>Here's what I've seen, in case you find it interesting.</p>\n<p>Metrics are shown as dice (one dimension, k2), surface dice (one dimension, k2), surface dice lb (3 dimensional TTA with 3,2,1 votes respectively, threshold for a vote is 0.5). </p>\n<p>k1 train only: 0.82, 0.68, (0.77, 0.77, 0.65) 3x TTA w/threshold of 3,2,1 votes.</p>\n<p>k1 &amp; k3 train together: 0.91, 0.86, (0.64, 0.75, 0.79) 3x TTA w/threshold of 3,2,1 votes.</p>\n<p>I feel like I'm missing something in this competition…trying different ways to preprocess, more augmentations, and a little bit with loss functions, but I really struggle to get lb scores better than these.</p>",
      "rawMarkdown": "Here's what I've seen, in case you find it interesting.\n\nMetrics are shown as dice (one dimension, k2), surface dice (one dimension, k2), surface dice lb (3 dimensional TTA with 3,2,1 votes respectively, threshold for a vote is 0.5). \n\nk1 train only: 0.82, 0.68, (0.77, 0.77, 0.65) 3x TTA w/threshold of 3,2,1 votes.\n\nk1 & k3 train together: 0.91, 0.86, (0.64, 0.75, 0.79) 3x TTA w/threshold of 3,2,1 votes.\n\nI feel like I'm missing something in this competition...trying different ways to preprocess, more augmentations, and a little bit with loss functions, but I really struggle to get lb scores better than these.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2605446,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "01/17/2024 04:38:21",
      "content": "<p>the possible problem is the public/private dataset distribution? if they are same type of data.  <br>\neither models trained on all data would be little bit lower on Public while stable on Private,<br>\nor models with 2 stages(classification and prediction) would be sustainable to shake up.</p>\n<p>However, without local computation power, I am lazy to try out my simple ideas.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2606272,
      "author_name": "optimo",
      "author_url": "",
      "post_date": "01/17/2024 14:29:25",
      "content": "<p>Can you please clarify what you mean by k2 and k2 sparse? So that I can share comparable results, thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2606278,
          "author_name": "sacuscreed",
          "author_url": "",
          "post_date": "01/17/2024 14:36:11",
          "content": "<p>May be k2 is semipseudolabeled</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2607125,
          "author_name": "lihaoweicvch",
          "author_url": "",
          "post_date": "01/18/2024 04:07:37",
          "content": "<p>I mean use kidney2 to train or kidney3 sparse to train, what’s the difference between using kidney1 dense in LB</p>",
          "votes": null,
          "replies": [
            {
              "id": 2607654,
              "author_name": "optimo",
              "author_url": "",
              "post_date": "01/18/2024 11:00:59",
              "content": "<p>is your score for the entire k2 or k2 on slices &gt; 900 as your teammate reports <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/468826\" target=\"_blank\">here</a> ?</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2616346,
              "author_name": "zz1zzi",
              "author_url": "",
              "post_date": "01/23/2024 15:32:30",
              "content": "<p>is your score for the entire k2 or k2 on slices &gt;= 900 ?</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2607884,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "01/18/2024 13:49:22",
      "content": "<p>Out of curiosity, have you tried training using different seeds and checking the LB vs CV correlation (for base models and averaged over seeds)?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2608618,
          "author_name": "lihaoweicvch",
          "author_url": "",
          "post_date": "01/19/2024 00:40:41",
          "content": "<p>I tried different seeds and also not found better correlation, have you found?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2609230,
              "author_name": "yassinealouini",
              "author_url": "",
              "post_date": "01/19/2024 10:25:23",
              "content": "<p>Not really. I need to do further experiments but I am lacking time. 😬</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2607904,
      "author_name": "optimo",
      "author_url": "",
      "post_date": "01/18/2024 13:58:17",
      "content": "<p>Here are the results of one experiment I just did (k2 here is the original kidney with all slices), LB threshold is 0.5:</p>\n<table>\n<thead>\n<tr>\n<th>Train</th>\n<th>Valid</th>\n<th>k1_dense</th>\n<th>k2</th>\n<th>k3_dense</th>\n<th>LB (Th 0.5)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>k1_dense, k3_dense</td>\n<td>k2</td>\n<td></td>\n<td>Thresh 0.10: 0.7446<br>Thresh 0.20: 0.7780<br>Thresh 0.30: 0.7909<br>Thresh 0.35: 0.7941<br>Thresh 0.45: 0.7968<br>Thresh 0.50: 0.7970<br>Thresh 0.55: 0.7963<br>Thresh 0.60: 0.7940<br>Thresh 0.70: 0.7839</td>\n<td></td>\n<td>0.799</td>\n</tr>\n<tr>\n<td>k2</td>\n<td>k1_dense, k3_dense</td>\n<td>Thresh 0.10: 0.6312<br>Thresh 0.20: 0.7462<br>Thresh 0.30: 0.8034<br>Thresh 0.35: 0.8215<br>Thresh 0.45: 0.8428<br>Thresh 0.50: 0.8477<br>Thresh 0.55: 0.8487<br>Thresh 0.60: 0.8459<br>Thresh 0.70: 0.8264</td>\n<td></td>\n<td>Thresh 0.10: 0.8047<br>Thresh 0.20: 0.8712<br>Thresh 0.30: 0.8991<br>Thresh 0.35: 0.9064<br>Thresh 0.45: 0.9122<br>Thresh 0.50: 0.9114<br>Thresh 0.55: 0.9084<br>Thresh 0.60: 0.9030<br>Thresh 0.70: 0.8835</td>\n<td>0.724</td>\n</tr>\n</tbody>\n</table>",
      "votes": null,
      "replies": [
        {
          "id": 2608274,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "01/18/2024 16:57:38",
          "content": "<p>Still same experiment, threshold adapted to CV (0.2)</p>\n<table>\n<thead>\n<tr>\n<th>Train</th>\n<th>Valid</th>\n<th>k1_dense</th>\n<th>k2</th>\n<th>k3_dense</th>\n<th>LB (Th 0.2)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>k1_dense</td>\n<td>k2, k3_dense</td>\n<td></td>\n<td>Thresh 0.10: 0.7884<br>Thresh 0.20: 0.8002<br>Thresh 0.30: 0.8008<br>Thresh 0.35: 0.7991<br>Thresh 0.45: 0.7930<br>Thresh 0.50: 0.7890<br>Thresh 0.55: 0.7842<br>Thresh 0.60: 0.7785<br>Thresh 0.70: 0.7629</td>\n<td>Thresh 0.10: 0.9005<br>Thresh 0.20: 0.9117<br>Thresh 0.30: 0.9089<br>Thresh 0.35: 0.9052<br>Thresh 0.45: 0.8945<br>Thresh 0.50: 0.8873<br>Thresh 0.55: 0.8785<br>Thresh 0.60: 0.8682<br>Thresh 0.70: 0.8406</td>\n<td>0.838</td>\n</tr>\n</tbody>\n</table>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2608617,
          "author_name": "lihaoweicvch",
          "author_url": "",
          "post_date": "01/19/2024 00:39:50",
          "content": "<p>Thanks for sharing! in my experiments, k1_dense, k3_dense train together is slightly higher in LB than only use k1 dense.<br>\nOnly use kidney2 train, LB is pretty low also around 0.72, similar with yours</p>",
          "votes": null,
          "replies": [
            {
              "id": 2609359,
              "author_name": "optimo",
              "author_url": "",
              "post_date": "01/19/2024 12:19:37",
              "content": "<p>my scores on kidney 2 seem much lower than yours… Do you report scores for the lower part of kindey 2 or the entire kidney 2? (on lower part of kidney 2 I have scores around 0.84-0.85)</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2610213,
                  "author_name": "lihaoweicvch",
                  "author_url": "",
                  "post_date": "01/20/2024 02:04:15",
                  "content": "<p>the result is on lower part of kidney 2, after 900 slice. my score is not stable range from 0.82--0.89</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2610228,
                      "author_name": "hengck23",
                      "author_url": "",
                      "post_date": "01/20/2024 02:53:45",
                      "content": "<p>if you perform TTA, you can easily go beyond 0.9+ for validation on kidney1,2(lower) or 3 !!</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        },
        {
          "id": 2612187,
          "author_name": "something4kag",
          "author_url": "",
          "post_date": "01/21/2024 08:49:51",
          "content": "<p><a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> - have you considered to use kidney 2 like for fine tuning on one of your kidney 1 or 1 and 3 models?  <br>\nwas considering doing that on the lower part of kidney 2 but not sure what to use for validation unless split part of it off, or use the rest but not sure how good it will be, if it is data leakage. etc.  may have a go, see what happens…</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2608640,
      "author_name": "lihaoweicvch",
      "author_url": "",
      "post_date": "01/19/2024 01:26:38",
      "content": "<p>Train a simple NN classifier online,  class1:  kidney1, public; class2:  kidney2; class3:  kidney3;   and then we can prob out the private?  but this is maybe classified by intensity, actually we need the metric domain.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2608646,
      "author_name": "lihaoweicvch",
      "author_url": "",
      "post_date": "01/19/2024 01:37:34",
      "content": "<p>Useful links:</p>\n<ul>\n<li><a href=\"https://towardsdatascience.com/modality-tests-and-kernel-density-estimations-3f349bb9e595\" target=\"_blank\">https://towardsdatascience.com/modality-tests-and-kernel-density-estimations-3f349bb9e595</a></li>\n<li><a href=\"https://towardsdatascience.com/how-to-find-the-best-theoretical-distribution-for-your-data-a26e5673b4bd\" target=\"_blank\">https://towardsdatascience.com/how-to-find-the-best-theoretical-distribution-for-your-data-a26e5673b4bd</a></li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2609098,
      "author_name": "zz1zzi",
      "author_url": "",
      "post_date": "01/19/2024 08:45:44",
      "content": "<p>May I inquire whether the metrics 'hit' and 'fp' are calculated based on a single axis or three axes?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2610212,
          "author_name": "lihaoweicvch",
          "author_url": "",
          "post_date": "01/20/2024 02:02:24",
          "content": "<p>three axes</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2616820,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "01/23/2024 21:14:21",
      "content": "<p>i think we may made a fundamental mistake.</p>\n<p>When the validation data is scare, you should use validation data + their augmentation for CV.<br>\nyou cannot validate with only \"one dataset\"<br>\nThis will have a more robust CV.</p>\n<p>Note: this require you to understand what is the different image characteristics of your test set (e.g. size, resolution, intensity ….)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2629724,
      "author_name": "chemdatafarmer",
      "author_url": "",
      "post_date": "02/01/2024 01:47:15",
      "content": "<p>Here's what I've seen, in case you find it interesting.</p>\n<p>Metrics are shown as dice (one dimension, k2), surface dice (one dimension, k2), surface dice lb (3 dimensional TTA with 3,2,1 votes respectively, threshold for a vote is 0.5). </p>\n<p>k1 train only: 0.82, 0.68, (0.77, 0.77, 0.65) 3x TTA w/threshold of 3,2,1 votes.</p>\n<p>k1 &amp; k3 train together: 0.91, 0.86, (0.64, 0.75, 0.79) 3x TTA w/threshold of 3,2,1 votes.</p>\n<p>I feel like I'm missing something in this competition…trying different ways to preprocess, more augmentations, and a little bit with loss functions, but I really struggle to get lb scores better than these.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2605231": "I trained on K1 and val on K3,  or  train K1,K3 val on K2,  both not found any correlation between CV and LB.\nHas anyone use K2 sparse, K3 sparse or other mixture of them  for train, what's the result?\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F856355%2F6ffe902ddc7efeaddfbc65dd31c2c3f7%2F5631705451322_.pic.jpg?generation=1705451397035794&alt=media)",
    "2605446": "the possible problem is the public/private dataset distribution? if they are same type of data.  \neither models trained on all data would be little bit lower on Public while stable on Private,\nor models with 2 stages(classification and prediction) would be sustainable to shake up.\n\nHowever, without local computation power, I am lazy to try out my simple ideas.",
    "2606272": "Can you please clarify what you mean by k2 and k2 sparse? So that I can share comparable results, thanks!",
    "2606278": "May be k2 is semipseudolabeled",
    "2607125": "I mean use kidney2 to train or kidney3 sparse to train, what’s the difference between using kidney1 dense in LB",
    "2607654": "is your score for the entire k2 or k2 on slices > 900 as your teammate reports [here](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/468826) ?",
    "2607884": "Out of curiosity, have you tried training using different seeds and checking the LB vs CV correlation (for base models and averaged over seeds)?",
    "2607904": "Here are the results of one experiment I just did (k2 here is the original kidney with all slices), LB threshold is 0.5:\n\n| Train         | Valid         | k1_dense |    k2      | k3_dense | LB (Th 0.5)     |\n|-----------|----------|--------------|----------|-------------|------|\n| k1_dense, k3_dense    |  k2  |          |   Thresh 0.10: 0.7446<br>Thresh 0.20: 0.7780<br>Thresh 0.30: 0.7909<br>Thresh 0.35: 0.7941<br>Thresh 0.45: 0.7968<br>Thresh 0.50: 0.7970<br>Thresh 0.55: 0.7963<br>Thresh 0.60: 0.7940<br>Thresh 0.70: 0.7839 |        |   0.799 |\n| k2    |  k1_dense, k3_dense  |  Thresh 0.10: 0.6312<br>Thresh 0.20: 0.7462<br>Thresh 0.30: 0.8034<br>Thresh 0.35: 0.8215<br>Thresh 0.45: 0.8428<br>Thresh 0.50: 0.8477<br>Thresh 0.55: 0.8487<br>Thresh 0.60: 0.8459<br>Thresh 0.70: 0.8264        |    |  Thresh 0.10: 0.8047<br>Thresh 0.20: 0.8712<br>Thresh 0.30: 0.8991<br>Thresh 0.35: 0.9064<br>Thresh 0.45: 0.9122<br>Thresh 0.50: 0.9114<br>Thresh 0.55: 0.9084<br>Thresh 0.60: 0.9030<br>Thresh 0.70: 0.8835      |   0.724 |",
    "2608274": "Still same experiment, threshold adapted to CV (0.2)\n\n| Train         | Valid         | k1_dense |    k2      | k3_dense | LB (Th 0.2)     |\n|-----------|----------|--------------|----------|-------------|------|\n| k1_dense   |  k2, k3_dense   |          |   Thresh 0.10: 0.7884<br>Thresh 0.20: 0.8002<br>Thresh 0.30: 0.8008<br>Thresh 0.35: 0.7991<br>Thresh 0.45: 0.7930<br>Thresh 0.50: 0.7890<br>Thresh 0.55: 0.7842<br>Thresh 0.60: 0.7785<br>Thresh 0.70: 0.7629  | Thresh 0.10: 0.9005<br>Thresh 0.20: 0.9117<br>Thresh 0.30: 0.9089<br>Thresh 0.35: 0.9052<br>Thresh 0.45: 0.8945<br>Thresh 0.50: 0.8873<br>Thresh 0.55: 0.8785<br>Thresh 0.60: 0.8682<br>Thresh 0.70: 0.8406       |   0.838 |",
    "2608617": "Thanks for sharing! in my experiments, k1_dense, k3_dense train together is slightly higher in LB than only use k1 dense.\nOnly use kidney2 train, LB is pretty low also around 0.72, similar with yours",
    "2608618": "I tried different seeds and also not found better correlation, have you found?",
    "2608640": "Train a simple NN classifier online,  class1:  kidney1, public; class2:  kidney2; class3:  kidney3;   and then we can prob out the private?  but this is maybe classified by intensity, actually we need the metric domain.",
    "2608646": "Useful links:\n- https://towardsdatascience.com/modality-tests-and-kernel-density-estimations-3f349bb9e595\n- https://towardsdatascience.com/how-to-find-the-best-theoretical-distribution-for-your-data-a26e5673b4bd",
    "2609098": "May I inquire whether the metrics 'hit' and 'fp' are calculated based on a single axis or three axes?",
    "2609230": "Not really. I need to do further experiments but I am lacking time. 😬",
    "2609359": "my scores on kidney 2 seem much lower than yours... Do you report scores for the lower part of kindey 2 or the entire kidney 2? (on lower part of kidney 2 I have scores around 0.84-0.85)",
    "2610212": "three axes",
    "2610213": "the result is on lower part of kidney 2, after 900 slice. my score is not stable range from 0.82--0.89",
    "2610228": "if you perform TTA, you can easily go beyond 0.9+ for validation on kidney1,2(lower) or 3 !!",
    "2612187": "optimo - have you considered to use kidney 2 like for fine tuning on one of your kidney 1 or 1 and 3 models?  \nwas considering doing that on the lower part of kidney 2 but not sure what to use for validation unless split part of it off, or use the rest but not sure how good it will be, if it is data leakage. etc.  may have a go, see what happens...",
    "2616346": "is your score for the entire k2 or k2 on slices >= 900 ?",
    "2616820": "i think we may made a fundamental mistake.\n\nWhen the validation data is scare, you should use validation data + their augmentation for CV.\nyou cannot validate with only \"one dataset\"\nThis will have a more robust CV.\n\nNote: this require you to understand what is the different image characteristics of your test set (e.g. size, resolution, intensity ....)",
    "2629724": "Here's what I've seen, in case you find it interesting.\n\nMetrics are shown as dice (one dimension, k2), surface dice (one dimension, k2), surface dice lb (3 dimensional TTA with 3,2,1 votes respectively, threshold for a vote is 0.5). \n\nk1 train only: 0.82, 0.68, (0.77, 0.77, 0.65) 3x TTA w/threshold of 3,2,1 votes.\n\nk1 & k3 train together: 0.91, 0.86, (0.64, 0.75, 0.79) 3x TTA w/threshold of 3,2,1 votes.\n\nI feel like I'm missing something in this competition...trying different ways to preprocess, more augmentations, and a little bit with loss functions, but I really struggle to get lb scores better than these."
  },
  "source": "meta"
}