{
  "id": 669994,
  "title": "CV0.5806/LB0.529？？？",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/669994",
  "author_name": "",
  "post_date": "2026-01-25T13:03:26.299173800Z",
  "votes": 4,
  "comment_count": 21,
  "views": 0,
  "content": "<p>Why do my CV and LB scores show no correlation? I used the official evaluation metric from \"https://www.kaggle.com/code/sohier/vesuvius-2025-metric-demo/\" and the training code from \"https://www.kaggle.com/code/choudharymanas/inference-baseline-transunet-lb-0-537\". I only changed the loss function and used the updated training data. Why is this happening?<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F21003169%2F125db5c8543f7d38592a466ae627b565%2F2026-01-25%20205953.jpg?generation=1769346196253531&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "3396621",
      "postDate": "01/25/2026 13:03:26",
      "content": "<p>Why do my CV and LB scores show no correlation? I used the official evaluation metric from \"https://www.kaggle.com/code/sohier/vesuvius-2025-metric-demo/\" and the training code from \"https://www.kaggle.com/code/choudharymanas/inference-baseline-transunet-lb-0-537\". I only changed the loss function and used the updated training data. Why is this happening?<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F21003169%2F125db5c8543f7d38592a466ae627b565%2F2026-01-25%20205953.jpg?generation=1769346196253531&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Why do my CV and LB scores show no correlation? I used the official evaluation metric from \"https://www.kaggle.com/code/sohier/vesuvius-2025-metric-demo/\" and the training code from \"https://www.kaggle.com/code/choudharymanas/inference-baseline-transunet-lb-0-537\". I only changed the loss function and used the updated training data. Why is this happening?![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F21003169%2F125db5c8543f7d38592a466ae627b565%2F2026-01-25%20205953.jpg?generation=1769346196253531&alt=media)",
      "votes": null
    },
    {
      "id": "3396623",
      "postDate": "01/25/2026 13:07:17",
      "content": "<p>Could training on the new dataset cause this?</p>",
      "rawMarkdown": "Could training on the new dataset cause this?",
      "votes": null
    },
    {
      "id": "3396656",
      "postDate": "01/25/2026 14:23:49",
      "content": "<p>It’s completely normal for the test data distribution to differ from the training data. A gap between CV and LB scores does not mean there is no correlation. You should run multiple experiments; as long as the CV–LB gap is relatively consistent across runs, it’s not an issue.</p>",
      "rawMarkdown": "It’s completely normal for the test data distribution to differ from the training data. A gap between CV and LB scores does not mean there is no correlation. You should run multiple experiments; as long as the CV–LB gap is relatively consistent across runs, it’s not an issue.",
      "votes": null
    },
    {
      "id": "3396658",
      "postDate": "01/25/2026 14:30:44",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/zejiezhang\" target=\"_blank\">@zejiezhang</a>, you shouldn't be having this problem from new dataset. But I didn't understand your cv result. Your leaving just one sample out?</p>",
      "rawMarkdown": "Hi @zejiezhang, you shouldn't be having this problem from new dataset. But I didn't understand your cv result. Your leaving just one sample out?",
      "votes": null
    },
    {
      "id": "3396662",
      "postDate": "01/25/2026 14:37:58",
      "content": "<p>only 6 validation samples</p>",
      "rawMarkdown": "only 6 validation samples",
      "votes": null
    },
    {
      "id": "3396668",
      "postDate": "01/25/2026 14:51:58",
      "content": "<p>Thanks — it looks like I need to retrain the model and reselect the validation samples, so I have to train again. I've wasted a lot of time.</p>",
      "rawMarkdown": "Thanks — it looks like I need to retrain the model and reselect the validation samples, so I have to train again. I've wasted a lot of time.",
      "votes": null
    },
    {
      "id": "3396670",
      "postDate": "01/25/2026 14:55:43",
      "content": "<p>Thanks, I'll run a few more sets of experiments to verify.</p>",
      "rawMarkdown": "Thanks, I'll run a few more sets of experiments to verify.",
      "votes": null
    },
    {
      "id": "3396675",
      "postDate": "01/25/2026 15:17:01",
      "content": "<p>Then that's why you're not seeing correlation. Try to keep 1/5 of data as validation, should work better than your current approach </p>",
      "rawMarkdown": "Then that's why you're not seeing correlation. Try to keep 1/5 of data as validation, should work better than your current approach",
      "votes": null
    },
    {
      "id": "3396739",
      "postDate": "01/25/2026 18:26:20",
      "content": "<p>I’ve run into this CV vs LB gap quite a few times myself.\nIn my case it usually wasn’t the model, but the validation setup that was too optimistic.</p>\n<p>I recently tried to break this down step by step in a short notebook — mostly to understand <em>when</em> CV can be trusted and when it can’t.\nSharing in case it’s useful to someone: <a href=\"https://www.kaggle.com/code/wsolcor/why-your-cross-validation-is-lying-cv-vs-lb\" target=\"_blank\">https://www.kaggle.com/code/wsolcor/why-your-cross-validation-is-lying-cv-vs-lb</a></p>",
      "rawMarkdown": "I’ve run into this CV vs LB gap quite a few times myself.\nIn my case it usually wasn’t the model, but the validation setup that was too optimistic.\n\nI recently tried to break this down step by step in a short notebook — mostly to understand *when* CV can be trusted and when it can’t.\nSharing in case it’s useful to someone: https://www.kaggle.com/code/wsolcor/why-your-cross-validation-is-lying-cv-vs-lb",
      "votes": null
    },
    {
      "id": "3396902",
      "postDate": "01/26/2026 05:46:27",
      "content": "<p>I am having similar issue,\nmodels doing well on local cv seems to be doing worse on lb. </p>",
      "rawMarkdown": "I am having similar issue,\nmodels doing well on local cv seems to be doing worse on lb.",
      "votes": null
    },
    {
      "id": "3396926",
      "postDate": "01/26/2026 06:17:56",
      "content": "<p>My local CV has reached 0.60+, but my public LB score is stuck at 0.567. I’m sure those ahead of me have even higher scores. What are your local CV scores looking like right now?</p>",
      "rawMarkdown": "My local CV has reached 0.60+, but my public LB score is stuck at 0.567. I’m sure those ahead of me have even higher scores. What are your local CV scores looking like right now?",
      "votes": null
    },
    {
      "id": "3396968",
      "postDate": "01/26/2026 08:01:34",
      "content": "<p>2026-01-26 12:38:12.588088: Mean Validation Dice:  0.6025550871361659\nDICE is also very high TT</p>",
      "rawMarkdown": "2026-01-26 12:38:12.588088: Mean Validation Dice:  0.6025550871361659\nDICE is also very high TT",
      "votes": null
    },
    {
      "id": "3396976",
      "postDate": "01/26/2026 08:19:49",
      "content": "<p>I am just reaching 0.57 on val 😭 best model and 0.545 on lb.\nVery confused as surfaces are cleaner visually compared to my \"best\" submission so far.</p>\n<p>Do you think this is an issue with just the public 20% data or with my cv batch?</p>",
      "rawMarkdown": "I am just reaching 0.57 on val 😭 best model and 0.545 on lb.\nVery confused as surfaces are cleaner visually compared to my \"best\" submission so far.\n\nDo you think this is an issue with just the public 20% data or with my cv batch?",
      "votes": null
    },
    {
      "id": "3397037",
      "postDate": "01/26/2026 10:54:30",
      "content": "<p>I checked our data, and a local CV of 0.57 actually matches up with a Public LB of about 0.54. I think it's just the 20% split causing this—after all, that's probably only &lt;= 40 samples we're looking at.</p>",
      "rawMarkdown": "I checked our data, and a local CV of 0.57 actually matches up with a Public LB of about 0.54. I think it's just the 20% split causing this—after all, that's probably only <= 40 samples we're looking at.",
      "votes": null
    },
    {
      "id": "3397051",
      "postDate": "01/26/2026 11:31:29",
      "content": "<p>thanks!! \nlooks like we will be looking at a significant lb shakeup</p>",
      "rawMarkdown": "thanks!! \nlooks like we will be looking at a significant lb shakeup",
      "votes": null
    },
    {
      "id": "3399248",
      "postDate": "01/30/2026 13:52:49",
      "content": "<p>I am a little late in this competition, and i have a question. How have you calculated the individual scores like toposcore, surfacedice and voi? is it through the same code that is defined <a href=\"https://www.kaggle.com/code/sohier/vesuvius-2025-metric-demo/\" target=\"_blank\">here</a>, or are you using any custom implementation from python libraries like scipy? Actually it is taking a lot of time for me on (160,160,160) dimension size (mora than 1.5 hrs on only 20% of validation data) using the official metric code.</p>",
      "rawMarkdown": "I am a little late in this competition, and i have a question. How have you calculated the individual scores like toposcore, surfacedice and voi? is it through the same code that is defined [here](https://www.kaggle.com/code/sohier/vesuvius-2025-metric-demo/), or are you using any custom implementation from python libraries like scipy? Actually it is taking a lot of time for me on (160,160,160) dimension size (mora than 1.5 hrs on only 20% of validation data) using the official metric code.",
      "votes": null
    },
    {
      "id": "3399821",
      "postDate": "01/31/2026 11:17:56",
      "content": "<p>I am using the Vesuvius 2025 Metric Demo for local verification.</p>",
      "rawMarkdown": "I am using the Vesuvius 2025 Metric Demo for local verification.",
      "votes": null
    },
    {
      "id": "3399840",
      "postDate": "01/31/2026 11:55:34",
      "content": "<p>Thanks for the help. can you tell how much time is it taking on local verification of only 20% data? </p>",
      "rawMarkdown": "Thanks for the help. can you tell how much time is it taking on local verification of only 20% data?",
      "votes": null
    },
    {
      "id": "3399854",
      "postDate": "01/31/2026 12:07:04",
      "content": "<p>Depending on your device, with an I9 processor and 64GB of RAM, it takes about 20 minutes.</p>",
      "rawMarkdown": "Depending on your device, with an I9 processor and 64GB of RAM, it takes about 20 minutes.",
      "votes": null
    },
    {
      "id": "3399867",
      "postDate": "01/31/2026 12:22:06",
      "content": "<p>I have a machine wiht H100 (80gbvram) with 1TB ram and 176 cpu cores, but still the validation on 20% train data is taking 1.5 to 2 hours. Let me refractor my code again. Thanks again for the help.</p>",
      "rawMarkdown": "I have a machine wiht H100 (80gbvram) with 1TB ram and 176 cpu cores, but still the validation on 20% train data is taking 1.5 to 2 hours. Let me refractor my code again. Thanks again for the help.",
      "votes": null
    },
    {
      "id": "3400105",
      "postDate": "02/01/2026 00:08:03",
      "content": "<p><a href=\"https://www.kaggle.com/chengtingyi\" target=\"_blank\">@chengtingyi</a> congrats on the crazy improvement\nlooks like i reached the position you were in with the 0.6+ val TT</p>",
      "rawMarkdown": "chengtingyi congrats on the crazy improvement\nlooks like i reached the position you were in with the 0.6+ val TT",
      "votes": null
    },
    {
      "id": "3400336",
      "postDate": "02/01/2026 11:41:50",
      "content": "<p>Thanks, but not so fast! Slow down a bit, or I'm afraid you're going to surpass me soon.</p>",
      "rawMarkdown": "Thanks, but not so fast! Slow down a bit, or I'm afraid you're going to surpass me soon.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3396623,
      "author_name": "zejiezhang",
      "author_url": "",
      "post_date": "01/25/2026 13:07:17",
      "content": "<p>Could training on the new dataset cause this?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3396658,
          "author_name": "sersasj",
          "author_url": "",
          "post_date": "01/25/2026 14:30:44",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/zejiezhang\" target=\"_blank\">@zejiezhang</a>, you shouldn't be having this problem from new dataset. But I didn't understand your cv result. Your leaving just one sample out?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3396662,
              "author_name": "zejiezhang",
              "author_url": "",
              "post_date": "01/25/2026 14:37:58",
              "content": "<p>only 6 validation samples</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3396668,
                  "author_name": "zejiezhang",
                  "author_url": "",
                  "post_date": "01/25/2026 14:51:58",
                  "content": "<p>Thanks — it looks like I need to retrain the model and reselect the validation samples, so I have to train again. I've wasted a lot of time.</p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 3396675,
                  "author_name": "sersasj",
                  "author_url": "",
                  "post_date": "01/25/2026 15:17:01",
                  "content": "<p>Then that's why you're not seeing correlation. Try to keep 1/5 of data as validation, should work better than your current approach </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3396656,
      "author_name": "wayne127",
      "author_url": "",
      "post_date": "01/25/2026 14:23:49",
      "content": "<p>It’s completely normal for the test data distribution to differ from the training data. A gap between CV and LB scores does not mean there is no correlation. You should run multiple experiments; as long as the CV–LB gap is relatively consistent across runs, it’s not an issue.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3396670,
          "author_name": "zejiezhang",
          "author_url": "",
          "post_date": "01/25/2026 14:55:43",
          "content": "<p>Thanks, I'll run a few more sets of experiments to verify.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3396739,
      "author_name": "wsolcor",
      "author_url": "",
      "post_date": "01/25/2026 18:26:20",
      "content": "<p>I’ve run into this CV vs LB gap quite a few times myself.\nIn my case it usually wasn’t the model, but the validation setup that was too optimistic.</p>\n<p>I recently tried to break this down step by step in a short notebook — mostly to understand <em>when</em> CV can be trusted and when it can’t.\nSharing in case it’s useful to someone: <a href=\"https://www.kaggle.com/code/wsolcor/why-your-cross-validation-is-lying-cv-vs-lb\" target=\"_blank\">https://www.kaggle.com/code/wsolcor/why-your-cross-validation-is-lying-cv-vs-lb</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3396902,
      "author_name": "arjunashokbhandary",
      "author_url": "",
      "post_date": "01/26/2026 05:46:27",
      "content": "<p>I am having similar issue,\nmodels doing well on local cv seems to be doing worse on lb. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3396926,
          "author_name": "chengtingyi",
          "author_url": "",
          "post_date": "01/26/2026 06:17:56",
          "content": "<p>My local CV has reached 0.60+, but my public LB score is stuck at 0.567. I’m sure those ahead of me have even higher scores. What are your local CV scores looking like right now?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3396968,
              "author_name": "ggayoayogg",
              "author_url": "",
              "post_date": "01/26/2026 08:01:34",
              "content": "<p>2026-01-26 12:38:12.588088: Mean Validation Dice:  0.6025550871361659\nDICE is also very high TT</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3396976,
              "author_name": "arjunashokbhandary",
              "author_url": "",
              "post_date": "01/26/2026 08:19:49",
              "content": "<p>I am just reaching 0.57 on val 😭 best model and 0.545 on lb.\nVery confused as surfaces are cleaner visually compared to my \"best\" submission so far.</p>\n<p>Do you think this is an issue with just the public 20% data or with my cv batch?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3397037,
                  "author_name": "chengtingyi",
                  "author_url": "",
                  "post_date": "01/26/2026 10:54:30",
                  "content": "<p>I checked our data, and a local CV of 0.57 actually matches up with a Public LB of about 0.54. I think it's just the 20% split causing this—after all, that's probably only &lt;= 40 samples we're looking at.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3397051,
                      "author_name": "arjunashokbhandary",
                      "author_url": "",
                      "post_date": "01/26/2026 11:31:29",
                      "content": "<p>thanks!! \nlooks like we will be looking at a significant lb shakeup</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3400105,
                          "author_name": "arjunashokbhandary",
                          "author_url": "",
                          "post_date": "02/01/2026 00:08:03",
                          "content": "<p><a href=\"https://www.kaggle.com/chengtingyi\" target=\"_blank\">@chengtingyi</a> congrats on the crazy improvement\nlooks like i reached the position you were in with the 0.6+ val TT</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 3400336,
                              "author_name": "chengtingyi",
                              "author_url": "",
                              "post_date": "02/01/2026 11:41:50",
                              "content": "<p>Thanks, but not so fast! Slow down a bit, or I'm afraid you're going to surpass me soon.</p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3399248,
      "author_name": "muhammadibrahim3093",
      "author_url": "",
      "post_date": "01/30/2026 13:52:49",
      "content": "<p>I am a little late in this competition, and i have a question. How have you calculated the individual scores like toposcore, surfacedice and voi? is it through the same code that is defined <a href=\"https://www.kaggle.com/code/sohier/vesuvius-2025-metric-demo/\" target=\"_blank\">here</a>, or are you using any custom implementation from python libraries like scipy? Actually it is taking a lot of time for me on (160,160,160) dimension size (mora than 1.5 hrs on only 20% of validation data) using the official metric code.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3399821,
          "author_name": "ggayoayogg",
          "author_url": "",
          "post_date": "01/31/2026 11:17:56",
          "content": "<p>I am using the Vesuvius 2025 Metric Demo for local verification.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3399840,
              "author_name": "muhammadibrahim3093",
              "author_url": "",
              "post_date": "01/31/2026 11:55:34",
              "content": "<p>Thanks for the help. can you tell how much time is it taking on local verification of only 20% data? </p>",
              "votes": null,
              "replies": [
                {
                  "id": 3399854,
                  "author_name": "ggayoayogg",
                  "author_url": "",
                  "post_date": "01/31/2026 12:07:04",
                  "content": "<p>Depending on your device, with an I9 processor and 64GB of RAM, it takes about 20 minutes.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3399867,
                      "author_name": "muhammadibrahim3093",
                      "author_url": "",
                      "post_date": "01/31/2026 12:22:06",
                      "content": "<p>I have a machine wiht H100 (80gbvram) with 1TB ram and 176 cpu cores, but still the validation on 20% train data is taking 1.5 to 2 hours. Let me refractor my code again. Thanks again for the help.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3396621": "Why do my CV and LB scores show no correlation? I used the official evaluation metric from \"https://www.kaggle.com/code/sohier/vesuvius-2025-metric-demo/\" and the training code from \"https://www.kaggle.com/code/choudharymanas/inference-baseline-transunet-lb-0-537\". I only changed the loss function and used the updated training data. Why is this happening?![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F21003169%2F125db5c8543f7d38592a466ae627b565%2F2026-01-25%20205953.jpg?generation=1769346196253531&alt=media)",
    "3396623": "Could training on the new dataset cause this?",
    "3396656": "It’s completely normal for the test data distribution to differ from the training data. A gap between CV and LB scores does not mean there is no correlation. You should run multiple experiments; as long as the CV–LB gap is relatively consistent across runs, it’s not an issue.",
    "3396658": "Hi @zejiezhang, you shouldn't be having this problem from new dataset. But I didn't understand your cv result. Your leaving just one sample out?",
    "3396662": "only 6 validation samples",
    "3396668": "Thanks — it looks like I need to retrain the model and reselect the validation samples, so I have to train again. I've wasted a lot of time.",
    "3396670": "Thanks, I'll run a few more sets of experiments to verify.",
    "3396675": "Then that's why you're not seeing correlation. Try to keep 1/5 of data as validation, should work better than your current approach",
    "3396739": "I’ve run into this CV vs LB gap quite a few times myself.\nIn my case it usually wasn’t the model, but the validation setup that was too optimistic.\n\nI recently tried to break this down step by step in a short notebook — mostly to understand *when* CV can be trusted and when it can’t.\nSharing in case it’s useful to someone: https://www.kaggle.com/code/wsolcor/why-your-cross-validation-is-lying-cv-vs-lb",
    "3396902": "I am having similar issue,\nmodels doing well on local cv seems to be doing worse on lb.",
    "3396926": "My local CV has reached 0.60+, but my public LB score is stuck at 0.567. I’m sure those ahead of me have even higher scores. What are your local CV scores looking like right now?",
    "3396968": "2026-01-26 12:38:12.588088: Mean Validation Dice:  0.6025550871361659\nDICE is also very high TT",
    "3396976": "I am just reaching 0.57 on val 😭 best model and 0.545 on lb.\nVery confused as surfaces are cleaner visually compared to my \"best\" submission so far.\n\nDo you think this is an issue with just the public 20% data or with my cv batch?",
    "3397037": "I checked our data, and a local CV of 0.57 actually matches up with a Public LB of about 0.54. I think it's just the 20% split causing this—after all, that's probably only <= 40 samples we're looking at.",
    "3397051": "thanks!! \nlooks like we will be looking at a significant lb shakeup",
    "3399248": "I am a little late in this competition, and i have a question. How have you calculated the individual scores like toposcore, surfacedice and voi? is it through the same code that is defined [here](https://www.kaggle.com/code/sohier/vesuvius-2025-metric-demo/), or are you using any custom implementation from python libraries like scipy? Actually it is taking a lot of time for me on (160,160,160) dimension size (mora than 1.5 hrs on only 20% of validation data) using the official metric code.",
    "3399821": "I am using the Vesuvius 2025 Metric Demo for local verification.",
    "3399840": "Thanks for the help. can you tell how much time is it taking on local verification of only 20% data?",
    "3399854": "Depending on your device, with an I9 processor and 64GB of RAM, it takes about 20 minutes.",
    "3399867": "I have a machine wiht H100 (80gbvram) with 1TB ram and 176 cpu cores, but still the validation on 20% train data is taking 1.5 to 2 hours. Let me refractor my code again. Thanks again for the help.",
    "3400105": "chengtingyi congrats on the crazy improvement\nlooks like i reached the position you were in with the 0.6+ val TT",
    "3400336": "Thanks, but not so fast! Slow down a bit, or I'm afraid you're going to surpass me soon."
  },
  "source": "meta"
}