{
  "id": 679262,
  "title": "Don't trust LB too much.  😭",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/679262",
  "author_name": "Zejun_",
  "post_date": "2026-02-28T08:45:06.528000",
  "votes": 4,
  "comment_count": 12,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F26336666%2Ff3a1c79f19d87f079191157f0f0253d8%2F33035207f48c3bb94543dacf67a22290.png?generation=1772268213211384&amp;alt=media\" alt=\"\">\nI gonna say I‘m very lucky for \"shake up\". But how can I choose the right one to submit? hahah🙃</p>",
  "messages": [
    {
      "id": 3415142,
      "postDate": "2026-02-28T08:45:06.530Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F26336666%2Ff3a1c79f19d87f079191157f0f0253d8%2F33035207f48c3bb94543dacf67a22290.png?generation=1772268213211384&amp;alt=media\" alt=\"\">\nI gonna say I‘m very lucky for \"shake up\". But how can I choose the right one to submit? hahah🙃</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F26336666%2Ff3a1c79f19d87f079191157f0f0253d8%2F33035207f48c3bb94543dacf67a22290.png?generation=1772268213211384&alt=media)\nI gonna say I‘m very lucky for \"shake up\". But how can I choose the right one to submit? hahah🙃",
      "votes": 4
    },
    {
      "id": 3415418,
      "postDate": "2026-02-28T21:57:54.970Z",
      "content": "<p>I agree. The public LB is 24 volumes. I observed locally that there is large variation in the competition metric using only 24 volumes. Therefore it is important in this competition to compute validation score using more than 24 volumes. When I used 130 volumes locally, my local CV score was the same as my private LB score (for all my submissions both selected and not selected). Note: private LB is 96 volumes.</p>",
      "rawMarkdown": "I agree. The public LB is 24 volumes. I observed locally that there is large variation in the competition metric using only 24 volumes. Therefore it is important in this competition to compute validation score using more than 24 volumes. When I used 130 volumes locally, my local CV score was the same as my private LB score (for all my submissions both selected and not selected). Note: private LB is 96 volumes.",
      "replies": [
        {
          "id": 3415424,
          "postDate": "2026-02-28T22:13:47.543Z",
          "content": "<p>I simulate some subet with multiple trials before deadline, actually lb seems just hit the rare subset.</p>",
          "rawMarkdown": "I simulate some subet with multiple trials before deadline, actually lb seems just hit the rare subset.",
          "votes": 1,
          "replies": [
            {
              "id": 3415427,
              "postDate": "2026-02-28T22:30:57.597Z",
              "content": "<p>This sounds interesting. Can you explain more, what does \"just hit the rare subset\" mean?</p>",
              "rawMarkdown": "This sounds interesting. Can you explain more, what does \"just hit the rare subset\" mean?"
            },
            {
              "id": 3415596,
              "postDate": "2026-03-01T03:03:37.253Z",
              "content": "<p>And it’s an especially unpredictable case.</p>",
              "rawMarkdown": "And it’s an especially unpredictable case."
            },
            {
              "id": 3415610,
              "postDate": "2026-03-01T03:31:51.897Z",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> The results show that, in most cases, the approach performs better (i.e., falls below the zero line in the plot), with only a few subsets showing worse scores.\nHowever, we observed the largest drop on the leaderboard so far, with the score decreasing from 0.595 to 0.588. The win rate of the 0.595 approach is only 22%, and it did not pass the hypothesis test for superiority. So at that time I confirm lb is rare cases I highlight in this plot.</p>",
              "rawMarkdown": "@cdeotte The results show that, in most cases, the approach performs better (i.e., falls below the zero line in the plot), with only a few subsets showing worse scores.\nHowever, we observed the largest drop on the leaderboard so far, with the score decreasing from 0.595 to 0.588. The win rate of the 0.595 approach is only 22%, and it did not pass the hypothesis test for superiority. So at that time I confirm lb is rare cases I highlight in this plot.",
              "votes": 1
            },
            {
              "id": 3415616,
              "postDate": "2026-03-01T03:44:40.663Z",
              "content": "<p>And seems public subset rewards more to the model good at predicting good sdice results.</p>",
              "rawMarkdown": "And seems public subset rewards more to the model good at predicting good sdice results.",
              "votes": 1
            },
            {
              "id": 3415624,
              "postDate": "2026-03-01T03:58:10.230Z",
              "content": "<p>The test set does have more samples which could be considered \"difficult\" than the train set, at least by percentage. I wanted these to be more in-line with eachother, but label creation on these is so, so time consuming. We wanted to ensure our test set was not too easy, and one which contained scroll data not publicly available,  because as i'm sure as you and others have noticed the models can grab the first 55 or so on the LB score with relative ease, and in the same fashion our unwrapping algorithms can get a large majority of the easy cases unwrapped well, so we wanted to focus more on these more difficult cases. </p>\n<p>if label creation was not so time consuming, i would have loved to have more of these cases in the train set. the split is reasonable as-is but i would have preferred many more so they represented a larger portion of train. </p>",
              "rawMarkdown": "The test set does have more samples which could be considered \"difficult\" than the train set, at least by percentage. I wanted these to be more in-line with eachother, but label creation on these is so, so time consuming. We wanted to ensure our test set was not too easy, and one which contained scroll data not publicly available,  because as i'm sure as you and others have noticed the models can grab the first 55 or so on the LB score with relative ease, and in the same fashion our unwrapping algorithms can get a large majority of the easy cases unwrapped well, so we wanted to focus more on these more difficult cases. \n\nif label creation was not so time consuming, i would have loved to have more of these cases in the train set. the split is reasonable as-is but i would have preferred many more so they represented a larger portion of train. ",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3415415,
      "postDate": "2026-02-28T21:55:34.750Z",
      "content": "<p>the reason is that the 3 metrics become contradictory when there is label noise with bg annotated as fg voxels. Then hole filling improves topo, but kills surface dice and voi. At such, any post-processing cannot be simply \"ad-hoc\". </p>\n<p>\". But how can I choose the right one to submit? \" when the uncertainty (aka shakeup) is large, one way to combat is to make many submissions and choose the \"median one\". now that post submission is opened, you can make 100 submissions per day to experiment and see the trend.</p>",
      "rawMarkdown": "the reason is that the 3 metrics become contradictory when there is label noise with bg annotated as fg voxels. Then hole filling improves topo, but kills surface dice and voi. At such, any post-processing cannot be simply \"ad-hoc\". \n\n\". But how can I choose the right one to submit? \" when the uncertainty (aka shakeup) is large, one way to combat is to make many submissions and choose the \"median one\". now that post submission is opened, you can make 100 submissions per day to experiment and see the trend."
    },
    {
      "id": 3415235,
      "postDate": "2026-02-28T13:44:30.430Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15255201%2F5278b540e914a10f93f718d69a2e74cb%2FScreenshot%202026-02-28%20213933.png?generation=1772286381948854&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15255201%2F5278b540e914a10f93f718d69a2e74cb%2FScreenshot%202026-02-28%20213933.png?generation=1772286381948854&alt=media)",
      "replies": [
        {
          "id": 3415269,
          "postDate": "2026-02-28T15:07:32.453Z",
          "content": "<p>that is very unfortunate \nthe difference is huge as well </p>",
          "rawMarkdown": "that is very unfortunate \nthe difference is huge as well "
        },
        {
          "id": 3415273,
          "postDate": "2026-02-28T15:12:20.033Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 3415844,
          "postDate": "2026-03-01T12:34:32.563Z",
          "content": "<p>Wow! ALLIN is the best choice for any gambler.</p>",
          "rawMarkdown": "Wow! ALLIN is the best choice for any gambler.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3415418,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2026-02-28T21:57:54.970000",
      "content": "<p>I agree. The public LB is 24 volumes. I observed locally that there is large variation in the competition metric using only 24 volumes. Therefore it is important in this competition to compute validation score using more than 24 volumes. When I used 130 volumes locally, my local CV score was the same as my private LB score (for all my submissions both selected and not selected). Note: private LB is 96 volumes.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3415424,
          "author_name": "Tom",
          "author_url": "",
          "post_date": "2026-02-28T22:13:47.543000",
          "content": "<p>I simulate some subet with multiple trials before deadline, actually lb seems just hit the rare subset.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3415427,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2026-02-28T22:30:57.597000",
              "content": "<p>This sounds interesting. Can you explain more, what does \"just hit the rare subset\" mean?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3415596,
              "author_name": "GG Ayo (AyoGG)",
              "author_url": "",
              "post_date": "2026-03-01T03:03:37.253000",
              "content": "<p>And it’s an especially unpredictable case.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3415610,
              "author_name": "Tom",
              "author_url": "",
              "post_date": "2026-03-01T03:31:51.897000",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> The results show that, in most cases, the approach performs better (i.e., falls below the zero line in the plot), with only a few subsets showing worse scores.\nHowever, we observed the largest drop on the leaderboard so far, with the score decreasing from 0.595 to 0.588. The win rate of the 0.595 approach is only 22%, and it did not pass the hypothesis test for superiority. So at that time I confirm lb is rare cases I highlight in this plot.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3415616,
              "author_name": "Tom",
              "author_url": "",
              "post_date": "2026-03-01T03:44:40.663000",
              "content": "<p>And seems public subset rewards more to the model good at predicting good sdice results.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3415624,
              "author_name": "Sean Johnson_SP",
              "author_url": "",
              "post_date": "2026-03-01T03:58:10.230000",
              "content": "<p>The test set does have more samples which could be considered \"difficult\" than the train set, at least by percentage. I wanted these to be more in-line with eachother, but label creation on these is so, so time consuming. We wanted to ensure our test set was not too easy, and one which contained scroll data not publicly available,  because as i'm sure as you and others have noticed the models can grab the first 55 or so on the LB score with relative ease, and in the same fashion our unwrapping algorithms can get a large majority of the easy cases unwrapped well, so we wanted to focus more on these more difficult cases. </p>\n<p>if label creation was not so time consuming, i would have loved to have more of these cases in the train set. the split is reasonable as-is but i would have preferred many more so they represented a larger portion of train. </p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3415415,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2026-02-28T21:55:34.750000",
      "content": "<p>the reason is that the 3 metrics become contradictory when there is label noise with bg annotated as fg voxels. Then hole filling improves topo, but kills surface dice and voi. At such, any post-processing cannot be simply \"ad-hoc\". </p>\n<p>\". But how can I choose the right one to submit? \" when the uncertainty (aka shakeup) is large, one way to combat is to make many submissions and choose the \"median one\". now that post submission is opened, you can make 100 submissions per day to experiment and see the trend.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3415235,
      "author_name": "GG Ayo (AyoGG)",
      "author_url": "",
      "post_date": "2026-02-28T13:44:30.430000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15255201%2F5278b540e914a10f93f718d69a2e74cb%2FScreenshot%202026-02-28%20213933.png?generation=1772286381948854&amp;alt=media\" alt=\"\"></p>",
      "votes": 0,
      "replies": [
        {
          "id": 3415269,
          "author_name": "ArjunB",
          "author_url": "",
          "post_date": "2026-02-28T15:07:32.453000",
          "content": "<p>that is very unfortunate \nthe difference is huge as well </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3415273,
          "author_name": "",
          "author_url": "",
          "post_date": "2026-02-28T15:12:20.033000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3415844,
          "author_name": "Zejun_",
          "author_url": "",
          "post_date": "2026-03-01T12:34:32.563000",
          "content": "<p>Wow! ALLIN is the best choice for any gambler.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3415142": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F26336666%2Ff3a1c79f19d87f079191157f0f0253d8%2F33035207f48c3bb94543dacf67a22290.png?generation=1772268213211384&alt=media)\nI gonna say I‘m very lucky for \"shake up\". But how can I choose the right one to submit? hahah🙃",
    "3415418": "I agree. The public LB is 24 volumes. I observed locally that there is large variation in the competition metric using only 24 volumes. Therefore it is important in this competition to compute validation score using more than 24 volumes. When I used 130 volumes locally, my local CV score was the same as my private LB score (for all my submissions both selected and not selected). Note: private LB is 96 volumes.",
    "3415415": "the reason is that the 3 metrics become contradictory when there is label noise with bg annotated as fg voxels. Then hole filling improves topo, but kills surface dice and voi. At such, any post-processing cannot be simply \"ad-hoc\". \n\n\". But how can I choose the right one to submit? \" when the uncertainty (aka shakeup) is large, one way to combat is to make many submissions and choose the \"median one\". now that post submission is opened, you can make 100 submissions per day to experiment and see the trend.",
    "3415235": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15255201%2F5278b540e914a10f93f718d69a2e74cb%2FScreenshot%202026-02-28%20213933.png?generation=1772286381948854&alt=media)"
  }
}