{
  "id": 679772,
  "title": "A costly mistake to learn from",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/679772",
  "author_name": "",
  "post_date": "2026-03-03T16:32:29.523881900Z",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>The sharing at the end of the competition has really helped me see why so many of the post processing ideas I had didn't work. More importantly, I see a crucial mistake that I made that allowed us to miss what could have been a large boost. </p>\n<p>For post processing for this competition I was primarily concerned with hole filling and sheet unmerging. Sheet unmerging, it seems was a lost cause for all of us but I do see some had success with hole filling. The only hole filling I get any success with is the diagonal hole filling (great idea, I didn't think of it). But the largest improvement for us would have been median filtering. And here is a story about how we just barely missed this metric hack. </p>\n<p>At some point during the competition we had assumed that people were getting higher scores through mass ensemble stacks. So we decided to try to convert all our models to TensorRT and begin a stack of our own. This did absolutely help our score, however, we noticed that the same pipeline but with TensorRT gave slightly better score on validation. More interesting, the only difference in prediction slices was slight blurs on the edges of sheets. This is where the mistake happened. I remember vividly telling my wife \"wow its so strange that converting to tensorrt made scores better, if anything it typically makes things a little worse, but im certainly not going to question it because its been a long week!\". I hadn't realized it at the time, but this blurring improving score was a key indicator of the fact that doing something like repeated median filtering would improve the score on masks. If I had only asked more questions on it, investigated other blurring and filtering, then there is a chance we would have found something like median filtering that significantly increases our score in both public and private lb. </p>\n<p>I was running out of steam, but this is a powerful reminder to stay curious, understand the why, and also probably a lesson in taking rest when you need it 😂</p>",
  "messages": [
    {
      "id": "3416752",
      "postDate": "03/03/2026 16:32:29",
      "content": "<p>The sharing at the end of the competition has really helped me see why so many of the post processing ideas I had didn't work. More importantly, I see a crucial mistake that I made that allowed us to miss what could have been a large boost. </p>\n<p>For post processing for this competition I was primarily concerned with hole filling and sheet unmerging. Sheet unmerging, it seems was a lost cause for all of us but I do see some had success with hole filling. The only hole filling I get any success with is the diagonal hole filling (great idea, I didn't think of it). But the largest improvement for us would have been median filtering. And here is a story about how we just barely missed this metric hack. </p>\n<p>At some point during the competition we had assumed that people were getting higher scores through mass ensemble stacks. So we decided to try to convert all our models to TensorRT and begin a stack of our own. This did absolutely help our score, however, we noticed that the same pipeline but with TensorRT gave slightly better score on validation. More interesting, the only difference in prediction slices was slight blurs on the edges of sheets. This is where the mistake happened. I remember vividly telling my wife \"wow its so strange that converting to tensorrt made scores better, if anything it typically makes things a little worse, but im certainly not going to question it because its been a long week!\". I hadn't realized it at the time, but this blurring improving score was a key indicator of the fact that doing something like repeated median filtering would improve the score on masks. If I had only asked more questions on it, investigated other blurring and filtering, then there is a chance we would have found something like median filtering that significantly increases our score in both public and private lb. </p>\n<p>I was running out of steam, but this is a powerful reminder to stay curious, understand the why, and also probably a lesson in taking rest when you need it 😂</p>",
      "rawMarkdown": "The sharing at the end of the competition has really helped me see why so many of the post processing ideas I had didn't work. More importantly, I see a crucial mistake that I made that allowed us to miss what could have been a large boost. \n\nFor post processing for this competition I was primarily concerned with hole filling and sheet unmerging. Sheet unmerging, it seems was a lost cause for all of us but I do see some had success with hole filling. The only hole filling I get any success with is the diagonal hole filling (great idea, I didn't think of it). But the largest improvement for us would have been median filtering. And here is a story about how we just barely missed this metric hack. \n\nAt some point during the competition we had assumed that people were getting higher scores through mass ensemble stacks. So we decided to try to convert all our models to TensorRT and begin a stack of our own. This did absolutely help our score, however, we noticed that the same pipeline but with TensorRT gave slightly better score on validation. More interesting, the only difference in prediction slices was slight blurs on the edges of sheets. This is where the mistake happened. I remember vividly telling my wife \"wow its so strange that converting to tensorrt made scores better, if anything it typically makes things a little worse, but im certainly not going to question it because its been a long week!\". I hadn't realized it at the time, but this blurring improving score was a key indicator of the fact that doing something like repeated median filtering would improve the score on masks. If I had only asked more questions on it, investigated other blurring and filtering, then there is a chance we would have found something like median filtering that significantly increases our score in both public and private lb. \n\nI was running out of steam, but this is a powerful reminder to stay curious, understand the why, and also probably a lesson in taking rest when you need it 😂",
      "votes": null
    },
    {
      "id": "3416769",
      "postDate": "03/03/2026 17:31:56",
      "content": "<p>If I didn't have my teammate <a href=\"https://www.kaggle.com/ggayoayogg\" target=\"_blank\">@ggayoayogg</a> , I probably wouldn't have known how effective the filter is. Although he doesn't fully understand why it boosts performance, this reminded me that I should try more seemingly unreasonable approaches—there might be surprises hidden within.</p>",
      "rawMarkdown": "If I didn't have my teammate @ggayoayogg , I probably wouldn't have known how effective the filter is. Although he doesn't fully understand why it boosts performance, this reminded me that I should try more seemingly unreasonable approaches—there might be surprises hidden within.",
      "votes": null
    },
    {
      "id": "3416773",
      "postDate": "03/03/2026 17:37:29",
      "content": "<p>I actually think our de-sticking solution is pretty good; it's just that the official baseline model is stuck way too severely—even high-confidence predictions sometimes get clumped together into a mess. 😀</p>",
      "rawMarkdown": "I actually think our de-sticking solution is pretty good; it's just that the official baseline model is stuck way too severely—even high-confidence predictions sometimes get clumped together into a mess. 😀",
      "votes": null
    },
    {
      "id": "3416798",
      "postDate": "03/03/2026 19:08:34",
      "content": "<p>Yeah I honestly didnt think of it because it seems really counter intuitive to the true goal of the competition. But however, if you did your post processing from a metric focus I think this would be more likely to find as it does make sense for the metric to have those softer boundaries. In theory if we were all at .8 then I think it would hurt but because we were all making fairly rough predictions anyway that bet hedge pays off a lot. At least that is the way I think about it. </p>",
      "rawMarkdown": "Yeah I honestly didnt think of it because it seems really counter intuitive to the true goal of the competition. But however, if you did your post processing from a metric focus I think this would be more likely to find as it does make sense for the metric to have those softer boundaries. In theory if we were all at .8 then I think it would hurt but because we were all making fairly rough predictions anyway that bet hedge pays off a lot. At least that is the way I think about it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3416769,
      "author_name": "chengtingyi",
      "author_url": "",
      "post_date": "03/03/2026 17:31:56",
      "content": "<p>If I didn't have my teammate <a href=\"https://www.kaggle.com/ggayoayogg\" target=\"_blank\">@ggayoayogg</a> , I probably wouldn't have known how effective the filter is. Although he doesn't fully understand why it boosts performance, this reminded me that I should try more seemingly unreasonable approaches—there might be surprises hidden within.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3416798,
          "author_name": "cody11null",
          "author_url": "",
          "post_date": "03/03/2026 19:08:34",
          "content": "<p>Yeah I honestly didnt think of it because it seems really counter intuitive to the true goal of the competition. But however, if you did your post processing from a metric focus I think this would be more likely to find as it does make sense for the metric to have those softer boundaries. In theory if we were all at .8 then I think it would hurt but because we were all making fairly rough predictions anyway that bet hedge pays off a lot. At least that is the way I think about it. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3416773,
      "author_name": "chengtingyi",
      "author_url": "",
      "post_date": "03/03/2026 17:37:29",
      "content": "<p>I actually think our de-sticking solution is pretty good; it's just that the official baseline model is stuck way too severely—even high-confidence predictions sometimes get clumped together into a mess. 😀</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3416752": "The sharing at the end of the competition has really helped me see why so many of the post processing ideas I had didn't work. More importantly, I see a crucial mistake that I made that allowed us to miss what could have been a large boost. \n\nFor post processing for this competition I was primarily concerned with hole filling and sheet unmerging. Sheet unmerging, it seems was a lost cause for all of us but I do see some had success with hole filling. The only hole filling I get any success with is the diagonal hole filling (great idea, I didn't think of it). But the largest improvement for us would have been median filtering. And here is a story about how we just barely missed this metric hack. \n\nAt some point during the competition we had assumed that people were getting higher scores through mass ensemble stacks. So we decided to try to convert all our models to TensorRT and begin a stack of our own. This did absolutely help our score, however, we noticed that the same pipeline but with TensorRT gave slightly better score on validation. More interesting, the only difference in prediction slices was slight blurs on the edges of sheets. This is where the mistake happened. I remember vividly telling my wife \"wow its so strange that converting to tensorrt made scores better, if anything it typically makes things a little worse, but im certainly not going to question it because its been a long week!\". I hadn't realized it at the time, but this blurring improving score was a key indicator of the fact that doing something like repeated median filtering would improve the score on masks. If I had only asked more questions on it, investigated other blurring and filtering, then there is a chance we would have found something like median filtering that significantly increases our score in both public and private lb. \n\nI was running out of steam, but this is a powerful reminder to stay curious, understand the why, and also probably a lesson in taking rest when you need it 😂",
    "3416769": "If I didn't have my teammate @ggayoayogg , I probably wouldn't have known how effective the filter is. Although he doesn't fully understand why it boosts performance, this reminded me that I should try more seemingly unreasonable approaches—there might be surprises hidden within.",
    "3416773": "I actually think our de-sticking solution is pretty good; it's just that the official baseline model is stuck way too severely—even high-confidence predictions sometimes get clumped together into a mess. 😀",
    "3416798": "Yeah I honestly didnt think of it because it seems really counter intuitive to the true goal of the competition. But however, if you did your post processing from a metric focus I think this would be more likely to find as it does make sense for the metric to have those softer boundaries. In theory if we were all at .8 then I think it would hurt but because we were all making fairly rough predictions anyway that bet hedge pays off a lot. At least that is the way I think about it."
  },
  "source": "meta"
}