{
  "id": 679220,
  "title": "18th : median filter x 7 post processing is very strong",
  "url": "/competitions/vesuvius-challenge-surface-detection/writeups/18th-median-filter-x-7-post-processing-is-very-s",
  "author_name": "",
  "post_date": "2026-02-28T00:43:45.663Z",
  "votes": 32,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Great competition—thanks to the Kaggle team for organizing and running everything so smoothly!</p>",
  "messages": [
    {
      "id": "3414934",
      "postDate": "02/28/2026 00:31:06",
      "content": "<p>Great competition—thanks to the Kaggle team for organizing and running everything so smoothly!</p>",
      "rawMarkdown": "Great competition—thanks to the Kaggle team for organizing and running everything so smoothly!",
      "votes": null
    },
    {
      "id": "3414940",
      "postDate": "02/28/2026 01:06:27",
      "content": "<p>6 models. runtime 6h. single thread</p>",
      "rawMarkdown": "6 models. runtime 6h. single thread",
      "votes": null
    },
    {
      "id": "3414975",
      "postDate": "02/28/2026 02:43:39",
      "content": "<p>code \n<a href=\"https://www.kaggle.com/code/sugupoko/v30-4-group-ensemble-blend-blur-bright-lowco?scriptVersionId=300017285\" target=\"_blank\">https://www.kaggle.com/code/sugupoko/v30-4-group-ensemble-blend-blur-bright-lowco?scriptVersionId=300017285</a></p>",
      "rawMarkdown": "code \nhttps://www.kaggle.com/code/sugupoko/v30-4-group-ensemble-blend-blur-bright-lowco?scriptVersionId=300017285",
      "votes": null
    },
    {
      "id": "3415131",
      "postDate": "02/28/2026 08:15:09",
      "content": "<p>Interesting solution!</p>\n<p>Did you compress in Z every input sample, or you decided the axis of compression depending on some analysis?</p>",
      "rawMarkdown": "Interesting solution!\n\nDid you compress in Z every input sample, or you decided the axis of compression depending on some analysis?",
      "votes": null
    },
    {
      "id": "3415146",
      "postDate": "02/28/2026 09:13:02",
      "content": "<p>Architecture : host model\nconfig : </p>\n<ul>\n<li>Pretrained model: 200 epochs, trained on ~700 pseudo-labeled samples + 400 external samples.</li>\n<li>Global-coverage model: 500 epochs, trained without pseudo-labeled data.</li>\n<li>Specialized model: 100 epochs, trained on selected data (~100–200 samples).</li>\n</ul>",
      "rawMarkdown": "Architecture : host model\nconfig : \n- Pretrained model: 200 epochs, trained on ~700 pseudo-labeled samples + 400 external samples.\n- Global-coverage model: 500 epochs, trained without pseudo-labeled data.\n- Specialized model: 100 epochs, trained on selected data (~100–200 samples).",
      "votes": null
    },
    {
      "id": "3415148",
      "postDate": "02/28/2026 09:16:15",
      "content": "<p>Thank you for your comment!!</p>\n<p>No, I didn’t choose it based on any analysis. Given the VRAM constraints of the RTX 5090, I just fixed it at around that range.</p>",
      "rawMarkdown": "Thank you for your comment!!\n\nNo, I didn’t choose it based on any analysis. Given the VRAM constraints of the RTX 5090, I just fixed it at around that range.",
      "votes": null
    },
    {
      "id": "3415362",
      "postDate": "02/28/2026 19:01:10",
      "content": "<p>How? :O\nChecking your shared notebook. </p>",
      "rawMarkdown": "How? :O\nChecking your shared notebook.",
      "votes": null
    },
    {
      "id": "3415441",
      "postDate": "02/28/2026 23:25:08",
      "content": "<p>Z compression is very fast. In most cases, Z is redundant.</p>",
      "rawMarkdown": "Z compression is very fast. In most cases, Z is redundant.",
      "votes": null
    },
    {
      "id": "3415444",
      "postDate": "02/28/2026 23:28:41",
      "content": "<p>Congratulations on strong 18th place finish. Can you explain the \"median filter x7 times\"? I'm curious how this works. Do you ensemble models and/or use TTA and then take the median?</p>",
      "rawMarkdown": "Congratulations on strong 18th place finish. Can you explain the \"median filter x7 times\"? I'm curious how this works. Do you ensemble models and/or use TTA and then take the median?",
      "votes": null
    },
    {
      "id": "3415448",
      "postDate": "02/28/2026 23:40:41",
      "content": "<p>Thank you for your comment!!!</p>\n<p>After ensemble and threshold-based binarization, I simply apply the following code.\nThe post-processing is minimal but very strong.</p>\n<pre><code>from scipy.ndimage import median_filter\n\nPP_MEDIAN_ITER = 7\nPP_MEDIAN_SIZE = 3\n\ndef postprocess(mask: np.ndarray) -&gt; np.ndarray:\n    mask = mask.copy()\n    for _ in range(PP_MEDIAN_ITER):\n        mask = (median_filter(mask, size=PP_MEDIAN_SIZE) &gt; 0).astype(np.uint8)\n    return mask.astype(np.uint8)\n\nmask = probs &gt;= THRESHOLD\nmask = postprocess(mask.astype(np.uint8))\n</code></pre>\n<p>I think the last cell will be helpful as a reference.\n<a href=\"https://www.kaggle.com/code/sugupoko/v30-4-group-ensemble-blend-blur-bright-lowco?scriptVersionId=300348690&amp;cellId=16\" target=\"_blank\">https://www.kaggle.com/code/sugupoko/v30-4-group-ensemble-blend-blur-bright-lowco?scriptVersionId=300348690&amp;cellId=16</a></p>",
      "rawMarkdown": "Thank you for your comment!!!\n\nAfter ensemble and threshold-based binarization, I simply apply the following code.\nThe post-processing is minimal but very strong.\n\n```python\nfrom scipy.ndimage import median_filter\n\nPP_MEDIAN_ITER = 7\nPP_MEDIAN_SIZE = 3\n\ndef postprocess(mask: np.ndarray) -> np.ndarray:\n    mask = mask.copy()\n    for _ in range(PP_MEDIAN_ITER):\n        mask = (median_filter(mask, size=PP_MEDIAN_SIZE) > 0).astype(np.uint8)\n    return mask.astype(np.uint8)\n\nmask = probs >= THRESHOLD\nmask = postprocess(mask.astype(np.uint8))\n\n```\n\nI think the last cell will be helpful as a reference.\nhttps://www.kaggle.com/code/sugupoko/v30-4-group-ensemble-blend-blur-bright-lowco?scriptVersionId=300348690&cellId=16",
      "votes": null
    },
    {
      "id": "3415453",
      "postDate": "02/28/2026 23:49:18",
      "content": "<p>Wow, awesome. This is a great post process. I see that you update each pixel prediction based on its 27 neighbors' median.</p>",
      "rawMarkdown": "Wow, awesome. This is a great post process. I see that you update each pixel prediction based on its 27 neighbors' median.",
      "votes": null
    },
    {
      "id": "3415532",
      "postDate": "03/01/2026 01:18:27",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a>  san</p>\n<p>I had forgotten about my experiment, I did analyze it.\nI wrote it up here — thank you!</p>\n<p><a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/679379\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/679379</a></p>",
      "rawMarkdown": "Hi @giorgioangelotti  san\n\nI had forgotten about my experiment, I did analyze it.\nI wrote it up here — thank you!\n\nhttps://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/679379",
      "votes": null
    },
    {
      "id": "3415705",
      "postDate": "03/01/2026 07:05:53",
      "content": "<p>Bravo for 18th place <a href=\"https://www.kaggle.com/sugupoko\" target=\"_blank\">@sugupoko</a> </p>",
      "rawMarkdown": "Bravo for 18th place @sugupoko",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3414940,
      "author_name": "sugupoko",
      "author_url": "",
      "post_date": "02/28/2026 01:06:27",
      "content": "<p>6 models. runtime 6h. single thread</p>",
      "votes": null,
      "replies": [
        {
          "id": 3415362,
          "author_name": "rajeshthevar",
          "author_url": "",
          "post_date": "02/28/2026 19:01:10",
          "content": "<p>How? :O\nChecking your shared notebook. </p>",
          "votes": null,
          "replies": [
            {
              "id": 3415441,
              "author_name": "sugupoko",
              "author_url": "",
              "post_date": "02/28/2026 23:25:08",
              "content": "<p>Z compression is very fast. In most cases, Z is redundant.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3414975,
      "author_name": "sugupoko",
      "author_url": "",
      "post_date": "02/28/2026 02:43:39",
      "content": "<p>code \n<a href=\"https://www.kaggle.com/code/sugupoko/v30-4-group-ensemble-blend-blur-bright-lowco?scriptVersionId=300017285\" target=\"_blank\">https://www.kaggle.com/code/sugupoko/v30-4-group-ensemble-blend-blur-bright-lowco?scriptVersionId=300017285</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3415131,
      "author_name": "giorgioangelotti",
      "author_url": "",
      "post_date": "02/28/2026 08:15:09",
      "content": "<p>Interesting solution!</p>\n<p>Did you compress in Z every input sample, or you decided the axis of compression depending on some analysis?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3415148,
          "author_name": "sugupoko",
          "author_url": "",
          "post_date": "02/28/2026 09:16:15",
          "content": "<p>Thank you for your comment!!</p>\n<p>No, I didn’t choose it based on any analysis. Given the VRAM constraints of the RTX 5090, I just fixed it at around that range.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3415532,
          "author_name": "sugupoko",
          "author_url": "",
          "post_date": "03/01/2026 01:18:27",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a>  san</p>\n<p>I had forgotten about my experiment, I did analyze it.\nI wrote it up here — thank you!</p>\n<p><a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/679379\" target=\"_blank\">https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/679379</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3415146,
      "author_name": "sugupoko",
      "author_url": "",
      "post_date": "02/28/2026 09:13:02",
      "content": "<p>Architecture : host model\nconfig : </p>\n<ul>\n<li>Pretrained model: 200 epochs, trained on ~700 pseudo-labeled samples + 400 external samples.</li>\n<li>Global-coverage model: 500 epochs, trained without pseudo-labeled data.</li>\n<li>Specialized model: 100 epochs, trained on selected data (~100–200 samples).</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3415444,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/28/2026 23:28:41",
      "content": "<p>Congratulations on strong 18th place finish. Can you explain the \"median filter x7 times\"? I'm curious how this works. Do you ensemble models and/or use TTA and then take the median?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3415448,
          "author_name": "sugupoko",
          "author_url": "",
          "post_date": "02/28/2026 23:40:41",
          "content": "<p>Thank you for your comment!!!</p>\n<p>After ensemble and threshold-based binarization, I simply apply the following code.\nThe post-processing is minimal but very strong.</p>\n<pre><code>from scipy.ndimage import median_filter\n\nPP_MEDIAN_ITER = 7\nPP_MEDIAN_SIZE = 3\n\ndef postprocess(mask: np.ndarray) -&gt; np.ndarray:\n    mask = mask.copy()\n    for _ in range(PP_MEDIAN_ITER):\n        mask = (median_filter(mask, size=PP_MEDIAN_SIZE) &gt; 0).astype(np.uint8)\n    return mask.astype(np.uint8)\n\nmask = probs &gt;= THRESHOLD\nmask = postprocess(mask.astype(np.uint8))\n</code></pre>\n<p>I think the last cell will be helpful as a reference.\n<a href=\"https://www.kaggle.com/code/sugupoko/v30-4-group-ensemble-blend-blur-bright-lowco?scriptVersionId=300348690&amp;cellId=16\" target=\"_blank\">https://www.kaggle.com/code/sugupoko/v30-4-group-ensemble-blend-blur-bright-lowco?scriptVersionId=300348690&amp;cellId=16</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 3415453,
              "author_name": "cdeotte",
              "author_url": "",
              "post_date": "02/28/2026 23:49:18",
              "content": "<p>Wow, awesome. This is a great post process. I see that you update each pixel prediction based on its 27 neighbors' median.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3415705,
      "author_name": "navneetbende",
      "author_url": "",
      "post_date": "03/01/2026 07:05:53",
      "content": "<p>Bravo for 18th place <a href=\"https://www.kaggle.com/sugupoko\" target=\"_blank\">@sugupoko</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3414934": "Great competition—thanks to the Kaggle team for organizing and running everything so smoothly!",
    "3414940": "6 models. runtime 6h. single thread",
    "3414975": "code \nhttps://www.kaggle.com/code/sugupoko/v30-4-group-ensemble-blend-blur-bright-lowco?scriptVersionId=300017285",
    "3415131": "Interesting solution!\n\nDid you compress in Z every input sample, or you decided the axis of compression depending on some analysis?",
    "3415146": "Architecture : host model\nconfig : \n- Pretrained model: 200 epochs, trained on ~700 pseudo-labeled samples + 400 external samples.\n- Global-coverage model: 500 epochs, trained without pseudo-labeled data.\n- Specialized model: 100 epochs, trained on selected data (~100–200 samples).",
    "3415148": "Thank you for your comment!!\n\nNo, I didn’t choose it based on any analysis. Given the VRAM constraints of the RTX 5090, I just fixed it at around that range.",
    "3415362": "How? :O\nChecking your shared notebook.",
    "3415441": "Z compression is very fast. In most cases, Z is redundant.",
    "3415444": "Congratulations on strong 18th place finish. Can you explain the \"median filter x7 times\"? I'm curious how this works. Do you ensemble models and/or use TTA and then take the median?",
    "3415448": "Thank you for your comment!!!\n\nAfter ensemble and threshold-based binarization, I simply apply the following code.\nThe post-processing is minimal but very strong.\n\n```python\nfrom scipy.ndimage import median_filter\n\nPP_MEDIAN_ITER = 7\nPP_MEDIAN_SIZE = 3\n\ndef postprocess(mask: np.ndarray) -> np.ndarray:\n    mask = mask.copy()\n    for _ in range(PP_MEDIAN_ITER):\n        mask = (median_filter(mask, size=PP_MEDIAN_SIZE) > 0).astype(np.uint8)\n    return mask.astype(np.uint8)\n\nmask = probs >= THRESHOLD\nmask = postprocess(mask.astype(np.uint8))\n\n```\n\nI think the last cell will be helpful as a reference.\nhttps://www.kaggle.com/code/sugupoko/v30-4-group-ensemble-blend-blur-bright-lowco?scriptVersionId=300348690&cellId=16",
    "3415453": "Wow, awesome. This is a great post process. I see that you update each pixel prediction based on its 27 neighbors' median.",
    "3415532": "Hi @giorgioangelotti  san\n\nI had forgotten about my experiment, I did analyze it.\nI wrote it up here — thank you!\n\nhttps://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/679379",
    "3415705": "Bravo for 18th place @sugupoko"
  },
  "source": "meta"
}