{
  "id": 587766,
  "title": "Cheap post-processing to fight on medal borders",
  "url": "/competitions/waveform-inversion/discussion/587766",
  "author_name": "",
  "post_date": "2025-07-02T15:39:17.135224100Z",
  "votes": 4,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Many great top solutions have already been published, but it is not easy for a beginner like me to replicate or apply them. On the other hand, solutions that are on the borderline of medals are not easily shared, and some of you may be struggling with the gap between the published notes and your own notes.(Like me.)<br>\nHere is a record of the minimal ingenuity and post-processing that actually earned me a bronze medal.</p>\n<p>■Assumptions and Strategies<br>\nThere were very good public notes and scores were densely packed in the medal borders (100th to 500th).<br>\nBased on this situation, we predicted that even a slight improvement would reach the medal range.<br>\nJudging that model improvement would be difficult in terms of resources, skills, and deadlines, we focused only on post-processing the output CSV.<br>\nAlthough we were aware of the risk of overfitting to the public Leaderboard, we accepted the evaluation on a submission basis because there were no other objective score indicators.</p>\n<p>■ Approach flow<br>\n1.Check CSV output<br>\nOutput velocity map contains many “close but slightly different values” and hypothesized to contain local noise.</p>\n<p>2.Introduce basic filter<br>\nImplement process of using a 3x3 moving window and averaging only values within ±1% of the center value → Slightly improved scores.</p>\n<p>3.Adjustment of threshold<br>\nTried multiple values around 0.5-1.0%, with 1.0% being the most stable and effective.<br>\nHypothesis of image structure and extension of filter shape<br>\nAssumed that information is compressed in the horizontal direction of the velocity map and that noise in the horizontal direction is noticeable.</p>\n<p>4.Switching to a 3 × 5 horizontal filter improves it further.<br>\nThreshold Optimization<br>\nRe-examined with a range of 0.5 to 1.0% for the 3×5 filter, with 0.65% being the best.</p>\n<p>5.Two-step application of filter<br>\nAfter applying 3×5 (threshold 0.65), reprocess with 5×3 (threshold 0.65) → Final score reached highest ever. Submitted this and received a bronze medal.</p>\n<p>■Trial that did not work<br>\nSmoothing only unstable regions by taking the difference between 25.6 and 27.1 predicted CSVs → No effect<br>\nSmoothing and edge enhancement near the boundary → Noise enhancement, which is counterproductive<br>\nExtract features from images and cluster them → Process segment by segment → Failed</p>\n<p>■ General Summary<br>\nWe are aware that this is a very cheap method. We have not made any modifications to the model learning or deep structure.<br>\nHowever, for those who say, “If I can improve 0.01 more, it will be a medal”, post-processing can be a realistic and effective means.<br>\nI hope this post will be helpful to those who, like me, are aiming for medals with limited resources and abilities.</p>\n<p><a href=\"https://www.kaggle.com/code/shunsukehayashi1993/simple-solution-for-medal-borders\" target=\"_blank\">https://www.kaggle.com/code/shunsukehayashi1993/simple-solution-for-medal-borders</a></p>",
  "messages": [
    {
      "id": "3239219",
      "postDate": "07/02/2025 15:39:17",
      "content": "<p>Many great top solutions have already been published, but it is not easy for a beginner like me to replicate or apply them. On the other hand, solutions that are on the borderline of medals are not easily shared, and some of you may be struggling with the gap between the published notes and your own notes.(Like me.)<br>\nHere is a record of the minimal ingenuity and post-processing that actually earned me a bronze medal.</p>\n<p>■Assumptions and Strategies<br>\nThere were very good public notes and scores were densely packed in the medal borders (100th to 500th).<br>\nBased on this situation, we predicted that even a slight improvement would reach the medal range.<br>\nJudging that model improvement would be difficult in terms of resources, skills, and deadlines, we focused only on post-processing the output CSV.<br>\nAlthough we were aware of the risk of overfitting to the public Leaderboard, we accepted the evaluation on a submission basis because there were no other objective score indicators.</p>\n<p>■ Approach flow<br>\n1.Check CSV output<br>\nOutput velocity map contains many “close but slightly different values” and hypothesized to contain local noise.</p>\n<p>2.Introduce basic filter<br>\nImplement process of using a 3x3 moving window and averaging only values within ±1% of the center value → Slightly improved scores.</p>\n<p>3.Adjustment of threshold<br>\nTried multiple values around 0.5-1.0%, with 1.0% being the most stable and effective.<br>\nHypothesis of image structure and extension of filter shape<br>\nAssumed that information is compressed in the horizontal direction of the velocity map and that noise in the horizontal direction is noticeable.</p>\n<p>4.Switching to a 3 × 5 horizontal filter improves it further.<br>\nThreshold Optimization<br>\nRe-examined with a range of 0.5 to 1.0% for the 3×5 filter, with 0.65% being the best.</p>\n<p>5.Two-step application of filter<br>\nAfter applying 3×5 (threshold 0.65), reprocess with 5×3 (threshold 0.65) → Final score reached highest ever. Submitted this and received a bronze medal.</p>\n<p>■Trial that did not work<br>\nSmoothing only unstable regions by taking the difference between 25.6 and 27.1 predicted CSVs → No effect<br>\nSmoothing and edge enhancement near the boundary → Noise enhancement, which is counterproductive<br>\nExtract features from images and cluster them → Process segment by segment → Failed</p>\n<p>■ General Summary<br>\nWe are aware that this is a very cheap method. We have not made any modifications to the model learning or deep structure.<br>\nHowever, for those who say, “If I can improve 0.01 more, it will be a medal”, post-processing can be a realistic and effective means.<br>\nI hope this post will be helpful to those who, like me, are aiming for medals with limited resources and abilities.</p>\n<p><a href=\"https://www.kaggle.com/code/shunsukehayashi1993/simple-solution-for-medal-borders\" target=\"_blank\">https://www.kaggle.com/code/shunsukehayashi1993/simple-solution-for-medal-borders</a></p>",
      "rawMarkdown": "Many great top solutions have already been published, but it is not easy for a beginner like me to replicate or apply them. On the other hand, solutions that are on the borderline of medals are not easily shared, and some of you may be struggling with the gap between the published notes and your own notes.(Like me.)\nHere is a record of the minimal ingenuity and post-processing that actually earned me a bronze medal.\n\n■Assumptions and Strategies\nThere were very good public notes and scores were densely packed in the medal borders (100th to 500th).\nBased on this situation, we predicted that even a slight improvement would reach the medal range.\nJudging that model improvement would be difficult in terms of resources, skills, and deadlines, we focused only on post-processing the output CSV.\nAlthough we were aware of the risk of overfitting to the public Leaderboard, we accepted the evaluation on a submission basis because there were no other objective score indicators.\n\n■ Approach flow\n1.Check CSV output\nOutput velocity map contains many “close but slightly different values” and hypothesized to contain local noise.\n\n2.Introduce basic filter\nImplement process of using a 3x3 moving window and averaging only values within ±1% of the center value → Slightly improved scores.\n\n3.Adjustment of threshold\nTried multiple values around 0.5-1.0%, with 1.0% being the most stable and effective.\nHypothesis of image structure and extension of filter shape\nAssumed that information is compressed in the horizontal direction of the velocity map and that noise in the horizontal direction is noticeable.\n\n4.Switching to a 3 × 5 horizontal filter improves it further.\nThreshold Optimization\nRe-examined with a range of 0.5 to 1.0% for the 3×5 filter, with 0.65% being the best.\n\n5.Two-step application of filter\nAfter applying 3×5 (threshold 0.65), reprocess with 5×3 (threshold 0.65) → Final score reached highest ever. Submitted this and received a bronze medal.\n\n■Trial that did not work\nSmoothing only unstable regions by taking the difference between 25.6 and 27.1 predicted CSVs → No effect\nSmoothing and edge enhancement near the boundary → Noise enhancement, which is counterproductive\nExtract features from images and cluster them → Process segment by segment → Failed\n\n■ General Summary\nWe are aware that this is a very cheap method. We have not made any modifications to the model learning or deep structure.\nHowever, for those who say, “If I can improve 0.01 more, it will be a medal”, post-processing can be a realistic and effective means.\nI hope this post will be helpful to those who, like me, are aiming for medals with limited resources and abilities.\n\nhttps://www.kaggle.com/code/shunsukehayashi1993/simple-solution-for-medal-borders",
      "votes": null
    },
    {
      "id": "3239275",
      "postDate": "07/02/2025 16:41:03",
      "content": "<p>I didn’t tried the latest best public notebook, but the simple following modifications to <a href=\"https://www.kaggle.com/code/gguillard/caformer-full-resolution-improved-bbe229\" target=\"_blank\">Overvalue Awareness’s public notebook</a> improved the LB score by 0.01579 (28.17865 against 28.19444, presumably), which allowed me to secure a solid ranking of 535th. (:</p>\n<pre><code>                    \n                    \n\n                    \n                    y_preds = torch.(outputs[:, ]).cpu().numpy()\n\n                    \n                     (\n                        np.allclose(inputs[,:,].cpu().numpy(), inputs[,:,].cpu().numpy(), atol=)  \n                        &amp; np.allclose(inputs[,:,].cpu().numpy(), inputs[,:,].cpu().numpy(), atol=)  \n                    ):\n                        \n                        mean_col = outputs[:,].mean(dim=-, keepdim=)  \n                        y_preds[:,:,:] = torch.(mean_col).cpu().numpy()\n</code></pre>",
      "rawMarkdown": "I didn’t tried the latest best public notebook, but the simple following modifications to [Overvalue Awareness’s public notebook](https://www.kaggle.com/code/gguillard/caformer-full-resolution-improved-bbe229) improved the LB score by 0.01579 (28.17865 against 28.19444, presumably), which allowed me to secure a solid ranking of 535th. (:\n\n```python\n                    # original code\n                    # y_preds = outputs[:, 0].cpu().numpy()\n\n                    # round velocities to integers\n                    y_preds = torch.round(outputs[:, 0]).cpu().numpy()\n\n                    # check FlatVel symmetry\n                    if (\n                        np.allclose(inputs[0,:,0].cpu().numpy(), inputs[4,:,69].cpu().numpy(), atol=1e-4)  # source 0 + détecteur 0 = source 4 + détecteur 69\n                        & np.allclose(inputs[1,:,17].cpu().numpy(), inputs[3,:,52].cpu().numpy(), atol=1e-4)  # source 1 + détecteur 17 = source 3 + détecteur 52\n                    ):\n                        # for FlatVel, set all channels to rounded average value (don’t remember if I tried with median)\n                        mean_col = outputs[:,0].mean(dim=-1, keepdim=True)  # (B, 1, 70, 1)\n                        y_preds[:,:,:] = torch.round(mean_col).cpu().numpy()\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3239275,
      "author_name": "gguillard",
      "author_url": "",
      "post_date": "07/02/2025 16:41:03",
      "content": "<p>I didn’t tried the latest best public notebook, but the simple following modifications to <a href=\"https://www.kaggle.com/code/gguillard/caformer-full-resolution-improved-bbe229\" target=\"_blank\">Overvalue Awareness’s public notebook</a> improved the LB score by 0.01579 (28.17865 against 28.19444, presumably), which allowed me to secure a solid ranking of 535th. (:</p>\n<pre><code>                    \n                    \n\n                    \n                    y_preds = torch.(outputs[:, ]).cpu().numpy()\n\n                    \n                     (\n                        np.allclose(inputs[,:,].cpu().numpy(), inputs[,:,].cpu().numpy(), atol=)  \n                        &amp; np.allclose(inputs[,:,].cpu().numpy(), inputs[,:,].cpu().numpy(), atol=)  \n                    ):\n                        \n                        mean_col = outputs[:,].mean(dim=-, keepdim=)  \n                        y_preds[:,:,:] = torch.(mean_col).cpu().numpy()\n</code></pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3239219": "Many great top solutions have already been published, but it is not easy for a beginner like me to replicate or apply them. On the other hand, solutions that are on the borderline of medals are not easily shared, and some of you may be struggling with the gap between the published notes and your own notes.(Like me.)\nHere is a record of the minimal ingenuity and post-processing that actually earned me a bronze medal.\n\n■Assumptions and Strategies\nThere were very good public notes and scores were densely packed in the medal borders (100th to 500th).\nBased on this situation, we predicted that even a slight improvement would reach the medal range.\nJudging that model improvement would be difficult in terms of resources, skills, and deadlines, we focused only on post-processing the output CSV.\nAlthough we were aware of the risk of overfitting to the public Leaderboard, we accepted the evaluation on a submission basis because there were no other objective score indicators.\n\n■ Approach flow\n1.Check CSV output\nOutput velocity map contains many “close but slightly different values” and hypothesized to contain local noise.\n\n2.Introduce basic filter\nImplement process of using a 3x3 moving window and averaging only values within ±1% of the center value → Slightly improved scores.\n\n3.Adjustment of threshold\nTried multiple values around 0.5-1.0%, with 1.0% being the most stable and effective.\nHypothesis of image structure and extension of filter shape\nAssumed that information is compressed in the horizontal direction of the velocity map and that noise in the horizontal direction is noticeable.\n\n4.Switching to a 3 × 5 horizontal filter improves it further.\nThreshold Optimization\nRe-examined with a range of 0.5 to 1.0% for the 3×5 filter, with 0.65% being the best.\n\n5.Two-step application of filter\nAfter applying 3×5 (threshold 0.65), reprocess with 5×3 (threshold 0.65) → Final score reached highest ever. Submitted this and received a bronze medal.\n\n■Trial that did not work\nSmoothing only unstable regions by taking the difference between 25.6 and 27.1 predicted CSVs → No effect\nSmoothing and edge enhancement near the boundary → Noise enhancement, which is counterproductive\nExtract features from images and cluster them → Process segment by segment → Failed\n\n■ General Summary\nWe are aware that this is a very cheap method. We have not made any modifications to the model learning or deep structure.\nHowever, for those who say, “If I can improve 0.01 more, it will be a medal”, post-processing can be a realistic and effective means.\nI hope this post will be helpful to those who, like me, are aiming for medals with limited resources and abilities.\n\nhttps://www.kaggle.com/code/shunsukehayashi1993/simple-solution-for-medal-borders",
    "3239275": "I didn’t tried the latest best public notebook, but the simple following modifications to [Overvalue Awareness’s public notebook](https://www.kaggle.com/code/gguillard/caformer-full-resolution-improved-bbe229) improved the LB score by 0.01579 (28.17865 against 28.19444, presumably), which allowed me to secure a solid ranking of 535th. (:\n\n```python\n                    # original code\n                    # y_preds = outputs[:, 0].cpu().numpy()\n\n                    # round velocities to integers\n                    y_preds = torch.round(outputs[:, 0]).cpu().numpy()\n\n                    # check FlatVel symmetry\n                    if (\n                        np.allclose(inputs[0,:,0].cpu().numpy(), inputs[4,:,69].cpu().numpy(), atol=1e-4)  # source 0 + détecteur 0 = source 4 + détecteur 69\n                        & np.allclose(inputs[1,:,17].cpu().numpy(), inputs[3,:,52].cpu().numpy(), atol=1e-4)  # source 1 + détecteur 17 = source 3 + détecteur 52\n                    ):\n                        # for FlatVel, set all channels to rounded average value (don’t remember if I tried with median)\n                        mean_col = outputs[:,0].mean(dim=-1, keepdim=True)  # (B, 1, 70, 1)\n                        y_preds[:,:,:] = torch.round(mean_col).cpu().numpy()\n```"
  },
  "source": "meta"
}