{
  "id": 475136,
  "title": "Inference Retraining Trick Boosted Private Score: 0.565 to 0.680 (Ranked 7th in Private LB)",
  "url": "/competitions/blood-vessel-segmentation/discussion/475136",
  "author_name": "yyykrk",
  "post_date": "2024-02-07T08:19:25.860000",
  "votes": 13,
  "comment_count": 13,
  "views": 0,
  "content": "<p>First and foremost, I would like to thank the organizers of the contest.</p>\n<p>I ranked 18th in the public LB and 34th in the private LB. However, I'd like to share that one of the ideas I tried during the contest turned out to be effective due to a late submission.</p>\n<p>It was confirmed that <strong>retraining using pseudo labels for the private test data (kidney_6) during inference can significantly boost the private score as below(0.565 -&gt; 0.680).</strong></p>\n<p>I attempted this idea with the public test data (kidney_5) during the contest period, but it didn't work well for the public score, so I did not pursue it further. However, it appears to have been impactful for the private test data.</p>\n<h1>Results</h1>\n<table>\n<thead>\n<tr>\n<th>Retraining</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>None</td>\n<td>0.877</td>\n<td>0.565</td>\n</tr>\n<tr>\n<td>With kidney_5 (public)</td>\n<td>0.871</td>\n<td>0.621</td>\n</tr>\n<tr>\n<td>With kidney_6 (private)</td>\n<td>0.870</td>\n<td>0.680</td>\n</tr>\n</tbody>\n</table>\n<h1>Model</h1>\n<ul>\n<li>network: custom 2D-UNet</li>\n<li>backbone: maxvit-small</li>\n<li>tiling size: 384x384</li>\n<li>train data: kidney_1, 50um_LADAF-2020-31_kidney_pag-0.01_0.02_jp2_(pseudo-labeling)</li>\n</ul>",
  "messages": [
    {
      "id": 2641009,
      "postDate": "2024-02-07T08:19:25.860Z",
      "content": "<p>First and foremost, I would like to thank the organizers of the contest.</p>\n<p>I ranked 18th in the public LB and 34th in the private LB. However, I'd like to share that one of the ideas I tried during the contest turned out to be effective due to a late submission.</p>\n<p>It was confirmed that <strong>retraining using pseudo labels for the private test data (kidney_6) during inference can significantly boost the private score as below(0.565 -&gt; 0.680).</strong></p>\n<p>I attempted this idea with the public test data (kidney_5) during the contest period, but it didn't work well for the public score, so I did not pursue it further. However, it appears to have been impactful for the private test data.</p>\n<h1>Results</h1>\n<table>\n<thead>\n<tr>\n<th>Retraining</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>None</td>\n<td>0.877</td>\n<td>0.565</td>\n</tr>\n<tr>\n<td>With kidney_5 (public)</td>\n<td>0.871</td>\n<td>0.621</td>\n</tr>\n<tr>\n<td>With kidney_6 (private)</td>\n<td>0.870</td>\n<td>0.680</td>\n</tr>\n</tbody>\n</table>\n<h1>Model</h1>\n<ul>\n<li>network: custom 2D-UNet</li>\n<li>backbone: maxvit-small</li>\n<li>tiling size: 384x384</li>\n<li>train data: kidney_1, 50um_LADAF-2020-31_kidney_pag-0.01_0.02_jp2_(pseudo-labeling)</li>\n</ul>",
      "rawMarkdown": "First and foremost, I would like to thank the organizers of the contest.\n\nI ranked 18th in the public LB and 34th in the private LB. However, I'd like to share that one of the ideas I tried during the contest turned out to be effective due to a late submission.\n\nIt was confirmed that **retraining using pseudo labels for the private test data (kidney_6) during inference can significantly boost the private score as below(0.565 -> 0.680).**\n\nI attempted this idea with the public test data (kidney_5) during the contest period, but it didn't work well for the public score, so I did not pursue it further. However, it appears to have been impactful for the private test data.\n\n# Results\n| Retraining| Public | Private |\n| --- | --- | --- |\n| None | 0.877 | 0.565 |\n| With kidney_5 (public)| 0.871 | 0.621 |\n| With kidney_6 (private) | 0.870 | 0.680 |\n\n# Model\n- network: custom 2D-UNet\n- backbone: maxvit-small\n- tiling size: 384x384\n- train data: kidney_1, 50um_LADAF-2020-31_kidney_pag-0.01_0.02_jp2_(pseudo-labeling)",
      "votes": 13
    },
    {
      "id": 2643891,
      "postDate": "2024-02-09T06:53:58.343Z",
      "content": "<p>Good work </p>",
      "rawMarkdown": "Good work "
    },
    {
      "id": 2641499,
      "postDate": "2024-02-07T14:19:22.890Z",
      "content": "<p>I, too, used something like this. Unfortunately this came to my mind just two days before the competition deadline, so I was not able to improve it much and in my case it improved CV, public and private LBs (0.93, 0.813, 0.621). In fact, my current rank of 19 is also using this trick.</p>",
      "rawMarkdown": "I, too, used something like this. Unfortunately this came to my mind just two days before the competition deadline, so I was not able to improve it much and in my case it improved CV, public and private LBs (0.93, 0.813, 0.621). In fact, my current rank of 19 is also using this trick.",
      "replies": [
        {
          "id": 2641508,
          "postDate": "2024-02-07T14:23:13.967Z",
          "content": "<p>I did not tweak my TH at all, and it remained 0.5 (after sigmoid) for all my submission. Perhaps I should have done some tweaks to it too.</p>",
          "rawMarkdown": "I did not tweak my TH at all, and it remained 0.5 (after sigmoid) for all my submission. Perhaps I should have done some tweaks to it too.",
          "replies": [
            {
              "id": 2644403,
              "postDate": "2024-02-09T13:03:10.093Z",
              "content": "<p>Since I also tried it right before the end of the competition, I didn't adjust the threshold, but depending on the threshold, I might have obtained good results even on the public dataset.</p>",
              "rawMarkdown": "Since I also tried it right before the end of the competition, I didn't adjust the threshold, but depending on the threshold, I might have obtained good results even on the public dataset."
            }
          ]
        }
      ]
    },
    {
      "id": 2641393,
      "postDate": "2024-02-07T13:17:12.867Z",
      "content": "<p>thanks for sharing, any thoughts on what could make public performance not boosted much (~0.87) but private data ends up bossted more (0.565 -&gt; 0.621 -&gt; 0.68)? I suppose for the pseudo labelling you simpling used the model threshold you got from training right? I wonder what your pub and private score difference might be if you do the pseudo labelling for the sparse data for training</p>",
      "rawMarkdown": "thanks for sharing, any thoughts on what could make public performance not boosted much (~0.87) but private data ends up bossted more (0.565 -> 0.621 -> 0.68)? I suppose for the pseudo labelling you simpling used the model threshold you got from training right? I wonder what your pub and private score difference might be if you do the pseudo labelling for the sparse data for training",
      "replies": [
        {
          "id": 2644401,
          "postDate": "2024-02-09T13:00:05.277Z",
          "content": "<p>I'm not entirely sure why there was a significant effect only on the private data. I used the threshold that resulted in the best cross-validation performance. I also tried pseudo-labeling on sparse data, but it didn't yield much improvement.</p>",
          "rawMarkdown": "I'm not entirely sure why there was a significant effect only on the private data. I used the threshold that resulted in the best cross-validation performance. I also tried pseudo-labeling on sparse data, but it didn't yield much improvement.",
          "replies": [
            {
              "id": 2645320,
              "postDate": "2024-02-10T06:27:09.073Z",
              "content": "<p>I was using a slightly customized dice loss which used only the values on the far right and far left [scale of 0-255] of the predicted mask for evaluation and update. I was not using values near 127 because these values were ambiguous. Even then, this technique quickly increased the false positives during the CV. So, may be the increase is just because of this increase of number of detected vessels. I think overall it's equivalent to lower thresholds.</p>",
              "rawMarkdown": "I was using a slightly customized dice loss which used only the values on the far right and far left [scale of 0-255] of the predicted mask for evaluation and update. I was not using values near 127 because these values were ambiguous. Even then, this technique quickly increased the false positives during the CV. So, may be the increase is just because of this increase of number of detected vessels. I think overall it's equivalent to lower thresholds."
            }
          ]
        }
      ]
    },
    {
      "id": 2641308,
      "postDate": "2024-02-07T12:13:41.340Z",
      "content": "<p>That's a really cool trick ! Especially for this kind of competition where this type of model will not be running on an iphone, and where it is important to end up with good preds but not necessarly in the minute the model is given those inputs. I am thinking about trying that idea out and do a late sub. Will let you know how it goes if I have time to do so. Congrats on the rank nevertheless</p>",
      "rawMarkdown": "That's a really cool trick ! Especially for this kind of competition where this type of model will not be running on an iphone, and where it is important to end up with good preds but not necessarly in the minute the model is given those inputs. I am thinking about trying that idea out and do a late sub. Will let you know how it goes if I have time to do so. Congrats on the rank nevertheless",
      "replies": [
        {
          "id": 2641310,
          "postDate": "2024-02-07T12:15:40.470Z",
          "content": "<p>Can you share more details on the \"finetuning\" params (epochs, lr and stuff like that) ? I really want to try this out as it can be used in any comp with a known public/private split but don't necessarl have the time to make 10 subs trying to optimise this. Thanks in advance</p>",
          "rawMarkdown": "Can you share more details on the \"finetuning\" params (epochs, lr and stuff like that) ? I really want to try this out as it can be used in any comp with a known public/private split but don't necessarl have the time to make 10 subs trying to optimise this. Thanks in advance",
          "replies": [
            {
              "id": 2641314,
              "postDate": "2024-02-07T12:24:10.533Z",
              "content": "<p>Thank you, and congrats on your medal too! I kept the training parameters the same but only ran one epoch because of time constraints during inference. Honestly, I think we'd get better results with a lower learning rate and more epochs.</p>",
              "rawMarkdown": "Thank you, and congrats on your medal too! I kept the training parameters the same but only ran one epoch because of time constraints during inference. Honestly, I think we'd get better results with a lower learning rate and more epochs.",
              "votes": 1
            },
            {
              "id": 2641350,
              "postDate": "2024-02-07T12:56:12.557Z",
              "content": "<p>Perfect thank you for the information.</p>",
              "rawMarkdown": "Perfect thank you for the information."
            },
            {
              "id": 2641617,
              "postDate": "2024-02-07T15:17:56.170Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2644405,
              "postDate": "2024-02-09T13:06:09.570Z",
              "content": "<p>I used my custom 2D-UNet model with the maxvit-small backbone. </p>",
              "rawMarkdown": "I used my custom 2D-UNet model with the maxvit-small backbone. "
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2643891,
      "author_name": "Arshpreet",
      "author_url": "",
      "post_date": "2024-02-09T06:53:58.343000",
      "content": "<p>Good work </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2641499,
      "author_name": "Zam",
      "author_url": "",
      "post_date": "2024-02-07T14:19:22.890000",
      "content": "<p>I, too, used something like this. Unfortunately this came to my mind just two days before the competition deadline, so I was not able to improve it much and in my case it improved CV, public and private LBs (0.93, 0.813, 0.621). In fact, my current rank of 19 is also using this trick.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2641508,
          "author_name": "Zam",
          "author_url": "",
          "post_date": "2024-02-07T14:23:13.967000",
          "content": "<p>I did not tweak my TH at all, and it remained 0.5 (after sigmoid) for all my submission. Perhaps I should have done some tweaks to it too.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2644403,
              "author_name": "yyykrk",
              "author_url": "",
              "post_date": "2024-02-09T13:03:10.093000",
              "content": "<p>Since I also tried it right before the end of the competition, I didn't adjust the threshold, but depending on the threshold, I might have obtained good results even on the public dataset.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2641393,
      "author_name": "samshipengs",
      "author_url": "",
      "post_date": "2024-02-07T13:17:12.867000",
      "content": "<p>thanks for sharing, any thoughts on what could make public performance not boosted much (~0.87) but private data ends up bossted more (0.565 -&gt; 0.621 -&gt; 0.68)? I suppose for the pseudo labelling you simpling used the model threshold you got from training right? I wonder what your pub and private score difference might be if you do the pseudo labelling for the sparse data for training</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2644401,
          "author_name": "yyykrk",
          "author_url": "",
          "post_date": "2024-02-09T13:00:05.277000",
          "content": "<p>I'm not entirely sure why there was a significant effect only on the private data. I used the threshold that resulted in the best cross-validation performance. I also tried pseudo-labeling on sparse data, but it didn't yield much improvement.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2645320,
              "author_name": "Zam",
              "author_url": "",
              "post_date": "2024-02-10T06:27:09.073000",
              "content": "<p>I was using a slightly customized dice loss which used only the values on the far right and far left [scale of 0-255] of the predicted mask for evaluation and update. I was not using values near 127 because these values were ambiguous. Even then, this technique quickly increased the false positives during the CV. So, may be the increase is just because of this increase of number of detected vessels. I think overall it's equivalent to lower thresholds.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2641308,
      "author_name": "JEANMPIA",
      "author_url": "",
      "post_date": "2024-02-07T12:13:41.340000",
      "content": "<p>That's a really cool trick ! Especially for this kind of competition where this type of model will not be running on an iphone, and where it is important to end up with good preds but not necessarly in the minute the model is given those inputs. I am thinking about trying that idea out and do a late sub. Will let you know how it goes if I have time to do so. Congrats on the rank nevertheless</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2641310,
          "author_name": "JEANMPIA",
          "author_url": "",
          "post_date": "2024-02-07T12:15:40.470000",
          "content": "<p>Can you share more details on the \"finetuning\" params (epochs, lr and stuff like that) ? I really want to try this out as it can be used in any comp with a known public/private split but don't necessarl have the time to make 10 subs trying to optimise this. Thanks in advance</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2641314,
              "author_name": "yyykrk",
              "author_url": "",
              "post_date": "2024-02-07T12:24:10.533000",
              "content": "<p>Thank you, and congrats on your medal too! I kept the training parameters the same but only ran one epoch because of time constraints during inference. Honestly, I think we'd get better results with a lower learning rate and more epochs.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2641350,
              "author_name": "JEANMPIA",
              "author_url": "",
              "post_date": "2024-02-07T12:56:12.557000",
              "content": "<p>Perfect thank you for the information.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2641617,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-02-07T15:17:56.170000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2644405,
              "author_name": "yyykrk",
              "author_url": "",
              "post_date": "2024-02-09T13:06:09.570000",
              "content": "<p>I used my custom 2D-UNet model with the maxvit-small backbone. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2641009": "First and foremost, I would like to thank the organizers of the contest.\n\nI ranked 18th in the public LB and 34th in the private LB. However, I'd like to share that one of the ideas I tried during the contest turned out to be effective due to a late submission.\n\nIt was confirmed that **retraining using pseudo labels for the private test data (kidney_6) during inference can significantly boost the private score as below(0.565 -> 0.680).**\n\nI attempted this idea with the public test data (kidney_5) during the contest period, but it didn't work well for the public score, so I did not pursue it further. However, it appears to have been impactful for the private test data.\n\n# Results\n| Retraining| Public | Private |\n| --- | --- | --- |\n| None | 0.877 | 0.565 |\n| With kidney_5 (public)| 0.871 | 0.621 |\n| With kidney_6 (private) | 0.870 | 0.680 |\n\n# Model\n- network: custom 2D-UNet\n- backbone: maxvit-small\n- tiling size: 384x384\n- train data: kidney_1, 50um_LADAF-2020-31_kidney_pag-0.01_0.02_jp2_(pseudo-labeling)",
    "2643891": "Good work ",
    "2641499": "I, too, used something like this. Unfortunately this came to my mind just two days before the competition deadline, so I was not able to improve it much and in my case it improved CV, public and private LBs (0.93, 0.813, 0.621). In fact, my current rank of 19 is also using this trick.",
    "2641393": "thanks for sharing, any thoughts on what could make public performance not boosted much (~0.87) but private data ends up bossted more (0.565 -> 0.621 -> 0.68)? I suppose for the pseudo labelling you simpling used the model threshold you got from training right? I wonder what your pub and private score difference might be if you do the pseudo labelling for the sparse data for training",
    "2641308": "That's a really cool trick ! Especially for this kind of competition where this type of model will not be running on an iphone, and where it is important to end up with good preds but not necessarly in the minute the model is given those inputs. I am thinking about trying that idea out and do a late sub. Will let you know how it goes if I have time to do so. Congrats on the rank nevertheless"
  }
}