{
  "id": 428392,
  "title": "public 9th / private 295th achieved private 0.575 with single modification",
  "url": "/competitions/hubmap-hacking-the-human-vasculature/writeups/yyama-public-9th-private-295th-achieved-private-0-",
  "author_name": "",
  "post_date": "2023-08-01T08:34:16.649463600Z",
  "votes": 32,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Despite achieving 9th place in the public leaderboard, I experienced a frightening shakeup, ending up at 295th place in the private leaderboard. However, many of the techniques I used proved effective in the private leaderboard, enabling me to achieve a score of 0.575 (equivalent to 4th place) by making just one modification to the 295th place code. Below, I will share the methods used and my reflections.</p>\n<h3>Method Overview:</h3>\n<p>I primarily utilized \"mmdet\" and employed ensemble learning with numerous models. I had hoped to use YOLO v8 and v7 as well, but the host did not disclose their availability, so I had to abandon the idea.</p>\n<p>For the ensemble, I made modifications to the \"WBF\" (Weighted Box Fusion) to use it with masks. Additionally, the major innovation was not limited to instance segmentation models but also involved using object detection models. I performed WBF based on bounding boxes and made mask predictions solely using the instance segmentation models. I used the object detection models only to calculate confidence scores. As you can see from the predictions, the models made far more predictions than needed. Given the nature of the evaluation metric, mAP, it was evident that properly aligning confidence scores could significantly improve the overall score.</p>\n<h3>Models</h3>\n<p>The following is a list of the models used:</p>\n<p>Instance Segmentation Models:</p>\n<ul>\n<li>Mask2Former</li>\n<li>MaskRCNN + various ConvNext v1 and v2</li>\n<li>Cascade RCNN + various ConvNext v1 and v2</li>\n<li>MaskRCNN + Swin Transformer<br>\nOther: MaskDino doesn’t work well.</li>\n</ul>\n<p>Object Detection Models:</p>\n<ul>\n<li>DINO</li>\n<li>Mask2Former (used only for object detection part)</li>\n<li>Diffusion Det</li>\n<li>ViT Det<br>\nOther: yolox doesn’t work well</li>\n</ul>\n<h3>Major Flaw:</h3>\n<p>I combined the outputs of the mask models using addition during post-processing. This resulted in improvements in ds1 (average IoU) and the public cross-validation (CV) score. However, as many of you may have noticed, when applying post-processing, the score on the private dataset degrades significantly. This indicates that the ds1 and private datasets have entirely different annotation tendencies. The annotations provided for ds1 were quite rough, and the dilation technique seemed effective due to this. One of the hosts mentioned that annotations were done in the same manner for both public and private datasets, but this information turned out to be inaccurate.</p>\n<h3>Single modification to achieve private 0.575:</h3>\n<p>Anticipating that the annotations for the private dataset might be more substantial, I submitted two versions: one with post-processing and another without. However, as mentioned earlier, during the ensemble mask generation, I was merely adding the masks instead of using averaging or voting. This method worked well for ds1 and the public dataset. On the other hand, I should have easily foreseen that if the private dataset had accurate and precise annotations, this approach could significantly reduce the score (which I overlooked).<br>\nTherefore, instead of adding the masks, I changed to using voting. This simple change dramatically improved the score from 0.397 to 0.575.</p>\n<div>\n  <img src=\"https://pbs.twimg.com/media/F2avc0eaEAEmdvL.jpg\" alt=\"score\">\n</div>\n<h3>Reflections:</h3>\n<p>One of the points I regret is being content with just having two submissions: one with post-processing and one without. Since I could have diversified the risks, I should have prepared submissions that could adapt to both small and accurate annotation masks.</p>",
  "messages": [
    {
      "id": "2368561",
      "postDate": "08/01/2023 08:34:16",
      "content": "<p>Despite achieving 9th place in the public leaderboard, I experienced a frightening shakeup, ending up at 295th place in the private leaderboard. However, many of the techniques I used proved effective in the private leaderboard, enabling me to achieve a score of 0.575 (equivalent to 4th place) by making just one modification to the 295th place code. Below, I will share the methods used and my reflections.</p>\n<h3>Method Overview:</h3>\n<p>I primarily utilized \"mmdet\" and employed ensemble learning with numerous models. I had hoped to use YOLO v8 and v7 as well, but the host did not disclose their availability, so I had to abandon the idea.</p>\n<p>For the ensemble, I made modifications to the \"WBF\" (Weighted Box Fusion) to use it with masks. Additionally, the major innovation was not limited to instance segmentation models but also involved using object detection models. I performed WBF based on bounding boxes and made mask predictions solely using the instance segmentation models. I used the object detection models only to calculate confidence scores. As you can see from the predictions, the models made far more predictions than needed. Given the nature of the evaluation metric, mAP, it was evident that properly aligning confidence scores could significantly improve the overall score.</p>\n<h3>Models</h3>\n<p>The following is a list of the models used:</p>\n<p>Instance Segmentation Models:</p>\n<ul>\n<li>Mask2Former</li>\n<li>MaskRCNN + various ConvNext v1 and v2</li>\n<li>Cascade RCNN + various ConvNext v1 and v2</li>\n<li>MaskRCNN + Swin Transformer<br>\nOther: MaskDino doesn’t work well.</li>\n</ul>\n<p>Object Detection Models:</p>\n<ul>\n<li>DINO</li>\n<li>Mask2Former (used only for object detection part)</li>\n<li>Diffusion Det</li>\n<li>ViT Det<br>\nOther: yolox doesn’t work well</li>\n</ul>\n<h3>Major Flaw:</h3>\n<p>I combined the outputs of the mask models using addition during post-processing. This resulted in improvements in ds1 (average IoU) and the public cross-validation (CV) score. However, as many of you may have noticed, when applying post-processing, the score on the private dataset degrades significantly. This indicates that the ds1 and private datasets have entirely different annotation tendencies. The annotations provided for ds1 were quite rough, and the dilation technique seemed effective due to this. One of the hosts mentioned that annotations were done in the same manner for both public and private datasets, but this information turned out to be inaccurate.</p>\n<h3>Single modification to achieve private 0.575:</h3>\n<p>Anticipating that the annotations for the private dataset might be more substantial, I submitted two versions: one with post-processing and another without. However, as mentioned earlier, during the ensemble mask generation, I was merely adding the masks instead of using averaging or voting. This method worked well for ds1 and the public dataset. On the other hand, I should have easily foreseen that if the private dataset had accurate and precise annotations, this approach could significantly reduce the score (which I overlooked).<br>\nTherefore, instead of adding the masks, I changed to using voting. This simple change dramatically improved the score from 0.397 to 0.575.</p>\n<div>\n  <img src=\"https://pbs.twimg.com/media/F2avc0eaEAEmdvL.jpg\" alt=\"score\">\n</div>\n<h3>Reflections:</h3>\n<p>One of the points I regret is being content with just having two submissions: one with post-processing and one without. Since I could have diversified the risks, I should have prepared submissions that could adapt to both small and accurate annotation masks.</p>",
      "rawMarkdown": "Despite achieving 9th place in the public leaderboard, I experienced a frightening shakeup, ending up at 295th place in the private leaderboard. However, many of the techniques I used proved effective in the private leaderboard, enabling me to achieve a score of 0.575 (equivalent to 4th place) by making just one modification to the 295th place code. Below, I will share the methods used and my reflections.\n\n### Method Overview:\nI primarily utilized \"mmdet\" and employed ensemble learning with numerous models. I had hoped to use YOLO v8 and v7 as well, but the host did not disclose their availability, so I had to abandon the idea.\n\nFor the ensemble, I made modifications to the \"WBF\" (Weighted Box Fusion) to use it with masks. Additionally, the major innovation was not limited to instance segmentation models but also involved using object detection models. I performed WBF based on bounding boxes and made mask predictions solely using the instance segmentation models. I used the object detection models only to calculate confidence scores. As you can see from the predictions, the models made far more predictions than needed. Given the nature of the evaluation metric, mAP, it was evident that properly aligning confidence scores could significantly improve the overall score.\n\n### Models\nThe following is a list of the models used:\n\nInstance Segmentation Models:\n* Mask2Former\n* MaskRCNN + various ConvNext v1 and v2\n* Cascade RCNN + various ConvNext v1 and v2\n* MaskRCNN + Swin Transformer\nOther: MaskDino doesn’t work well.\n\nObject Detection Models:\n* DINO\n* Mask2Former (used only for object detection part)\n* Diffusion Det\n* ViT Det\nOther: yolox doesn’t work well\n\n### Major Flaw:\nI combined the outputs of the mask models using addition during post-processing. This resulted in improvements in ds1 (average IoU) and the public cross-validation (CV) score. However, as many of you may have noticed, when applying post-processing, the score on the private dataset degrades significantly. This indicates that the ds1 and private datasets have entirely different annotation tendencies. The annotations provided for ds1 were quite rough, and the dilation technique seemed effective due to this. One of the hosts mentioned that annotations were done in the same manner for both public and private datasets, but this information turned out to be inaccurate.\n\n### Single modification to achieve private 0.575:\nAnticipating that the annotations for the private dataset might be more substantial, I submitted two versions: one with post-processing and another without. However, as mentioned earlier, during the ensemble mask generation, I was merely adding the masks instead of using averaging or voting. This method worked well for ds1 and the public dataset. On the other hand, I should have easily foreseen that if the private dataset had accurate and precise annotations, this approach could significantly reduce the score (which I overlooked).\nTherefore, instead of adding the masks, I changed to using voting. This simple change dramatically improved the score from 0.397 to 0.575.\n\n<div style=\"text-align: center;\">\n  <img src=\"https://pbs.twimg.com/media/F2avc0eaEAEmdvL.jpg\" style=\"max-width: 450px; display: block; margin: 0 auto;\" alt=\"score\">\n</div>\n\n### Reflections:\nOne of the points I regret is being content with just having two submissions: one with post-processing and one without. Since I could have diversified the risks, I should have prepared submissions that could adapt to both small and accurate annotation masks.",
      "votes": null
    },
    {
      "id": "2368652",
      "postDate": "08/01/2023 09:48:56",
      "content": "<p>thanks for sharing!! 　<br>\nI could not get a good local cv with mask2former. can you tell me what config you used to learn mask2former?</p>",
      "rawMarkdown": "thanks for sharing!! 　\nI could not get a good local cv with mask2former. can you tell me what config you used to learn mask2former?",
      "votes": null
    },
    {
      "id": "2368702",
      "postDate": "08/01/2023 10:21:55",
      "content": "<p>Of course it's OK!</p>\n<p>However, this was written by me when I was a mmdetection beginner, so there are many mistakes. For example, the size of the inference is improperly 1333x800 because it is taken from the base, it should be 1024x1024.<br>\nThe scheduler is also unchanged.</p>\n<p>However, the basic setting is this one, so I think it will give some good results!</p>\n<p><a href=\"https://www.kaggle.com/code/yosukeyama/hubmap23-mmdet-train-mask2former-1024/notebook\" target=\"_blank\">https://www.kaggle.com/code/yosukeyama/hubmap23-mmdet-train-mask2former-1024/notebook</a></p>\n<p>EDIT: I have checked past scores. Using this training note, even a single-fold and improperly size inference can earn a silver medal.</p>",
      "rawMarkdown": "Of course it's OK!\n\nHowever, this was written by me when I was a mmdetection beginner, so there are many mistakes. For example, the size of the inference is improperly 1333x800 because it is taken from the base, it should be 1024x1024.\nThe scheduler is also unchanged.\n\nHowever, the basic setting is this one, so I think it will give some good results!\n\nhttps://www.kaggle.com/code/yosukeyama/hubmap23-mmdet-train-mask2former-1024/notebook\n\nEDIT: I have checked past scores. Using this training note, even a single-fold and improperly size inference can earn a silver medal.",
      "votes": null
    },
    {
      "id": "2368855",
      "postDate": "08/01/2023 12:21:42",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.",
      "votes": null
    },
    {
      "id": "2370433",
      "postDate": "08/02/2023 12:04:59",
      "content": "<p>Thank you, <a href=\"https://www.kaggle.com/yosukeyama\" target=\"_blank\">@yosukeyama</a> for the report! This is quite an elegant solution with a very strong private result after switching to voting. Btw what voting method did you use in the corrected ensembling of masks?</p>\n<ul>\n<li>np.median(masks &gt; threshold, axis=0) ?  #&nbsp; majority after thresholding</li>\n<li>np.mean(masks, axis=0) &gt; threshold ?                # or thresholding after averaging?</li>\n</ul>",
      "rawMarkdown": "Thank you, @yosukeyama for the report! This is quite an elegant solution with a very strong private result after switching to voting. Btw what voting method did you use in the corrected ensembling of masks?\n- np.median(masks > threshold, axis=0) ?  #  majority after thresholding\n- np.mean(masks, axis=0) > threshold ?                # or thresholding after averaging?",
      "votes": null
    },
    {
      "id": "2375070",
      "postDate": "08/05/2023 11:35:43",
      "content": "<p>Thank you for your question!</p>\n<p>The \"voting\" here is that once the mask is obtained in binary and the majority of the model predicts the pixel to be a binary mask. The process is called.<br>\nSo the former of the codes is my method.</p>",
      "rawMarkdown": "Thank you for your question!\n\nThe \"voting\" here is that once the mask is obtained in binary and the majority of the model predicts the pixel to be a binary mask. The process is called.\nSo the former of the codes is my method.",
      "votes": null
    },
    {
      "id": "2377088",
      "postDate": "08/07/2023 00:02:12",
      "content": "<p><a href=\"https://www.kaggle.com/yosukeyama\" target=\"_blank\">@yosukeyama</a> It was unfortunate to miss out 4th place.</p>\n<blockquote>\n  <p>For the ensemble, I made modifications to the \"WBF\" (Weighted Box Fusion) to use it with masks.</p>\n</blockquote>\n<p>Could you share the detail of WDF adaptation for mask prediction and CV strategy?</p>",
      "rawMarkdown": "yosukeyama It was unfortunate to miss out 4th place.\n\n> For the ensemble, I made modifications to the \"WBF\" (Weighted Box Fusion) to use it with masks.\n\nCould you share the detail of WDF adaptation for mask prediction and CV strategy?",
      "votes": null
    },
    {
      "id": "2377771",
      "postDate": "08/07/2023 10:17:57",
      "content": "<p>Thank you for your question!</p>\n<p>The base is utilizing the WBF library.<br>\nWhen clustering boxes, the boxes and scores are processed together, but in the original code, masks could not be used as inputs.<br>\nTherefore, when acquiring the combinations, I've made it so that masks are clustered along with them.<br>\nI thought it might be good to calculate the IoU using masks, but like WBF, I calculated the IoU using boxes.<br>\nUsually, a single model will output multiple overlapping boxes. Generally, after a model's output, techniques like NMS are used to leave only the high-scoring ones before performing WBF.<br>\nAt the time, I didn't think to apply NMS, so I decided to use the same number of top-scoring boxes as the number of models.<br>\nI believe this is essentially similar, but I think this way might result in lower confidence scores when less accurate masks are outputted in some cases. I have not been able to examine the superiority compared to using NMS. I would like to try it when I have time, but it seems like it will be difficult for a while.</p>\n<p>CV strategy. This is my biggest failure.<br>\nBasically, I divided the fold into six (although I intended to divide the folds into five, due to typo, it was divided into six…), and essentially calculated only with fold0.<br>\nI mixed up dataset1 and dataset2 in the CV, and I now think this was a significant mistake.<br>\nFurthermore, I also conducted studies with a small number of ds1 only in fold0.<br>\nThis was because I had much work privately, and I had to choose whether to experiment with fold0-5 or reduce the number of experiments.<br>\nIn retrospect, I should have prioritized robustness even if reducing experiments, carried out training with all folds in all experiments, and calculated the CV. <br>\nThat would have prevented this situation…</p>",
      "rawMarkdown": "Thank you for your question!\n\nThe base is utilizing the WBF library.\nWhen clustering boxes, the boxes and scores are processed together, but in the original code, masks could not be used as inputs.\nTherefore, when acquiring the combinations, I've made it so that masks are clustered along with them.\nI thought it might be good to calculate the IoU using masks, but like WBF, I calculated the IoU using boxes.\nUsually, a single model will output multiple overlapping boxes. Generally, after a model's output, techniques like NMS are used to leave only the high-scoring ones before performing WBF.\nAt the time, I didn't think to apply NMS, so I decided to use the same number of top-scoring boxes as the number of models.\nI believe this is essentially similar, but I think this way might result in lower confidence scores when less accurate masks are outputted in some cases. I have not been able to examine the superiority compared to using NMS. I would like to try it when I have time, but it seems like it will be difficult for a while.\n\nCV strategy. This is my biggest failure.\nBasically, I divided the fold into six (although I intended to divide the folds into five, due to typo, it was divided into six…), and essentially calculated only with fold0.\nI mixed up dataset1 and dataset2 in the CV, and I now think this was a significant mistake.\nFurthermore, I also conducted studies with a small number of ds1 only in fold0.\nThis was because I had much work privately, and I had to choose whether to experiment with fold0-5 or reduce the number of experiments.\nIn retrospect, I should have prioritized robustness even if reducing experiments, carried out training with all folds in all experiments, and calculated the CV. \nThat would have prevented this situation…",
      "votes": null
    },
    {
      "id": "2378019",
      "postDate": "08/07/2023 12:41:58",
      "content": "<p><a href=\"https://www.kaggle.com/yosukeyama\" target=\"_blank\">@yosukeyama</a> </p>\n<p>Thank you for additional explanation. Below is my understanding, please correct them if they are wrong.</p>\n<ol>\n<li>WBF is used as to suppress duplicated boxes instead of NMS, and hopefully that could avoid the over-confident boxes generated.</li>\n</ol>\n<p>(By the way, it reminds me max pooling vs mean pooling on CV tasks.)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fe6faf525d2ba296a54a327f93c2632a0%2FScreenshot%202023-08-07%20at%2021.25.14.png?generation=1691411164747665&amp;alt=media\" alt=\"\"></p>\n<ol>\n<li>For the second part, I am not sure if I understand correctly, but if you had problem with imbalanced statistics of DS1 and DS2 on your validation set, I think <code>StratifiedKFold</code> might mitigate the issue.</li>\n</ol>",
      "rawMarkdown": "yosukeyama \n\nThank you for additional explanation. Below is my understanding, please correct them if they are wrong.\n\n1. WBF is used as to suppress duplicated boxes instead of NMS, and hopefully that could avoid the over-confident boxes generated.\n\n(By the way, it reminds me max pooling vs mean pooling on CV tasks.)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fe6faf525d2ba296a54a327f93c2632a0%2FScreenshot%202023-08-07%20at%2021.25.14.png?generation=1691411164747665&alt=media)\n\n2. For the second part, I am not sure if I understand correctly, but if you had problem with imbalanced statistics of DS1 and DS2 on your validation set, I think `StratifiedKFold` might mitigate the issue.",
      "votes": null
    },
    {
      "id": "2378055",
      "postDate": "08/07/2023 13:08:25",
      "content": "<p>About WBF.<br>\nThis is the same as my understanding.<br>\nYes, it is similar to max and mean pooling lol.<br>\nHowever, mAP can be improved by reordering the forecasts correctly. Therefore, I think it is quite nonsense to use max to aggregate predictions. This is because the implication is the same as taking max when performing an ensemble of many models in image classification. In such cases, taking mean or median is the most basic and robust method, isn't it?<br>\nIn object detection, as in image classification, it seems to be a good idea to calculate the confidence score using mean or weighted mean.</p>\n<p>You are right about the second. <br>\nFurthermore, we were informed in advance that only ds1 was used in the test data, and as ds1 and ds2 are annotated differently, it was more efficient to calculate CVs according to ds1 only. So I should have spent more time considering how to fit ds1, as in the top solutions…</p>",
      "rawMarkdown": "About WBF.\nThis is the same as my understanding.\nYes, it is similar to max and mean pooling lol.\nHowever, mAP can be improved by reordering the forecasts correctly. Therefore, I think it is quite nonsense to use max to aggregate predictions. This is because the implication is the same as taking max when performing an ensemble of many models in image classification. In such cases, taking mean or median is the most basic and robust method, isn't it?\nIn object detection, as in image classification, it seems to be a good idea to calculate the confidence score using mean or weighted mean.\n\nYou are right about the second. \nFurthermore, we were informed in advance that only ds1 was used in the test data, and as ds1 and ds2 are annotated differently, it was more efficient to calculate CVs according to ds1 only. So I should have spent more time considering how to fit ds1, as in the top solutions...",
      "votes": null
    },
    {
      "id": "2378096",
      "postDate": "08/07/2023 13:26:56",
      "content": "<blockquote>\n  <p>In such cases, taking mean or median is the most basic and robust method, isn't it?</p>\n</blockquote>\n<p>Yeah, I agree with that for the ensemble method of different models' predictions.</p>\n<p></p>\n<p>Sorry, I was misunderstood. I misunderstood NMS is applied before calculating loss. As NMS is applied as a post process, it would surely tend to generate over-confident boxes. So in the situation where accurate confidence value is required, NMS could be bad choice.</p>",
      "rawMarkdown": "> In such cases, taking mean or median is the most basic and robust method, isn't it?\n\nYeah, I agree with that for the ensemble method of different models' predictions.\n\n~~However, I am not sure if WBF is always effective as a suppression method for duplicated boxes from a single model. As global max pooling is successful on CV tasks, I believe NMS also could learn correct confidence values from the statistics of the dataset. I think try and error is required per dataset/tasks which aggregation strategy is optimal.~~\n\nSorry, I was misunderstood. I misunderstood NMS is applied before calculating loss. As NMS is applied as a post process, it would surely tend to generate over-confident boxes. So in the situation where accurate confidence value is required, NMS could be bad choice.",
      "votes": null
    },
    {
      "id": "2378939",
      "postDate": "08/07/2023 21:44:28",
      "content": "<p>After looking at the revised comments, I understood what you meant. Indeed, if we assume the process before loss calculation, it becomes almost the same argument as max-pooling, and I don't know which one is more suitable.</p>\n<p>That's right, in cases where mAP is used as an evaluation metric (requring confidence scores), one can say that WBF is a superior process to NMS. If we aggregate scores from multiple models like stacking and use gradient boosting, we might be able to predict even better scores, but the process becomes too complicated. So calculating mean values is reasonable.</p>",
      "rawMarkdown": "After looking at the revised comments, I understood what you meant. Indeed, if we assume the process before loss calculation, it becomes almost the same argument as max-pooling, and I don't know which one is more suitable.\n\nThat's right, in cases where mAP is used as an evaluation metric (requring confidence scores), one can say that WBF is a superior process to NMS. If we aggregate scores from multiple models like stacking and use gradient boosting, we might be able to predict even better scores, but the process becomes too complicated. So calculating mean values is reasonable.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2368652,
      "author_name": "abebe9849",
      "author_url": "",
      "post_date": "08/01/2023 09:48:56",
      "content": "<p>thanks for sharing!! 　<br>\nI could not get a good local cv with mask2former. can you tell me what config you used to learn mask2former?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2368702,
          "author_name": "yosukeyama",
          "author_url": "",
          "post_date": "08/01/2023 10:21:55",
          "content": "<p>Of course it's OK!</p>\n<p>However, this was written by me when I was a mmdetection beginner, so there are many mistakes. For example, the size of the inference is improperly 1333x800 because it is taken from the base, it should be 1024x1024.<br>\nThe scheduler is also unchanged.</p>\n<p>However, the basic setting is this one, so I think it will give some good results!</p>\n<p><a href=\"https://www.kaggle.com/code/yosukeyama/hubmap23-mmdet-train-mask2former-1024/notebook\" target=\"_blank\">https://www.kaggle.com/code/yosukeyama/hubmap23-mmdet-train-mask2former-1024/notebook</a></p>\n<p>EDIT: I have checked past scores. Using this training note, even a single-fold and improperly size inference can earn a silver medal.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2368855,
      "author_name": "shigengtian",
      "author_url": "",
      "post_date": "08/01/2023 12:21:42",
      "content": "<p>Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2370433,
      "author_name": "vslaykovsky",
      "author_url": "",
      "post_date": "08/02/2023 12:04:59",
      "content": "<p>Thank you, <a href=\"https://www.kaggle.com/yosukeyama\" target=\"_blank\">@yosukeyama</a> for the report! This is quite an elegant solution with a very strong private result after switching to voting. Btw what voting method did you use in the corrected ensembling of masks?</p>\n<ul>\n<li>np.median(masks &gt; threshold, axis=0) ?  #&nbsp; majority after thresholding</li>\n<li>np.mean(masks, axis=0) &gt; threshold ?                # or thresholding after averaging?</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 2375070,
          "author_name": "yosukeyama",
          "author_url": "",
          "post_date": "08/05/2023 11:35:43",
          "content": "<p>Thank you for your question!</p>\n<p>The \"voting\" here is that once the mask is obtained in binary and the majority of the model predicts the pixel to be a binary mask. The process is called.<br>\nSo the former of the codes is my method.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2377088,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "08/07/2023 00:02:12",
      "content": "<p><a href=\"https://www.kaggle.com/yosukeyama\" target=\"_blank\">@yosukeyama</a> It was unfortunate to miss out 4th place.</p>\n<blockquote>\n  <p>For the ensemble, I made modifications to the \"WBF\" (Weighted Box Fusion) to use it with masks.</p>\n</blockquote>\n<p>Could you share the detail of WDF adaptation for mask prediction and CV strategy?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2377771,
          "author_name": "yosukeyama",
          "author_url": "",
          "post_date": "08/07/2023 10:17:57",
          "content": "<p>Thank you for your question!</p>\n<p>The base is utilizing the WBF library.<br>\nWhen clustering boxes, the boxes and scores are processed together, but in the original code, masks could not be used as inputs.<br>\nTherefore, when acquiring the combinations, I've made it so that masks are clustered along with them.<br>\nI thought it might be good to calculate the IoU using masks, but like WBF, I calculated the IoU using boxes.<br>\nUsually, a single model will output multiple overlapping boxes. Generally, after a model's output, techniques like NMS are used to leave only the high-scoring ones before performing WBF.<br>\nAt the time, I didn't think to apply NMS, so I decided to use the same number of top-scoring boxes as the number of models.<br>\nI believe this is essentially similar, but I think this way might result in lower confidence scores when less accurate masks are outputted in some cases. I have not been able to examine the superiority compared to using NMS. I would like to try it when I have time, but it seems like it will be difficult for a while.</p>\n<p>CV strategy. This is my biggest failure.<br>\nBasically, I divided the fold into six (although I intended to divide the folds into five, due to typo, it was divided into six…), and essentially calculated only with fold0.<br>\nI mixed up dataset1 and dataset2 in the CV, and I now think this was a significant mistake.<br>\nFurthermore, I also conducted studies with a small number of ds1 only in fold0.<br>\nThis was because I had much work privately, and I had to choose whether to experiment with fold0-5 or reduce the number of experiments.<br>\nIn retrospect, I should have prioritized robustness even if reducing experiments, carried out training with all folds in all experiments, and calculated the CV. <br>\nThat would have prevented this situation…</p>",
          "votes": null,
          "replies": [
            {
              "id": 2378019,
              "author_name": "tatamikenn",
              "author_url": "",
              "post_date": "08/07/2023 12:41:58",
              "content": "<p><a href=\"https://www.kaggle.com/yosukeyama\" target=\"_blank\">@yosukeyama</a> </p>\n<p>Thank you for additional explanation. Below is my understanding, please correct them if they are wrong.</p>\n<ol>\n<li>WBF is used as to suppress duplicated boxes instead of NMS, and hopefully that could avoid the over-confident boxes generated.</li>\n</ol>\n<p>(By the way, it reminds me max pooling vs mean pooling on CV tasks.)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fe6faf525d2ba296a54a327f93c2632a0%2FScreenshot%202023-08-07%20at%2021.25.14.png?generation=1691411164747665&amp;alt=media\" alt=\"\"></p>\n<ol>\n<li>For the second part, I am not sure if I understand correctly, but if you had problem with imbalanced statistics of DS1 and DS2 on your validation set, I think <code>StratifiedKFold</code> might mitigate the issue.</li>\n</ol>",
              "votes": null,
              "replies": [
                {
                  "id": 2378055,
                  "author_name": "yosukeyama",
                  "author_url": "",
                  "post_date": "08/07/2023 13:08:25",
                  "content": "<p>About WBF.<br>\nThis is the same as my understanding.<br>\nYes, it is similar to max and mean pooling lol.<br>\nHowever, mAP can be improved by reordering the forecasts correctly. Therefore, I think it is quite nonsense to use max to aggregate predictions. This is because the implication is the same as taking max when performing an ensemble of many models in image classification. In such cases, taking mean or median is the most basic and robust method, isn't it?<br>\nIn object detection, as in image classification, it seems to be a good idea to calculate the confidence score using mean or weighted mean.</p>\n<p>You are right about the second. <br>\nFurthermore, we were informed in advance that only ds1 was used in the test data, and as ds1 and ds2 are annotated differently, it was more efficient to calculate CVs according to ds1 only. So I should have spent more time considering how to fit ds1, as in the top solutions…</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2378096,
                      "author_name": "tatamikenn",
                      "author_url": "",
                      "post_date": "08/07/2023 13:26:56",
                      "content": "<blockquote>\n  <p>In such cases, taking mean or median is the most basic and robust method, isn't it?</p>\n</blockquote>\n<p>Yeah, I agree with that for the ensemble method of different models' predictions.</p>\n<p></p>\n<p>Sorry, I was misunderstood. I misunderstood NMS is applied before calculating loss. As NMS is applied as a post process, it would surely tend to generate over-confident boxes. So in the situation where accurate confidence value is required, NMS could be bad choice.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2378939,
                          "author_name": "yosukeyama",
                          "author_url": "",
                          "post_date": "08/07/2023 21:44:28",
                          "content": "<p>After looking at the revised comments, I understood what you meant. Indeed, if we assume the process before loss calculation, it becomes almost the same argument as max-pooling, and I don't know which one is more suitable.</p>\n<p>That's right, in cases where mAP is used as an evaluation metric (requring confidence scores), one can say that WBF is a superior process to NMS. If we aggregate scores from multiple models like stacking and use gradient boosting, we might be able to predict even better scores, but the process becomes too complicated. So calculating mean values is reasonable.</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2368561": "Despite achieving 9th place in the public leaderboard, I experienced a frightening shakeup, ending up at 295th place in the private leaderboard. However, many of the techniques I used proved effective in the private leaderboard, enabling me to achieve a score of 0.575 (equivalent to 4th place) by making just one modification to the 295th place code. Below, I will share the methods used and my reflections.\n\n### Method Overview:\nI primarily utilized \"mmdet\" and employed ensemble learning with numerous models. I had hoped to use YOLO v8 and v7 as well, but the host did not disclose their availability, so I had to abandon the idea.\n\nFor the ensemble, I made modifications to the \"WBF\" (Weighted Box Fusion) to use it with masks. Additionally, the major innovation was not limited to instance segmentation models but also involved using object detection models. I performed WBF based on bounding boxes and made mask predictions solely using the instance segmentation models. I used the object detection models only to calculate confidence scores. As you can see from the predictions, the models made far more predictions than needed. Given the nature of the evaluation metric, mAP, it was evident that properly aligning confidence scores could significantly improve the overall score.\n\n### Models\nThe following is a list of the models used:\n\nInstance Segmentation Models:\n* Mask2Former\n* MaskRCNN + various ConvNext v1 and v2\n* Cascade RCNN + various ConvNext v1 and v2\n* MaskRCNN + Swin Transformer\nOther: MaskDino doesn’t work well.\n\nObject Detection Models:\n* DINO\n* Mask2Former (used only for object detection part)\n* Diffusion Det\n* ViT Det\nOther: yolox doesn’t work well\n\n### Major Flaw:\nI combined the outputs of the mask models using addition during post-processing. This resulted in improvements in ds1 (average IoU) and the public cross-validation (CV) score. However, as many of you may have noticed, when applying post-processing, the score on the private dataset degrades significantly. This indicates that the ds1 and private datasets have entirely different annotation tendencies. The annotations provided for ds1 were quite rough, and the dilation technique seemed effective due to this. One of the hosts mentioned that annotations were done in the same manner for both public and private datasets, but this information turned out to be inaccurate.\n\n### Single modification to achieve private 0.575:\nAnticipating that the annotations for the private dataset might be more substantial, I submitted two versions: one with post-processing and another without. However, as mentioned earlier, during the ensemble mask generation, I was merely adding the masks instead of using averaging or voting. This method worked well for ds1 and the public dataset. On the other hand, I should have easily foreseen that if the private dataset had accurate and precise annotations, this approach could significantly reduce the score (which I overlooked).\nTherefore, instead of adding the masks, I changed to using voting. This simple change dramatically improved the score from 0.397 to 0.575.\n\n<div style=\"text-align: center;\">\n  <img src=\"https://pbs.twimg.com/media/F2avc0eaEAEmdvL.jpg\" style=\"max-width: 450px; display: block; margin: 0 auto;\" alt=\"score\">\n</div>\n\n### Reflections:\nOne of the points I regret is being content with just having two submissions: one with post-processing and one without. Since I could have diversified the risks, I should have prepared submissions that could adapt to both small and accurate annotation masks.",
    "2368652": "thanks for sharing!! 　\nI could not get a good local cv with mask2former. can you tell me what config you used to learn mask2former?",
    "2368702": "Of course it's OK!\n\nHowever, this was written by me when I was a mmdetection beginner, so there are many mistakes. For example, the size of the inference is improperly 1333x800 because it is taken from the base, it should be 1024x1024.\nThe scheduler is also unchanged.\n\nHowever, the basic setting is this one, so I think it will give some good results!\n\nhttps://www.kaggle.com/code/yosukeyama/hubmap23-mmdet-train-mask2former-1024/notebook\n\nEDIT: I have checked past scores. Using this training note, even a single-fold and improperly size inference can earn a silver medal.",
    "2368855": "Thanks for sharing.",
    "2370433": "Thank you, @yosukeyama for the report! This is quite an elegant solution with a very strong private result after switching to voting. Btw what voting method did you use in the corrected ensembling of masks?\n- np.median(masks > threshold, axis=0) ?  #  majority after thresholding\n- np.mean(masks, axis=0) > threshold ?                # or thresholding after averaging?",
    "2375070": "Thank you for your question!\n\nThe \"voting\" here is that once the mask is obtained in binary and the majority of the model predicts the pixel to be a binary mask. The process is called.\nSo the former of the codes is my method.",
    "2377088": "yosukeyama It was unfortunate to miss out 4th place.\n\n> For the ensemble, I made modifications to the \"WBF\" (Weighted Box Fusion) to use it with masks.\n\nCould you share the detail of WDF adaptation for mask prediction and CV strategy?",
    "2377771": "Thank you for your question!\n\nThe base is utilizing the WBF library.\nWhen clustering boxes, the boxes and scores are processed together, but in the original code, masks could not be used as inputs.\nTherefore, when acquiring the combinations, I've made it so that masks are clustered along with them.\nI thought it might be good to calculate the IoU using masks, but like WBF, I calculated the IoU using boxes.\nUsually, a single model will output multiple overlapping boxes. Generally, after a model's output, techniques like NMS are used to leave only the high-scoring ones before performing WBF.\nAt the time, I didn't think to apply NMS, so I decided to use the same number of top-scoring boxes as the number of models.\nI believe this is essentially similar, but I think this way might result in lower confidence scores when less accurate masks are outputted in some cases. I have not been able to examine the superiority compared to using NMS. I would like to try it when I have time, but it seems like it will be difficult for a while.\n\nCV strategy. This is my biggest failure.\nBasically, I divided the fold into six (although I intended to divide the folds into five, due to typo, it was divided into six…), and essentially calculated only with fold0.\nI mixed up dataset1 and dataset2 in the CV, and I now think this was a significant mistake.\nFurthermore, I also conducted studies with a small number of ds1 only in fold0.\nThis was because I had much work privately, and I had to choose whether to experiment with fold0-5 or reduce the number of experiments.\nIn retrospect, I should have prioritized robustness even if reducing experiments, carried out training with all folds in all experiments, and calculated the CV. \nThat would have prevented this situation…",
    "2378019": "yosukeyama \n\nThank you for additional explanation. Below is my understanding, please correct them if they are wrong.\n\n1. WBF is used as to suppress duplicated boxes instead of NMS, and hopefully that could avoid the over-confident boxes generated.\n\n(By the way, it reminds me max pooling vs mean pooling on CV tasks.)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fe6faf525d2ba296a54a327f93c2632a0%2FScreenshot%202023-08-07%20at%2021.25.14.png?generation=1691411164747665&alt=media)\n\n2. For the second part, I am not sure if I understand correctly, but if you had problem with imbalanced statistics of DS1 and DS2 on your validation set, I think `StratifiedKFold` might mitigate the issue.",
    "2378055": "About WBF.\nThis is the same as my understanding.\nYes, it is similar to max and mean pooling lol.\nHowever, mAP can be improved by reordering the forecasts correctly. Therefore, I think it is quite nonsense to use max to aggregate predictions. This is because the implication is the same as taking max when performing an ensemble of many models in image classification. In such cases, taking mean or median is the most basic and robust method, isn't it?\nIn object detection, as in image classification, it seems to be a good idea to calculate the confidence score using mean or weighted mean.\n\nYou are right about the second. \nFurthermore, we were informed in advance that only ds1 was used in the test data, and as ds1 and ds2 are annotated differently, it was more efficient to calculate CVs according to ds1 only. So I should have spent more time considering how to fit ds1, as in the top solutions...",
    "2378096": "> In such cases, taking mean or median is the most basic and robust method, isn't it?\n\nYeah, I agree with that for the ensemble method of different models' predictions.\n\n~~However, I am not sure if WBF is always effective as a suppression method for duplicated boxes from a single model. As global max pooling is successful on CV tasks, I believe NMS also could learn correct confidence values from the statistics of the dataset. I think try and error is required per dataset/tasks which aggregation strategy is optimal.~~\n\nSorry, I was misunderstood. I misunderstood NMS is applied before calculating loss. As NMS is applied as a post process, it would surely tend to generate over-confident boxes. So in the situation where accurate confidence value is required, NMS could be bad choice.",
    "2378939": "After looking at the revised comments, I understood what you meant. Indeed, if we assume the process before loss calculation, it becomes almost the same argument as max-pooling, and I don't know which one is more suitable.\n\nThat's right, in cases where mAP is used as an evaluation metric (requring confidence scores), one can say that WBF is a superior process to NMS. If we aggregate scores from multiple models like stacking and use gradient boosting, we might be able to predict even better scores, but the process becomes too complicated. So calculating mean values is reasonable."
  },
  "source": "meta"
}