{
  "id": 238024,
  "title": "4th place solution",
  "url": "/competitions/hubmap-kidney-segmentation/writeups/deeplive-exe-4th-place-solution",
  "author_name": "",
  "post_date": "2021-05-12T13:45:49.570Z",
  "votes": 71,
  "comment_count": 35,
  "views": 0,
  "content": "<p><strong>Please consider the following links:</strong> <a href=\"https://www.kaggle.com/theoviel/hubmap-final-methodology-submission\" target=\"_blank\"><strong>writeup</strong></a>, <a href=\"https://github.com/Optimox/HubMap\" target=\"_blank\"><strong>code</strong></a>. Below I provide just a short description.</p>\n<h3>The key:</h3>\n<ul>\n<li>proper consideration of healthy and unhealthy glomeruli and trust in CV</li>\n<li>2 class approach</li>\n<li>proper annotation of the data</li>\n<li>be careful with d488c759a when check LB</li>\n</ul>\n<h3>Introduction</h3>\n<p>Our team would like to thank organizers and kaggle for making this competition possible. And a really big thanks to my amazing teammates <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> and <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> for their incredible efforts in working on the competition and putting everything into writeup for the judge prize. Also <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> received his first gold medal in this competition, big congratulation to him! Finally, congratulations to all participants of this challenge.</p>\n<p>It was quite a long journey for some of us, since the competition started back in 2020. There were a number of challenges we needed to overcome, and there were many learning opportunities. For me, as well as for some participants, it was the opportunity to <strong>learn labeling medical data</strong>. It may sound funny, but doing this job I realized the amount of work organizers put into this challenge.</p>\n<h3>Our approach</h3>\n<p><strong>In core of our approach is proper treatment of healthy and unhealthy glomeruli.</strong><br>\nIn short, we quickly realized that d488c759a with FC annotated is destabilizing LB and creating a gap between CV and LB. Use of external data with <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/208972\" target=\"_blank\">this annotation</a> improved LB but did not affect or even degraded CV. The explanation is simple: it contained annotation of both healthy and unhealthy glomeruli (at least large portion of them), while the train set contained only annotation of healthy glomeruli. So <strong>how can we make a model to predict FC: train it on images with FC… But it was not a right thing</strong>: organizers were looking for healthy glomeruli only, and all images except d488c759a did not include FC into annotation. So, use the external data with the shared hand labeling was dangerous, and some people, I may expect, got affected by shake up at the end of the competition because of this.</p>\n<p><strong>Our strategy was proper hand annotation of</strong> <a href=\"https://data.mendeley.com/datasets/k7nvtgn2x6/3\" target=\"_blank\"><strong>this external dataset</strong></a> <strong>with 2 classes</strong>: healthy and unhealthy glomeruli. There are many kinds of unhealthy glomeruli, but we decided to keep the things simple, combining them into a single class. We also hand annotated unhealthy glomeruli class in test images, which combined with PL for healthy glomeruli, were used in our training pipeline. We also used <a href=\"https://zenodo.org/record/4299694\" target=\"_blank\">zondo dataset</a>, which naturally includes two classes. However, there was quite a number of missed glomeruli, especially at the image edges, and we hand annotated them. Finally, we used 2 images from HuBMAP portal, not included into train/test datasets with PL (both healthy and unhealthy) generated by our model. Using of multiple data sources helped us to ensure that the model is able to generalize well on new cases. We also used a combined human+AI guided method to add missing annotation to the train data (there were just a few instances tough). </p>\n<p><strong>Secret sauce: ignore d488c759a from LB evaluation.</strong> LB with fixed predictions on this image correlated quite well with CV</p>\n<p>From the model perspective of view, there is nothing really special, EfficientNet, Unet… Though, we tried more advanced things, like BoT middle layer, Swin transformer backbone, FPN skip connections, but it was not giving any noticeable benefit. This competition really was not about it. The thing that helped was using 2 classes during training. The level of injury of glomeruli is on a continuous scale. So instead of confusing the model by providing nearly identical glomeruli included/excluded from masks, we made the model to learn that unhealthy glomeruli are also glomeruli, and at the end model decides whether a particular instance is healthy enough. To be honest, this competition was not really about segmentation but rather about classification. The uncertainty of including a particular glomerulus in the mask predicted by the model is the major factor degrading the model performance.</p>\n<p>Our best submission (ensemble of two models) got 0.9439 CV, 0.941 LB (with including hand annotation of unhealthy glomeruli for d488c759a image), and 0.950 private LB. There were some technical difficulties with making submissions on the last day, so we could not make a submission with all our final models(</p>\n<h4>Links: <a href=\"https://www.kaggle.com/theoviel/hubmap-final-methodology-submission\" target=\"_blank\"><strong>writeup</strong></a> &lt;- please take a look (it provides additional details and results of the analysis we made), <a href=\"https://github.com/Optimox/HubMap\" target=\"_blank\"><strong>code</strong></a></h4>",
  "messages": [
    {
      "id": "1301236",
      "postDate": "05/11/2021 01:52:43",
      "content": "<p><strong>Please consider the following links:</strong> <a href=\"https://www.kaggle.com/theoviel/hubmap-final-methodology-submission\" target=\"_blank\"><strong>writeup</strong></a>, <a href=\"https://github.com/Optimox/HubMap\" target=\"_blank\"><strong>code</strong></a>. Below I provide just a short description.</p>\n<h3>The key:</h3>\n<ul>\n<li>proper consideration of healthy and unhealthy glomeruli and trust in CV</li>\n<li>2 class approach</li>\n<li>proper annotation of the data</li>\n<li>be careful with d488c759a when check LB</li>\n</ul>\n<h3>Introduction</h3>\n<p>Our team would like to thank organizers and kaggle for making this competition possible. And a really big thanks to my amazing teammates <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> and <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> for their incredible efforts in working on the competition and putting everything into writeup for the judge prize. Also <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> received his first gold medal in this competition, big congratulation to him! Finally, congratulations to all participants of this challenge.</p>\n<p>It was quite a long journey for some of us, since the competition started back in 2020. There were a number of challenges we needed to overcome, and there were many learning opportunities. For me, as well as for some participants, it was the opportunity to <strong>learn labeling medical data</strong>. It may sound funny, but doing this job I realized the amount of work organizers put into this challenge.</p>\n<h3>Our approach</h3>\n<p><strong>In core of our approach is proper treatment of healthy and unhealthy glomeruli.</strong><br>\nIn short, we quickly realized that d488c759a with FC annotated is destabilizing LB and creating a gap between CV and LB. Use of external data with <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/208972\" target=\"_blank\">this annotation</a> improved LB but did not affect or even degraded CV. The explanation is simple: it contained annotation of both healthy and unhealthy glomeruli (at least large portion of them), while the train set contained only annotation of healthy glomeruli. So <strong>how can we make a model to predict FC: train it on images with FC… But it was not a right thing</strong>: organizers were looking for healthy glomeruli only, and all images except d488c759a did not include FC into annotation. So, use the external data with the shared hand labeling was dangerous, and some people, I may expect, got affected by shake up at the end of the competition because of this.</p>\n<p><strong>Our strategy was proper hand annotation of</strong> <a href=\"https://data.mendeley.com/datasets/k7nvtgn2x6/3\" target=\"_blank\"><strong>this external dataset</strong></a> <strong>with 2 classes</strong>: healthy and unhealthy glomeruli. There are many kinds of unhealthy glomeruli, but we decided to keep the things simple, combining them into a single class. We also hand annotated unhealthy glomeruli class in test images, which combined with PL for healthy glomeruli, were used in our training pipeline. We also used <a href=\"https://zenodo.org/record/4299694\" target=\"_blank\">zondo dataset</a>, which naturally includes two classes. However, there was quite a number of missed glomeruli, especially at the image edges, and we hand annotated them. Finally, we used 2 images from HuBMAP portal, not included into train/test datasets with PL (both healthy and unhealthy) generated by our model. Using of multiple data sources helped us to ensure that the model is able to generalize well on new cases. We also used a combined human+AI guided method to add missing annotation to the train data (there were just a few instances tough). </p>\n<p><strong>Secret sauce: ignore d488c759a from LB evaluation.</strong> LB with fixed predictions on this image correlated quite well with CV</p>\n<p>From the model perspective of view, there is nothing really special, EfficientNet, Unet… Though, we tried more advanced things, like BoT middle layer, Swin transformer backbone, FPN skip connections, but it was not giving any noticeable benefit. This competition really was not about it. The thing that helped was using 2 classes during training. The level of injury of glomeruli is on a continuous scale. So instead of confusing the model by providing nearly identical glomeruli included/excluded from masks, we made the model to learn that unhealthy glomeruli are also glomeruli, and at the end model decides whether a particular instance is healthy enough. To be honest, this competition was not really about segmentation but rather about classification. The uncertainty of including a particular glomerulus in the mask predicted by the model is the major factor degrading the model performance.</p>\n<p>Our best submission (ensemble of two models) got 0.9439 CV, 0.941 LB (with including hand annotation of unhealthy glomeruli for d488c759a image), and 0.950 private LB. There were some technical difficulties with making submissions on the last day, so we could not make a submission with all our final models(</p>\n<h4>Links: <a href=\"https://www.kaggle.com/theoviel/hubmap-final-methodology-submission\" target=\"_blank\"><strong>writeup</strong></a> &lt;- please take a look (it provides additional details and results of the analysis we made), <a href=\"https://github.com/Optimox/HubMap\" target=\"_blank\"><strong>code</strong></a></h4>",
      "rawMarkdown": "**Please consider the following links:** [**writeup**](https://www.kaggle.com/theoviel/hubmap-final-methodology-submission), [**code**] (https://github.com/Optimox/HubMap). Below I provide just a short description.\n\n### The key: \n- proper consideration of healthy and unhealthy glomeruli and trust in CV\n- 2 class approach\n- proper annotation of the data\n- be careful with d488c759a when check LB\n\n### Introduction \nOur team would like to thank organizers and kaggle for making this competition possible. And a really big thanks to my amazing teammates @theoviel and @optimo for their incredible efforts in working on the competition and putting everything into writeup for the judge prize. Also @optimo received his first gold medal in this competition, big congratulation to him! Finally, congratulations to all participants of this challenge.\n\nIt was quite a long journey for some of us, since the competition started back in 2020. There were a number of challenges we needed to overcome, and there were many learning opportunities. For me, as well as for some participants, it was the opportunity to **learn labeling medical data**. It may sound funny, but doing this job I realized the amount of work organizers put into this challenge.\n\n### Our approach\n**In core of our approach is proper treatment of healthy and unhealthy glomeruli.**\nIn short, we quickly realized that d488c759a with FC annotated is destabilizing LB and creating a gap between CV and LB. Use of external data with [this annotation](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/208972) improved LB but did not affect or even degraded CV. The explanation is simple: it contained annotation of both healthy and unhealthy glomeruli (at least large portion of them), while the train set contained only annotation of healthy glomeruli. So **how can we make a model to predict FC: train it on images with FC... But it was not a right thing**: organizers were looking for healthy glomeruli only, and all images except d488c759a did not include FC into annotation. So, use the external data with the shared hand labeling was dangerous, and some people, I may expect, got affected by shake up at the end of the competition because of this.\n\n**Our strategy was proper hand annotation of** [**this external dataset**](https://data.mendeley.com/datasets/k7nvtgn2x6/3) **with 2 classes**: healthy and unhealthy glomeruli. There are many kinds of unhealthy glomeruli, but we decided to keep the things simple, combining them into a single class. We also hand annotated unhealthy glomeruli class in test images, which combined with PL for healthy glomeruli, were used in our training pipeline. We also used [zondo dataset](https://zenodo.org/record/4299694), which naturally includes two classes. However, there was quite a number of missed glomeruli, especially at the image edges, and we hand annotated them. Finally, we used 2 images from HuBMAP portal, not included into train/test datasets with PL (both healthy and unhealthy) generated by our model. Using of multiple data sources helped us to ensure that the model is able to generalize well on new cases. We also used a combined human+AI guided method to add missing annotation to the train data (there were just a few instances tough). \n\n**Secret sauce: ignore d488c759a from LB evaluation.** LB with fixed predictions on this image correlated quite well with CV\n\nFrom the model perspective of view, there is nothing really special, EfficientNet, Unet... Though, we tried more advanced things, like BoT middle layer, Swin transformer backbone, FPN skip connections, but it was not giving any noticeable benefit. This competition really was not about it. The thing that helped was using 2 classes during training. The level of injury of glomeruli is on a continuous scale. So instead of confusing the model by providing nearly identical glomeruli included/excluded from masks, we made the model to learn that unhealthy glomeruli are also glomeruli, and at the end model decides whether a particular instance is healthy enough. To be honest, this competition was not really about segmentation but rather about classification. The uncertainty of including a particular glomerulus in the mask predicted by the model is the major factor degrading the model performance.\n\nOur best submission (ensemble of two models) got 0.9439 CV, 0.941 LB (with including hand annotation of unhealthy glomeruli for d488c759a image), and 0.950 private LB. There were some technical difficulties with making submissions on the last day, so we could not make a submission with all our final models(\n####Links: [**writeup**](https://www.kaggle.com/theoviel/hubmap-final-methodology-submission) <- please take a look (it provides additional details and results of the analysis we made), [**code**] (https://github.com/Optimox/HubMap)",
      "votes": null
    },
    {
      "id": "1301253",
      "postDate": "05/11/2021 01:59:33",
      "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> Congratulations on getting 4th place. This is the 3rd competition I have participated on Kaggle and in each of them, your code/discussions have really been of help to me.</p>",
      "rawMarkdown": "iafoss Congratulations on getting 4th place. This is the 3rd competition I have participated on Kaggle and in each of them, your code/discussions have really been of help to me.",
      "votes": null
    },
    {
      "id": "1301255",
      "postDate": "05/11/2021 02:02:20",
      "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> Congratulations on 4 th Place and thanks for sharing the solution </p>",
      "rawMarkdown": "iafoss Congratulations on 4 th Place and thanks for sharing the solution",
      "votes": null
    },
    {
      "id": "1301274",
      "postDate": "05/11/2021 02:14:53",
      "content": "<p><a href=\"https://www.kaggle.com/nauyan\" target=\"_blank\">@nauyan</a> thank you, I'm glad that you found my kernels helpful</p>",
      "rawMarkdown": "nauyan thank you, I'm glad that you found my kernels helpful",
      "votes": null
    },
    {
      "id": "1301275",
      "postDate": "05/11/2021 02:15:39",
      "content": "<p><a href=\"https://www.kaggle.com/usharengaraju\" target=\"_blank\">@usharengaraju</a> Thanks so much, and you are very welcome.</p>",
      "rawMarkdown": "usharengaraju Thanks so much, and you are very welcome.",
      "votes": null
    },
    {
      "id": "1301328",
      "postDate": "05/11/2021 03:04:46",
      "content": "<p>Congratulations a lot ! your code was truly helped. and I learned a lot from on it. <br>\nAnd thanks for sharing the solution :) </p>",
      "rawMarkdown": "Congratulations a lot ! your code was truly helped. and I learned a lot from on it. \nAnd thanks for sharing the solution :)",
      "votes": null
    },
    {
      "id": "1301348",
      "postDate": "05/11/2021 03:21:14",
      "content": "<p>You are very welcome</p>",
      "rawMarkdown": "You are very welcome",
      "votes": null
    },
    {
      "id": "1301367",
      "postDate": "05/11/2021 03:43:33",
      "content": "<p>Congrats! I've learned a lot from your code both in the early and final step of this competition!</p>",
      "rawMarkdown": "Congrats! I've learned a lot from your code both in the early and final step of this competition!",
      "votes": null
    },
    {
      "id": "1301374",
      "postDate": "05/11/2021 03:48:32",
      "content": "<p>Thanks, you are very welcome</p>",
      "rawMarkdown": "Thanks, you are very welcome",
      "votes": null
    },
    {
      "id": "1301402",
      "postDate": "05/11/2021 04:09:42",
      "content": "<p>Very great solution indeed !</p>",
      "rawMarkdown": "Very great solution indeed !",
      "votes": null
    },
    {
      "id": "1301408",
      "postDate": "05/11/2021 04:15:35",
      "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> Thank you so much</p>",
      "rawMarkdown": "jaideepvalani Thank you so much",
      "votes": null
    },
    {
      "id": "1301418",
      "postDate": "05/11/2021 04:25:09",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> and <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a>. Nice job and thanks for the writeup. </p>",
      "rawMarkdown": "Congrats @iafoss @theoviel and @optimo. Nice job and thanks for the writeup.",
      "votes": null
    },
    {
      "id": "1301440",
      "postDate": "05/11/2021 04:45:09",
      "content": "<p>Thank you so much!</p>",
      "rawMarkdown": "Thank you so much!",
      "votes": null
    },
    {
      "id": "1301461",
      "postDate": "05/11/2021 05:07:19",
      "content": "<p>It was my first time competing and I thought I knew what was important. I thought I knew what was important, but I didn't realize how important preprocessing was. I tried to cross-validate (eefficientnet) and your wonderful stuff, but I couldn't get it done in time. I realized it was 4 hours before the end of the day, so I was done. But I didn't realize that if I wanted to go higher, I should spend more time on the pre-processing at the beginning. I'll keep that in mind next time. I'd like to thank you for all the help you've given me, from the code to the content. I have been referring to you for everything, from the code to the content. You are the only one who writes so clearly and concisely. I've only been doing this for about 7 days, but I've packed 3 years worth of experience. Let's meet at the leaderboard next time!<br>\nSorry for the long post.<br>\nThank you very much!!!!!!!!!!!!!!!!!!!!</p>\n<p>Translated with <a href=\"http://www.DeepL.com/Translator\" target=\"_blank\">www.DeepL.com/Translator</a> (free version)</p>",
      "rawMarkdown": "It was my first time competing and I thought I knew what was important. I thought I knew what was important, but I didn't realize how important preprocessing was. I tried to cross-validate (eefficientnet) and your wonderful stuff, but I couldn't get it done in time. I realized it was 4 hours before the end of the day, so I was done. But I didn't realize that if I wanted to go higher, I should spend more time on the pre-processing at the beginning. I'll keep that in mind next time. I'd like to thank you for all the help you've given me, from the code to the content. I have been referring to you for everything, from the code to the content. You are the only one who writes so clearly and concisely. I've only been doing this for about 7 days, but I've packed 3 years worth of experience. Let's meet at the leaderboard next time!\nSorry for the long post.\nThank you very much!!!!!!!!!!!!!!!!!!!!\n\n\n\nTranslated with www.DeepL.com/Translator (free version)",
      "votes": null
    },
    {
      "id": "1301462",
      "postDate": "05/11/2021 05:07:39",
      "content": "<p>The secret sauce is really on point. I also realized that but 3 days before the competition ended so there was no time to experiment based on that. </p>\n<p>Aggressive augmentations are extremely important as you mentioned in the write-up. I actually had a vague idea about it during the PANDA conference, where my models did not generalize well to external data. I think that was because my PANDA models were trained on simple augmentations such as rotation and cutout, so now I think augmentations probably should be as aggressive as possible without hurting cv/lb </p>",
      "rawMarkdown": "The secret sauce is really on point. I also realized that but 3 days before the competition ended so there was no time to experiment based on that. \n\nAggressive augmentations are extremely important as you mentioned in the write-up. I actually had a vague idea about it during the PANDA conference, where my models did not generalize well to external data. I think that was because my PANDA models were trained on simple augmentations such as rotation and cutout, so now I think augmentations probably should be as aggressive as possible without hurting cv/lb",
      "votes": null
    },
    {
      "id": "1301464",
      "postDate": "05/11/2021 05:08:47",
      "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>  wcm .learning is always mine.<br>\n Did u you use post processing also .  In our case we came up with two <br>\n   a) Area based -Used Pneumothorax competition .It gave slight improvement.<br>\n   b) To be tested :Filtering Extra  borderers based on HSV. This was some thing that i learnt off lately  by that time it was late to use and test it  but looked promising. I will see if it works  in late submissions</p>",
      "rawMarkdown": "iafoss  wcm .learning is always mine.\n Did u you use post processing also .  In our case we came up with two \n   a) Area based -Used Pneumothorax competition .It gave slight improvement.\n   b) To be tested :Filtering Extra  borderers based on HSV. This was some thing that i learnt off lately  by that time it was late to use and test it  but looked promising. I will see if it works  in late submissions",
      "votes": null
    },
    {
      "id": "1301476",
      "postDate": "05/11/2021 05:21:19",
      "content": "<p>Thank you so much, and congratulations with completing your first competition.</p>",
      "rawMarkdown": "Thank you so much, and congratulations with completing your first competition.",
      "votes": null
    },
    {
      "id": "1301483",
      "postDate": "05/11/2021 05:30:04",
      "content": "<p>Yes, the augmentation may be sometimes quite tricky, and there could be a trade of between CV and the ability of the model to work on a completely new data. Doing some analysis based on activation and/or attention maps is often quite helpful to get an idea if model really good, or just pretending to be good by capturing some irrelevant things of the particular set… And based on that one can adjust the augmentation.</p>",
      "rawMarkdown": "Yes, the augmentation may be sometimes quite tricky, and there could be a trade of between CV and the ability of the model to work on a completely new data. Doing some analysis based on activation and/or attention maps is often quite helpful to get an idea if model really good, or just pretending to be good by capturing some irrelevant things of the particular set... And based on that one can adjust the augmentation.",
      "votes": null
    },
    {
      "id": "1301502",
      "postDate": "05/11/2021 05:40:55",
      "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> super congratulations for you :) you deserve a gold !</p>",
      "rawMarkdown": "theoviel super congratulations for you :) you deserve a gold !",
      "votes": null
    },
    {
      "id": "1301538",
      "postDate": "05/11/2021 06:11:14",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/jinssaa\" target=\"_blank\">@jinssaa</a>, it wouldn't have been possible without my teammates though :)</p>",
      "rawMarkdown": "Thanks @jinssaa, it wouldn't have been possible without my teammates though :)",
      "votes": null
    },
    {
      "id": "1301542",
      "postDate": "05/11/2021 06:13:50",
      "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> I tried a lot of post-processing but couldn't find a reliable strategy. <br>\nCRFs, remove small components, remove components with a small max proba, use the convex hull of components… Nothing really worked.</p>",
      "rawMarkdown": "jaideepvalani I tried a lot of post-processing but couldn't find a reliable strategy. \nCRFs, remove small components, remove components with a small max proba, use the convex hull of components... Nothing really worked.",
      "votes": null
    },
    {
      "id": "1301578",
      "postDate": "05/11/2021 06:36:43",
      "content": "<p>removing small components seems worked for me. since my model is overfit and aggresive to predict the target.</p>",
      "rawMarkdown": "removing small components seems worked for me. since my model is overfit and aggresive to predict the target.",
      "votes": null
    },
    {
      "id": "1301618",
      "postDate": "05/11/2021 07:01:06",
      "content": "<p>Congratulations! Without your code I would never get a silver. Everytime I had problems in my competition, I would always go throught your code to get some ideas.</p>",
      "rawMarkdown": "Congratulations! Without your code I would never get a silver. Everytime I had problems in my competition, I would always go throught your code to get some ideas.",
      "votes": null
    },
    {
      "id": "1301651",
      "postDate": "05/11/2021 07:17:54",
      "content": "<p>Yes its difficult to set any generalized  threshold to select or unselect contour . I checked some of my predictions , found that some masks always had an extra border or better to say dilated predictions  i suppose there may at least 1 to 2 in every 100 prediction that one could find such adding to FPS. </p>",
      "rawMarkdown": "Yes its difficult to set any generalized  threshold to select or unselect contour . I checked some of my predictions , found that some masks always had an extra border or better to say dilated predictions  i suppose there may at least 1 to 2 in every 100 prediction that one could find such adding to FPS.",
      "votes": null
    },
    {
      "id": "1301864",
      "postDate": "05/11/2021 09:26:01",
      "content": "<p>Well done!</p>",
      "rawMarkdown": "Well done!",
      "votes": null
    },
    {
      "id": "1301955",
      "postDate": "05/11/2021 10:30:57",
      "content": "<p>Congrats, Elegant solution </p>",
      "rawMarkdown": "Congrats, Elegant solution",
      "votes": null
    },
    {
      "id": "1302137",
      "postDate": "05/11/2021 12:02:49",
      "content": "<p>great work <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> </p>\n<p>I also experimented with the zenodo dataset (provided no boost for me),the mendeley dataset (used pseudo labels and used those 227x227 glomes as mixup; no boost for me) and pseudo labels on the two tiffs from the hubmap portal (again no boost). I didn't try to hand label these datasets tho.</p>\n<p>On the other hand, wrt to ignoring d488, I'm struggling to understand why one model I trained with the d488 hand labels score the same (0.947) as 4 other models with very similar config but no d488 hand labels.</p>\n<p><img src=\"https://i.ibb.co/rbBCL4c/Capture.png\" alt=\"\"></p>",
      "rawMarkdown": "great work @iafoss \n\nI also experimented with the zenodo dataset (provided no boost for me),the mendeley dataset (used pseudo labels and used those 227x227 glomes as mixup; no boost for me) and pseudo labels on the two tiffs from the hubmap portal (again no boost). I didn't try to hand label these datasets tho.\n\nOn the other hand, wrt to ignoring d488, I'm struggling to understand why one model I trained with the d488 hand labels score the same (0.947) as 4 other models with very similar config but no d488 hand labels.\n\n![](https://i.ibb.co/rbBCL4c/Capture.png)",
      "votes": null
    },
    {
      "id": "1303207",
      "postDate": "05/12/2021 00:51:24",
      "content": "<p>Thanks     </p>",
      "rawMarkdown": "Thanks",
      "votes": null
    },
    {
      "id": "1303211",
      "postDate": "05/12/2021 00:55:41",
      "content": "<p>Thanks so much</p>",
      "rawMarkdown": "Thanks so much",
      "votes": null
    },
    {
      "id": "1303231",
      "postDate": "05/12/2021 01:18:54",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/rosuluc\" target=\"_blank\">@rosuluc</a>, quite interesting observation. My best explanation could be that tissue is extremally different from once patient to another (I spent really a lot of time annotating and looking into the images). It may happen that for a specific subset of images the impact of adding FC to d488c759a is not significant, like, for example, there may be no images with FC in the private test set (adding FC to d488c759a does not generate FP at private set). It is difficult to say without looking into images and actual predictions. Meanwhile, it could be another private subset of images where the effect will be game changing.</p>\n<p>In our work we tried to minimize the risks and create a model that generalizes well on diverse images. It is obvious that there may be some randomness in the LB, but we tried to follow a methodology driven approach rather than at the end, if get lucky enough, write about a success story.</p>",
      "rawMarkdown": "Thanks @rosuluc, quite interesting observation. My best explanation could be that tissue is extremally different from once patient to another (I spent really a lot of time annotating and looking into the images). It may happen that for a specific subset of images the impact of adding FC to d488c759a is not significant, like, for example, there may be no images with FC in the private test set (adding FC to d488c759a does not generate FP at private set). It is difficult to say without looking into images and actual predictions. Meanwhile, it could be another private subset of images where the effect will be game changing.\n\nIn our work we tried to minimize the risks and create a model that generalizes well on diverse images. It is obvious that there may be some randomness in the LB, but we tried to follow a methodology driven approach rather than at the end, if get lucky enough, write about a success story.",
      "votes": null
    },
    {
      "id": "1303533",
      "postDate": "05/12/2021 05:49:38",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/duykhanh99\" target=\"_blank\">@duykhanh99</a> !</p>",
      "rawMarkdown": "Thanks @duykhanh99 !",
      "votes": null
    },
    {
      "id": "1304225",
      "postDate": "05/12/2021 13:44:12",
      "content": "<p>Thank you so much, I'm very happy.</p>",
      "rawMarkdown": "Thank you so much, I'm very happy.",
      "votes": null
    },
    {
      "id": "1304636",
      "postDate": "05/12/2021 18:46:22",
      "content": "<blockquote>\n  <p>On the other hand, wrt to ignoring d488, I'm struggling to understand why one model I trained with the d488 hand labels score the same (0.947) as 4 other models with very similar config but no d488 hand labels.</p>\n</blockquote>\n<p>The same thing happened to me described <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/237999\" target=\"_blank\">here</a>. If i added my very accurate public test d488 labels to model training, it did not change my CV score nor did it change my private LB score. (It did increase my public LB score and it would have increased private LB, if private LB contained the same outliers).</p>",
      "rawMarkdown": "> On the other hand, wrt to ignoring d488, I'm struggling to understand why one model I trained with the d488 hand labels score the same (0.947) as 4 other models with very similar config but no d488 hand labels.\n\nThe same thing happened to me described [here][1]. If i added my very accurate public test d488 labels to model training, it did not change my CV score nor did it change my private LB score. (It did increase my public LB score and it would have increased private LB, if private LB contained the same outliers).\n\n[1]: https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/237999",
      "votes": null
    },
    {
      "id": "1304642",
      "postDate": "05/12/2021 18:49:57",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> This is a wonderful solution, i like your idea of 2 class segmentation and beneficial use of external data. </p>\n<p>Your secret sauce is very smart. The presence of d488c759a made public LB score meaningless, but your trick allowed you to collect meaningful feedback from public LB score.</p>\n<blockquote>\n  <p>Secret sauce: ignore d488c759a from LB evaluation. LB with fixed predictions on this image correlated quite well with CV</p>\n</blockquote>",
      "rawMarkdown": "Congratulations @iafoss @theoviel @optimo This is a wonderful solution, i like your idea of 2 class segmentation and beneficial use of external data. \n\nYour secret sauce is very smart. The presence of d488c759a made public LB score meaningless, but your trick allowed you to collect meaningful feedback from public LB score.\n\n> Secret sauce: ignore d488c759a from LB evaluation. LB with fixed predictions on this image correlated quite well with CV",
      "votes": null
    },
    {
      "id": "1304687",
      "postDate": "05/12/2021 19:18:05",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , this competition was quite confusing because of incorrect annotation of 'd488c759a', and host never clarified it. There is a number of other images where FC instances are not annotated, so we assumed that 'd488c759a' is just and exception annotated by some external person, who didn't know what to include. We also had a plan B, though, where we used PL and hand annotation of FC for 'd488c759a' (but not as elaborated as yours), in case if the private test set had this kind of mistakes as train.</p>",
      "rawMarkdown": "Thank you @cdeotte , this competition was quite confusing because of incorrect annotation of 'd488c759a', and host never clarified it. There is a number of other images where FC instances are not annotated, so we assumed that 'd488c759a' is just and exception annotated by some external person, who didn't know what to include. We also had a plan B, though, where we used PL and hand annotation of FC for 'd488c759a' (but not as elaborated as yours), in case if the private test set had this kind of mistakes as train.",
      "votes": null
    },
    {
      "id": "1304719",
      "postDate": "05/12/2021 19:50:55",
      "content": "<p>Thanks Chris ! <br>\nCongratz to you as well, grabbing a solo gold and a solo silver the same day is far from easy :)</p>",
      "rawMarkdown": "Thanks Chris ! \nCongratz to you as well, grabbing a solo gold and a solo silver the same day is far from easy :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1301253,
      "author_name": "nauyan",
      "author_url": "",
      "post_date": "05/11/2021 01:59:33",
      "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> Congratulations on getting 4th place. This is the 3rd competition I have participated on Kaggle and in each of them, your code/discussions have really been of help to me.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1301274,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "05/11/2021 02:14:53",
          "content": "<p><a href=\"https://www.kaggle.com/nauyan\" target=\"_blank\">@nauyan</a> thank you, I'm glad that you found my kernels helpful</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1301255,
      "author_name": "usharengaraju",
      "author_url": "",
      "post_date": "05/11/2021 02:02:20",
      "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> Congratulations on 4 th Place and thanks for sharing the solution </p>",
      "votes": null,
      "replies": [
        {
          "id": 1301275,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "05/11/2021 02:15:39",
          "content": "<p><a href=\"https://www.kaggle.com/usharengaraju\" target=\"_blank\">@usharengaraju</a> Thanks so much, and you are very welcome.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1301328,
      "author_name": "jinssaa",
      "author_url": "",
      "post_date": "05/11/2021 03:04:46",
      "content": "<p>Congratulations a lot ! your code was truly helped. and I learned a lot from on it. <br>\nAnd thanks for sharing the solution :) </p>",
      "votes": null,
      "replies": [
        {
          "id": 1301348,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "05/11/2021 03:21:14",
          "content": "<p>You are very welcome</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1301367,
      "author_name": "southsakura",
      "author_url": "",
      "post_date": "05/11/2021 03:43:33",
      "content": "<p>Congrats! I've learned a lot from your code both in the early and final step of this competition!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1301374,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "05/11/2021 03:48:32",
          "content": "<p>Thanks, you are very welcome</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1301464,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "05/11/2021 05:08:47",
          "content": "<p><a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a>  wcm .learning is always mine.<br>\n Did u you use post processing also .  In our case we came up with two <br>\n   a) Area based -Used Pneumothorax competition .It gave slight improvement.<br>\n   b) To be tested :Filtering Extra  borderers based on HSV. This was some thing that i learnt off lately  by that time it was late to use and test it  but looked promising. I will see if it works  in late submissions</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1301542,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "05/11/2021 06:13:50",
          "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> I tried a lot of post-processing but couldn't find a reliable strategy. <br>\nCRFs, remove small components, remove components with a small max proba, use the convex hull of components… Nothing really worked.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1301578,
          "author_name": "southsakura",
          "author_url": "",
          "post_date": "05/11/2021 06:36:43",
          "content": "<p>removing small components seems worked for me. since my model is overfit and aggresive to predict the target.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1301651,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "05/11/2021 07:17:54",
          "content": "<p>Yes its difficult to set any generalized  threshold to select or unselect contour . I checked some of my predictions , found that some masks always had an extra border or better to say dilated predictions  i suppose there may at least 1 to 2 in every 100 prediction that one could find such adding to FPS. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1301402,
      "author_name": "jaideepvalani",
      "author_url": "",
      "post_date": "05/11/2021 04:09:42",
      "content": "<p>Very great solution indeed !</p>",
      "votes": null,
      "replies": [
        {
          "id": 1301408,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "05/11/2021 04:15:35",
          "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> Thank you so much</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1301418,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "05/11/2021 04:25:09",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> and <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a>. Nice job and thanks for the writeup. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1301440,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "05/11/2021 04:45:09",
          "content": "<p>Thank you so much!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1303533,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "05/12/2021 05:49:38",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/duykhanh99\" target=\"_blank\">@duykhanh99</a> !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1301461,
      "author_name": "taiyofuckyou",
      "author_url": "",
      "post_date": "05/11/2021 05:07:19",
      "content": "<p>It was my first time competing and I thought I knew what was important. I thought I knew what was important, but I didn't realize how important preprocessing was. I tried to cross-validate (eefficientnet) and your wonderful stuff, but I couldn't get it done in time. I realized it was 4 hours before the end of the day, so I was done. But I didn't realize that if I wanted to go higher, I should spend more time on the pre-processing at the beginning. I'll keep that in mind next time. I'd like to thank you for all the help you've given me, from the code to the content. I have been referring to you for everything, from the code to the content. You are the only one who writes so clearly and concisely. I've only been doing this for about 7 days, but I've packed 3 years worth of experience. Let's meet at the leaderboard next time!<br>\nSorry for the long post.<br>\nThank you very much!!!!!!!!!!!!!!!!!!!!</p>\n<p>Translated with <a href=\"http://www.DeepL.com/Translator\" target=\"_blank\">www.DeepL.com/Translator</a> (free version)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1301476,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "05/11/2021 05:21:19",
          "content": "<p>Thank you so much, and congratulations with completing your first competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1301462,
      "author_name": "shujun717",
      "author_url": "",
      "post_date": "05/11/2021 05:07:39",
      "content": "<p>The secret sauce is really on point. I also realized that but 3 days before the competition ended so there was no time to experiment based on that. </p>\n<p>Aggressive augmentations are extremely important as you mentioned in the write-up. I actually had a vague idea about it during the PANDA conference, where my models did not generalize well to external data. I think that was because my PANDA models were trained on simple augmentations such as rotation and cutout, so now I think augmentations probably should be as aggressive as possible without hurting cv/lb </p>",
      "votes": null,
      "replies": [
        {
          "id": 1301483,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "05/11/2021 05:30:04",
          "content": "<p>Yes, the augmentation may be sometimes quite tricky, and there could be a trade of between CV and the ability of the model to work on a completely new data. Doing some analysis based on activation and/or attention maps is often quite helpful to get an idea if model really good, or just pretending to be good by capturing some irrelevant things of the particular set… And based on that one can adjust the augmentation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1301502,
      "author_name": "jinssaa",
      "author_url": "",
      "post_date": "05/11/2021 05:40:55",
      "content": "<p><a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> super congratulations for you :) you deserve a gold !</p>",
      "votes": null,
      "replies": [
        {
          "id": 1301538,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "05/11/2021 06:11:14",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/jinssaa\" target=\"_blank\">@jinssaa</a>, it wouldn't have been possible without my teammates though :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1301618,
      "author_name": "luciusk",
      "author_url": "",
      "post_date": "05/11/2021 07:01:06",
      "content": "<p>Congratulations! Without your code I would never get a silver. Everytime I had problems in my competition, I would always go throught your code to get some ideas.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1304225,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "05/12/2021 13:44:12",
          "content": "<p>Thank you so much, I'm very happy.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1301864,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "05/11/2021 09:26:01",
      "content": "<p>Well done!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1303207,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "05/12/2021 00:51:24",
          "content": "<p>Thanks     </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1301955,
      "author_name": "salimkhazem",
      "author_url": "",
      "post_date": "05/11/2021 10:30:57",
      "content": "<p>Congrats, Elegant solution </p>",
      "votes": null,
      "replies": [
        {
          "id": 1303211,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "05/12/2021 00:55:41",
          "content": "<p>Thanks so much</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1302137,
      "author_name": "rosuluc",
      "author_url": "",
      "post_date": "05/11/2021 12:02:49",
      "content": "<p>great work <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> </p>\n<p>I also experimented with the zenodo dataset (provided no boost for me),the mendeley dataset (used pseudo labels and used those 227x227 glomes as mixup; no boost for me) and pseudo labels on the two tiffs from the hubmap portal (again no boost). I didn't try to hand label these datasets tho.</p>\n<p>On the other hand, wrt to ignoring d488, I'm struggling to understand why one model I trained with the d488 hand labels score the same (0.947) as 4 other models with very similar config but no d488 hand labels.</p>\n<p><img src=\"https://i.ibb.co/rbBCL4c/Capture.png\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1303231,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "05/12/2021 01:18:54",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/rosuluc\" target=\"_blank\">@rosuluc</a>, quite interesting observation. My best explanation could be that tissue is extremally different from once patient to another (I spent really a lot of time annotating and looking into the images). It may happen that for a specific subset of images the impact of adding FC to d488c759a is not significant, like, for example, there may be no images with FC in the private test set (adding FC to d488c759a does not generate FP at private set). It is difficult to say without looking into images and actual predictions. Meanwhile, it could be another private subset of images where the effect will be game changing.</p>\n<p>In our work we tried to minimize the risks and create a model that generalizes well on diverse images. It is obvious that there may be some randomness in the LB, but we tried to follow a methodology driven approach rather than at the end, if get lucky enough, write about a success story.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1304636,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "05/12/2021 18:46:22",
          "content": "<blockquote>\n  <p>On the other hand, wrt to ignoring d488, I'm struggling to understand why one model I trained with the d488 hand labels score the same (0.947) as 4 other models with very similar config but no d488 hand labels.</p>\n</blockquote>\n<p>The same thing happened to me described <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/237999\" target=\"_blank\">here</a>. If i added my very accurate public test d488 labels to model training, it did not change my CV score nor did it change my private LB score. (It did increase my public LB score and it would have increased private LB, if private LB contained the same outliers).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1304642,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "05/12/2021 18:49:57",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> This is a wonderful solution, i like your idea of 2 class segmentation and beneficial use of external data. </p>\n<p>Your secret sauce is very smart. The presence of d488c759a made public LB score meaningless, but your trick allowed you to collect meaningful feedback from public LB score.</p>\n<blockquote>\n  <p>Secret sauce: ignore d488c759a from LB evaluation. LB with fixed predictions on this image correlated quite well with CV</p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 1304687,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "05/12/2021 19:18:05",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , this competition was quite confusing because of incorrect annotation of 'd488c759a', and host never clarified it. There is a number of other images where FC instances are not annotated, so we assumed that 'd488c759a' is just and exception annotated by some external person, who didn't know what to include. We also had a plan B, though, where we used PL and hand annotation of FC for 'd488c759a' (but not as elaborated as yours), in case if the private test set had this kind of mistakes as train.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1304719,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "05/12/2021 19:50:55",
          "content": "<p>Thanks Chris ! <br>\nCongratz to you as well, grabbing a solo gold and a solo silver the same day is far from easy :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1301236": "**Please consider the following links:** [**writeup**](https://www.kaggle.com/theoviel/hubmap-final-methodology-submission), [**code**] (https://github.com/Optimox/HubMap). Below I provide just a short description.\n\n### The key: \n- proper consideration of healthy and unhealthy glomeruli and trust in CV\n- 2 class approach\n- proper annotation of the data\n- be careful with d488c759a when check LB\n\n### Introduction \nOur team would like to thank organizers and kaggle for making this competition possible. And a really big thanks to my amazing teammates @theoviel and @optimo for their incredible efforts in working on the competition and putting everything into writeup for the judge prize. Also @optimo received his first gold medal in this competition, big congratulation to him! Finally, congratulations to all participants of this challenge.\n\nIt was quite a long journey for some of us, since the competition started back in 2020. There were a number of challenges we needed to overcome, and there were many learning opportunities. For me, as well as for some participants, it was the opportunity to **learn labeling medical data**. It may sound funny, but doing this job I realized the amount of work organizers put into this challenge.\n\n### Our approach\n**In core of our approach is proper treatment of healthy and unhealthy glomeruli.**\nIn short, we quickly realized that d488c759a with FC annotated is destabilizing LB and creating a gap between CV and LB. Use of external data with [this annotation](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/208972) improved LB but did not affect or even degraded CV. The explanation is simple: it contained annotation of both healthy and unhealthy glomeruli (at least large portion of them), while the train set contained only annotation of healthy glomeruli. So **how can we make a model to predict FC: train it on images with FC... But it was not a right thing**: organizers were looking for healthy glomeruli only, and all images except d488c759a did not include FC into annotation. So, use the external data with the shared hand labeling was dangerous, and some people, I may expect, got affected by shake up at the end of the competition because of this.\n\n**Our strategy was proper hand annotation of** [**this external dataset**](https://data.mendeley.com/datasets/k7nvtgn2x6/3) **with 2 classes**: healthy and unhealthy glomeruli. There are many kinds of unhealthy glomeruli, but we decided to keep the things simple, combining them into a single class. We also hand annotated unhealthy glomeruli class in test images, which combined with PL for healthy glomeruli, were used in our training pipeline. We also used [zondo dataset](https://zenodo.org/record/4299694), which naturally includes two classes. However, there was quite a number of missed glomeruli, especially at the image edges, and we hand annotated them. Finally, we used 2 images from HuBMAP portal, not included into train/test datasets with PL (both healthy and unhealthy) generated by our model. Using of multiple data sources helped us to ensure that the model is able to generalize well on new cases. We also used a combined human+AI guided method to add missing annotation to the train data (there were just a few instances tough). \n\n**Secret sauce: ignore d488c759a from LB evaluation.** LB with fixed predictions on this image correlated quite well with CV\n\nFrom the model perspective of view, there is nothing really special, EfficientNet, Unet... Though, we tried more advanced things, like BoT middle layer, Swin transformer backbone, FPN skip connections, but it was not giving any noticeable benefit. This competition really was not about it. The thing that helped was using 2 classes during training. The level of injury of glomeruli is on a continuous scale. So instead of confusing the model by providing nearly identical glomeruli included/excluded from masks, we made the model to learn that unhealthy glomeruli are also glomeruli, and at the end model decides whether a particular instance is healthy enough. To be honest, this competition was not really about segmentation but rather about classification. The uncertainty of including a particular glomerulus in the mask predicted by the model is the major factor degrading the model performance.\n\nOur best submission (ensemble of two models) got 0.9439 CV, 0.941 LB (with including hand annotation of unhealthy glomeruli for d488c759a image), and 0.950 private LB. There were some technical difficulties with making submissions on the last day, so we could not make a submission with all our final models(\n####Links: [**writeup**](https://www.kaggle.com/theoviel/hubmap-final-methodology-submission) <- please take a look (it provides additional details and results of the analysis we made), [**code**] (https://github.com/Optimox/HubMap)",
    "1301253": "iafoss Congratulations on getting 4th place. This is the 3rd competition I have participated on Kaggle and in each of them, your code/discussions have really been of help to me.",
    "1301255": "iafoss Congratulations on 4 th Place and thanks for sharing the solution",
    "1301274": "nauyan thank you, I'm glad that you found my kernels helpful",
    "1301275": "usharengaraju Thanks so much, and you are very welcome.",
    "1301328": "Congratulations a lot ! your code was truly helped. and I learned a lot from on it. \nAnd thanks for sharing the solution :)",
    "1301348": "You are very welcome",
    "1301367": "Congrats! I've learned a lot from your code both in the early and final step of this competition!",
    "1301374": "Thanks, you are very welcome",
    "1301402": "Very great solution indeed !",
    "1301408": "jaideepvalani Thank you so much",
    "1301418": "Congrats @iafoss @theoviel and @optimo. Nice job and thanks for the writeup.",
    "1301440": "Thank you so much!",
    "1301461": "It was my first time competing and I thought I knew what was important. I thought I knew what was important, but I didn't realize how important preprocessing was. I tried to cross-validate (eefficientnet) and your wonderful stuff, but I couldn't get it done in time. I realized it was 4 hours before the end of the day, so I was done. But I didn't realize that if I wanted to go higher, I should spend more time on the pre-processing at the beginning. I'll keep that in mind next time. I'd like to thank you for all the help you've given me, from the code to the content. I have been referring to you for everything, from the code to the content. You are the only one who writes so clearly and concisely. I've only been doing this for about 7 days, but I've packed 3 years worth of experience. Let's meet at the leaderboard next time!\nSorry for the long post.\nThank you very much!!!!!!!!!!!!!!!!!!!!\n\n\n\nTranslated with www.DeepL.com/Translator (free version)",
    "1301462": "The secret sauce is really on point. I also realized that but 3 days before the competition ended so there was no time to experiment based on that. \n\nAggressive augmentations are extremely important as you mentioned in the write-up. I actually had a vague idea about it during the PANDA conference, where my models did not generalize well to external data. I think that was because my PANDA models were trained on simple augmentations such as rotation and cutout, so now I think augmentations probably should be as aggressive as possible without hurting cv/lb",
    "1301464": "iafoss  wcm .learning is always mine.\n Did u you use post processing also .  In our case we came up with two \n   a) Area based -Used Pneumothorax competition .It gave slight improvement.\n   b) To be tested :Filtering Extra  borderers based on HSV. This was some thing that i learnt off lately  by that time it was late to use and test it  but looked promising. I will see if it works  in late submissions",
    "1301476": "Thank you so much, and congratulations with completing your first competition.",
    "1301483": "Yes, the augmentation may be sometimes quite tricky, and there could be a trade of between CV and the ability of the model to work on a completely new data. Doing some analysis based on activation and/or attention maps is often quite helpful to get an idea if model really good, or just pretending to be good by capturing some irrelevant things of the particular set... And based on that one can adjust the augmentation.",
    "1301502": "theoviel super congratulations for you :) you deserve a gold !",
    "1301538": "Thanks @jinssaa, it wouldn't have been possible without my teammates though :)",
    "1301542": "jaideepvalani I tried a lot of post-processing but couldn't find a reliable strategy. \nCRFs, remove small components, remove components with a small max proba, use the convex hull of components... Nothing really worked.",
    "1301578": "removing small components seems worked for me. since my model is overfit and aggresive to predict the target.",
    "1301618": "Congratulations! Without your code I would never get a silver. Everytime I had problems in my competition, I would always go throught your code to get some ideas.",
    "1301651": "Yes its difficult to set any generalized  threshold to select or unselect contour . I checked some of my predictions , found that some masks always had an extra border or better to say dilated predictions  i suppose there may at least 1 to 2 in every 100 prediction that one could find such adding to FPS.",
    "1301864": "Well done!",
    "1301955": "Congrats, Elegant solution",
    "1302137": "great work @iafoss \n\nI also experimented with the zenodo dataset (provided no boost for me),the mendeley dataset (used pseudo labels and used those 227x227 glomes as mixup; no boost for me) and pseudo labels on the two tiffs from the hubmap portal (again no boost). I didn't try to hand label these datasets tho.\n\nOn the other hand, wrt to ignoring d488, I'm struggling to understand why one model I trained with the d488 hand labels score the same (0.947) as 4 other models with very similar config but no d488 hand labels.\n\n![](https://i.ibb.co/rbBCL4c/Capture.png)",
    "1303207": "Thanks",
    "1303211": "Thanks so much",
    "1303231": "Thanks @rosuluc, quite interesting observation. My best explanation could be that tissue is extremally different from once patient to another (I spent really a lot of time annotating and looking into the images). It may happen that for a specific subset of images the impact of adding FC to d488c759a is not significant, like, for example, there may be no images with FC in the private test set (adding FC to d488c759a does not generate FP at private set). It is difficult to say without looking into images and actual predictions. Meanwhile, it could be another private subset of images where the effect will be game changing.\n\nIn our work we tried to minimize the risks and create a model that generalizes well on diverse images. It is obvious that there may be some randomness in the LB, but we tried to follow a methodology driven approach rather than at the end, if get lucky enough, write about a success story.",
    "1303533": "Thanks @duykhanh99 !",
    "1304225": "Thank you so much, I'm very happy.",
    "1304636": "> On the other hand, wrt to ignoring d488, I'm struggling to understand why one model I trained with the d488 hand labels score the same (0.947) as 4 other models with very similar config but no d488 hand labels.\n\nThe same thing happened to me described [here][1]. If i added my very accurate public test d488 labels to model training, it did not change my CV score nor did it change my private LB score. (It did increase my public LB score and it would have increased private LB, if private LB contained the same outliers).\n\n[1]: https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/237999",
    "1304642": "Congratulations @iafoss @theoviel @optimo This is a wonderful solution, i like your idea of 2 class segmentation and beneficial use of external data. \n\nYour secret sauce is very smart. The presence of d488c759a made public LB score meaningless, but your trick allowed you to collect meaningful feedback from public LB score.\n\n> Secret sauce: ignore d488c759a from LB evaluation. LB with fixed predictions on this image correlated quite well with CV",
    "1304687": "Thank you @cdeotte , this competition was quite confusing because of incorrect annotation of 'd488c759a', and host never clarified it. There is a number of other images where FC instances are not annotated, so we assumed that 'd488c759a' is just and exception annotated by some external person, who didn't know what to include. We also had a plan B, though, where we used PL and hand annotation of FC for 'd488c759a' (but not as elaborated as yours), in case if the private test set had this kind of mistakes as train.",
    "1304719": "Thanks Chris ! \nCongratz to you as well, grabbing a solo gold and a solo silver the same day is far from easy :)"
  },
  "source": "meta"
}