{
  "id": 238552,
  "title": "16th Place Solution: CNN, Yolov5, and use image-wise model to predict the cell",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/238552",
  "author_name": "Alien",
  "post_date": "2021-05-12T15:00:03.410000",
  "votes": 24,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Congrats to everyone! We are glad to get 16th in this competition. We want to say thank to our team members <br>\n<a href=\"https://www.kaggle.com/louieshao\" target=\"_blank\">@louieshao</a> <a href=\"https://www.kaggle.com/lucamtb\" target=\"_blank\">@lucamtb</a> <a href=\"https://www.kaggle.com/joven1997\" target=\"_blank\">@joven1997</a> <a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> and all participants in this competitions! Especially thank host to celebrate such an amazing competition! <a href=\"https://www.kaggle.com/emmalumpan\" target=\"_blank\">@emmalumpan</a> <a href=\"https://www.kaggle.com/lnhtrang\" target=\"_blank\">@lnhtrang</a> <a href=\"https://www.kaggle.com/cwinsnes\" target=\"_blank\">@cwinsnes</a> <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> <a href=\"https://www.kaggle.com/maggiemd\" target=\"_blank\">@maggiemd</a> It’s my first time to see hosts are so nice to help people facing issues in discussion, also post some great kernels to help us dive into this competition! Great job everyone!</p>\n<p>Pipeline:<br>\n5 cell-wise model, 3 image-wise model, 1 yolov5 model. All image-wise model receive green channels only. </p>\n<p><img src=\"https://imgur.com/rkpVAAi.jpg\" alt=\"\"></p>\n<p>Weights: <br>\n0.5 cell-wise model + 0.3 image-wise model + 0.1 yolov5 model + 0.1 image-wise model(predict the cell)</p>\n<p>Cell-wise models:<br>\n1.Efficientnet b0, external data included, pseudo labeling, epoch 10;<br>\n2.Efficientnet b0, external data included, pseudo labeling, epoch 1;<br>\n3.Efficientnet b7, external data included, epoch 1;<br>\n4.Efficientnet b7, epoch 1;<br>\n5.Efficientnet b7, external data included, green channel only, epoch 1.</p>\n<p>Image-wise models:<br>\n1.Efficientnet b7, external data included, green channel only, epoch 20;<br>\n2.Efficientnet b7, green channel only, epoch 20; <br>\nMy <a href=\"https://www.kaggle.com/h053473666/hpa-classification-efnb7-train\" target=\"_blank\">public kernel</a><br>\n3.Efficientnet b7, green channel only, input size 720, epoch 20; <br>\n<a href=\"https://www.kaggle.com/aristotelisch\" target=\"_blank\">@aristotelisch</a> ‘s <a href=\"https://www.kaggle.com/aristotelisch/hpa-classification-efnb7-train-13cc0d\" target=\"_blank\">public kernel</a></p>\n<p>Cell tile resize trick:<br>\nAdd padding to keep the cell tile’s width equals height, then resize to the input size, which can keep the cell tile’s width-height ratio unchanged, highly boosting our score. </p>\n<p>Pseudo Labeling:<br>\nConsidering the cell-wise tile has unreliable class, we think pseudo labeling on cell may help. Here is our strategy:<br>\n1)Split datasets to two part;<br>\n2)Use one predict the other;<br>\n3)For the possibility of each class:<br>\ni)If confidence value&gt;0.8, put this class into class list;<br>\nii)If confidence value0.2, if the answer is YES, we will put class 18 into class list;<br>\nv)Drop cell tiles without any class.</p>\n<p>Some postscripts by me:<br>\nI have posted my public kernel<a href=\"https://www.kaggle.com/h053473666/0-354-efnb7-classification-weights-0-4-0-6\" target=\"_blank\">https://kaggle.com/h053473666/0-354-efnb7-classification-weights-0-4-0-6</a> about the ensemble of cell-wise and image-wise models. I’m glad to see that it helps. Nevertheless, in some cases, adding image-wise models may not improve the performance. Basically, if it does not work, you can try to tune down the weights of image-wise models. I joined VinBigData competition before, in which i use the weight of 0.95<em>detection + 0.05</em>image-wise model. Even the score went down when i tuned the weight of detection to 0.9. In most cases, more powerful the cell-wise model is, less improvement you can get by ensemble with image-wise models. It’s normal because the mAP focus on the cell-wise confidence anyway. Basically, if one want to do ensemble of cell-wise and image-wise models. Maybe you want to use multiplication first. Nevertheless, I think weighted average may help and easy to conduct. This idea is from my combine of detection and CNN. I think detection is to get bbox, but CNN may predict more accurate confidence value than detection. However, False positives becomes a main issue here. We will expect that image-wise model can help the performance a little bit and that is enough. In this competition, however, we do not have reliable cell-wise labels, so image-wise models may help a lot. If a person can solve the weak-label issue, they will get less improvement from image-level models, but the upper limit of score is relatively high</p>\n<p>Here comes the interesting part--</p>\n<p>Yolov5 model:<br>\nConsidering we don’t have the GT of the mask( they can only be obtained by CellSegmentator ), we think it is not easy to make our segmentation models more accurate than CellSegmentator . Nevertheless, yolov5 can not only detect the image into cell bbox but also give each bbox a confidence value. Considering all of these, we select yolov5 to our model zoo. However, considering the mAP is calculated by the sorting of confidence value, different location of cells in images may influence the confidence value for yolov5. So it is necessary to train other models to predict cell-tile data to do the ensemble.</p>\n<p>Training: Green channel only, no external data, 20 epochs for basic training, 20 epochs for fine tuning.<br>\nPrediction: <br>\n1)Get the mid points of each cell by Cellsegmentator;<br>\n2)Do the KNN(N=1) with mid points predicted by yolov5, making cells and confidence values compared. When a cell mask has two bbox confidence values, choose the larger one. (like NMS)<br>\n3)Assign each cell with the confidence value given by yolov5. If KNN does not find bbox of yolov5 to compare, set the confidence value to 0. Finally, weighted average with our cell-level model. (Like WBF)</p>\n<p>New prediction approach:<br>\nIn this competition, the labels of cell tiles are not quite stable. Meanwhile, the image-wise model seems more robust because the training data of image-wise label is reliable. So, we mainly want to find a way to use image-wise model to predict the cell. </p>\n<p>So we came up with a way, it goes like this:</p>\n<p>1)Label a whole image using Host’s Segmentator, so we will get different cells in the image(1,2,3 …);<br>\n2)Choose a cell in interest, set all area except that to 0;<br>\n3)Resize the image to the size that image-wise model need;<br>\n4)Predict. Image-wise models(My public kernel )<br>\nimg example<br>\n<img src=\"https://imgur.com/kj8vk8t.png\" alt=\"img\"></p>\n<p>I think it is a ‘magic’ part of our solution.</p>\n<p>Thanks all from participants to host. It’s a such amazing journey for us. Hope to see you guys in HPAv3 :)</p>",
  "messages": [
    {
      "id": 1304331,
      "postDate": "2021-05-12T15:00:03.410Z",
      "content": "<p>Congrats to everyone! We are glad to get 16th in this competition. We want to say thank to our team members <br>\n<a href=\"https://www.kaggle.com/louieshao\" target=\"_blank\">@louieshao</a> <a href=\"https://www.kaggle.com/lucamtb\" target=\"_blank\">@lucamtb</a> <a href=\"https://www.kaggle.com/joven1997\" target=\"_blank\">@joven1997</a> <a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> and all participants in this competitions! Especially thank host to celebrate such an amazing competition! <a href=\"https://www.kaggle.com/emmalumpan\" target=\"_blank\">@emmalumpan</a> <a href=\"https://www.kaggle.com/lnhtrang\" target=\"_blank\">@lnhtrang</a> <a href=\"https://www.kaggle.com/cwinsnes\" target=\"_blank\">@cwinsnes</a> <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> <a href=\"https://www.kaggle.com/maggiemd\" target=\"_blank\">@maggiemd</a> It’s my first time to see hosts are so nice to help people facing issues in discussion, also post some great kernels to help us dive into this competition! Great job everyone!</p>\n<p>Pipeline:<br>\n5 cell-wise model, 3 image-wise model, 1 yolov5 model. All image-wise model receive green channels only. </p>\n<p><img src=\"https://imgur.com/rkpVAAi.jpg\" alt=\"\"></p>\n<p>Weights: <br>\n0.5 cell-wise model + 0.3 image-wise model + 0.1 yolov5 model + 0.1 image-wise model(predict the cell)</p>\n<p>Cell-wise models:<br>\n1.Efficientnet b0, external data included, pseudo labeling, epoch 10;<br>\n2.Efficientnet b0, external data included, pseudo labeling, epoch 1;<br>\n3.Efficientnet b7, external data included, epoch 1;<br>\n4.Efficientnet b7, epoch 1;<br>\n5.Efficientnet b7, external data included, green channel only, epoch 1.</p>\n<p>Image-wise models:<br>\n1.Efficientnet b7, external data included, green channel only, epoch 20;<br>\n2.Efficientnet b7, green channel only, epoch 20; <br>\nMy <a href=\"https://www.kaggle.com/h053473666/hpa-classification-efnb7-train\" target=\"_blank\">public kernel</a><br>\n3.Efficientnet b7, green channel only, input size 720, epoch 20; <br>\n<a href=\"https://www.kaggle.com/aristotelisch\" target=\"_blank\">@aristotelisch</a> ‘s <a href=\"https://www.kaggle.com/aristotelisch/hpa-classification-efnb7-train-13cc0d\" target=\"_blank\">public kernel</a></p>\n<p>Cell tile resize trick:<br>\nAdd padding to keep the cell tile’s width equals height, then resize to the input size, which can keep the cell tile’s width-height ratio unchanged, highly boosting our score. </p>\n<p>Pseudo Labeling:<br>\nConsidering the cell-wise tile has unreliable class, we think pseudo labeling on cell may help. Here is our strategy:<br>\n1)Split datasets to two part;<br>\n2)Use one predict the other;<br>\n3)For the possibility of each class:<br>\ni)If confidence value&gt;0.8, put this class into class list;<br>\nii)If confidence value0.2, if the answer is YES, we will put class 18 into class list;<br>\nv)Drop cell tiles without any class.</p>\n<p>Some postscripts by me:<br>\nI have posted my public kernel<a href=\"https://www.kaggle.com/h053473666/0-354-efnb7-classification-weights-0-4-0-6\" target=\"_blank\">https://kaggle.com/h053473666/0-354-efnb7-classification-weights-0-4-0-6</a> about the ensemble of cell-wise and image-wise models. I’m glad to see that it helps. Nevertheless, in some cases, adding image-wise models may not improve the performance. Basically, if it does not work, you can try to tune down the weights of image-wise models. I joined VinBigData competition before, in which i use the weight of 0.95<em>detection + 0.05</em>image-wise model. Even the score went down when i tuned the weight of detection to 0.9. In most cases, more powerful the cell-wise model is, less improvement you can get by ensemble with image-wise models. It’s normal because the mAP focus on the cell-wise confidence anyway. Basically, if one want to do ensemble of cell-wise and image-wise models. Maybe you want to use multiplication first. Nevertheless, I think weighted average may help and easy to conduct. This idea is from my combine of detection and CNN. I think detection is to get bbox, but CNN may predict more accurate confidence value than detection. However, False positives becomes a main issue here. We will expect that image-wise model can help the performance a little bit and that is enough. In this competition, however, we do not have reliable cell-wise labels, so image-wise models may help a lot. If a person can solve the weak-label issue, they will get less improvement from image-level models, but the upper limit of score is relatively high</p>\n<p>Here comes the interesting part--</p>\n<p>Yolov5 model:<br>\nConsidering we don’t have the GT of the mask( they can only be obtained by CellSegmentator ), we think it is not easy to make our segmentation models more accurate than CellSegmentator . Nevertheless, yolov5 can not only detect the image into cell bbox but also give each bbox a confidence value. Considering all of these, we select yolov5 to our model zoo. However, considering the mAP is calculated by the sorting of confidence value, different location of cells in images may influence the confidence value for yolov5. So it is necessary to train other models to predict cell-tile data to do the ensemble.</p>\n<p>Training: Green channel only, no external data, 20 epochs for basic training, 20 epochs for fine tuning.<br>\nPrediction: <br>\n1)Get the mid points of each cell by Cellsegmentator;<br>\n2)Do the KNN(N=1) with mid points predicted by yolov5, making cells and confidence values compared. When a cell mask has two bbox confidence values, choose the larger one. (like NMS)<br>\n3)Assign each cell with the confidence value given by yolov5. If KNN does not find bbox of yolov5 to compare, set the confidence value to 0. Finally, weighted average with our cell-level model. (Like WBF)</p>\n<p>New prediction approach:<br>\nIn this competition, the labels of cell tiles are not quite stable. Meanwhile, the image-wise model seems more robust because the training data of image-wise label is reliable. So, we mainly want to find a way to use image-wise model to predict the cell. </p>\n<p>So we came up with a way, it goes like this:</p>\n<p>1)Label a whole image using Host’s Segmentator, so we will get different cells in the image(1,2,3 …);<br>\n2)Choose a cell in interest, set all area except that to 0;<br>\n3)Resize the image to the size that image-wise model need;<br>\n4)Predict. Image-wise models(My public kernel )<br>\nimg example<br>\n<img src=\"https://imgur.com/kj8vk8t.png\" alt=\"img\"></p>\n<p>I think it is a ‘magic’ part of our solution.</p>\n<p>Thanks all from participants to host. It’s a such amazing journey for us. Hope to see you guys in HPAv3 :)</p>",
      "rawMarkdown": "Congrats to everyone! We are glad to get 16th in this competition. We want to say thank to our team members \n@louieshao @lucamtb @joven1997 @morizin and all participants in this competitions! Especially thank host to celebrate such an amazing competition! @emmalumpan @lnhtrang @cwinsnes @philculliton @maggiemd It’s my first time to see hosts are so nice to help people facing issues in discussion, also post some great kernels to help us dive into this competition! Great job everyone!\n\nPipeline:\n5 cell-wise model, 3 image-wise model, 1 yolov5 model. All image-wise model receive green channels only. \n\n![](https://imgur.com/rkpVAAi.jpg)\n\nWeights: \n0.5 cell-wise model + 0.3 image-wise model + 0.1 yolov5 model + 0.1 image-wise model(predict the cell)\n\nCell-wise models:\n1.Efficientnet b0, external data included, pseudo labeling, epoch 10;\n2.Efficientnet b0, external data included, pseudo labeling, epoch 1;\n3.Efficientnet b7, external data included, epoch 1;\n4.Efficientnet b7, epoch 1;\n5.Efficientnet b7, external data included, green channel only, epoch 1.\n\nImage-wise models:\n1.Efficientnet b7, external data included, green channel only, epoch 20;\n2.Efficientnet b7, green channel only, epoch 20; \nMy [public kernel](https://www.kaggle.com/h053473666/hpa-classification-efnb7-train)\n3.Efficientnet b7, green channel only, input size 720, epoch 20; \n@aristotelisch ‘s [public kernel](https://www.kaggle.com/aristotelisch/hpa-classification-efnb7-train-13cc0d)\n\nCell tile resize trick:\nAdd padding to keep the cell tile’s width equals height, then resize to the input size, which can keep the cell tile’s width-height ratio unchanged, highly boosting our score. \n\nPseudo Labeling:\nConsidering the cell-wise tile has unreliable class, we think pseudo labeling on cell may help. Here is our strategy:\n1)Split datasets to two part;\n2)Use one predict the other;\n3)For the possibility of each class:\ni)If confidence value>0.8, put this class into class list;\nii)If confidence value<0.2, drop this class;\niii)If confidence is between 0.2 and 0.8, keep the class unchanged.\niv)If a cell tile does not belong to any class, then considering whether confidence value of class 18>0.2, if the answer is YES, we will put class 18 into class list;\nv)Drop cell tiles without any class.\n\nSome postscripts by me:\nI have posted my public kernel[https://kaggle.com/h053473666/0-354-efnb7-classification-weights-0-4-0-6](https://www.kaggle.com/h053473666/0-354-efnb7-classification-weights-0-4-0-6) about the ensemble of cell-wise and image-wise models. I’m glad to see that it helps. Nevertheless, in some cases, adding image-wise models may not improve the performance. Basically, if it does not work, you can try to tune down the weights of image-wise models. I joined VinBigData competition before, in which i use the weight of 0.95*detection + 0.05*image-wise model. Even the score went down when i tuned the weight of detection to 0.9. In most cases, more powerful the cell-wise model is, less improvement you can get by ensemble with image-wise models. It’s normal because the mAP focus on the cell-wise confidence anyway. Basically, if one want to do ensemble of cell-wise and image-wise models. Maybe you want to use multiplication first. Nevertheless, I think weighted average may help and easy to conduct. This idea is from my combine of detection and CNN. I think detection is to get bbox, but CNN may predict more accurate confidence value than detection. However, False positives becomes a main issue here. We will expect that image-wise model can help the performance a little bit and that is enough. In this competition, however, we do not have reliable cell-wise labels, so image-wise models may help a lot. If a person can solve the weak-label issue, they will get less improvement from image-level models, but the upper limit of score is relatively high\n\nHere comes the interesting part--\n\nYolov5 model:\nConsidering we don’t have the GT of the mask( they can only be obtained by CellSegmentator ), we think it is not easy to make our segmentation models more accurate than CellSegmentator . Nevertheless, yolov5 can not only detect the image into cell bbox but also give each bbox a confidence value. Considering all of these, we select yolov5 to our model zoo. However, considering the mAP is calculated by the sorting of confidence value, different location of cells in images may influence the confidence value for yolov5. So it is necessary to train other models to predict cell-tile data to do the ensemble.\n\nTraining: Green channel only, no external data, 20 epochs for basic training, 20 epochs for fine tuning.\nPrediction: \n1)Get the mid points of each cell by Cellsegmentator;\n2)Do the KNN(N=1) with mid points predicted by yolov5, making cells and confidence values compared. When a cell mask has two bbox confidence values, choose the larger one. (like NMS)\n3)Assign each cell with the confidence value given by yolov5. If KNN does not find bbox of yolov5 to compare, set the confidence value to 0. Finally, weighted average with our cell-level model. (Like WBF)\n\nNew prediction approach:\nIn this competition, the labels of cell tiles are not quite stable. Meanwhile, the image-wise model seems more robust because the training data of image-wise label is reliable. So, we mainly want to find a way to use image-wise model to predict the cell. \n\nSo we came up with a way, it goes like this:\n\n1)Label a whole image using Host’s Segmentator, so we will get different cells in the image(1,2,3 ...);\n2)Choose a cell in interest, set all area except that to 0;\n3)Resize the image to the size that image-wise model need;\n4)Predict. Image-wise models(My public kernel )\nimg example\n![img](https://imgur.com/kj8vk8t.png)\n\nI think it is a ‘magic’ part of our solution.\n\nThanks all from participants to host. It’s a such amazing journey for us. Hope to see you guys in HPAv3 :)",
      "votes": 24
    },
    {
      "id": 1305437,
      "postDate": "2021-05-13T09:13:46.823Z",
      "content": "<p>Congratulations to your team and thanks for the write-up! Very interesting! Nice idea to use YOLO to get detection confidence! Do you by any chance remember the boost YOLO gave you? Thanks!</p>",
      "rawMarkdown": "Congratulations to your team and thanks for the write-up! Very interesting! Nice idea to use YOLO to get detection confidence! Do you by any chance remember the boost YOLO gave you? Thanks!",
      "votes": 3,
      "replies": [
        {
          "id": 1305464,
          "postDate": "2021-05-13T09:45:51.977Z",
          "content": "<p>0.005~0.009. There is an advantage that yolo can increase robustness</p>",
          "rawMarkdown": "0.005~0.009. There is an advantage that yolo can increase robustness",
          "votes": 2
        },
        {
          "id": 1305643,
          "postDate": "2021-05-13T12:06:04.337Z",
          "content": "<p>Thank you for the reply, <a href=\"https://www.kaggle.com/h053473666\" target=\"_blank\">@h053473666</a> !</p>",
          "rawMarkdown": "Thank you for the reply, @h053473666 !",
          "votes": 2
        }
      ]
    },
    {
      "id": 1304436,
      "postDate": "2021-05-12T16:07:28.960Z",
      "content": "<p><a href=\"https://www.kaggle.com/h053473666\" target=\"_blank\">@h053473666</a> 's yolov5 really help us boost score, no need to say his nice public kernel. Thanks bro!</p>",
      "rawMarkdown": "@h053473666 's yolov5 really help us boost score, no need to say his nice public kernel. Thanks bro!",
      "votes": 3
    },
    {
      "id": 1304411,
      "postDate": "2021-05-12T15:52:24.760Z",
      "content": "<p>It was fun to work with you 😃</p>",
      "rawMarkdown": "It was fun to work with you 😃",
      "votes": 3,
      "replies": [
        {
          "id": 1304424,
          "postDate": "2021-05-12T15:57:11.193Z",
          "content": "<p>Thanks! It’s also fun to work with you! I learned a lot from your kernel. 👍</p>",
          "rawMarkdown": "Thanks! It’s also fun to work with you! I learned a lot from your kernel. 👍",
          "votes": 2
        }
      ]
    },
    {
      "id": 1305023,
      "postDate": "2021-05-13T04:25:27.513Z",
      "content": "<p><a href=\"https://www.kaggle.com/h053473666\" target=\"_blank\">@h053473666</a> Congratulations  and Thanks for sharing the approach</p>",
      "rawMarkdown": "@h053473666 Congratulations  and Thanks for sharing the approach",
      "votes": 1
    },
    {
      "id": 1305618,
      "postDate": "2021-05-13T11:49:48.570Z",
      "content": "<p>Wow, I like the way you do Pseudo Labeling in a way to address the cell-wise label problem. And I also like the work you drew out the pipeline for us to understand 😂 That's sweet.</p>\n<p>Congratulations!</p>",
      "rawMarkdown": "Wow, I like the way you do Pseudo Labeling in a way to address the cell-wise label problem. And I also like the work you drew out the pipeline for us to understand 😂 That's sweet.\n\nCongratulations!",
      "votes": 2
    },
    {
      "id": 1304688,
      "postDate": "2021-05-12T19:18:18.150Z",
      "content": "<p>I would like to say about one more part in our solution which got into notebook timeout so we avoided it in the final submission</p>\n<p>It was inspired from the idea of this repo<br>\n<a href=\"https://github.com/jfhealthcare/Chexpert\" target=\"_blank\">https://github.com/jfhealthcare/Chexpert</a></p>\n<p>The idea was to create a heat map for each disease for each images <br>\nAnd this heat map was probabilistic. The heatmap was in the same image size as of the image. So this helped us in cropping the cell ( we made indices of cropping with segmentor ). </p>\n<p>This cropped cell would be a numpy array containing values between 0 and 1. We global averaged this cropped cell heatmap pixel values for each classes to generate final confidence for that class and weighted average them with our existing models.</p>\n<p>Here is a sample of heatmap<img src=\"https://i.postimg.cc/sXKVr9dX/plot-3000.png\" alt=\"\"></p>\n<p>This sample was created in earlier epochs the heatmaps much more good in future epochs</p>\n<p>In the minimal implementation ( <br>\nbecause of checking memory) of it, it scored 0.541 in the existing model. But i hope it would improve our model.</p>\n<p>Thanks to my amazing teammates. Through out the Competition they were supporting me and helping me.</p>",
      "rawMarkdown": "I would like to say about one more part in our solution which got into notebook timeout so we avoided it in the final submission\n\nIt was inspired from the idea of this repo\nhttps://github.com/jfhealthcare/Chexpert\n\nThe idea was to create a heat map for each disease for each images \nAnd this heat map was probabilistic. The heatmap was in the same image size as of the image. So this helped us in cropping the cell ( we made indices of cropping with segmentor ). \n\nThis cropped cell would be a numpy array containing values between 0 and 1. We global averaged this cropped cell heatmap pixel values for each classes to generate final confidence for that class and weighted average them with our existing models.\n\nHere is a sample of heatmap![](https://i.postimg.cc/sXKVr9dX/plot-3000.png)\n\nThis sample was created in earlier epochs the heatmaps much more good in future epochs\n\nIn the minimal implementation ( \nbecause of checking memory) of it, it scored 0.541 in the existing model. But i hope it would improve our model.\n\nThanks to my amazing teammates. Through out the Competition they were supporting me and helping me.",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 1305437,
      "author_name": "Raman",
      "author_url": "",
      "post_date": "2021-05-13T09:13:46.823000",
      "content": "<p>Congratulations to your team and thanks for the write-up! Very interesting! Nice idea to use YOLO to get detection confidence! Do you by any chance remember the boost YOLO gave you? Thanks!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1305464,
          "author_name": "Alien",
          "author_url": "",
          "post_date": "2021-05-13T09:45:51.977000",
          "content": "<p>0.005~0.009. There is an advantage that yolo can increase robustness</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1305643,
          "author_name": "Raman",
          "author_url": "",
          "post_date": "2021-05-13T12:06:04.337000",
          "content": "<p>Thank you for the reply, <a href=\"https://www.kaggle.com/h053473666\" target=\"_blank\">@h053473666</a> !</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1304436,
      "author_name": "Shihao Shao",
      "author_url": "",
      "post_date": "2021-05-12T16:07:28.960000",
      "content": "<p><a href=\"https://www.kaggle.com/h053473666\" target=\"_blank\">@h053473666</a> 's yolov5 really help us boost score, no need to say his nice public kernel. Thanks bro!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1304411,
      "author_name": "LucaMTB",
      "author_url": "",
      "post_date": "2021-05-12T15:52:24.760000",
      "content": "<p>It was fun to work with you 😃</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1304424,
          "author_name": "Alien",
          "author_url": "",
          "post_date": "2021-05-12T15:57:11.193000",
          "content": "<p>Thanks! It’s also fun to work with you! I learned a lot from your kernel. 👍</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1305023,
      "author_name": "Tensor Girl",
      "author_url": "",
      "post_date": "2021-05-13T04:25:27.513000",
      "content": "<p><a href=\"https://www.kaggle.com/h053473666\" target=\"_blank\">@h053473666</a> Congratulations  and Thanks for sharing the approach</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1305618,
      "author_name": "Chien-Hsiang Hung",
      "author_url": "",
      "post_date": "2021-05-13T11:49:48.570000",
      "content": "<p>Wow, I like the way you do Pseudo Labeling in a way to address the cell-wise label problem. And I also like the work you drew out the pipeline for us to understand 😂 That's sweet.</p>\n<p>Congratulations!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1304688,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2021-05-12T19:18:18.150000",
      "content": "<p>I would like to say about one more part in our solution which got into notebook timeout so we avoided it in the final submission</p>\n<p>It was inspired from the idea of this repo<br>\n<a href=\"https://github.com/jfhealthcare/Chexpert\" target=\"_blank\">https://github.com/jfhealthcare/Chexpert</a></p>\n<p>The idea was to create a heat map for each disease for each images <br>\nAnd this heat map was probabilistic. The heatmap was in the same image size as of the image. So this helped us in cropping the cell ( we made indices of cropping with segmentor ). </p>\n<p>This cropped cell would be a numpy array containing values between 0 and 1. We global averaged this cropped cell heatmap pixel values for each classes to generate final confidence for that class and weighted average them with our existing models.</p>\n<p>Here is a sample of heatmap<img src=\"https://i.postimg.cc/sXKVr9dX/plot-3000.png\" alt=\"\"></p>\n<p>This sample was created in earlier epochs the heatmaps much more good in future epochs</p>\n<p>In the minimal implementation ( <br>\nbecause of checking memory) of it, it scored 0.541 in the existing model. But i hope it would improve our model.</p>\n<p>Thanks to my amazing teammates. Through out the Competition they were supporting me and helping me.</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1304331": "Congrats to everyone! We are glad to get 16th in this competition. We want to say thank to our team members \n@louieshao @lucamtb @joven1997 @morizin and all participants in this competitions! Especially thank host to celebrate such an amazing competition! @emmalumpan @lnhtrang @cwinsnes @philculliton @maggiemd It’s my first time to see hosts are so nice to help people facing issues in discussion, also post some great kernels to help us dive into this competition! Great job everyone!\n\nPipeline:\n5 cell-wise model, 3 image-wise model, 1 yolov5 model. All image-wise model receive green channels only. \n\n![](https://imgur.com/rkpVAAi.jpg)\n\nWeights: \n0.5 cell-wise model + 0.3 image-wise model + 0.1 yolov5 model + 0.1 image-wise model(predict the cell)\n\nCell-wise models:\n1.Efficientnet b0, external data included, pseudo labeling, epoch 10;\n2.Efficientnet b0, external data included, pseudo labeling, epoch 1;\n3.Efficientnet b7, external data included, epoch 1;\n4.Efficientnet b7, epoch 1;\n5.Efficientnet b7, external data included, green channel only, epoch 1.\n\nImage-wise models:\n1.Efficientnet b7, external data included, green channel only, epoch 20;\n2.Efficientnet b7, green channel only, epoch 20; \nMy [public kernel](https://www.kaggle.com/h053473666/hpa-classification-efnb7-train)\n3.Efficientnet b7, green channel only, input size 720, epoch 20; \n@aristotelisch ‘s [public kernel](https://www.kaggle.com/aristotelisch/hpa-classification-efnb7-train-13cc0d)\n\nCell tile resize trick:\nAdd padding to keep the cell tile’s width equals height, then resize to the input size, which can keep the cell tile’s width-height ratio unchanged, highly boosting our score. \n\nPseudo Labeling:\nConsidering the cell-wise tile has unreliable class, we think pseudo labeling on cell may help. Here is our strategy:\n1)Split datasets to two part;\n2)Use one predict the other;\n3)For the possibility of each class:\ni)If confidence value>0.8, put this class into class list;\nii)If confidence value<0.2, drop this class;\niii)If confidence is between 0.2 and 0.8, keep the class unchanged.\niv)If a cell tile does not belong to any class, then considering whether confidence value of class 18>0.2, if the answer is YES, we will put class 18 into class list;\nv)Drop cell tiles without any class.\n\nSome postscripts by me:\nI have posted my public kernel[https://kaggle.com/h053473666/0-354-efnb7-classification-weights-0-4-0-6](https://www.kaggle.com/h053473666/0-354-efnb7-classification-weights-0-4-0-6) about the ensemble of cell-wise and image-wise models. I’m glad to see that it helps. Nevertheless, in some cases, adding image-wise models may not improve the performance. Basically, if it does not work, you can try to tune down the weights of image-wise models. I joined VinBigData competition before, in which i use the weight of 0.95*detection + 0.05*image-wise model. Even the score went down when i tuned the weight of detection to 0.9. In most cases, more powerful the cell-wise model is, less improvement you can get by ensemble with image-wise models. It’s normal because the mAP focus on the cell-wise confidence anyway. Basically, if one want to do ensemble of cell-wise and image-wise models. Maybe you want to use multiplication first. Nevertheless, I think weighted average may help and easy to conduct. This idea is from my combine of detection and CNN. I think detection is to get bbox, but CNN may predict more accurate confidence value than detection. However, False positives becomes a main issue here. We will expect that image-wise model can help the performance a little bit and that is enough. In this competition, however, we do not have reliable cell-wise labels, so image-wise models may help a lot. If a person can solve the weak-label issue, they will get less improvement from image-level models, but the upper limit of score is relatively high\n\nHere comes the interesting part--\n\nYolov5 model:\nConsidering we don’t have the GT of the mask( they can only be obtained by CellSegmentator ), we think it is not easy to make our segmentation models more accurate than CellSegmentator . Nevertheless, yolov5 can not only detect the image into cell bbox but also give each bbox a confidence value. Considering all of these, we select yolov5 to our model zoo. However, considering the mAP is calculated by the sorting of confidence value, different location of cells in images may influence the confidence value for yolov5. So it is necessary to train other models to predict cell-tile data to do the ensemble.\n\nTraining: Green channel only, no external data, 20 epochs for basic training, 20 epochs for fine tuning.\nPrediction: \n1)Get the mid points of each cell by Cellsegmentator;\n2)Do the KNN(N=1) with mid points predicted by yolov5, making cells and confidence values compared. When a cell mask has two bbox confidence values, choose the larger one. (like NMS)\n3)Assign each cell with the confidence value given by yolov5. If KNN does not find bbox of yolov5 to compare, set the confidence value to 0. Finally, weighted average with our cell-level model. (Like WBF)\n\nNew prediction approach:\nIn this competition, the labels of cell tiles are not quite stable. Meanwhile, the image-wise model seems more robust because the training data of image-wise label is reliable. So, we mainly want to find a way to use image-wise model to predict the cell. \n\nSo we came up with a way, it goes like this:\n\n1)Label a whole image using Host’s Segmentator, so we will get different cells in the image(1,2,3 ...);\n2)Choose a cell in interest, set all area except that to 0;\n3)Resize the image to the size that image-wise model need;\n4)Predict. Image-wise models(My public kernel )\nimg example\n![img](https://imgur.com/kj8vk8t.png)\n\nI think it is a ‘magic’ part of our solution.\n\nThanks all from participants to host. It’s a such amazing journey for us. Hope to see you guys in HPAv3 :)",
    "1305437": "Congratulations to your team and thanks for the write-up! Very interesting! Nice idea to use YOLO to get detection confidence! Do you by any chance remember the boost YOLO gave you? Thanks!",
    "1304436": "@h053473666 's yolov5 really help us boost score, no need to say his nice public kernel. Thanks bro!",
    "1304411": "It was fun to work with you 😃",
    "1305023": "@h053473666 Congratulations  and Thanks for sharing the approach",
    "1305618": "Wow, I like the way you do Pseudo Labeling in a way to address the cell-wise label problem. And I also like the work you drew out the pipeline for us to understand 😂 That's sweet.\n\nCongratulations!",
    "1304688": "I would like to say about one more part in our solution which got into notebook timeout so we avoided it in the final submission\n\nIt was inspired from the idea of this repo\nhttps://github.com/jfhealthcare/Chexpert\n\nThe idea was to create a heat map for each disease for each images \nAnd this heat map was probabilistic. The heatmap was in the same image size as of the image. So this helped us in cropping the cell ( we made indices of cropping with segmentor ). \n\nThis cropped cell would be a numpy array containing values between 0 and 1. We global averaged this cropped cell heatmap pixel values for each classes to generate final confidence for that class and weighted average them with our existing models.\n\nHere is a sample of heatmap![](https://i.postimg.cc/sXKVr9dX/plot-3000.png)\n\nThis sample was created in earlier epochs the heatmaps much more good in future epochs\n\nIn the minimal implementation ( \nbecause of checking memory) of it, it scored 0.541 in the existing model. But i hope it would improve our model.\n\nThanks to my amazing teammates. Through out the Competition they were supporting me and helping me."
  }
}