{
  "id": 227103,
  "title": "10th Place Solution",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/227103",
  "author_name": "toxu",
  "post_date": "2021-03-19T02:04:22.113000",
  "votes": 30,
  "comment_count": 10,
  "views": 0,
  "content": "<p>First of all, I would like to thank Kaggle and the organizers for hosting this great competitions. Also, I would like to thank my teammates - <a href=\"https://www.kaggle.com/ludovick\" target=\"_blank\">@ludovick</a> and <a href=\"https://www.kaggle.com/woshifym\" target=\"_blank\">@woshifym</a> . Due to the time difference, we can work 24 hours a day.</p>\n<h1>Solutions</h1>\n<p>Our solution can be divided into 2 parts, pytorch part and tensorflow part.</p>\n<h2>Pytorch part</h2>\n<p>For Pytorch part I would like to thank <a href=\"https://www.kaggle.com/ammarali32\" target=\"_blank\">@ammarali32</a> for the 3-step method first. All our models based on 3-step method. We trained models with different backbones and different image size to get better results.</p>\n<ul>\n<li>resnet200d + image size: 600  CV: 0.9578</li>\n<li>ecaresnet269d + image size: 600 CV: 0.9584</li>\n<li>resnest200e + image size: 640 CV:0.9625</li>\n</ul>\n<p>After that we found <a href=\"https://github.com/alipay/cvpr2020-plant-pathology\" target=\"_blank\">Soft label method</a> performs great in this competition. So we trained all the models above using this method to form our 4-step method.</p>\n<h2>Tensorflow part</h2>\n<p>My teammates <a href=\"https://www.kaggle.com/ludovick\" target=\"_blank\">@ludovick</a> did a great job using tensorflow. I would like to invite him to introduce his method in details later.<br>\nFor brief summary:</p>\n<ul>\n<li>Efficient B7 + 1024 image size</li>\n<li>Efficient L2 + 768 image size</li>\n<li>Soft label method</li>\n<li>Pesudo Label</li>\n</ul>\n<h2>Ensemble</h2>\n<p>For ensemble, we use the OOF files to train a linear regression model. Then we use this linear model to get the weight for each model. We select our best CV and best public score in our final submissions. Unfortunately, we just missed out best ensemble which is the second best in public score.</p>\n<p>Final submission one <br>\nCV: 0.9692 Public LB: 0.971 Private LB: 0.973</p>\n<ul>\n<li>resnet200d + 600</li>\n<li>ecaresnet269d + 600</li>\n<li>resnest200e + 640</li>\n<li>Efficient B7 + 1024 + soft label + PL</li>\n<li>Efficient L2 + 768 + soft label</li>\n</ul>\n<p>Final submission two<br>\nCV: 0.9689 Public LB: 0.971 Private LB: 0.974</p>\n<ul>\n<li>resnet200d + 600</li>\n<li>resnest200e + 640</li>\n<li>Efficient B7 + 1024 + soft label + PL</li>\n<li>Efficient B7 + 1024 + PL</li>\n<li>Efficient L2 + 768 + soft label</li>\n</ul>\n<h2>Tricks</h2>\n<ul>\n<li>Larger model</li>\n<li>Higher resolution</li>\n<li>Soft Label</li>\n<li>Pesudo Label</li>\n</ul>\n<h2>Acknowledge</h2>\n<p>Thanks to TFRC team for providing us free TPU.</p>",
  "messages": [
    {
      "id": 1244443,
      "postDate": "2021-03-19T02:04:22.113Z",
      "content": "<p>First of all, I would like to thank Kaggle and the organizers for hosting this great competitions. Also, I would like to thank my teammates - <a href=\"https://www.kaggle.com/ludovick\" target=\"_blank\">@ludovick</a> and <a href=\"https://www.kaggle.com/woshifym\" target=\"_blank\">@woshifym</a> . Due to the time difference, we can work 24 hours a day.</p>\n<h1>Solutions</h1>\n<p>Our solution can be divided into 2 parts, pytorch part and tensorflow part.</p>\n<h2>Pytorch part</h2>\n<p>For Pytorch part I would like to thank <a href=\"https://www.kaggle.com/ammarali32\" target=\"_blank\">@ammarali32</a> for the 3-step method first. All our models based on 3-step method. We trained models with different backbones and different image size to get better results.</p>\n<ul>\n<li>resnet200d + image size: 600  CV: 0.9578</li>\n<li>ecaresnet269d + image size: 600 CV: 0.9584</li>\n<li>resnest200e + image size: 640 CV:0.9625</li>\n</ul>\n<p>After that we found <a href=\"https://github.com/alipay/cvpr2020-plant-pathology\" target=\"_blank\">Soft label method</a> performs great in this competition. So we trained all the models above using this method to form our 4-step method.</p>\n<h2>Tensorflow part</h2>\n<p>My teammates <a href=\"https://www.kaggle.com/ludovick\" target=\"_blank\">@ludovick</a> did a great job using tensorflow. I would like to invite him to introduce his method in details later.<br>\nFor brief summary:</p>\n<ul>\n<li>Efficient B7 + 1024 image size</li>\n<li>Efficient L2 + 768 image size</li>\n<li>Soft label method</li>\n<li>Pesudo Label</li>\n</ul>\n<h2>Ensemble</h2>\n<p>For ensemble, we use the OOF files to train a linear regression model. Then we use this linear model to get the weight for each model. We select our best CV and best public score in our final submissions. Unfortunately, we just missed out best ensemble which is the second best in public score.</p>\n<p>Final submission one <br>\nCV: 0.9692 Public LB: 0.971 Private LB: 0.973</p>\n<ul>\n<li>resnet200d + 600</li>\n<li>ecaresnet269d + 600</li>\n<li>resnest200e + 640</li>\n<li>Efficient B7 + 1024 + soft label + PL</li>\n<li>Efficient L2 + 768 + soft label</li>\n</ul>\n<p>Final submission two<br>\nCV: 0.9689 Public LB: 0.971 Private LB: 0.974</p>\n<ul>\n<li>resnet200d + 600</li>\n<li>resnest200e + 640</li>\n<li>Efficient B7 + 1024 + soft label + PL</li>\n<li>Efficient B7 + 1024 + PL</li>\n<li>Efficient L2 + 768 + soft label</li>\n</ul>\n<h2>Tricks</h2>\n<ul>\n<li>Larger model</li>\n<li>Higher resolution</li>\n<li>Soft Label</li>\n<li>Pesudo Label</li>\n</ul>\n<h2>Acknowledge</h2>\n<p>Thanks to TFRC team for providing us free TPU.</p>",
      "rawMarkdown": "First of all, I would like to thank Kaggle and the organizers for hosting this great competitions. Also, I would like to thank my teammates - @ludovick and @woshifym . Due to the time difference, we can work 24 hours a day.\n\n# Solutions\nOur solution can be divided into 2 parts, pytorch part and tensorflow part.\n\n## Pytorch part\nFor Pytorch part I would like to thank @ammarali32 for the 3-step method first. All our models based on 3-step method. We trained models with different backbones and different image size to get better results.\n\n* resnet200d + image size: 600  CV: 0.9578\n* ecaresnet269d + image size: 600 CV: 0.9584\n* resnest200e + image size: 640 CV:0.9625\n\nAfter that we found [Soft label method](https://github.com/alipay/cvpr2020-plant-pathology) performs great in this competition. So we trained all the models above using this method to form our 4-step method.\n\n## Tensorflow part\nMy teammates @ludovick did a great job using tensorflow. I would like to invite him to introduce his method in details later.\nFor brief summary:\n\n* Efficient B7 + 1024 image size\n* Efficient L2 + 768 image size\n* Soft label method\n* Pesudo Label\n\n## Ensemble\nFor ensemble, we use the OOF files to train a linear regression model. Then we use this linear model to get the weight for each model. We select our best CV and best public score in our final submissions. Unfortunately, we just missed out best ensemble which is the second best in public score.\n\nFinal submission one \nCV: 0.9692 Public LB: 0.971 Private LB: 0.973\n* resnet200d + 600\n* ecaresnet269d + 600\n* resnest200e + 640\n* Efficient B7 + 1024 + soft label + PL\n* Efficient L2 + 768 + soft label\n\nFinal submission two\nCV: 0.9689 Public LB: 0.971 Private LB: 0.974\n* resnet200d + 600\n* resnest200e + 640\n* Efficient B7 + 1024 + soft label + PL\n* Efficient B7 + 1024 + PL\n* Efficient L2 + 768 + soft label\n\n## Tricks\n* Larger model\n* Higher resolution\n* Soft Label\n* Pesudo Label\n\n## Acknowledge\nThanks to TFRC team for providing us free TPU.",
      "votes": 30
    },
    {
      "id": 1245610,
      "postDate": "2021-03-20T00:15:24.033Z",
      "content": "<p>Thank you to Kaggle and the organizers for hosting this competition. </p>\n<p>regarding the Tensorflow part :</p>\n<p>We use for our solution only EfficientNet architectures for the TF solution. In fact, we are quite limited regarding the different strong models in TF, only inceptionresnetV2 was giving interesting results, but it was still lower than EfficientNet B6/B7/L2.<br>\nWe use 3 resolutions : 768, 1024, 1280 and lights data aug : rotation, translation, shear, zoom, color (contrast etc)</p>\n<ul>\n<li><p>1st step : Train EfficientNetB6/B7 on image size of 1024 gives respectively a LB of 0.966 and 0.965. EfficientNetB7 on 1280 gives LB of 0.966. <br>\nEfficientNet L2 with an image size of 768 gives 0.967. It would have been interesting to increase the image size for EffL2 unfortunately it seems impossible with TPUv3.<br>\nHowever, EfficientNetB6 was not adding enought diversity in our ensembling (no big increase for CV/LB) so we discard it for the next steps.</p></li>\n<li><p>2nd step : self distillation. We include the softlabel of the models trained in the 1st step, such as the softlabel represent 30% of the \"final label\".<br>\nFor example, the label of the image is 1.0 and the model predict a probability of 0.8. The final label will be equal to : 0.7*hard label + 0.3 * softlabel = 0.7 + 0.24 = 0.94<br>\nThat  gives a boost of 0.001 and 0.002 depending of the model. The ensembling of these 3 models gives an LB of 0.970.</p></li>\n<li><p>3rd step: This final step use the pseudo label from the NIH dataset. We average the prediction of our 3 models from steps 2 to build the pseudo label. We train only EfficientNetB7 on image size of 1024 as we were out of time. (We also trained an Xception model but the submission randomly failed the last day :/ )<br>\nThe results of the EffB7 was 0.969 on LB.</p></li>\n</ul>\n<p>Our final submission was ensembling the EffB7 of Step3 and effL2 of step2 and the pytorch model described by <a href=\"https://www.kaggle.com/tonyxu\" target=\"_blank\">@tonyxu</a> </p>",
      "rawMarkdown": "Thank you to Kaggle and the organizers for hosting this competition. \n\nregarding the Tensorflow part :\n\nWe use for our solution only EfficientNet architectures for the TF solution. In fact, we are quite limited regarding the different strong models in TF, only inceptionresnetV2 was giving interesting results, but it was still lower than EfficientNet B6/B7/L2.\nWe use 3 resolutions : 768, 1024, 1280 and lights data aug : rotation, translation, shear, zoom, color (contrast etc)\n\n- 1st step : Train EfficientNetB6/B7 on image size of 1024 gives respectively a LB of 0.966 and 0.965. EfficientNetB7 on 1280 gives LB of 0.966. \nEfficientNet L2 with an image size of 768 gives 0.967. It would have been interesting to increase the image size for EffL2 unfortunately it seems impossible with TPUv3.\nHowever, EfficientNetB6 was not adding enought diversity in our ensembling (no big increase for CV/LB) so we discard it for the next steps.\n\n- 2nd step : self distillation. We include the softlabel of the models trained in the 1st step, such as the softlabel represent 30% of the \"final label\".\nFor example, the label of the image is 1.0 and the model predict a probability of 0.8. The final label will be equal to : 0.7*hard label + 0.3 * softlabel = 0.7 + 0.24 = 0.94\nThat  gives a boost of 0.001 and 0.002 depending of the model. The ensembling of these 3 models gives an LB of 0.970.\n\n- 3rd step: This final step use the pseudo label from the NIH dataset. We average the prediction of our 3 models from steps 2 to build the pseudo label. We train only EfficientNetB7 on image size of 1024 as we were out of time. (We also trained an Xception model but the submission randomly failed the last day :/ )\nThe results of the EffB7 was 0.969 on LB.\n\nOur final submission was ensembling the EffB7 of Step3 and effL2 of step2 and the pytorch model described by @tonyxu ",
      "votes": 6
    },
    {
      "id": 1244491,
      "postDate": "2021-03-19T02:54:13.527Z",
      "content": "<p>Congrats on our new Grandmaster!!! Really happy to work with you. Thanks so much for teaching me how to use TPU in GCP.</p>",
      "rawMarkdown": "Congrats on our new Grandmaster!!! Really happy to work with you. Thanks so much for teaching me how to use TPU in GCP.",
      "votes": 3
    },
    {
      "id": 1248639,
      "postDate": "2021-03-22T18:13:46.827Z",
      "content": "<p>Congrats on 10th place <a href=\"https://www.kaggle.com/ludovick\" target=\"_blank\">@ludovick</a> <a href=\"https://www.kaggle.com/woshifym\" target=\"_blank\">@woshifym</a>.<br>\nAnd competitions GM! <a href=\"https://www.kaggle.com/tonyxu\" target=\"_blank\">@tonyxu</a></p>\n<p>It is amazing to achieve that 10th place with only the classification model.</p>\n<p>It seems 3-stage trainings.<br>\n<strong>1) Pytorch Part</strong> - It also has sub 3 or 4-stage training.<br>\n<strong>2) TF Part</strong> - Using 1) models and generating soft labels. And train student models (self-distillation)<br>\n<strong>3) Ensemble All</strong></p>\n<p>Could you explain more about <strong>2)</strong>?<br>\n<strong>2-1)</strong> Generating soft labels using 1) models and retraining all of 1) models?<br>\n<strong>2-2)</strong> When is the right time to use the pseudo label?  (already used in sub 3 or 4-stage training?)</p>\n<p>In other competitions, the ensemble didn't work out well with distillation models together.<br>\nI am curious if the ensemble worked well in self-distillation using soft labels.<br>\n(Any tips when ensemble in this cases?)</p>\n<p>p.s. cvpr2020-plant-pathology was my first competition in kaggle. It's nice to see great solution again in this competition.<br>\np.s2. My questions may be lacking because I am not participating deeply in the competition. </p>\n<p>Thanks for sharing! :)</p>",
      "rawMarkdown": "Congrats on 10th place @ludovick @woshifym.\nAnd competitions GM! @tonyxu\n\nIt is amazing to achieve that 10th place with only the classification model.\n\nIt seems 3-stage trainings.\n**1) Pytorch Part** - It also has sub 3 or 4-stage training.\n**2) TF Part** - Using 1) models and generating soft labels. And train student models (self-distillation)\n**3) Ensemble All**\n\nCould you explain more about **2)**?\n**2-1)** Generating soft labels using 1) models and retraining all of 1) models?\n**2-2)** When is the right time to use the pseudo label?  (already used in sub 3 or 4-stage training?)\n\nIn other competitions, the ensemble didn't work out well with distillation models together.\nI am curious if the ensemble worked well in self-distillation using soft labels.\n(Any tips when ensemble in this cases?)\n\np.s. cvpr2020-plant-pathology was my first competition in kaggle. It's nice to see great solution again in this competition.\np.s2. My questions may be lacking because I am not participating deeply in the competition. \n\nThanks for sharing! :)",
      "votes": 1,
      "replies": [
        {
          "id": 1250060,
          "postDate": "2021-03-23T18:38:16.040Z",
          "content": "<p>hello,</p>\n<p>the pytorch and TF part are completely independant, we merge the different models in the part 3) in order to boost our final submission.</p>\n<p>For the TF part it can be decomposed into 3 steps:</p>\n<ul>\n<li><p>the first step is a regular training using only light data augmentation such as rotation/translation/shear/color etc. We use EfficientNetB6/B7/L2 on different image size such as 768</p></li>\n<li><p>the second step will use the softlabel from the first step to train new EfficientNet model (same architecture and same split of the data (5 folds) [ I give an exemple in an other reply in this discussion]</p></li>\n<li><p>the third step will use Pseudo Label. PL is obtained from the NIH dataset and the soft prediction (probabilities) of all the models trained in step 2. We trained only one model for PL as we were out of time. (EfficientNet B7 with an image size of 1024)</p></li>\n</ul>\n<p>Hope it answers your question  :) </p>",
          "rawMarkdown": "hello,\n\nthe pytorch and TF part are completely independant, we merge the different models in the part 3) in order to boost our final submission.\n\nFor the TF part it can be decomposed into 3 steps:\n- the first step is a regular training using only light data augmentation such as rotation/translation/shear/color etc. We use EfficientNetB6/B7/L2 on different image size such as 768\n\n- the second step will use the softlabel from the first step to train new EfficientNet model (same architecture and same split of the data (5 folds) [ I give an exemple in an other reply in this discussion]\n\n- the third step will use Pseudo Label. PL is obtained from the NIH dataset and the soft prediction (probabilities) of all the models trained in step 2. We trained only one model for PL as we were out of time. (EfficientNet B7 with an image size of 1024)\n\nHope it answers your question  :) ",
          "votes": 1
        },
        {
          "id": 1250094,
          "postDate": "2021-03-23T19:20:57.080Z",
          "content": "<p>Thanks for your answer. :) <a href=\"https://www.kaggle.com/ludovick\" target=\"_blank\">@ludovick</a> </p>",
          "rawMarkdown": "Thanks for your answer. :) @ludovick ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1244920,
      "postDate": "2021-03-19T10:13:50.990Z",
      "content": "<p>Congrats on becoming GM <a href=\"https://www.kaggle.com/tonyxu\" target=\"_blank\">@tonyxu</a> </p>",
      "rawMarkdown": "Congrats on becoming GM @tonyxu ",
      "votes": 1,
      "replies": [
        {
          "id": 1245821,
          "postDate": "2021-03-20T07:42:57.493Z",
          "content": "<p>Thank you. You can also be a GM sooner or later😃</p>",
          "rawMarkdown": "Thank you. You can also be a GM sooner or later😃"
        }
      ]
    },
    {
      "id": 1244579,
      "postDate": "2021-03-19T04:51:26.097Z",
      "content": "<p>Congratulation in Gold Win and to the new Grandmaster Candidate<br>\nI think PseudoLabelling were the main breakthrough for this competition.</p>",
      "rawMarkdown": "Congratulation in Gold Win and to the new Grandmaster Candidate\nI think PseudoLabelling were the main breakthrough for this competition.",
      "replies": [
        {
          "id": 1244634,
          "postDate": "2021-03-19T05:34:20.423Z",
          "content": "<p>Thanks. We didn't try segmentation method in our solution. Soft label and Pesudo Label is the most important tricks in out solution.</p>",
          "rawMarkdown": "Thanks. We didn't try segmentation method in our solution. Soft label and Pesudo Label is the most important tricks in out solution.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1246811,
      "postDate": "2021-03-21T06:13:32.630Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1245610,
      "author_name": "Shiro",
      "author_url": "",
      "post_date": "2021-03-20T00:15:24.033000",
      "content": "<p>Thank you to Kaggle and the organizers for hosting this competition. </p>\n<p>regarding the Tensorflow part :</p>\n<p>We use for our solution only EfficientNet architectures for the TF solution. In fact, we are quite limited regarding the different strong models in TF, only inceptionresnetV2 was giving interesting results, but it was still lower than EfficientNet B6/B7/L2.<br>\nWe use 3 resolutions : 768, 1024, 1280 and lights data aug : rotation, translation, shear, zoom, color (contrast etc)</p>\n<ul>\n<li><p>1st step : Train EfficientNetB6/B7 on image size of 1024 gives respectively a LB of 0.966 and 0.965. EfficientNetB7 on 1280 gives LB of 0.966. <br>\nEfficientNet L2 with an image size of 768 gives 0.967. It would have been interesting to increase the image size for EffL2 unfortunately it seems impossible with TPUv3.<br>\nHowever, EfficientNetB6 was not adding enought diversity in our ensembling (no big increase for CV/LB) so we discard it for the next steps.</p></li>\n<li><p>2nd step : self distillation. We include the softlabel of the models trained in the 1st step, such as the softlabel represent 30% of the \"final label\".<br>\nFor example, the label of the image is 1.0 and the model predict a probability of 0.8. The final label will be equal to : 0.7*hard label + 0.3 * softlabel = 0.7 + 0.24 = 0.94<br>\nThat  gives a boost of 0.001 and 0.002 depending of the model. The ensembling of these 3 models gives an LB of 0.970.</p></li>\n<li><p>3rd step: This final step use the pseudo label from the NIH dataset. We average the prediction of our 3 models from steps 2 to build the pseudo label. We train only EfficientNetB7 on image size of 1024 as we were out of time. (We also trained an Xception model but the submission randomly failed the last day :/ )<br>\nThe results of the EffB7 was 0.969 on LB.</p></li>\n</ul>\n<p>Our final submission was ensembling the EffB7 of Step3 and effL2 of step2 and the pytorch model described by <a href=\"https://www.kaggle.com/tonyxu\" target=\"_blank\">@tonyxu</a> </p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 1244491,
      "author_name": "LeoF",
      "author_url": "",
      "post_date": "2021-03-19T02:54:13.527000",
      "content": "<p>Congrats on our new Grandmaster!!! Really happy to work with you. Thanks so much for teaching me how to use TPU in GCP.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1248639,
      "author_name": "Heroseo",
      "author_url": "",
      "post_date": "2021-03-22T18:13:46.827000",
      "content": "<p>Congrats on 10th place <a href=\"https://www.kaggle.com/ludovick\" target=\"_blank\">@ludovick</a> <a href=\"https://www.kaggle.com/woshifym\" target=\"_blank\">@woshifym</a>.<br>\nAnd competitions GM! <a href=\"https://www.kaggle.com/tonyxu\" target=\"_blank\">@tonyxu</a></p>\n<p>It is amazing to achieve that 10th place with only the classification model.</p>\n<p>It seems 3-stage trainings.<br>\n<strong>1) Pytorch Part</strong> - It also has sub 3 or 4-stage training.<br>\n<strong>2) TF Part</strong> - Using 1) models and generating soft labels. And train student models (self-distillation)<br>\n<strong>3) Ensemble All</strong></p>\n<p>Could you explain more about <strong>2)</strong>?<br>\n<strong>2-1)</strong> Generating soft labels using 1) models and retraining all of 1) models?<br>\n<strong>2-2)</strong> When is the right time to use the pseudo label?  (already used in sub 3 or 4-stage training?)</p>\n<p>In other competitions, the ensemble didn't work out well with distillation models together.<br>\nI am curious if the ensemble worked well in self-distillation using soft labels.<br>\n(Any tips when ensemble in this cases?)</p>\n<p>p.s. cvpr2020-plant-pathology was my first competition in kaggle. It's nice to see great solution again in this competition.<br>\np.s2. My questions may be lacking because I am not participating deeply in the competition. </p>\n<p>Thanks for sharing! :)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1250060,
          "author_name": "Shiro",
          "author_url": "",
          "post_date": "2021-03-23T18:38:16.040000",
          "content": "<p>hello,</p>\n<p>the pytorch and TF part are completely independant, we merge the different models in the part 3) in order to boost our final submission.</p>\n<p>For the TF part it can be decomposed into 3 steps:</p>\n<ul>\n<li><p>the first step is a regular training using only light data augmentation such as rotation/translation/shear/color etc. We use EfficientNetB6/B7/L2 on different image size such as 768</p></li>\n<li><p>the second step will use the softlabel from the first step to train new EfficientNet model (same architecture and same split of the data (5 folds) [ I give an exemple in an other reply in this discussion]</p></li>\n<li><p>the third step will use Pseudo Label. PL is obtained from the NIH dataset and the soft prediction (probabilities) of all the models trained in step 2. We trained only one model for PL as we were out of time. (EfficientNet B7 with an image size of 1024)</p></li>\n</ul>\n<p>Hope it answers your question  :) </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1250094,
          "author_name": "Heroseo",
          "author_url": "",
          "post_date": "2021-03-23T19:20:57.080000",
          "content": "<p>Thanks for your answer. :) <a href=\"https://www.kaggle.com/ludovick\" target=\"_blank\">@ludovick</a> </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1244920,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2021-03-19T10:13:50.990000",
      "content": "<p>Congrats on becoming GM <a href=\"https://www.kaggle.com/tonyxu\" target=\"_blank\">@tonyxu</a> </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1245821,
          "author_name": "toxu",
          "author_url": "",
          "post_date": "2021-03-20T07:42:57.493000",
          "content": "<p>Thank you. You can also be a GM sooner or later😃</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1244579,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2021-03-19T04:51:26.097000",
      "content": "<p>Congratulation in Gold Win and to the new Grandmaster Candidate<br>\nI think PseudoLabelling were the main breakthrough for this competition.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1244634,
          "author_name": "toxu",
          "author_url": "",
          "post_date": "2021-03-19T05:34:20.423000",
          "content": "<p>Thanks. We didn't try segmentation method in our solution. Soft label and Pesudo Label is the most important tricks in out solution.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1246811,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-21T06:13:32.630000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1244443": "First of all, I would like to thank Kaggle and the organizers for hosting this great competitions. Also, I would like to thank my teammates - @ludovick and @woshifym . Due to the time difference, we can work 24 hours a day.\n\n# Solutions\nOur solution can be divided into 2 parts, pytorch part and tensorflow part.\n\n## Pytorch part\nFor Pytorch part I would like to thank @ammarali32 for the 3-step method first. All our models based on 3-step method. We trained models with different backbones and different image size to get better results.\n\n* resnet200d + image size: 600  CV: 0.9578\n* ecaresnet269d + image size: 600 CV: 0.9584\n* resnest200e + image size: 640 CV:0.9625\n\nAfter that we found [Soft label method](https://github.com/alipay/cvpr2020-plant-pathology) performs great in this competition. So we trained all the models above using this method to form our 4-step method.\n\n## Tensorflow part\nMy teammates @ludovick did a great job using tensorflow. I would like to invite him to introduce his method in details later.\nFor brief summary:\n\n* Efficient B7 + 1024 image size\n* Efficient L2 + 768 image size\n* Soft label method\n* Pesudo Label\n\n## Ensemble\nFor ensemble, we use the OOF files to train a linear regression model. Then we use this linear model to get the weight for each model. We select our best CV and best public score in our final submissions. Unfortunately, we just missed out best ensemble which is the second best in public score.\n\nFinal submission one \nCV: 0.9692 Public LB: 0.971 Private LB: 0.973\n* resnet200d + 600\n* ecaresnet269d + 600\n* resnest200e + 640\n* Efficient B7 + 1024 + soft label + PL\n* Efficient L2 + 768 + soft label\n\nFinal submission two\nCV: 0.9689 Public LB: 0.971 Private LB: 0.974\n* resnet200d + 600\n* resnest200e + 640\n* Efficient B7 + 1024 + soft label + PL\n* Efficient B7 + 1024 + PL\n* Efficient L2 + 768 + soft label\n\n## Tricks\n* Larger model\n* Higher resolution\n* Soft Label\n* Pesudo Label\n\n## Acknowledge\nThanks to TFRC team for providing us free TPU.",
    "1245610": "Thank you to Kaggle and the organizers for hosting this competition. \n\nregarding the Tensorflow part :\n\nWe use for our solution only EfficientNet architectures for the TF solution. In fact, we are quite limited regarding the different strong models in TF, only inceptionresnetV2 was giving interesting results, but it was still lower than EfficientNet B6/B7/L2.\nWe use 3 resolutions : 768, 1024, 1280 and lights data aug : rotation, translation, shear, zoom, color (contrast etc)\n\n- 1st step : Train EfficientNetB6/B7 on image size of 1024 gives respectively a LB of 0.966 and 0.965. EfficientNetB7 on 1280 gives LB of 0.966. \nEfficientNet L2 with an image size of 768 gives 0.967. It would have been interesting to increase the image size for EffL2 unfortunately it seems impossible with TPUv3.\nHowever, EfficientNetB6 was not adding enought diversity in our ensembling (no big increase for CV/LB) so we discard it for the next steps.\n\n- 2nd step : self distillation. We include the softlabel of the models trained in the 1st step, such as the softlabel represent 30% of the \"final label\".\nFor example, the label of the image is 1.0 and the model predict a probability of 0.8. The final label will be equal to : 0.7*hard label + 0.3 * softlabel = 0.7 + 0.24 = 0.94\nThat  gives a boost of 0.001 and 0.002 depending of the model. The ensembling of these 3 models gives an LB of 0.970.\n\n- 3rd step: This final step use the pseudo label from the NIH dataset. We average the prediction of our 3 models from steps 2 to build the pseudo label. We train only EfficientNetB7 on image size of 1024 as we were out of time. (We also trained an Xception model but the submission randomly failed the last day :/ )\nThe results of the EffB7 was 0.969 on LB.\n\nOur final submission was ensembling the EffB7 of Step3 and effL2 of step2 and the pytorch model described by @tonyxu ",
    "1244491": "Congrats on our new Grandmaster!!! Really happy to work with you. Thanks so much for teaching me how to use TPU in GCP.",
    "1248639": "Congrats on 10th place @ludovick @woshifym.\nAnd competitions GM! @tonyxu\n\nIt is amazing to achieve that 10th place with only the classification model.\n\nIt seems 3-stage trainings.\n**1) Pytorch Part** - It also has sub 3 or 4-stage training.\n**2) TF Part** - Using 1) models and generating soft labels. And train student models (self-distillation)\n**3) Ensemble All**\n\nCould you explain more about **2)**?\n**2-1)** Generating soft labels using 1) models and retraining all of 1) models?\n**2-2)** When is the right time to use the pseudo label?  (already used in sub 3 or 4-stage training?)\n\nIn other competitions, the ensemble didn't work out well with distillation models together.\nI am curious if the ensemble worked well in self-distillation using soft labels.\n(Any tips when ensemble in this cases?)\n\np.s. cvpr2020-plant-pathology was my first competition in kaggle. It's nice to see great solution again in this competition.\np.s2. My questions may be lacking because I am not participating deeply in the competition. \n\nThanks for sharing! :)",
    "1244920": "Congrats on becoming GM @tonyxu ",
    "1244579": "Congratulation in Gold Win and to the new Grandmaster Candidate\nI think PseudoLabelling were the main breakthrough for this competition.",
    "1246811": ""
  }
}