{
  "id": 226715,
  "title": "23rd Place Solution ",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/226715",
  "author_name": "Tawara",
  "post_date": "2021-03-17T12:59:47.006000",
  "votes": 42,
  "comment_count": 31,
  "views": 0,
  "content": "<p>First of all, congrats to all the teams got the medal and participants finished this competition, and thanks to Kaggle team and Competition host.<br>\nI learned many things especially teacher-student learning through this competition.</p>\n<p>I'd like to say special thanks to <a href=\"https://www.kaggle.com/ammarali32\" target=\"_blank\">@ammarali32</a> . I was about to give up in the middle of competition, but reading <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/215910\" target=\"_blank\">his interesting approach</a> made me resume.</p>\n<p>My approach is not so special . Then I show brief summary here.</p>\n<h3>Summary</h3>\n<p>Final submission(Public: 0.970, Private: 0.972) is averaging of the following 6 models.<br>\nI used image size of 640x640 for all models.</p>\n<ul>\n<li>model1: ResNeSt200 (Public: 0.965, Private: 0.966)<ul>\n<li>pretrained model: ImageNet</li>\n<li>Classification Head: Single MLP</li>\n<li>training: fine-tuning on competition data<br>\n<br></li></ul></li>\n<li>model2: ResNet200d (Public: 0.968, Private: 0.971)<ul>\n<li>pretrained model: <a href=\"https://www.kaggle.com/ammarali32/startingpointschestx\" target=\"_blank\">@ammarali32 's starting points</a></li>\n<li>Classification Head:  Multi Spatial-Attention Head</li>\n<li>training: fine-tuning on competition data<br>\n<br></li></ul></li>\n<li>model3: EfficientNetB5-NoisyStudnet (Public: 0.966, Private: 0.969)<ul>\n<li>pretrained model: <a href=\"https://www.kaggle.com/ammarali32/startingpointschestx\" target=\"_blank\">@ammarali32 's starting points</a></li>\n<li>Classification Head:  Multi Spatial-Attention Head</li>\n<li>training: fine-tuning on competition data<br>\n<br></li></ul></li>\n<li>model4: SE-ResNet152d (Public: 0.967, Private: 0.968)<ul>\n<li>pretrained model: <a href=\"https://www.kaggle.com/ammarali32/startingpointschestx\" target=\"_blank\">@ammarali32 's starting points</a></li>\n<li>Classification Head:  Multi Spatial-Attention Head</li>\n<li>training: fine-tuning on competition data<br>\n<br></li></ul></li>\n<li>model5: ResNeSt200e (Public: 0.966, Private: 0.967)<ul>\n<li>pretrained model: trained model on <a href=\"https://www.kaggle.com/nih-chest-xrays/data\" target=\"_blank\">NIH Chest X-rays</a> by Teacher-Student Training </li>\n<li>Classification Head:  Multi Spatial-Attention Head</li>\n<li>training: fine-tuning on competition data<br>\n<br></li></ul></li>\n<li>model6: ResNet200d (Public: 0.966, Private: 0.970)<ul>\n<li>pretrained model: <a href=\"https://www.kaggle.com/ammarali32/startingpointschestx\" target=\"_blank\">@ammarali32 's starting points</a></li>\n<li>Classification Head:  Multi-Head Attention with <strong>Segmentation</strong> Branch</li>\n<li>training:<ul>\n<li>stage1: trained on only annotated data</li>\n<li>stage2: trained on only <strong>non</strong>-annotated data</li></ul></li></ul></li>\n</ul>\n<p>I really wanted to do teacher-student training on <a href=\"https://www.kaggle.com/nih-chest-xrays/data\" target=\"_blank\">NIH Chest X-rays</a> for all the models. But I used <a href=\"https://www.kaggle.com/ammarali32/startingpointschestx\" target=\"_blank\">@ammarali32 's starting points</a> due to lack of time and computing resources.</p>\n<p>As a matter of fact, you can achieve Private 0.972(silver?) by averaging model 2,3,4 (<a href=\"https://www.kaggle.com/ttahara/ranzcr-ensemble-fine-tuned-multi-head-models\" target=\"_blank\">submission notebook</a>), which is combination of <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207230\" target=\"_blank\">my multi-head approach</a> and <a href=\"https://www.kaggle.com/ammarali32/startingpointschestx\" target=\"_blank\">@ammarali32 's starting points</a>.  <br>\nI wanted to exceed this score with a single model, but could not 😓</p>",
  "messages": [
    {
      "id": 1242170,
      "postDate": "2021-03-17T12:59:47.007Z",
      "content": "<p>First of all, congrats to all the teams got the medal and participants finished this competition, and thanks to Kaggle team and Competition host.<br>\nI learned many things especially teacher-student learning through this competition.</p>\n<p>I'd like to say special thanks to <a href=\"https://www.kaggle.com/ammarali32\" target=\"_blank\">@ammarali32</a> . I was about to give up in the middle of competition, but reading <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/215910\" target=\"_blank\">his interesting approach</a> made me resume.</p>\n<p>My approach is not so special . Then I show brief summary here.</p>\n<h3>Summary</h3>\n<p>Final submission(Public: 0.970, Private: 0.972) is averaging of the following 6 models.<br>\nI used image size of 640x640 for all models.</p>\n<ul>\n<li>model1: ResNeSt200 (Public: 0.965, Private: 0.966)<ul>\n<li>pretrained model: ImageNet</li>\n<li>Classification Head: Single MLP</li>\n<li>training: fine-tuning on competition data<br>\n<br></li></ul></li>\n<li>model2: ResNet200d (Public: 0.968, Private: 0.971)<ul>\n<li>pretrained model: <a href=\"https://www.kaggle.com/ammarali32/startingpointschestx\" target=\"_blank\">@ammarali32 's starting points</a></li>\n<li>Classification Head:  Multi Spatial-Attention Head</li>\n<li>training: fine-tuning on competition data<br>\n<br></li></ul></li>\n<li>model3: EfficientNetB5-NoisyStudnet (Public: 0.966, Private: 0.969)<ul>\n<li>pretrained model: <a href=\"https://www.kaggle.com/ammarali32/startingpointschestx\" target=\"_blank\">@ammarali32 's starting points</a></li>\n<li>Classification Head:  Multi Spatial-Attention Head</li>\n<li>training: fine-tuning on competition data<br>\n<br></li></ul></li>\n<li>model4: SE-ResNet152d (Public: 0.967, Private: 0.968)<ul>\n<li>pretrained model: <a href=\"https://www.kaggle.com/ammarali32/startingpointschestx\" target=\"_blank\">@ammarali32 's starting points</a></li>\n<li>Classification Head:  Multi Spatial-Attention Head</li>\n<li>training: fine-tuning on competition data<br>\n<br></li></ul></li>\n<li>model5: ResNeSt200e (Public: 0.966, Private: 0.967)<ul>\n<li>pretrained model: trained model on <a href=\"https://www.kaggle.com/nih-chest-xrays/data\" target=\"_blank\">NIH Chest X-rays</a> by Teacher-Student Training </li>\n<li>Classification Head:  Multi Spatial-Attention Head</li>\n<li>training: fine-tuning on competition data<br>\n<br></li></ul></li>\n<li>model6: ResNet200d (Public: 0.966, Private: 0.970)<ul>\n<li>pretrained model: <a href=\"https://www.kaggle.com/ammarali32/startingpointschestx\" target=\"_blank\">@ammarali32 's starting points</a></li>\n<li>Classification Head:  Multi-Head Attention with <strong>Segmentation</strong> Branch</li>\n<li>training:<ul>\n<li>stage1: trained on only annotated data</li>\n<li>stage2: trained on only <strong>non</strong>-annotated data</li></ul></li></ul></li>\n</ul>\n<p>I really wanted to do teacher-student training on <a href=\"https://www.kaggle.com/nih-chest-xrays/data\" target=\"_blank\">NIH Chest X-rays</a> for all the models. But I used <a href=\"https://www.kaggle.com/ammarali32/startingpointschestx\" target=\"_blank\">@ammarali32 's starting points</a> due to lack of time and computing resources.</p>\n<p>As a matter of fact, you can achieve Private 0.972(silver?) by averaging model 2,3,4 (<a href=\"https://www.kaggle.com/ttahara/ranzcr-ensemble-fine-tuned-multi-head-models\" target=\"_blank\">submission notebook</a>), which is combination of <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207230\" target=\"_blank\">my multi-head approach</a> and <a href=\"https://www.kaggle.com/ammarali32/startingpointschestx\" target=\"_blank\">@ammarali32 's starting points</a>.  <br>\nI wanted to exceed this score with a single model, but could not 😓</p>",
      "rawMarkdown": "First of all, congrats to all the teams got the medal and participants finished this competition, and thanks to Kaggle team and Competition host.\nI learned many things especially teacher-student learning through this competition.\n\nI'd like to say special thanks to @ammarali32 . I was about to give up in the middle of competition, but reading [his interesting approach](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/215910) made me resume.\n\nMy approach is not so special . Then I show brief summary here.\n\n### Summary\n\nFinal submission(Public: 0.970, Private: 0.972) is averaging of the following 6 models.\nI used image size of 640x640 for all models.\n\n* model1: ResNeSt200 (Public: 0.965, Private: 0.966)\n    * pretrained model: ImageNet\n    * Classification Head: Single MLP\n    * training: fine-tuning on competition data\n<br>\n* model2: ResNet200d (Public: 0.968, Private: 0.971)\n    * pretrained model: [@ammarali32 's starting points](https://www.kaggle.com/ammarali32/startingpointschestx)\n    * Classification Head:  Multi Spatial-Attention Head\n    * training: fine-tuning on competition data\n<br>\n* model3: EfficientNetB5-NoisyStudnet (Public: 0.966, Private: 0.969)\n    * pretrained model: [@ammarali32 's starting points](https://www.kaggle.com/ammarali32/startingpointschestx)\n    * Classification Head:  Multi Spatial-Attention Head\n    * training: fine-tuning on competition data\n<br>\n* model4: SE-ResNet152d (Public: 0.967, Private: 0.968)\n    * pretrained model: [@ammarali32 's starting points](https://www.kaggle.com/ammarali32/startingpointschestx)\n    * Classification Head:  Multi Spatial-Attention Head\n    * training: fine-tuning on competition data\n<br>\n* model5: ResNeSt200e (Public: 0.966, Private: 0.967)\n    * pretrained model: trained model on [NIH Chest X-rays](https://www.kaggle.com/nih-chest-xrays/data) by Teacher-Student Training \n    * Classification Head:  Multi Spatial-Attention Head\n    * training: fine-tuning on competition data\n<br>\n* model6: ResNet200d (Public: 0.966, Private: 0.970)\n    * pretrained model: [@ammarali32 's starting points](https://www.kaggle.com/ammarali32/startingpointschestx)\n    * Classification Head:  Multi-Head Attention with **Segmentation** Branch\n    * training:\n        * stage1: trained on only annotated data\n        * stage2: trained on only **non**-annotated data\n\nI really wanted to do teacher-student training on [NIH Chest X-rays](https://www.kaggle.com/nih-chest-xrays/data) for all the models. But I used [@ammarali32 's starting points](https://www.kaggle.com/ammarali32/startingpointschestx) due to lack of time and computing resources.\n\nAs a matter of fact, you can achieve Private 0.972(silver?) by averaging model 2,3,4 ([submission notebook](https://www.kaggle.com/ttahara/ranzcr-ensemble-fine-tuned-multi-head-models)), which is combination of [my multi-head approach](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207230) and [@ammarali32 's starting points](https://www.kaggle.com/ammarali32/startingpointschestx).  \nI wanted to exceed this score with a single model, but could not 😓",
      "votes": 42
    },
    {
      "id": 1244276,
      "postDate": "2021-03-18T21:22:57.667Z",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> for the thought-provoking presentation. Congratulations on achieving 23rd on the leadership board.<br>\nWhat variance did you notice with size? You used  640x640. I was thinking if 1024x1024 would have given an edge.</p>",
      "rawMarkdown": "Thank you @ttahara for the thought-provoking presentation. Congratulations on achieving 23rd on the leadership board.\nWhat variance did you notice with size? You used  640x640. I was thinking if 1024x1024 would have given an edge.",
      "votes": 1,
      "replies": [
        {
          "id": 1244944,
          "postDate": "2021-03-19T10:37:34.130Z",
          "content": "<p>Thanks.</p>\n<p>I tried 768x768 with small models, but not 1024x1024.</p>\n<p>I don't have time and computation resources for large models such as ResNet200d with 1024x1024 😂</p>",
          "rawMarkdown": "Thanks.\n\nI tried 768x768 with small models, but not 1024x1024.\n\nI don't have time and computation resources for large models such as ResNet200d with 1024x1024 😂",
          "votes": 2
        }
      ]
    },
    {
      "id": 1243535,
      "postDate": "2021-03-18T09:47:36.810Z",
      "content": "<p>Thank you for the write up and sharing the multi head approach.<br>\nMay I know what is the different between EfficientNetB5-NoisyStudnet  and normal EfficientNetB5 ?</p>\n<blockquote>\n  <p>As a matter of fact, you can achieve Private 0.972(silver?) by averaging model 2,3,4</p>\n</blockquote>\n<p>This is what I did, but my EfficientNet gets lower Public Score (0.964) compared to Seresnet (0.967), although on Private score it is higher (0.969 vs 0.968). </p>",
      "rawMarkdown": "Thank you for the write up and sharing the multi head approach.\nMay I know what is the different between EfficientNetB5-NoisyStudnet  and normal EfficientNetB5 ?\n\n> As a matter of fact, you can achieve Private 0.972(silver?) by averaging model 2,3,4\n\nThis is what I did, but my EfficientNet gets lower Public Score (0.964) compared to Seresnet (0.967), although on Private score it is higher (0.969 vs 0.968). ",
      "votes": 1,
      "replies": [
        {
          "id": 1243563,
          "postDate": "2021-03-18T10:13:11.223Z",
          "content": "<p>I think there is no difference between them in terms of the model architecture. But pre-trained models of them we can use by <a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">timm</a> are trained by different ways.</p>",
          "rawMarkdown": "I think there is no difference between them in terms of the model architecture. But pre-trained models of them we can use by [timm](https://github.com/rwightman/pytorch-image-models) are trained by different ways.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1243150,
      "postDate": "2021-03-18T03:43:02.563Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a>  on your finish!  You were also a very valuable contributor to this competition. Multi Spatial-Attention Head was pretty special here and helped a lot of people.  So many great ideas here. Thanks. </p>",
      "rawMarkdown": "Congratulations @ttahara  on your finish!  You were also a very valuable contributor to this competition. Multi Spatial-Attention Head was pretty special here and helped a lot of people.  So many great ideas here. Thanks. ",
      "votes": 1,
      "replies": [
        {
          "id": 1243158,
          "postDate": "2021-03-18T04:00:31.247Z",
          "content": "<p>Thanks! </p>\n<p>I'm very happy that my concepts and baseline are helpful for participants.</p>",
          "rawMarkdown": "Thanks! \n\nI'm very happy that my concepts and baseline are helpful for participants.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1242739,
      "postDate": "2021-03-17T19:02:57.247Z",
      "content": "<p>Thanks for your special Spatial Attention approach , a lot of people would not have won a medal without it , beautiful concept . I also tried self attention on top of it I wonder why it didn't work</p>",
      "rawMarkdown": "Thanks for your special Spatial Attention approach , a lot of people would not have won a medal without it , beautiful concept . I also tried self attention on top of it I wonder why it didn't work",
      "votes": 1,
      "replies": [
        {
          "id": 1243163,
          "postDate": "2021-03-18T04:07:45.207Z",
          "content": "<p>Thanks!</p>\n<p>I didn't try self-attention but, IMO, calculating attention weight by neighborhood features(Spatial-Attention) may be better than by all the features(Self-Attention).</p>",
          "rawMarkdown": "Thanks!\n\nI didn't try self-attention but, IMO, calculating attention weight by neighborhood features(Spatial-Attention) may be better than by all the features(Self-Attention).",
          "votes": 2
        },
        {
          "id": 1243635,
          "postDate": "2021-03-18T11:31:33.850Z",
          "content": "<p>Yeah it makes sense <br>\nThanks</p>",
          "rawMarkdown": "Yeah it makes sense \nThanks"
        }
      ]
    },
    {
      "id": 1242291,
      "postDate": "2021-03-17T14:23:47.887Z",
      "content": "<p>Congrats on 23th place <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> </p>",
      "rawMarkdown": "Congrats on 23th place @ttahara ",
      "votes": 1,
      "replies": [
        {
          "id": 1242292,
          "postDate": "2021-03-17T14:24:18.173Z",
          "content": "<p>Thanks!    </p>",
          "rawMarkdown": "Thanks!    "
        }
      ]
    },
    {
      "id": 1242214,
      "postDate": "2021-03-17T13:32:21.643Z",
      "content": "<p>Congrats!</p>\n<p>Thank you for your great training notebook. I used it to train one of our ensemble models. <br>\nI couldn't get 0.972 with single model too. I have similar result to your resnet 200d model 2 (CV: 0.966 Public: 0.967, Private: 0.971).</p>\n<p>I think it should be able to get 0.972 single model if you add semgment layers like other top competitors.</p>",
      "rawMarkdown": "Congrats!\n\nThank you for your great training notebook. I used it to train one of our ensemble models. \nI couldn't get 0.972 with single model too. I have similar result to your resnet 200d model 2 (CV: 0.966 Public: 0.967, Private: 0.971).\n\nI think it should be able to get 0.972 single model if you add semgment layers like other top competitors.\n",
      "votes": 1,
      "replies": [
        {
          "id": 1242229,
          "postDate": "2021-03-17T13:39:37.353Z",
          "content": "<p>Thanks.</p>\n<p>I tried model 6(with segmentation) in last few days but there was no sufficient training time.</p>",
          "rawMarkdown": "Thanks.\n\nI tried model 6(with segmentation) in last few days but there was no sufficient training time."
        }
      ]
    },
    {
      "id": 1243116,
      "postDate": "2021-03-18T03:05:02.423Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> on your position.<br>\nBTW, I achieved 0.972 on private lb just by using ResNet200d with Multi Spatial-Attention Head only (5 fold ensemble of img size 640) + 1 fold of Img size 1024.</p>",
      "rawMarkdown": "Congrats @ttahara on your position.\nBTW, I achieved 0.972 on private lb just by using ResNet200d with Multi Spatial-Attention Head only (5 fold ensemble of img size 640) + 1 fold of Img size 1024.\n ",
      "votes": 2,
      "replies": [
        {
          "id": 1243153,
          "postDate": "2021-03-18T03:54:52.533Z",
          "content": "<p>Thanks!</p>\n<blockquote>\n  <p>I achieved 0.972 on private lb just by using ResNet200d with Multi Spatial-Attention Head only</p>\n</blockquote>\n<p>That's great! </p>\n<p>I don't understand \"+ 1 fold of Img size 1024.\" You mean doing ensemble 6 predictions (<strong>5</strong> folds whose training and inference was done with img size 640 and <strong>1</strong> folds with img size 1024), right?</p>",
          "rawMarkdown": "Thanks!\n\n\n>  I achieved 0.972 on private lb just by using ResNet200d with Multi Spatial-Attention Head only\n\nThat's great! \n\nI don't understand \"+ 1 fold of Img size 1024.\" You mean doing ensemble 6 predictions (**5** folds whose training and inference was done with img size 640 and **1** folds with img size 1024), right?",
          "votes": 2
        },
        {
          "id": 1243238,
          "postDate": "2021-03-18T05:26:35.043Z",
          "content": "<p>yup , thats right.</p>",
          "rawMarkdown": "yup , thats right.",
          "votes": 2
        },
        {
          "id": 1243244,
          "postDate": "2021-03-18T05:31:26.137Z",
          "content": "<p>I see. Thanks.</p>",
          "rawMarkdown": "I see. Thanks."
        }
      ]
    },
    {
      "id": 1242303,
      "postDate": "2021-03-17T14:33:37.243Z",
      "content": "<p>Congrats on your position! well deserved!! And thank you for your insight about Multi Head Spatial Attention.</p>",
      "rawMarkdown": "Congrats on your position! well deserved!! And thank you for your insight about Multi Head Spatial Attention.",
      "votes": 2,
      "replies": [
        {
          "id": 1242348,
          "postDate": "2021-03-17T15:01:23.277Z",
          "content": "<p>Thanks!    </p>",
          "rawMarkdown": "Thanks!    "
        }
      ]
    },
    {
      "id": 1242202,
      "postDate": "2021-03-17T13:25:40.693Z",
      "content": "<p>Congratulations !! I would like to thank you for your amazing work and contribution.  </p>",
      "rawMarkdown": "Congratulations !! I would like to thank you for your amazing work and contribution.  \n",
      "votes": 2,
      "replies": [
        {
          "id": 1242224,
          "postDate": "2021-03-17T13:38:17.850Z",
          "content": "<p>Thanks!</p>\n<p>I think your contribution is also great. Thank you so much again.</p>",
          "rawMarkdown": "Thanks!\n\nI think your contribution is also great. Thank you so much again.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1242188,
      "postDate": "2021-03-17T13:13:33.303Z",
      "content": "<p>May I know your training pipeline for efficientnet. I tried but cannot get a good results at all. </p>",
      "rawMarkdown": "May I know your training pipeline for efficientnet. I tried but cannot get a good results at all. ",
      "votes": 2,
      "replies": [
        {
          "id": 1242193,
          "postDate": "2021-03-17T13:21:42.220Z",
          "content": "<p>I trained EfficientNet-B5 by the same pipeline with other models. The different settings among models were batch size and learning rate because of model size.</p>",
          "rawMarkdown": "I trained EfficientNet-B5 by the same pipeline with other models. The different settings among models were batch size and learning rate because of model size.",
          "votes": 1
        },
        {
          "id": 1242268,
          "postDate": "2021-03-17T14:03:42.260Z",
          "content": "<p>What batch size did you use and did you need to use gradient accumulation/freezing BatchNorm layers? If not, what hardware did you use?</p>",
          "rawMarkdown": "What batch size did you use and did you need to use gradient accumulation/freezing BatchNorm layers? If not, what hardware did you use?"
        },
        {
          "id": 1242288,
          "postDate": "2021-03-17T14:18:07.017Z",
          "content": "<p>My GPU resource is a single TitanRTX(24GB) and I didn't use gradient accumulation/freezing BatchNorm layers.</p>\n<p>Batch size I used for each models is as follows:</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>image size</th>\n<th>batch size</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ResNet200d</td>\n<td>640x640</td>\n<td>16</td>\n</tr>\n<tr>\n<td>EfficientNet-B5</td>\n<td>640x640</td>\n<td>16</td>\n</tr>\n<tr>\n<td>SE-ResNet152d</td>\n<td>640x640</td>\n<td>16</td>\n</tr>\n<tr>\n<td>ResNeSt200e</td>\n<td>640x640</td>\n<td>15</td>\n</tr>\n</tbody>\n</table>\n<p>Sorry, I remember that I was trying to keep their batch size as same as possible. You can increase batch size of ResNet200d and SE-ResNet152d on TitanRTX.</p>",
          "rawMarkdown": "My GPU resource is a single TitanRTX(24GB) and I didn't use gradient accumulation/freezing BatchNorm layers.\n\nBatch size I used for each models is as follows:\n\n| Model | image size | batch size |\n|:------:|:-----------:|:-----------:|\n| ResNet200d | 640x640 | 16 |\n| EfficientNet-B5 | 640x640 | 16 |\n| SE-ResNet152d | 640x640 | 16 |\n| ResNeSt200e | 640x640 | 15 |\n\nSorry, I remember that I was trying to keep their batch size as same as possible. You can increase batch size of ResNet200d and SE-ResNet152d on TitanRTX.\n\n\n \n  ",
          "votes": 2
        }
      ]
    },
    {
      "id": 1243162,
      "postDate": "2021-03-18T04:04:39.293Z",
      "content": "<p>Hello,do they all trained on the whole data?</p>",
      "rawMarkdown": "Hello,do they all trained on the whole data?",
      "replies": [
        {
          "id": 1243168,
          "postDate": "2021-03-18T04:15:12.353Z",
          "content": "<p>Yes. I used whole the data for fine-tuning but did cross validation.</p>\n<p>I  split the competition data by Multi-Label Stratified Group K-Fold manner(K=5).</p>",
          "rawMarkdown": "Yes. I used whole the data for fine-tuning but did cross validation.\n\nI  split the competition data by Multi-Label Stratified Group K-Fold manner(K=5).",
          "votes": 1
        }
      ]
    },
    {
      "id": 1242249,
      "postDate": "2021-03-17T13:48:47.060Z",
      "content": "<p><a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> thank you for summarizing your solution! And also thanks for sharing very useful results/ideas during the competition 🙂</p>",
      "rawMarkdown": "@ttahara thank you for summarizing your solution! And also thanks for sharing very useful results/ideas during the competition 🙂",
      "replies": [
        {
          "id": 1242252,
          "postDate": "2021-03-17T13:51:24.833Z",
          "content": "<p>Thanks. I'm glad that my discussions/baselines are helpful to you 🙂</p>",
          "rawMarkdown": "Thanks. I'm glad that my discussions/baselines are helpful to you 🙂"
        }
      ]
    },
    {
      "id": 1246822,
      "postDate": "2021-03-21T06:20:54.760Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 1246832,
          "postDate": "2021-03-21T06:33:44.363Z",
          "content": "<p>Thanks 😃  </p>",
          "rawMarkdown": "Thanks 😃  "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1244276,
      "author_name": "Olusesi Adebisi",
      "author_url": "",
      "post_date": "2021-03-18T21:22:57.667000",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> for the thought-provoking presentation. Congratulations on achieving 23rd on the leadership board.<br>\nWhat variance did you notice with size? You used  640x640. I was thinking if 1024x1024 would have given an edge.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1244944,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-03-19T10:37:34.130000",
          "content": "<p>Thanks.</p>\n<p>I tried 768x768 with small models, but not 1024x1024.</p>\n<p>I don't have time and computation resources for large models such as ResNet200d with 1024x1024 😂</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1243535,
      "author_name": "yukiya",
      "author_url": "",
      "post_date": "2021-03-18T09:47:36.810000",
      "content": "<p>Thank you for the write up and sharing the multi head approach.<br>\nMay I know what is the different between EfficientNetB5-NoisyStudnet  and normal EfficientNetB5 ?</p>\n<blockquote>\n  <p>As a matter of fact, you can achieve Private 0.972(silver?) by averaging model 2,3,4</p>\n</blockquote>\n<p>This is what I did, but my EfficientNet gets lower Public Score (0.964) compared to Seresnet (0.967), although on Private score it is higher (0.969 vs 0.968). </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1243563,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-03-18T10:13:11.223000",
          "content": "<p>I think there is no difference between them in terms of the model architecture. But pre-trained models of them we can use by <a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">timm</a> are trained by different ways.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1243150,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "2021-03-18T03:43:02.563000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a>  on your finish!  You were also a very valuable contributor to this competition. Multi Spatial-Attention Head was pretty special here and helped a lot of people.  So many great ideas here. Thanks. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1243158,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-03-18T04:00:31.247000",
          "content": "<p>Thanks! </p>\n<p>I'm very happy that my concepts and baseline are helpful for participants.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1242739,
      "author_name": "Mr_KnowNothing",
      "author_url": "",
      "post_date": "2021-03-17T19:02:57.247000",
      "content": "<p>Thanks for your special Spatial Attention approach , a lot of people would not have won a medal without it , beautiful concept . I also tried self attention on top of it I wonder why it didn't work</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1243163,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-03-18T04:07:45.207000",
          "content": "<p>Thanks!</p>\n<p>I didn't try self-attention but, IMO, calculating attention weight by neighborhood features(Spatial-Attention) may be better than by all the features(Self-Attention).</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1243635,
          "author_name": "Mr_KnowNothing",
          "author_url": "",
          "post_date": "2021-03-18T11:31:33.850000",
          "content": "<p>Yeah it makes sense <br>\nThanks</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1242291,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2021-03-17T14:23:47.887000",
      "content": "<p>Congrats on 23th place <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1242292,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-03-17T14:24:18.173000",
          "content": "<p>Thanks!    </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1242214,
      "author_name": "SHih Chieh Lai",
      "author_url": "",
      "post_date": "2021-03-17T13:32:21.643000",
      "content": "<p>Congrats!</p>\n<p>Thank you for your great training notebook. I used it to train one of our ensemble models. <br>\nI couldn't get 0.972 with single model too. I have similar result to your resnet 200d model 2 (CV: 0.966 Public: 0.967, Private: 0.971).</p>\n<p>I think it should be able to get 0.972 single model if you add semgment layers like other top competitors.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1242229,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-03-17T13:39:37.353000",
          "content": "<p>Thanks.</p>\n<p>I tried model 6(with segmentation) in last few days but there was no sufficient training time.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1243116,
      "author_name": "RAHUL SINGH INDA",
      "author_url": "",
      "post_date": "2021-03-18T03:05:02.423000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> on your position.<br>\nBTW, I achieved 0.972 on private lb just by using ResNet200d with Multi Spatial-Attention Head only (5 fold ensemble of img size 640) + 1 fold of Img size 1024.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1243153,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-03-18T03:54:52.533000",
          "content": "<p>Thanks!</p>\n<blockquote>\n  <p>I achieved 0.972 on private lb just by using ResNet200d with Multi Spatial-Attention Head only</p>\n</blockquote>\n<p>That's great! </p>\n<p>I don't understand \"+ 1 fold of Img size 1024.\" You mean doing ensemble 6 predictions (<strong>5</strong> folds whose training and inference was done with img size 640 and <strong>1</strong> folds with img size 1024), right?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1243238,
          "author_name": "RAHUL SINGH INDA",
          "author_url": "",
          "post_date": "2021-03-18T05:26:35.043000",
          "content": "<p>yup , thats right.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1243244,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-03-18T05:31:26.137000",
          "content": "<p>I see. Thanks.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1242303,
      "author_name": "Mohammed Rizin V K",
      "author_url": "",
      "post_date": "2021-03-17T14:33:37.243000",
      "content": "<p>Congrats on your position! well deserved!! And thank you for your insight about Multi Head Spatial Attention.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1242348,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-03-17T15:01:23.277000",
          "content": "<p>Thanks!    </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1242202,
      "author_name": "ammarali32",
      "author_url": "",
      "post_date": "2021-03-17T13:25:40.693000",
      "content": "<p>Congratulations !! I would like to thank you for your amazing work and contribution.  </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1242224,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-03-17T13:38:17.850000",
          "content": "<p>Thanks!</p>\n<p>I think your contribution is also great. Thank you so much again.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1242188,
      "author_name": "gao-hongnan",
      "author_url": "",
      "post_date": "2021-03-17T13:13:33.303000",
      "content": "<p>May I know your training pipeline for efficientnet. I tried but cannot get a good results at all. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1242193,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-03-17T13:21:42.220000",
          "content": "<p>I trained EfficientNet-B5 by the same pipeline with other models. The different settings among models were batch size and learning rate because of model size.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1242268,
          "author_name": "Björn",
          "author_url": "",
          "post_date": "2021-03-17T14:03:42.260000",
          "content": "<p>What batch size did you use and did you need to use gradient accumulation/freezing BatchNorm layers? If not, what hardware did you use?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1242288,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-03-17T14:18:07.017000",
          "content": "<p>My GPU resource is a single TitanRTX(24GB) and I didn't use gradient accumulation/freezing BatchNorm layers.</p>\n<p>Batch size I used for each models is as follows:</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>image size</th>\n<th>batch size</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ResNet200d</td>\n<td>640x640</td>\n<td>16</td>\n</tr>\n<tr>\n<td>EfficientNet-B5</td>\n<td>640x640</td>\n<td>16</td>\n</tr>\n<tr>\n<td>SE-ResNet152d</td>\n<td>640x640</td>\n<td>16</td>\n</tr>\n<tr>\n<td>ResNeSt200e</td>\n<td>640x640</td>\n<td>15</td>\n</tr>\n</tbody>\n</table>\n<p>Sorry, I remember that I was trying to keep their batch size as same as possible. You can increase batch size of ResNet200d and SE-ResNet152d on TitanRTX.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1243162,
      "author_name": "Zekun",
      "author_url": "",
      "post_date": "2021-03-18T04:04:39.293000",
      "content": "<p>Hello,do they all trained on the whole data?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1243168,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-03-18T04:15:12.353000",
          "content": "<p>Yes. I used whole the data for fine-tuning but did cross validation.</p>\n<p>I  split the competition data by Multi-Label Stratified Group K-Fold manner(K=5).</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1242249,
      "author_name": "Nikita Kozodoi",
      "author_url": "",
      "post_date": "2021-03-17T13:48:47.060000",
      "content": "<p><a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> thank you for summarizing your solution! And also thanks for sharing very useful results/ideas during the competition 🙂</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1242252,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-03-17T13:51:24.833000",
          "content": "<p>Thanks. I'm glad that my discussions/baselines are helpful to you 🙂</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1246822,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-21T06:20:54.760000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1246832,
          "author_name": "Tawara",
          "author_url": "",
          "post_date": "2021-03-21T06:33:44.363000",
          "content": "<p>Thanks 😃  </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1242170": "First of all, congrats to all the teams got the medal and participants finished this competition, and thanks to Kaggle team and Competition host.\nI learned many things especially teacher-student learning through this competition.\n\nI'd like to say special thanks to @ammarali32 . I was about to give up in the middle of competition, but reading [his interesting approach](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/215910) made me resume.\n\nMy approach is not so special . Then I show brief summary here.\n\n### Summary\n\nFinal submission(Public: 0.970, Private: 0.972) is averaging of the following 6 models.\nI used image size of 640x640 for all models.\n\n* model1: ResNeSt200 (Public: 0.965, Private: 0.966)\n    * pretrained model: ImageNet\n    * Classification Head: Single MLP\n    * training: fine-tuning on competition data\n<br>\n* model2: ResNet200d (Public: 0.968, Private: 0.971)\n    * pretrained model: [@ammarali32 's starting points](https://www.kaggle.com/ammarali32/startingpointschestx)\n    * Classification Head:  Multi Spatial-Attention Head\n    * training: fine-tuning on competition data\n<br>\n* model3: EfficientNetB5-NoisyStudnet (Public: 0.966, Private: 0.969)\n    * pretrained model: [@ammarali32 's starting points](https://www.kaggle.com/ammarali32/startingpointschestx)\n    * Classification Head:  Multi Spatial-Attention Head\n    * training: fine-tuning on competition data\n<br>\n* model4: SE-ResNet152d (Public: 0.967, Private: 0.968)\n    * pretrained model: [@ammarali32 's starting points](https://www.kaggle.com/ammarali32/startingpointschestx)\n    * Classification Head:  Multi Spatial-Attention Head\n    * training: fine-tuning on competition data\n<br>\n* model5: ResNeSt200e (Public: 0.966, Private: 0.967)\n    * pretrained model: trained model on [NIH Chest X-rays](https://www.kaggle.com/nih-chest-xrays/data) by Teacher-Student Training \n    * Classification Head:  Multi Spatial-Attention Head\n    * training: fine-tuning on competition data\n<br>\n* model6: ResNet200d (Public: 0.966, Private: 0.970)\n    * pretrained model: [@ammarali32 's starting points](https://www.kaggle.com/ammarali32/startingpointschestx)\n    * Classification Head:  Multi-Head Attention with **Segmentation** Branch\n    * training:\n        * stage1: trained on only annotated data\n        * stage2: trained on only **non**-annotated data\n\nI really wanted to do teacher-student training on [NIH Chest X-rays](https://www.kaggle.com/nih-chest-xrays/data) for all the models. But I used [@ammarali32 's starting points](https://www.kaggle.com/ammarali32/startingpointschestx) due to lack of time and computing resources.\n\nAs a matter of fact, you can achieve Private 0.972(silver?) by averaging model 2,3,4 ([submission notebook](https://www.kaggle.com/ttahara/ranzcr-ensemble-fine-tuned-multi-head-models)), which is combination of [my multi-head approach](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207230) and [@ammarali32 's starting points](https://www.kaggle.com/ammarali32/startingpointschestx).  \nI wanted to exceed this score with a single model, but could not 😓",
    "1244276": "Thank you @ttahara for the thought-provoking presentation. Congratulations on achieving 23rd on the leadership board.\nWhat variance did you notice with size? You used  640x640. I was thinking if 1024x1024 would have given an edge.",
    "1243535": "Thank you for the write up and sharing the multi head approach.\nMay I know what is the different between EfficientNetB5-NoisyStudnet  and normal EfficientNetB5 ?\n\n> As a matter of fact, you can achieve Private 0.972(silver?) by averaging model 2,3,4\n\nThis is what I did, but my EfficientNet gets lower Public Score (0.964) compared to Seresnet (0.967), although on Private score it is higher (0.969 vs 0.968). ",
    "1243150": "Congratulations @ttahara  on your finish!  You were also a very valuable contributor to this competition. Multi Spatial-Attention Head was pretty special here and helped a lot of people.  So many great ideas here. Thanks. ",
    "1242739": "Thanks for your special Spatial Attention approach , a lot of people would not have won a medal without it , beautiful concept . I also tried self attention on top of it I wonder why it didn't work",
    "1242291": "Congrats on 23th place @ttahara ",
    "1242214": "Congrats!\n\nThank you for your great training notebook. I used it to train one of our ensemble models. \nI couldn't get 0.972 with single model too. I have similar result to your resnet 200d model 2 (CV: 0.966 Public: 0.967, Private: 0.971).\n\nI think it should be able to get 0.972 single model if you add semgment layers like other top competitors.\n",
    "1243116": "Congrats @ttahara on your position.\nBTW, I achieved 0.972 on private lb just by using ResNet200d with Multi Spatial-Attention Head only (5 fold ensemble of img size 640) + 1 fold of Img size 1024.\n ",
    "1242303": "Congrats on your position! well deserved!! And thank you for your insight about Multi Head Spatial Attention.",
    "1242202": "Congratulations !! I would like to thank you for your amazing work and contribution.  \n",
    "1242188": "May I know your training pipeline for efficientnet. I tried but cannot get a good results at all. ",
    "1243162": "Hello,do they all trained on the whole data?",
    "1242249": "@ttahara thank you for summarizing your solution! And also thanks for sharing very useful results/ideas during the competition 🙂",
    "1246822": ""
  }
}