{
  "id": 226616,
  "title": "6th place solution",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/writeups/datnt-6th-place-solution",
  "author_name": "",
  "post_date": "2021-03-18T05:44:15.337Z",
  "votes": 79,
  "comment_count": 36,
  "views": 0,
  "content": "<p>I will publish the kernel first then write the solution later today when I have time ;)</p>\n<p>key factors: image size, additional classes, auxiliary head training, self-distillation, and noisy student.</p>\n<p><a href=\"https://www.kaggle.com/moewie94/975-with-fixes\" target=\"_blank\">https://www.kaggle.com/moewie94/975-with-fixes</a></p>\n<p>For this competition, I did not use segmentation mask information because I got quite a good local CV and public LB scores from early without it, and training with segmentation masks cost a lot of time and computation resources so I decided to leave it. I'm still wondering if I can have any boost if I dedicate time and efforts to train models with segmentation masks 😂</p>\n<ol>\n<li>Image size: To be honest, I doubt if anyone can finish in the gold medal zone with anything less than 1024-by-1024. Firstly I do experiments with small image sizes (448-by-448) but after 1 month, I focused on 1024-by-1024 or bigger image sizes to get high scores. My final models use and collections of sizes (1024, 1280, 1344, and 1408) to improve diversification. I also try to fit as much model as possible into the ensemble and use big boys like efficient-net b6 and b7 😂</li>\n<li>Additional classes: I do not train with no-ETT class and cross-entropy loss but I create 3 more classes: CVC present, NGT present, and ETT present. This improves the local CV score a little bit.</li>\n<li>Auxilliary heads: To make training more efficient, I also add output layers to the intermediate blocks of the classification model and train them with the same ground truth as the final output layer (some guys call this the Supervision technique).</li>\n<li>Self-distillation: I detach the output of deeper blocks and use them as extra ground-truth to train the shallower blocks (weight 0.5)</li>\n<li>Noisy student: Every time I finished training k-fold, I generate OOF prediction soft-labels and use the soft-labels to train the model in the next cycle. This boost my local CV and leaderboard score quite a lot (0.004-0.005) and I think this is the reason I can get myself a high position finish :D</li>\n</ol>\n<p>P/S: Thank <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> for the quick submission tricks and <a href=\"https://www.kaggle.com/roydatascience\" target=\"_blank\">@roydatascience</a>, <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> for the multi head implementation. I benefitted a lot from you guys's ideas and tricks :)</p>",
  "messages": [
    {
      "id": "1241561",
      "postDate": "03/17/2021 05:20:21",
      "content": "<p>I will publish the kernel first then write the solution later today when I have time ;)</p>\n<p>key factors: image size, additional classes, auxiliary head training, self-distillation, and noisy student.</p>\n<p><a href=\"https://www.kaggle.com/moewie94/975-with-fixes\" target=\"_blank\">https://www.kaggle.com/moewie94/975-with-fixes</a></p>\n<p>For this competition, I did not use segmentation mask information because I got quite a good local CV and public LB scores from early without it, and training with segmentation masks cost a lot of time and computation resources so I decided to leave it. I'm still wondering if I can have any boost if I dedicate time and efforts to train models with segmentation masks 😂</p>\n<ol>\n<li>Image size: To be honest, I doubt if anyone can finish in the gold medal zone with anything less than 1024-by-1024. Firstly I do experiments with small image sizes (448-by-448) but after 1 month, I focused on 1024-by-1024 or bigger image sizes to get high scores. My final models use and collections of sizes (1024, 1280, 1344, and 1408) to improve diversification. I also try to fit as much model as possible into the ensemble and use big boys like efficient-net b6 and b7 😂</li>\n<li>Additional classes: I do not train with no-ETT class and cross-entropy loss but I create 3 more classes: CVC present, NGT present, and ETT present. This improves the local CV score a little bit.</li>\n<li>Auxilliary heads: To make training more efficient, I also add output layers to the intermediate blocks of the classification model and train them with the same ground truth as the final output layer (some guys call this the Supervision technique).</li>\n<li>Self-distillation: I detach the output of deeper blocks and use them as extra ground-truth to train the shallower blocks (weight 0.5)</li>\n<li>Noisy student: Every time I finished training k-fold, I generate OOF prediction soft-labels and use the soft-labels to train the model in the next cycle. This boost my local CV and leaderboard score quite a lot (0.004-0.005) and I think this is the reason I can get myself a high position finish :D</li>\n</ol>\n<p>P/S: Thank <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> for the quick submission tricks and <a href=\"https://www.kaggle.com/roydatascience\" target=\"_blank\">@roydatascience</a>, <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a> for the multi head implementation. I benefitted a lot from you guys's ideas and tricks :)</p>",
      "rawMarkdown": "I will publish the kernel first then write the solution later today when I have time ;)\n\nkey factors: image size, additional classes, auxiliary head training, self-distillation, and noisy student.\n\nhttps://www.kaggle.com/moewie94/975-with-fixes\n\n\nFor this competition, I did not use segmentation mask information because I got quite a good local CV and public LB scores from early without it, and training with segmentation masks cost a lot of time and computation resources so I decided to leave it. I'm still wondering if I can have any boost if I dedicate time and efforts to train models with segmentation masks 😂\n\n1. Image size: To be honest, I doubt if anyone can finish in the gold medal zone with anything less than 1024-by-1024. Firstly I do experiments with small image sizes (448-by-448) but after 1 month, I focused on 1024-by-1024 or bigger image sizes to get high scores. My final models use and collections of sizes (1024, 1280, 1344, and 1408) to improve diversification. I also try to fit as much model as possible into the ensemble and use big boys like efficient-net b6 and b7 😂\n2. Additional classes: I do not train with no-ETT class and cross-entropy loss but I create 3 more classes: CVC present, NGT present, and ETT present. This improves the local CV score a little bit.\n3. Auxilliary heads: To make training more efficient, I also add output layers to the intermediate blocks of the classification model and train them with the same ground truth as the final output layer (some guys call this the Supervision technique).\n4. Self-distillation: I detach the output of deeper blocks and use them as extra ground-truth to train the shallower blocks (weight 0.5)\n5. Noisy student: Every time I finished training k-fold, I generate OOF prediction soft-labels and use the soft-labels to train the model in the next cycle. This boost my local CV and leaderboard score quite a lot (0.004-0.005) and I think this is the reason I can get myself a high position finish :D\n\nP/S: Thank @underwearfitting for the quick submission tricks and @roydatascience, @ttahara for the multi head implementation. I benefitted a lot from you guys's ideas and tricks :)",
      "votes": null
    },
    {
      "id": "1241600",
      "postDate": "03/17/2021 05:48:49",
      "content": "<p>Thanks for sharing, are large model and img size help you boost CV and LB <a href=\"https://www.kaggle.com/moewie94\" target=\"_blank\">@moewie94</a>? </p>",
      "rawMarkdown": "Thanks for sharing, are large model and img size help you boost CV and LB @moewie94?",
      "votes": null
    },
    {
      "id": "1241783",
      "postDate": "03/17/2021 07:39:20",
      "content": "<p>thanks and congrats! </p>\n<p>i am interested in self-distillation. which paper did you use?<br>\nthis one?<br>\nBe Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self Distillation</p>",
      "rawMarkdown": "thanks and congrats! \n\ni am interested in self-distillation. which paper did you use?\nthis one?\nBe Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self Distillation",
      "votes": null
    },
    {
      "id": "1241847",
      "postDate": "03/17/2021 08:37:17",
      "content": "<p>yes. I used the idea from that paper. also, instead of training with KL-divergence, I just simple train with binary cross-entropy with sigmoid-ed predictions from the last layer.</p>",
      "rawMarkdown": "yes. I used the idea from that paper. also, instead of training with KL-divergence, I just simple train with binary cross-entropy with sigmoid-ed predictions from the last layer.",
      "votes": null
    },
    {
      "id": "1241867",
      "postDate": "03/17/2021 08:53:25",
      "content": "<p>I use 1024-by-1024 b7, 1408-by-1408 b7, and 1344-by-1344 b6. So… yes 😂</p>",
      "rawMarkdown": "I use 1024-by-1024 b7, 1408-by-1408 b7, and 1344-by-1344 b6. So... yes 😂",
      "votes": null
    },
    {
      "id": "1241931",
      "postDate": "03/17/2021 09:33:05",
      "content": "<p>Lấy thịt đè người quá bro ;)</p>",
      "rawMarkdown": "Lấy thịt đè người quá bro ;)",
      "votes": null
    },
    {
      "id": "1241936",
      "postDate": "03/17/2021 09:37:26",
      "content": "<p>The noisy student approach is amazing. May I ask how many cycles did you train to boost that much? Congrats on solo gold.</p>",
      "rawMarkdown": "The noisy student approach is amazing. May I ask how many cycles did you train to boost that much? Congrats on solo gold.",
      "votes": null
    },
    {
      "id": "1241946",
      "postDate": "03/17/2021 09:50:43",
      "content": "<p>First I train normally and do experiments normally. When I finish experimenting single-fold and start doing k fold for submission I call this cycle 0.<br>\nThen I keep all configs and only change csv file. I do it twice on my experiment arch (b4). So it can be cycle 1&amp;2.<br>\nThen when I do bigger arch or move to bigger image size, I continuously do this. After each time I have done a k-fold, I blend the new OOF with the old OOF to get new csv file. <br>\nSo it can be 7-8 cycles or 2-3 cycles depends on how you define a cycle :D</p>",
      "rawMarkdown": "First I train normally and do experiments normally. When I finish experimenting single-fold and start doing k fold for submission I call this cycle 0.\nThen I keep all configs and only change csv file. I do it twice on my experiment arch (b4). So it can be cycle 1&2.\nThen when I do bigger arch or move to bigger image size, I continuously do this. After each time I have done a k-fold, I blend the new OOF with the old OOF to get new csv file. \nSo it can be 7-8 cycles or 2-3 cycles depends on how you define a cycle :D",
      "votes": null
    },
    {
      "id": "1241947",
      "postDate": "03/17/2021 09:51:47",
      "content": "<p>May I ask how much improvement did the self distillation bring, if you have an estimate.</p>",
      "rawMarkdown": "May I ask how much improvement did the self distillation bring, if you have an estimate.",
      "votes": null
    },
    {
      "id": "1241951",
      "postDate": "03/17/2021 09:53:53",
      "content": "<p>Thanks for the explanation. This is something I have never tried. I will def try it.</p>",
      "rawMarkdown": "Thanks for the explanation. This is something I have never tried. I will def try it.",
      "votes": null
    },
    {
      "id": "1241957",
      "postDate": "03/17/2021 09:55:51",
      "content": "<p>I tried this from the early experiments, it improves my local validation score ~0.003 iirc (for b4 size 448)</p>",
      "rawMarkdown": "I tried this from the early experiments, it improves my local validation score ~0.003 iirc (for b4 size 448)",
      "votes": null
    },
    {
      "id": "1241959",
      "postDate": "03/17/2021 09:56:31",
      "content": "<p>may I ask when you say \"blend the new OOF with the old OOF\" you mean averaging the prediction from cycle -2 and cycle -1 to produce a new soft label? </p>",
      "rawMarkdown": "may I ask when you say \"blend the new OOF with the old OOF\" you mean averaging the prediction from cycle -2 and cycle -1 to produce a new soft label?",
      "votes": null
    },
    {
      "id": "1241966",
      "postDate": "03/17/2021 10:02:39",
      "content": "<p>Congratulations amazing work !! </p>",
      "rawMarkdown": "Congratulations amazing work !!",
      "votes": null
    },
    {
      "id": "1242027",
      "postDate": "03/17/2021 10:57:12",
      "content": "<p><a href=\"https://www.kaggle.com/moewie94\" target=\"_blank\">@moewie94</a> Congratulations on Solo Gold Finish . Great work !!!</p>",
      "rawMarkdown": "moewie94 Congratulations on Solo Gold Finish . Great work !!!",
      "votes": null
    },
    {
      "id": "1242039",
      "postDate": "03/17/2021 11:05:13",
      "content": "<p>Congratulations! Your approach, especially self-distillation and noisy student part are very insightful.</p>\n<p><code>To be honest, I doubt if anyone can finish in the gold medal zone with anything less than 1024-by-1024.</code><br>\nWe used 768 by 768 only :)</p>",
      "rawMarkdown": "Congratulations! Your approach, especially self-distillation and noisy student part are very insightful.\n\n`To be honest, I doubt if anyone can finish in the gold medal zone with anything less than 1024-by-1024. `\nWe used 768 by 768 only :)",
      "votes": null
    },
    {
      "id": "1242343",
      "postDate": "03/17/2021 14:58:31",
      "content": "<p>Wow. Lot of insights from your approach. Congrats on the win!</p>",
      "rawMarkdown": "Wow. Lot of insights from your approach. Congrats on the win!",
      "votes": null
    },
    {
      "id": "1242360",
      "postDate": "03/17/2021 15:08:35",
      "content": "<p>yes. for example: assume first I have oof_b4 as b4 prediction. then I train b5 with it and predict oof_b5. I will take (oof_b4 + oof_b5) / 2. as the new csv and train b6 and so on.</p>",
      "rawMarkdown": "yes. for example: assume first I have oof_b4 as b4 prediction. then I train b5 with it and predict oof_b5. I will take (oof_b4 + oof_b5) / 2. as the new csv and train b6 and so on.",
      "votes": null
    },
    {
      "id": "1242497",
      "postDate": "03/17/2021 16:41:40",
      "content": "<p>Congratulations! May I ask: how many gpu resources did you have?</p>",
      "rawMarkdown": "Congratulations! May I ask: how many gpu resources did you have?",
      "votes": null
    },
    {
      "id": "1242711",
      "postDate": "03/17/2021 18:50:04",
      "content": "<p>Congrats on Money Medal! But, it seems that you did not use any of InceptionV3 and Densenet121. Why did u suggest to <code>add inceptionV3 and densenet121 to your ensembles</code> <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/212856#1162676\" target=\"_blank\">here</a>?</p>",
      "rawMarkdown": "Congrats on Money Medal! But, it seems that you did not use any of InceptionV3 and Densenet121. Why did u suggest to `add inceptionV3 and densenet121 to your ensembles` [here](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/212856#1162676)?",
      "votes": null
    },
    {
      "id": "1243078",
      "postDate": "03/18/2021 02:15:21",
      "content": "<p>I tried but densenet can not beat big efficientnets :(</p>",
      "rawMarkdown": "I tried but densenet can not beat big efficientnets :(",
      "votes": null
    },
    {
      "id": "1243079",
      "postDate": "03/18/2021 02:16:42",
      "content": "<p>sometimes 1, sometimes 2, sometimes 4 V100 and I used some cloud computing too. I have to prioritize company works so no really dedicated GPU</p>",
      "rawMarkdown": "sometimes 1, sometimes 2, sometimes 4 V100 and I used some cloud computing too. I have to prioritize company works so no really dedicated GPU",
      "votes": null
    },
    {
      "id": "1243091",
      "postDate": "03/18/2021 02:37:17",
      "content": "<p>Congrats! please correct that the original multi-head implementation must be credited to <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a>. </p>",
      "rawMarkdown": "Congrats! please correct that the original multi-head implementation must be credited to @ttahara.",
      "votes": null
    },
    {
      "id": "1243198",
      "postDate": "03/18/2021 04:41:16",
      "content": "<p>Same 768 for us as well. Our least scored model in the ensemble scores 0.974 in private LB.</p>",
      "rawMarkdown": "Same 768 for us as well. Our least scored model in the ensemble scores 0.974 in private LB.",
      "votes": null
    },
    {
      "id": "1243254",
      "postDate": "03/18/2021 05:40:51",
      "content": "<p>Congrats!!!!!!!!!!!!!!! </p>\n<p>As I am trying to learn from the greats, would you mind explaining a little more in depth on the idea of point 5? I cannot quite catch what it meant~</p>",
      "rawMarkdown": "Congrats!!!!!!!!!!!!!!! \n\nAs I am trying to learn from the greats, would you mind explaining a little more in depth on the idea of point 5? I cannot quite catch what it meant~",
      "votes": null
    },
    {
      "id": "1243257",
      "postDate": "03/18/2021 05:43:18",
      "content": "<p>This is awesome work. Congratulations <a href=\"https://www.kaggle.com/moewie94\" target=\"_blank\">@moewie94</a>! I cannot wait for the details. Please pass it on!</p>",
      "rawMarkdown": "This is awesome work. Congratulations @moewie94! I cannot wait for the details. Please pass it on!",
      "votes": null
    },
    {
      "id": "1243264",
      "postDate": "03/18/2021 05:44:51",
      "content": "<p>you can search the noisy student paper from Quoc Le and his team at Google Brain :)</p>",
      "rawMarkdown": "you can search the noisy student paper from Quoc Le and his team at Google Brain :)",
      "votes": null
    },
    {
      "id": "1243265",
      "postDate": "03/18/2021 05:45:36",
      "content": "<p><a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> Congratulations )). Will you publish your solution details ???</p>",
      "rawMarkdown": "underwearfitting Congratulations )). Will you publish your solution details ???",
      "votes": null
    },
    {
      "id": "1243355",
      "postDate": "03/18/2021 06:44:32",
      "content": "<p>Working on that. We will also release the full code.</p>",
      "rawMarkdown": "Working on that. We will also release the full code.",
      "votes": null
    },
    {
      "id": "1243390",
      "postDate": "03/18/2021 07:19:37",
      "content": "<p><a href=\"https://www.kaggle.com/moewie94\" target=\"_blank\">@moewie94</a> Thanks!!!</p>",
      "rawMarkdown": "moewie94 Thanks!!!",
      "votes": null
    },
    {
      "id": "1243391",
      "postDate": "03/18/2021 07:19:54",
      "content": "<p><a href=\"https://www.kaggle.com/reighns\" target=\"_blank\">@reighns</a> <br>\n<a href=\"https://arxiv.org/pdf/1911.04252.pdf\" target=\"_blank\">https://arxiv.org/pdf/1911.04252.pdf</a><br>\nSelf-training with Noisy Student improves ImageNet classification</p>",
      "rawMarkdown": "reighns \nhttps://arxiv.org/pdf/1911.04252.pdf\nSelf-training with Noisy Student improves ImageNet classification",
      "votes": null
    },
    {
      "id": "1243394",
      "postDate": "03/18/2021 07:22:55",
      "content": "<p>Congratulations on gold!  Great idea noisy student.  </p>",
      "rawMarkdown": "Congratulations on gold!  Great idea noisy student.",
      "votes": null
    },
    {
      "id": "1243525",
      "postDate": "03/18/2021 09:36:19",
      "content": "<p>On the \"Self-distillation: I detach the output of deeper blocks and use them as extra ground-truth to train the shallower blocks (weight 0.5)\" bit, do you mean that you that you do something like the following:</p>\n<ul>\n<li>Fit model one epoch primarily with main binary competition labels.</li>\n<li>Take some hidden layer (perhaps the flattened one at the very end after some form of pooling), save it for every training image.</li>\n<li>In the next epoch, you now can use this as an additional target, but you branch the  model at an earlier layer (e.g. if the model consists of 10 blocks of some kind, you add an extra output with pooling + flattening after the 5th block - obviously you probably had it there from the start, you just now have some actual input for it).</li>\n</ul>\n<p>Did I get that approximately right? If so, I assume the idea is to try to get the earlier layers to already have a stronger representation?</p>",
      "rawMarkdown": "On the \"Self-distillation: I detach the output of deeper blocks and use them as extra ground-truth to train the shallower blocks (weight 0.5)\" bit, do you mean that you that you do something like the following:\n* Fit model one epoch primarily with main binary competition labels.\n* Take some hidden layer (perhaps the flattened one at the very end after some form of pooling), save it for every training image.\n* In the next epoch, you now can use this as an additional target, but you branch the  model at an earlier layer (e.g. if the model consists of 10 blocks of some kind, you add an extra output with pooling + flattening after the 5th block - obviously you probably had it there from the start, you just now have some actual input for it).\n\nDid I get that approximately right? If so, I assume the idea is to try to get the earlier layers to already have a stronger representation?",
      "votes": null
    },
    {
      "id": "1243540",
      "postDate": "03/18/2021 09:51:24",
      "content": "<p>No, not the epoch thing. First, I already have auxiliary heads, so my models can have up to 3 outputs from different blocks. And in the same iteration, I just call loss(output_weak, output_strong.detach())</p>",
      "rawMarkdown": "No, not the epoch thing. First, I already have auxiliary heads, so my models can have up to 3 outputs from different blocks. And in the same iteration, I just call loss(output_weak, output_strong.detach())",
      "votes": null
    },
    {
      "id": "1243607",
      "postDate": "03/18/2021 11:05:06",
      "content": "<p>Ah! Thanks. Very interesting idea.</p>",
      "rawMarkdown": "Ah! Thanks. Very interesting idea.",
      "votes": null
    },
    {
      "id": "1245221",
      "postDate": "03/19/2021 15:15:16",
      "content": "<p>Thanks for sharing, really interesting stuff!<br>\nI have a question regarding the noisy student approach: For example, if we have 2 folds, model_1_1 is trained on fold_1 (ground truth) and validated on fold_2 (model_2_1 vice versa), then model_1_1 is used to predict the soft labels for model 2_2, and model2_1 predicts the soft labels for model1_2. Is this correct? If so, wouldn't this introduce some kind of leakage? The predicted fold1 (soft) labels of model2_1 include somekind of information of fold2, because model2_1 is trained on fold2. Do I miss something, or is this negligible? </p>",
      "rawMarkdown": "Thanks for sharing, really interesting stuff!\nI have a question regarding the noisy student approach: For example, if we have 2 folds, model_1_1 is trained on fold_1 (ground truth) and validated on fold_2 (model_2_1 vice versa), then model_1_1 is used to predict the soft labels for model 2_2, and model2_1 predicts the soft labels for model1_2. Is this correct? If so, wouldn't this introduce some kind of leakage? The predicted fold1 (soft) labels of model2_1 include somekind of information of fold2, because model2_1 is trained on fold2. Do I miss something, or is this negligible?",
      "votes": null
    },
    {
      "id": "1246530",
      "postDate": "03/20/2021 20:11:09",
      "content": "<p>Congratulations!  Thanks for sharing your approach and your kernel.  Did you try different ensembling approaches, and any comments on what did/didn't help?  Thanks!</p>",
      "rawMarkdown": "Congratulations!  Thanks for sharing your approach and your kernel.  Did you try different ensembling approaches, and any comments on what did/didn't help?  Thanks!",
      "votes": null
    },
    {
      "id": "1478218",
      "postDate": "08/17/2021 21:26:06",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/moewie94\" target=\"_blank\">@moewie94</a></p>\n<p>I am also interested in the noisy student approach.</p>\n<p>In the paper, they train a teacher on labelled data, predict on unlabelled data, then train a bigger student on labels + pseudo labels. That's one iteration.</p>\n<p>In your case, you're saying you predict OOF, then use only that to train the student? You do not use labelled data at all after you train your first teacher?</p>",
      "rawMarkdown": "Hi @moewie94\n\nI am also interested in the noisy student approach.\n\nIn the paper, they train a teacher on labelled data, predict on unlabelled data, then train a bigger student on labels + pseudo labels. That's one iteration.\n\nIn your case, you're saying you predict OOF, then use only that to train the student? You do not use labelled data at all after you train your first teacher?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1241600,
      "author_name": "duykhanh99",
      "author_url": "",
      "post_date": "03/17/2021 05:48:49",
      "content": "<p>Thanks for sharing, are large model and img size help you boost CV and LB <a href=\"https://www.kaggle.com/moewie94\" target=\"_blank\">@moewie94</a>? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1241867,
          "author_name": "moewie94",
          "author_url": "",
          "post_date": "03/17/2021 08:53:25",
          "content": "<p>I use 1024-by-1024 b7, 1408-by-1408 b7, and 1344-by-1344 b6. So… yes 😂</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1241931,
          "author_name": "duykhanh99",
          "author_url": "",
          "post_date": "03/17/2021 09:33:05",
          "content": "<p>Lấy thịt đè người quá bro ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1241783,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/17/2021 07:39:20",
      "content": "<p>thanks and congrats! </p>\n<p>i am interested in self-distillation. which paper did you use?<br>\nthis one?<br>\nBe Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self Distillation</p>",
      "votes": null,
      "replies": [
        {
          "id": 1241847,
          "author_name": "moewie94",
          "author_url": "",
          "post_date": "03/17/2021 08:37:17",
          "content": "<p>yes. I used the idea from that paper. also, instead of training with KL-divergence, I just simple train with binary cross-entropy with sigmoid-ed predictions from the last layer.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1241936,
      "author_name": "underwearfitting",
      "author_url": "",
      "post_date": "03/17/2021 09:37:26",
      "content": "<p>The noisy student approach is amazing. May I ask how many cycles did you train to boost that much? Congrats on solo gold.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1241946,
          "author_name": "moewie94",
          "author_url": "",
          "post_date": "03/17/2021 09:50:43",
          "content": "<p>First I train normally and do experiments normally. When I finish experimenting single-fold and start doing k fold for submission I call this cycle 0.<br>\nThen I keep all configs and only change csv file. I do it twice on my experiment arch (b4). So it can be cycle 1&amp;2.<br>\nThen when I do bigger arch or move to bigger image size, I continuously do this. After each time I have done a k-fold, I blend the new OOF with the old OOF to get new csv file. <br>\nSo it can be 7-8 cycles or 2-3 cycles depends on how you define a cycle :D</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1241951,
          "author_name": "underwearfitting",
          "author_url": "",
          "post_date": "03/17/2021 09:53:53",
          "content": "<p>Thanks for the explanation. This is something I have never tried. I will def try it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1241959,
          "author_name": "yl1202",
          "author_url": "",
          "post_date": "03/17/2021 09:56:31",
          "content": "<p>may I ask when you say \"blend the new OOF with the old OOF\" you mean averaging the prediction from cycle -2 and cycle -1 to produce a new soft label? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1242360,
          "author_name": "moewie94",
          "author_url": "",
          "post_date": "03/17/2021 15:08:35",
          "content": "<p>yes. for example: assume first I have oof_b4 as b4 prediction. then I train b5 with it and predict oof_b5. I will take (oof_b4 + oof_b5) / 2. as the new csv and train b6 and so on.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1478218,
          "author_name": "yousof9",
          "author_url": "",
          "post_date": "08/17/2021 21:26:06",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/moewie94\" target=\"_blank\">@moewie94</a></p>\n<p>I am also interested in the noisy student approach.</p>\n<p>In the paper, they train a teacher on labelled data, predict on unlabelled data, then train a bigger student on labels + pseudo labels. That's one iteration.</p>\n<p>In your case, you're saying you predict OOF, then use only that to train the student? You do not use labelled data at all after you train your first teacher?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1241947,
      "author_name": "yl1202",
      "author_url": "",
      "post_date": "03/17/2021 09:51:47",
      "content": "<p>May I ask how much improvement did the self distillation bring, if you have an estimate.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1241957,
          "author_name": "moewie94",
          "author_url": "",
          "post_date": "03/17/2021 09:55:51",
          "content": "<p>I tried this from the early experiments, it improves my local validation score ~0.003 iirc (for b4 size 448)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1241966,
      "author_name": "ammarali32",
      "author_url": "",
      "post_date": "03/17/2021 10:02:39",
      "content": "<p>Congratulations amazing work !! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1242027,
      "author_name": "usharengaraju",
      "author_url": "",
      "post_date": "03/17/2021 10:57:12",
      "content": "<p><a href=\"https://www.kaggle.com/moewie94\" target=\"_blank\">@moewie94</a> Congratulations on Solo Gold Finish . Great work !!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1242039,
      "author_name": "analokamus",
      "author_url": "",
      "post_date": "03/17/2021 11:05:13",
      "content": "<p>Congratulations! Your approach, especially self-distillation and noisy student part are very insightful.</p>\n<p><code>To be honest, I doubt if anyone can finish in the gold medal zone with anything less than 1024-by-1024.</code><br>\nWe used 768 by 768 only :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1243198,
          "author_name": "underwearfitting",
          "author_url": "",
          "post_date": "03/18/2021 04:41:16",
          "content": "<p>Same 768 for us as well. Our least scored model in the ensemble scores 0.974 in private LB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1243265,
          "author_name": "ammarali32",
          "author_url": "",
          "post_date": "03/18/2021 05:45:36",
          "content": "<p><a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> Congratulations )). Will you publish your solution details ???</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1243355,
          "author_name": "underwearfitting",
          "author_url": "",
          "post_date": "03/18/2021 06:44:32",
          "content": "<p>Working on that. We will also release the full code.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1242343,
      "author_name": "drcodikpollonny",
      "author_url": "",
      "post_date": "03/17/2021 14:58:31",
      "content": "<p>Wow. Lot of insights from your approach. Congrats on the win!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1242497,
      "author_name": "trytolose",
      "author_url": "",
      "post_date": "03/17/2021 16:41:40",
      "content": "<p>Congratulations! May I ask: how many gpu resources did you have?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1243079,
          "author_name": "moewie94",
          "author_url": "",
          "post_date": "03/18/2021 02:16:42",
          "content": "<p>sometimes 1, sometimes 2, sometimes 4 V100 and I used some cloud computing too. I have to prioritize company works so no really dedicated GPU</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1242711,
      "author_name": "woshifym",
      "author_url": "",
      "post_date": "03/17/2021 18:50:04",
      "content": "<p>Congrats on Money Medal! But, it seems that you did not use any of InceptionV3 and Densenet121. Why did u suggest to <code>add inceptionV3 and densenet121 to your ensembles</code> <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/212856#1162676\" target=\"_blank\">here</a>?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1243078,
          "author_name": "moewie94",
          "author_url": "",
          "post_date": "03/18/2021 02:15:21",
          "content": "<p>I tried but densenet can not beat big efficientnets :(</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1243091,
      "author_name": "projdev",
      "author_url": "",
      "post_date": "03/18/2021 02:37:17",
      "content": "<p>Congrats! please correct that the original multi-head implementation must be credited to <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a>. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1243254,
      "author_name": "reighns",
      "author_url": "",
      "post_date": "03/18/2021 05:40:51",
      "content": "<p>Congrats!!!!!!!!!!!!!!! </p>\n<p>As I am trying to learn from the greats, would you mind explaining a little more in depth on the idea of point 5? I cannot quite catch what it meant~</p>",
      "votes": null,
      "replies": [
        {
          "id": 1243264,
          "author_name": "moewie94",
          "author_url": "",
          "post_date": "03/18/2021 05:44:51",
          "content": "<p>you can search the noisy student paper from Quoc Le and his team at Google Brain :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1243390,
          "author_name": "reighns",
          "author_url": "",
          "post_date": "03/18/2021 07:19:37",
          "content": "<p><a href=\"https://www.kaggle.com/moewie94\" target=\"_blank\">@moewie94</a> Thanks!!!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1243391,
          "author_name": "something4kag",
          "author_url": "",
          "post_date": "03/18/2021 07:19:54",
          "content": "<p><a href=\"https://www.kaggle.com/reighns\" target=\"_blank\">@reighns</a> <br>\n<a href=\"https://arxiv.org/pdf/1911.04252.pdf\" target=\"_blank\">https://arxiv.org/pdf/1911.04252.pdf</a><br>\nSelf-training with Noisy Student improves ImageNet classification</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1243257,
      "author_name": "olusesiadebisi",
      "author_url": "",
      "post_date": "03/18/2021 05:43:18",
      "content": "<p>This is awesome work. Congratulations <a href=\"https://www.kaggle.com/moewie94\" target=\"_blank\">@moewie94</a>! I cannot wait for the details. Please pass it on!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1243394,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "03/18/2021 07:22:55",
      "content": "<p>Congratulations on gold!  Great idea noisy student.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1243525,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "03/18/2021 09:36:19",
      "content": "<p>On the \"Self-distillation: I detach the output of deeper blocks and use them as extra ground-truth to train the shallower blocks (weight 0.5)\" bit, do you mean that you that you do something like the following:</p>\n<ul>\n<li>Fit model one epoch primarily with main binary competition labels.</li>\n<li>Take some hidden layer (perhaps the flattened one at the very end after some form of pooling), save it for every training image.</li>\n<li>In the next epoch, you now can use this as an additional target, but you branch the  model at an earlier layer (e.g. if the model consists of 10 blocks of some kind, you add an extra output with pooling + flattening after the 5th block - obviously you probably had it there from the start, you just now have some actual input for it).</li>\n</ul>\n<p>Did I get that approximately right? If so, I assume the idea is to try to get the earlier layers to already have a stronger representation?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1243540,
          "author_name": "moewie94",
          "author_url": "",
          "post_date": "03/18/2021 09:51:24",
          "content": "<p>No, not the epoch thing. First, I already have auxiliary heads, so my models can have up to 3 outputs from different blocks. And in the same iteration, I just call loss(output_weak, output_strong.detach())</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1243607,
          "author_name": "bjoernholzhauer",
          "author_url": "",
          "post_date": "03/18/2021 11:05:06",
          "content": "<p>Ah! Thanks. Very interesting idea.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1245221,
      "author_name": "maschwa",
      "author_url": "",
      "post_date": "03/19/2021 15:15:16",
      "content": "<p>Thanks for sharing, really interesting stuff!<br>\nI have a question regarding the noisy student approach: For example, if we have 2 folds, model_1_1 is trained on fold_1 (ground truth) and validated on fold_2 (model_2_1 vice versa), then model_1_1 is used to predict the soft labels for model 2_2, and model2_1 predicts the soft labels for model1_2. Is this correct? If so, wouldn't this introduce some kind of leakage? The predicted fold1 (soft) labels of model2_1 include somekind of information of fold2, because model2_1 is trained on fold2. Do I miss something, or is this negligible? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1246530,
      "author_name": "socated",
      "author_url": "",
      "post_date": "03/20/2021 20:11:09",
      "content": "<p>Congratulations!  Thanks for sharing your approach and your kernel.  Did you try different ensembling approaches, and any comments on what did/didn't help?  Thanks!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1241561": "I will publish the kernel first then write the solution later today when I have time ;)\n\nkey factors: image size, additional classes, auxiliary head training, self-distillation, and noisy student.\n\nhttps://www.kaggle.com/moewie94/975-with-fixes\n\n\nFor this competition, I did not use segmentation mask information because I got quite a good local CV and public LB scores from early without it, and training with segmentation masks cost a lot of time and computation resources so I decided to leave it. I'm still wondering if I can have any boost if I dedicate time and efforts to train models with segmentation masks 😂\n\n1. Image size: To be honest, I doubt if anyone can finish in the gold medal zone with anything less than 1024-by-1024. Firstly I do experiments with small image sizes (448-by-448) but after 1 month, I focused on 1024-by-1024 or bigger image sizes to get high scores. My final models use and collections of sizes (1024, 1280, 1344, and 1408) to improve diversification. I also try to fit as much model as possible into the ensemble and use big boys like efficient-net b6 and b7 😂\n2. Additional classes: I do not train with no-ETT class and cross-entropy loss but I create 3 more classes: CVC present, NGT present, and ETT present. This improves the local CV score a little bit.\n3. Auxilliary heads: To make training more efficient, I also add output layers to the intermediate blocks of the classification model and train them with the same ground truth as the final output layer (some guys call this the Supervision technique).\n4. Self-distillation: I detach the output of deeper blocks and use them as extra ground-truth to train the shallower blocks (weight 0.5)\n5. Noisy student: Every time I finished training k-fold, I generate OOF prediction soft-labels and use the soft-labels to train the model in the next cycle. This boost my local CV and leaderboard score quite a lot (0.004-0.005) and I think this is the reason I can get myself a high position finish :D\n\nP/S: Thank @underwearfitting for the quick submission tricks and @roydatascience, @ttahara for the multi head implementation. I benefitted a lot from you guys's ideas and tricks :)",
    "1241600": "Thanks for sharing, are large model and img size help you boost CV and LB @moewie94?",
    "1241783": "thanks and congrats! \n\ni am interested in self-distillation. which paper did you use?\nthis one?\nBe Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self Distillation",
    "1241847": "yes. I used the idea from that paper. also, instead of training with KL-divergence, I just simple train with binary cross-entropy with sigmoid-ed predictions from the last layer.",
    "1241867": "I use 1024-by-1024 b7, 1408-by-1408 b7, and 1344-by-1344 b6. So... yes 😂",
    "1241931": "Lấy thịt đè người quá bro ;)",
    "1241936": "The noisy student approach is amazing. May I ask how many cycles did you train to boost that much? Congrats on solo gold.",
    "1241946": "First I train normally and do experiments normally. When I finish experimenting single-fold and start doing k fold for submission I call this cycle 0.\nThen I keep all configs and only change csv file. I do it twice on my experiment arch (b4). So it can be cycle 1&2.\nThen when I do bigger arch or move to bigger image size, I continuously do this. After each time I have done a k-fold, I blend the new OOF with the old OOF to get new csv file. \nSo it can be 7-8 cycles or 2-3 cycles depends on how you define a cycle :D",
    "1241947": "May I ask how much improvement did the self distillation bring, if you have an estimate.",
    "1241951": "Thanks for the explanation. This is something I have never tried. I will def try it.",
    "1241957": "I tried this from the early experiments, it improves my local validation score ~0.003 iirc (for b4 size 448)",
    "1241959": "may I ask when you say \"blend the new OOF with the old OOF\" you mean averaging the prediction from cycle -2 and cycle -1 to produce a new soft label?",
    "1241966": "Congratulations amazing work !!",
    "1242027": "moewie94 Congratulations on Solo Gold Finish . Great work !!!",
    "1242039": "Congratulations! Your approach, especially self-distillation and noisy student part are very insightful.\n\n`To be honest, I doubt if anyone can finish in the gold medal zone with anything less than 1024-by-1024. `\nWe used 768 by 768 only :)",
    "1242343": "Wow. Lot of insights from your approach. Congrats on the win!",
    "1242360": "yes. for example: assume first I have oof_b4 as b4 prediction. then I train b5 with it and predict oof_b5. I will take (oof_b4 + oof_b5) / 2. as the new csv and train b6 and so on.",
    "1242497": "Congratulations! May I ask: how many gpu resources did you have?",
    "1242711": "Congrats on Money Medal! But, it seems that you did not use any of InceptionV3 and Densenet121. Why did u suggest to `add inceptionV3 and densenet121 to your ensembles` [here](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/212856#1162676)?",
    "1243078": "I tried but densenet can not beat big efficientnets :(",
    "1243079": "sometimes 1, sometimes 2, sometimes 4 V100 and I used some cloud computing too. I have to prioritize company works so no really dedicated GPU",
    "1243091": "Congrats! please correct that the original multi-head implementation must be credited to @ttahara.",
    "1243198": "Same 768 for us as well. Our least scored model in the ensemble scores 0.974 in private LB.",
    "1243254": "Congrats!!!!!!!!!!!!!!! \n\nAs I am trying to learn from the greats, would you mind explaining a little more in depth on the idea of point 5? I cannot quite catch what it meant~",
    "1243257": "This is awesome work. Congratulations @moewie94! I cannot wait for the details. Please pass it on!",
    "1243264": "you can search the noisy student paper from Quoc Le and his team at Google Brain :)",
    "1243265": "underwearfitting Congratulations )). Will you publish your solution details ???",
    "1243355": "Working on that. We will also release the full code.",
    "1243390": "moewie94 Thanks!!!",
    "1243391": "reighns \nhttps://arxiv.org/pdf/1911.04252.pdf\nSelf-training with Noisy Student improves ImageNet classification",
    "1243394": "Congratulations on gold!  Great idea noisy student.",
    "1243525": "On the \"Self-distillation: I detach the output of deeper blocks and use them as extra ground-truth to train the shallower blocks (weight 0.5)\" bit, do you mean that you that you do something like the following:\n* Fit model one epoch primarily with main binary competition labels.\n* Take some hidden layer (perhaps the flattened one at the very end after some form of pooling), save it for every training image.\n* In the next epoch, you now can use this as an additional target, but you branch the  model at an earlier layer (e.g. if the model consists of 10 blocks of some kind, you add an extra output with pooling + flattening after the 5th block - obviously you probably had it there from the start, you just now have some actual input for it).\n\nDid I get that approximately right? If so, I assume the idea is to try to get the earlier layers to already have a stronger representation?",
    "1243540": "No, not the epoch thing. First, I already have auxiliary heads, so my models can have up to 3 outputs from different blocks. And in the same iteration, I just call loss(output_weak, output_strong.detach())",
    "1243607": "Ah! Thanks. Very interesting idea.",
    "1245221": "Thanks for sharing, really interesting stuff!\nI have a question regarding the noisy student approach: For example, if we have 2 folds, model_1_1 is trained on fold_1 (ground truth) and validated on fold_2 (model_2_1 vice versa), then model_1_1 is used to predict the soft labels for model 2_2, and model2_1 predicts the soft labels for model1_2. Is this correct? If so, wouldn't this introduce some kind of leakage? The predicted fold1 (soft) labels of model2_1 include somekind of information of fold2, because model2_1 is trained on fold2. Do I miss something, or is this negligible?",
    "1246530": "Congratulations!  Thanks for sharing your approach and your kernel.  Did you try different ensembling approaches, and any comments on what did/didn't help?  Thanks!",
    "1478218": "Hi @moewie94\n\nI am also interested in the noisy student approach.\n\nIn the paper, they train a teacher on labelled data, predict on unlabelled data, then train a bigger student on labels + pseudo labels. That's one iteration.\n\nIn your case, you're saying you predict OOF, then use only that to train the student? You do not use labelled data at all after you train your first teacher?"
  },
  "source": "meta"
}