{
  "id": 207577,
  "title": "3-stage training with additional annotation [CV:0.95x]",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/207577",
  "author_name": "Y.Nakama",
  "post_date": "2020-12-30T11:29:47.003000",
  "votes": 149,
  "comment_count": 26,
  "views": 0,
  "content": "<p>The simplest way to use additional annotation is already discussed in <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/205243\" target=\"_blank\">this thread</a> by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>.<br>\nI implemented this kind of approach on kaggle notebooks (you can select GPU/TPU) and local validation score can get 0.95x. In my experience, there is 0.01 improvement with this approach.</p>\n<h2>Training strategy</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/yasufuminakama/ranzcr-resnet200d-3-stage-training-step1\" target=\"_blank\">1st-stage training</a><ul>\n<li>teacher model training for annotated image<ul>\n<li>data: annotated data</li>\n<li>pretrained weight: imagenet weight</li>\n<li><code>BCEWithLogitsLoss(y_preds, labels)</code></li>\n<li><code>y_preds: teacher model predictions for annotated image</code></li></ul></li></ul></li>\n<li><a href=\"https://www.kaggle.com/yasufuminakama/ranzcr-resnet200d-3-stage-training-step2\" target=\"_blank\">2nd-stage training</a><ul>\n<li>student model training with teacher model features<ul>\n<li>data: annotated data</li>\n<li>student model pretrained weight: imagenet weight</li>\n<li>teacher model pretrained weight: 1st-stage weight</li>\n<li><code>BCEWithLogitsLoss(y_preds, labels) + w * MSELoss(student_features, teacher_features)</code></li>\n<li><code>y_preds: student model predictions for normal image</code></li>\n<li><code>student_features: student model features for normal image</code></li>\n<li><code>teacher_features: teacher model features for annotated image</code></li></ul></li></ul></li>\n<li><a href=\"https://www.kaggle.com/yasufuminakama/ranzcr-resnet200d-3-stage-training-step3\" target=\"_blank\">3rd-stage training</a><ul>\n<li>model training<ul>\n<li>data: all data</li>\n<li>pretrained weight: 2nd-stage weight</li>\n<li><code>BCEWithLogitsLoss(y_preds, labels)</code></li>\n<li><code>y_preds: student model predictions for normal image</code></li></ul></li></ul></li>\n<li><a href=\"https://www.kaggle.com/yasufuminakama/ranzcr-resnet200d-3-stage-training-sub\" target=\"_blank\">inference notebook</a></li>\n</ul>\n<p>Hope this helps, happy kaggling :)</p>",
  "messages": [
    {
      "id": 1132394,
      "postDate": "2020-12-30T11:29:47.003Z",
      "content": "<p>The simplest way to use additional annotation is already discussed in <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/205243\" target=\"_blank\">this thread</a> by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>.<br>\nI implemented this kind of approach on kaggle notebooks (you can select GPU/TPU) and local validation score can get 0.95x. In my experience, there is 0.01 improvement with this approach.</p>\n<h2>Training strategy</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/yasufuminakama/ranzcr-resnet200d-3-stage-training-step1\" target=\"_blank\">1st-stage training</a><ul>\n<li>teacher model training for annotated image<ul>\n<li>data: annotated data</li>\n<li>pretrained weight: imagenet weight</li>\n<li><code>BCEWithLogitsLoss(y_preds, labels)</code></li>\n<li><code>y_preds: teacher model predictions for annotated image</code></li></ul></li></ul></li>\n<li><a href=\"https://www.kaggle.com/yasufuminakama/ranzcr-resnet200d-3-stage-training-step2\" target=\"_blank\">2nd-stage training</a><ul>\n<li>student model training with teacher model features<ul>\n<li>data: annotated data</li>\n<li>student model pretrained weight: imagenet weight</li>\n<li>teacher model pretrained weight: 1st-stage weight</li>\n<li><code>BCEWithLogitsLoss(y_preds, labels) + w * MSELoss(student_features, teacher_features)</code></li>\n<li><code>y_preds: student model predictions for normal image</code></li>\n<li><code>student_features: student model features for normal image</code></li>\n<li><code>teacher_features: teacher model features for annotated image</code></li></ul></li></ul></li>\n<li><a href=\"https://www.kaggle.com/yasufuminakama/ranzcr-resnet200d-3-stage-training-step3\" target=\"_blank\">3rd-stage training</a><ul>\n<li>model training<ul>\n<li>data: all data</li>\n<li>pretrained weight: 2nd-stage weight</li>\n<li><code>BCEWithLogitsLoss(y_preds, labels)</code></li>\n<li><code>y_preds: student model predictions for normal image</code></li></ul></li></ul></li>\n<li><a href=\"https://www.kaggle.com/yasufuminakama/ranzcr-resnet200d-3-stage-training-sub\" target=\"_blank\">inference notebook</a></li>\n</ul>\n<p>Hope this helps, happy kaggling :)</p>",
      "rawMarkdown": "The simplest way to use additional annotation is already discussed in [this thread](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/205243) by @hengck23.\nI implemented this kind of approach on kaggle notebooks (you can select GPU/TPU) and local validation score can get 0.95x. In my experience, there is 0.01 improvement with this approach.\n\n## Training strategy\n- [1st-stage training](https://www.kaggle.com/yasufuminakama/ranzcr-resnet200d-3-stage-training-step1)\n    - teacher model training for annotated image\n        - data: annotated data\n        - pretrained weight: imagenet weight\n        - `BCEWithLogitsLoss(y_preds, labels)`\n        - `y_preds: teacher model predictions for annotated image`\n- [2nd-stage training](https://www.kaggle.com/yasufuminakama/ranzcr-resnet200d-3-stage-training-step2)\n    - student model training with teacher model features\n        - data: annotated data\n        - student model pretrained weight: imagenet weight\n        - teacher model pretrained weight: 1st-stage weight\n        - `BCEWithLogitsLoss(y_preds, labels) + w * MSELoss(student_features, teacher_features)`\n        - `y_preds: student model predictions for normal image`\n        - `student_features: student model features for normal image`\n        - `teacher_features: teacher model features for annotated image`\n- [3rd-stage training](https://www.kaggle.com/yasufuminakama/ranzcr-resnet200d-3-stage-training-step3)\n    - model training\n        - data: all data\n        - pretrained weight: 2nd-stage weight\n        - `BCEWithLogitsLoss(y_preds, labels)`\n        - `y_preds: student model predictions for normal image`\n- [inference notebook](https://www.kaggle.com/yasufuminakama/ranzcr-resnet200d-3-stage-training-sub)\n\nHope this helps, happy kaggling :)",
      "votes": 149
    },
    {
      "id": 1132504,
      "postDate": "2020-12-30T12:58:31.920Z",
      "content": "<p>When I first joined Kaggle few years ago, I was amazed to see a brilliant 15 or (16) y.o guy <a href=\"https://www.kaggle.com/anokas\" target=\"_blank\">@anokas</a>  who used to wrote quickly very robust starter kernels at almost every competition. </p>\n<p>It seems now <a href=\"https://www.kaggle.com/xhlulu\" target=\"_blank\">@xhlulu</a> and you have brilliantly taken it over for TF and Pytorch respectively. </p>\n<p>Even though I have my own workflow and Pipeline. Your kernels are always a great inspiration to adapt mines on new competitions. </p>\n<p>Many thanks and keep up the great work ! </p>\n<p>And of course huge thanks to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for giving too much inspirations and ideas to all of us. </p>",
      "rawMarkdown": "When I first joined Kaggle few years ago, I was amazed to see a brilliant 15 or (16) y.o guy @anokas  who used to wrote quickly very robust starter kernels at almost every competition. \n\nIt seems now @xhlulu and you have brilliantly taken it over for TF and Pytorch respectively. \n\nEven though I have my own workflow and Pipeline. Your kernels are always a great inspiration to adapt mines on new competitions. \n\nMany thanks and keep up the great work ! \n\nAnd of course huge thanks to @hengck23 for giving too much inspirations and ideas to all of us. ",
      "votes": 7,
      "replies": [
        {
          "id": 1132511,
          "postDate": "2020-12-30T13:05:14Z",
          "content": "<p>Thank you! :)</p>",
          "rawMarkdown": "Thank you! :)"
        }
      ]
    },
    {
      "id": 1132554,
      "postDate": "2020-12-30T13:44:41.333Z",
      "content": "<p><a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> </p>\n<p>thanks for the results and experiments.<br>\n I only verify that it would work based on my experiments. but I haven't implemented and submitted it yet.<br>\nbased on my experiments, the potential increase can be about 1%.</p>\n<p>you can check the CAM map difference before and after the approach.</p>",
      "rawMarkdown": "@yasufuminakama \n\nthanks for the results and experiments.\n I only verify that it would work based on my experiments. but I haven't implemented and submitted it yet.\nbased on my experiments, the potential increase can be about 1%.\n\nyou can check the CAM map difference before and after the approach.\n\n",
      "votes": 5
    },
    {
      "id": 1135016,
      "postDate": "2021-01-01T19:45:37.923Z",
      "content": "<p>It's always amazing how ideas turn into reality. :)</p>\n<p>Thanks <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a>.</p>",
      "rawMarkdown": "It's always amazing how ideas turn into reality. :)\n\nThanks @hengck23 @yasufuminakama.",
      "votes": 1
    },
    {
      "id": 1133846,
      "postDate": "2020-12-31T15:26:01.370Z",
      "content": "<p>Is this <strong>Knowledge Distillation</strong>? or another technique?</p>",
      "rawMarkdown": "Is this **Knowledge Distillation**? or another technique?",
      "votes": 2
    },
    {
      "id": 1132438,
      "postDate": "2020-12-30T12:08:28.443Z",
      "content": "<p>Unfortunately can upvote only once all these kernels :( hats off <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> !</p>",
      "rawMarkdown": "Unfortunately can upvote only once all these kernels :( hats off @yasufuminakama !",
      "votes": 2,
      "replies": [
        {
          "id": 1132441,
          "postDate": "2020-12-30T12:13:12.513Z",
          "content": "<p>Thanks for your upvote <a href=\"https://www.kaggle.com/atanasova\" target=\"_blank\">@atanasova</a> :)</p>",
          "rawMarkdown": "Thanks for your upvote @atanasova :)"
        }
      ]
    },
    {
      "id": 1230264,
      "postDate": "2021-03-08T00:22:59.380Z",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a>, thanks for this. What is 'w' in the 2nd stage loss. A constant?</p>",
      "rawMarkdown": "Hey @yasufuminakama, thanks for this. What is 'w' in the 2nd stage loss. A constant?",
      "replies": [
        {
          "id": 1236032,
          "postDate": "2021-03-12T17:31:20.207Z",
          "content": "<p>Yes, this gives weightage to  MSELoss. <br>\nFor me 1 for 'BCEWithLogitsLoss' and 0.5 for 'MSELoss' worked best.</p>",
          "rawMarkdown": "Yes, this gives weightage to  MSELoss. \nFor me 1 for 'BCEWithLogitsLoss' and 0.5 for 'MSELoss' worked best."
        }
      ]
    },
    {
      "id": 1225263,
      "postDate": "2021-03-03T13:38:29.990Z",
      "content": "<p>Great work, thanx</p>",
      "rawMarkdown": "Great work, thanx"
    },
    {
      "id": 1224982,
      "postDate": "2021-03-03T08:09:17.320Z",
      "content": "<p>Thanks, Y.Nakama for your insightful and helpful posts.</p>\n<p>May I ask a newbie's question?</p>\n<p>How do you do training in stages 1 and 2? I mean, In understand that for stage 3 a conventional CV Stratified Group K-fold strategy can be applied, but for stages 1 and 2 I'm not sure. Maybe in stage 1 train the teacher model with all the data, without validation, and for stage 2 reserving a holdout set for validation for training the student model?</p>\n<p>Thanks, and sorry for my dumb question.</p>",
      "rawMarkdown": "Thanks, Y.Nakama for your insightful and helpful posts.\n\nMay I ask a newbie's question?\n\nHow do you do training in stages 1 and 2? I mean, In understand that for stage 3 a conventional CV Stratified Group K-fold strategy can be applied, but for stages 1 and 2 I'm not sure. Maybe in stage 1 train the teacher model with all the data, without validation, and for stage 2 reserving a holdout set for validation for training the student model?\n\nThanks, and sorry for my dumb question."
    },
    {
      "id": 1148861,
      "postDate": "2021-01-11T12:40:52.253Z",
      "content": "<p>Hi,I have a question.How do we adjust the params 'weights' in step2 training,you gaved is [0.5,1.0].Thanks!</p>",
      "rawMarkdown": "Hi,I have a question.How do we adjust the params 'weights' in step2 training,you gaved is [0.5,1.0].Thanks!"
    },
    {
      "id": 1148629,
      "postDate": "2021-01-11T09:19:17.293Z",
      "content": "<p>Thanks for sharing such great work!<br>\nCould you tell me why the COLOR MAP trick works on this project? I mean it is likely that the teacher model would be supervised by the COLOR signals rather than the semantic information of the X-ray. </p>",
      "rawMarkdown": "Thanks for sharing such great work!\nCould you tell me why the COLOR MAP trick works on this project? I mean it is likely that the teacher model would be supervised by the COLOR signals rather than the semantic information of the X-ray. ",
      "replies": [
        {
          "id": 1148705,
          "postDate": "2021-01-11T10:29:54.037Z",
          "content": "<p>in summary:<br>\n1) human input supervision by color<br>\n2)teacher convert supervision to feature map<br>\n3)student must find real signals to produce the same feature map as the teacher<br>\n(i.e. find the image features that is equivalent as if the image are colored)</p>\n<p>the role of the teacher is to convert human supervision to feature map supervision</p>",
          "rawMarkdown": "in summary:\n1) human input supervision by color\n2)teacher convert supervision to feature map\n3)student must find real signals to produce the same feature map as the teacher\n(i.e. find the image features that is equivalent as if the image are colored)\n\nthe role of the teacher is to convert human supervision to feature map supervision",
          "votes": 11
        }
      ]
    },
    {
      "id": 1147692,
      "postDate": "2021-01-10T16:43:04.007Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> thanks for sharing an excellent series of notebooks, I've learned a lot from them. In stage 2, have you tried other losses here? Seems like KL divergence would make a good fit.</p>",
      "rawMarkdown": "Hi @yasufuminakama thanks for sharing an excellent series of notebooks, I've learned a lot from them. In stage 2, have you tried other losses here? Seems like KL divergence would make a good fit.",
      "replies": [
        {
          "id": 1148696,
          "postDate": "2021-01-11T10:20:45.480Z",
          "content": "<p>You're welcome :) I haven't tried other losses yet.</p>",
          "rawMarkdown": "You're welcome :) I haven't tried other losses yet."
        },
        {
          "id": 1158524,
          "postDate": "2021-01-18T16:02:48.667Z",
          "content": "<p>Hi,did you test the KL loss,I test it ,but I get a lower score in stage2.It is only get 0.776.</p>",
          "rawMarkdown": "Hi,did you test the KL loss,I test it ,but I get a lower score in stage2.It is only get 0.776."
        }
      ]
    },
    {
      "id": 1141302,
      "postDate": "2021-01-06T15:59:36.410Z",
      "content": "<p>Hey, great work and thank you for ideas and the code. Side question though. I noticed you are training on TPU with <code>nproc=1</code>. Have you seen speed or batch size improvement comparing pytorch xla on TPU and regular Pytorch on GPU? Have you managed to make it work on all 8 cores of TPU accelerator?</p>\n<p>Thanks in advance)<br>\nCheers,<br>\nAlexey</p>",
      "rawMarkdown": "Hey, great work and thank you for ideas and the code. Side question though. I noticed you are training on TPU with `nproc=1`. Have you seen speed or batch size improvement comparing pytorch xla on TPU and regular Pytorch on GPU? Have you managed to make it work on all 8 cores of TPU accelerator?\n\nThanks in advance)\nCheers,\nAlexey",
      "replies": [
        {
          "id": 1148702,
          "postDate": "2021-01-11T10:27:29.253Z",
          "content": "<p>In my experiment, GPU was faster than single core of TPU.</p>\n<blockquote>\n  <p>Have you managed to make it work on all 8 cores of TPU accelerator?</p>\n</blockquote>\n<p>No, I'm not sure but maybe something is wrong…</p>",
          "rawMarkdown": "In my experiment, GPU was faster than single core of TPU.\n\n> Have you managed to make it work on all 8 cores of TPU accelerator?\n\nNo, I'm not sure but maybe something is wrong..."
        }
      ]
    },
    {
      "id": 1135664,
      "postDate": "2021-01-02T12:28:50.667Z",
      "content": "<p>Can someone please explain the rationale of this 3 step process? What added benefit it has over the normal training?<br>\nAlso, what is a Teacher model and Student model - is this in context to Transfer learning? </p>",
      "rawMarkdown": "Can someone please explain the rationale of this 3 step process? What added benefit it has over the normal training?\nAlso, what is a Teacher model and Student model - is this in context to Transfer learning? \n ",
      "replies": [
        {
          "id": 1135666,
          "postDate": "2021-01-02T12:30:04.460Z",
          "content": "<p><a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a>, <a href=\"https://www.kaggle.com/piantic\" target=\"_blank\">@piantic</a>, <a href=\"https://www.kaggle.com/hiramcho\" target=\"_blank\">@hiramcho</a> please shed some light.</p>",
          "rawMarkdown": " @yasufuminakama, @piantic, @hiramcho please shed some light."
        },
        {
          "id": 1135667,
          "postDate": "2021-01-02T12:31:01.287Z",
          "content": "<p>Also, goes without saying, <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> you for sharing your approach.</p>",
          "rawMarkdown": "Also, goes without saying, @yasufuminakama you for sharing your approach.",
          "votes": -2
        },
        {
          "id": 1135932,
          "postDate": "2021-01-02T15:53:13.033Z",
          "content": "<p>Please check this thread <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/205243\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/205243</a></p>",
          "rawMarkdown": "Please check this thread https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/205243",
          "votes": 1
        },
        {
          "id": 1136579,
          "postDate": "2021-01-03T08:10:50.673Z",
          "content": "<p><a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> Thank you!</p>",
          "rawMarkdown": "@yasufuminakama Thank you!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1244437,
      "postDate": "2021-03-19T01:58:29.510Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1132903,
      "postDate": "2020-12-30T19:15:52.580Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1132504,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2020-12-30T12:58:31.920000",
      "content": "<p>When I first joined Kaggle few years ago, I was amazed to see a brilliant 15 or (16) y.o guy <a href=\"https://www.kaggle.com/anokas\" target=\"_blank\">@anokas</a>  who used to wrote quickly very robust starter kernels at almost every competition. </p>\n<p>It seems now <a href=\"https://www.kaggle.com/xhlulu\" target=\"_blank\">@xhlulu</a> and you have brilliantly taken it over for TF and Pytorch respectively. </p>\n<p>Even though I have my own workflow and Pipeline. Your kernels are always a great inspiration to adapt mines on new competitions. </p>\n<p>Many thanks and keep up the great work ! </p>\n<p>And of course huge thanks to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for giving too much inspirations and ideas to all of us. </p>",
      "votes": 7,
      "replies": [
        {
          "id": 1132511,
          "author_name": "Y.Nakama",
          "author_url": "",
          "post_date": "2020-12-30T13:05:14",
          "content": "<p>Thank you! :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1132554,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-12-30T13:44:41.333000",
      "content": "<p><a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> </p>\n<p>thanks for the results and experiments.<br>\n I only verify that it would work based on my experiments. but I haven't implemented and submitted it yet.<br>\nbased on my experiments, the potential increase can be about 1%.</p>\n<p>you can check the CAM map difference before and after the approach.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1135016,
      "author_name": "Heroseo",
      "author_url": "",
      "post_date": "2021-01-01T19:45:37.923000",
      "content": "<p>It's always amazing how ideas turn into reality. :)</p>\n<p>Thanks <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a>.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1133846,
      "author_name": "Hiram Coria 🧬",
      "author_url": "",
      "post_date": "2020-12-31T15:26:01.370000",
      "content": "<p>Is this <strong>Knowledge Distillation</strong>? or another technique?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1132438,
      "author_name": "Atanas Atanasov",
      "author_url": "",
      "post_date": "2020-12-30T12:08:28.443000",
      "content": "<p>Unfortunately can upvote only once all these kernels :( hats off <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> !</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1132441,
          "author_name": "Y.Nakama",
          "author_url": "",
          "post_date": "2020-12-30T12:13:12.513000",
          "content": "<p>Thanks for your upvote <a href=\"https://www.kaggle.com/atanasova\" target=\"_blank\">@atanasova</a> :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1230264,
      "author_name": "James Condon",
      "author_url": "",
      "post_date": "2021-03-08T00:22:59.380000",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a>, thanks for this. What is 'w' in the 2nd stage loss. A constant?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1236032,
          "author_name": "Aman Deep Gupta",
          "author_url": "",
          "post_date": "2021-03-12T17:31:20.207000",
          "content": "<p>Yes, this gives weightage to  MSELoss. <br>\nFor me 1 for 'BCEWithLogitsLoss' and 0.5 for 'MSELoss' worked best.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1225263,
      "author_name": "Evgeny Sidorov",
      "author_url": "",
      "post_date": "2021-03-03T13:38:29.990000",
      "content": "<p>Great work, thanx</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1224982,
      "author_name": "jcesquivel",
      "author_url": "",
      "post_date": "2021-03-03T08:09:17.320000",
      "content": "<p>Thanks, Y.Nakama for your insightful and helpful posts.</p>\n<p>May I ask a newbie's question?</p>\n<p>How do you do training in stages 1 and 2? I mean, In understand that for stage 3 a conventional CV Stratified Group K-fold strategy can be applied, but for stages 1 and 2 I'm not sure. Maybe in stage 1 train the teacher model with all the data, without validation, and for stage 2 reserving a holdout set for validation for training the student model?</p>\n<p>Thanks, and sorry for my dumb question.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1148861,
      "author_name": "Bcw93",
      "author_url": "",
      "post_date": "2021-01-11T12:40:52.253000",
      "content": "<p>Hi,I have a question.How do we adjust the params 'weights' in step2 training,you gaved is [0.5,1.0].Thanks!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1148629,
      "author_name": "BIRDAD",
      "author_url": "",
      "post_date": "2021-01-11T09:19:17.293000",
      "content": "<p>Thanks for sharing such great work!<br>\nCould you tell me why the COLOR MAP trick works on this project? I mean it is likely that the teacher model would be supervised by the COLOR signals rather than the semantic information of the X-ray. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1148705,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-01-11T10:29:54.037000",
          "content": "<p>in summary:<br>\n1) human input supervision by color<br>\n2)teacher convert supervision to feature map<br>\n3)student must find real signals to produce the same feature map as the teacher<br>\n(i.e. find the image features that is equivalent as if the image are colored)</p>\n<p>the role of the teacher is to convert human supervision to feature map supervision</p>",
          "votes": 11,
          "replies": []
        }
      ]
    },
    {
      "id": 1147692,
      "author_name": "datasaurus",
      "author_url": "",
      "post_date": "2021-01-10T16:43:04.007000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> thanks for sharing an excellent series of notebooks, I've learned a lot from them. In stage 2, have you tried other losses here? Seems like KL divergence would make a good fit.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1148696,
          "author_name": "Y.Nakama",
          "author_url": "",
          "post_date": "2021-01-11T10:20:45.480000",
          "content": "<p>You're welcome :) I haven't tried other losses yet.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1158524,
          "author_name": "Bcw93",
          "author_url": "",
          "post_date": "2021-01-18T16:02:48.667000",
          "content": "<p>Hi,did you test the KL loss,I test it ,but I get a lower score in stage2.It is only get 0.776.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1141302,
      "author_name": "A.Demyanchuk",
      "author_url": "",
      "post_date": "2021-01-06T15:59:36.410000",
      "content": "<p>Hey, great work and thank you for ideas and the code. Side question though. I noticed you are training on TPU with <code>nproc=1</code>. Have you seen speed or batch size improvement comparing pytorch xla on TPU and regular Pytorch on GPU? Have you managed to make it work on all 8 cores of TPU accelerator?</p>\n<p>Thanks in advance)<br>\nCheers,<br>\nAlexey</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1148702,
          "author_name": "Y.Nakama",
          "author_url": "",
          "post_date": "2021-01-11T10:27:29.253000",
          "content": "<p>In my experiment, GPU was faster than single core of TPU.</p>\n<blockquote>\n  <p>Have you managed to make it work on all 8 cores of TPU accelerator?</p>\n</blockquote>\n<p>No, I'm not sure but maybe something is wrong…</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1135664,
      "author_name": "Dr. Amritpal Singh",
      "author_url": "",
      "post_date": "2021-01-02T12:28:50.667000",
      "content": "<p>Can someone please explain the rationale of this 3 step process? What added benefit it has over the normal training?<br>\nAlso, what is a Teacher model and Student model - is this in context to Transfer learning? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1135666,
          "author_name": "Dr. Amritpal Singh",
          "author_url": "",
          "post_date": "2021-01-02T12:30:04.460000",
          "content": "<p><a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a>, <a href=\"https://www.kaggle.com/piantic\" target=\"_blank\">@piantic</a>, <a href=\"https://www.kaggle.com/hiramcho\" target=\"_blank\">@hiramcho</a> please shed some light.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1135667,
          "author_name": "Dr. Amritpal Singh",
          "author_url": "",
          "post_date": "2021-01-02T12:31:01.287000",
          "content": "<p>Also, goes without saying, <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> you for sharing your approach.</p>",
          "votes": -2,
          "replies": []
        },
        {
          "id": 1135932,
          "author_name": "Y.Nakama",
          "author_url": "",
          "post_date": "2021-01-02T15:53:13.033000",
          "content": "<p>Please check this thread <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/205243\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/205243</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1136579,
          "author_name": "Dr. Amritpal Singh",
          "author_url": "",
          "post_date": "2021-01-03T08:10:50.673000",
          "content": "<p><a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> Thank you!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1244437,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-19T01:58:29.510000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1132903,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-30T19:15:52.580000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1132394": "The simplest way to use additional annotation is already discussed in [this thread](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/205243) by @hengck23.\nI implemented this kind of approach on kaggle notebooks (you can select GPU/TPU) and local validation score can get 0.95x. In my experience, there is 0.01 improvement with this approach.\n\n## Training strategy\n- [1st-stage training](https://www.kaggle.com/yasufuminakama/ranzcr-resnet200d-3-stage-training-step1)\n    - teacher model training for annotated image\n        - data: annotated data\n        - pretrained weight: imagenet weight\n        - `BCEWithLogitsLoss(y_preds, labels)`\n        - `y_preds: teacher model predictions for annotated image`\n- [2nd-stage training](https://www.kaggle.com/yasufuminakama/ranzcr-resnet200d-3-stage-training-step2)\n    - student model training with teacher model features\n        - data: annotated data\n        - student model pretrained weight: imagenet weight\n        - teacher model pretrained weight: 1st-stage weight\n        - `BCEWithLogitsLoss(y_preds, labels) + w * MSELoss(student_features, teacher_features)`\n        - `y_preds: student model predictions for normal image`\n        - `student_features: student model features for normal image`\n        - `teacher_features: teacher model features for annotated image`\n- [3rd-stage training](https://www.kaggle.com/yasufuminakama/ranzcr-resnet200d-3-stage-training-step3)\n    - model training\n        - data: all data\n        - pretrained weight: 2nd-stage weight\n        - `BCEWithLogitsLoss(y_preds, labels)`\n        - `y_preds: student model predictions for normal image`\n- [inference notebook](https://www.kaggle.com/yasufuminakama/ranzcr-resnet200d-3-stage-training-sub)\n\nHope this helps, happy kaggling :)",
    "1132504": "When I first joined Kaggle few years ago, I was amazed to see a brilliant 15 or (16) y.o guy @anokas  who used to wrote quickly very robust starter kernels at almost every competition. \n\nIt seems now @xhlulu and you have brilliantly taken it over for TF and Pytorch respectively. \n\nEven though I have my own workflow and Pipeline. Your kernels are always a great inspiration to adapt mines on new competitions. \n\nMany thanks and keep up the great work ! \n\nAnd of course huge thanks to @hengck23 for giving too much inspirations and ideas to all of us. ",
    "1132554": "@yasufuminakama \n\nthanks for the results and experiments.\n I only verify that it would work based on my experiments. but I haven't implemented and submitted it yet.\nbased on my experiments, the potential increase can be about 1%.\n\nyou can check the CAM map difference before and after the approach.\n\n",
    "1135016": "It's always amazing how ideas turn into reality. :)\n\nThanks @hengck23 @yasufuminakama.",
    "1133846": "Is this **Knowledge Distillation**? or another technique?",
    "1132438": "Unfortunately can upvote only once all these kernels :( hats off @yasufuminakama !",
    "1230264": "Hey @yasufuminakama, thanks for this. What is 'w' in the 2nd stage loss. A constant?",
    "1225263": "Great work, thanx",
    "1224982": "Thanks, Y.Nakama for your insightful and helpful posts.\n\nMay I ask a newbie's question?\n\nHow do you do training in stages 1 and 2? I mean, In understand that for stage 3 a conventional CV Stratified Group K-fold strategy can be applied, but for stages 1 and 2 I'm not sure. Maybe in stage 1 train the teacher model with all the data, without validation, and for stage 2 reserving a holdout set for validation for training the student model?\n\nThanks, and sorry for my dumb question.",
    "1148861": "Hi,I have a question.How do we adjust the params 'weights' in step2 training,you gaved is [0.5,1.0].Thanks!",
    "1148629": "Thanks for sharing such great work!\nCould you tell me why the COLOR MAP trick works on this project? I mean it is likely that the teacher model would be supervised by the COLOR signals rather than the semantic information of the X-ray. ",
    "1147692": "Hi @yasufuminakama thanks for sharing an excellent series of notebooks, I've learned a lot from them. In stage 2, have you tried other losses here? Seems like KL divergence would make a good fit.",
    "1141302": "Hey, great work and thank you for ideas and the code. Side question though. I noticed you are training on TPU with `nproc=1`. Have you seen speed or batch size improvement comparing pytorch xla on TPU and regular Pytorch on GPU? Have you managed to make it work on all 8 cores of TPU accelerator?\n\nThanks in advance)\nCheers,\nAlexey",
    "1135664": "Can someone please explain the rationale of this 3 step process? What added benefit it has over the normal training?\nAlso, what is a Teacher model and Student model - is this in context to Transfer learning? \n ",
    "1244437": "",
    "1132903": ""
  }
}