{
  "id": 205243,
  "title": "simplest way to use additional annotation to train classifier without segmentation",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/205243",
  "author_name": "hengck23",
  "post_date": "2020-12-19T06:27:46.523000",
  "votes": 170,
  "comment_count": 34,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd8e62984343f73764b13346d5d16aa71%2FSelection_089.png?generation=1608359259306689&amp;alt=media\" alt=\"\"></p>\n<p>the key idea is to modify the activation (activation of the feature map in this case) by changing the input so that the network can pick up extra supervision and decide if they want to use them or not.</p>\n<p>There are many ways to implement this.  if you want to mix those with and without annotations, you can add a flag \"is additional annotation available\" as input at training. this is probably the easiest \"human in the loop for training\"</p>\n<p>note: if the kaggle annotation is not enough, add your own annotation, etc. part of lung, etc</p>\n<p>in my experiment, I estimated the addition annotation supervision  signals will cause an improvement of about 2.0% to 2.5% for the AUC metric</p>",
  "messages": [
    {
      "id": 1118515,
      "postDate": "2020-12-19T06:27:46.523Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd8e62984343f73764b13346d5d16aa71%2FSelection_089.png?generation=1608359259306689&amp;alt=media\" alt=\"\"></p>\n<p>the key idea is to modify the activation (activation of the feature map in this case) by changing the input so that the network can pick up extra supervision and decide if they want to use them or not.</p>\n<p>There are many ways to implement this.  if you want to mix those with and without annotations, you can add a flag \"is additional annotation available\" as input at training. this is probably the easiest \"human in the loop for training\"</p>\n<p>note: if the kaggle annotation is not enough, add your own annotation, etc. part of lung, etc</p>\n<p>in my experiment, I estimated the addition annotation supervision  signals will cause an improvement of about 2.0% to 2.5% for the AUC metric</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd8e62984343f73764b13346d5d16aa71%2FSelection_089.png?generation=1608359259306689&alt=media)\n\nthe key idea is to modify the activation (activation of the feature map in this case) by changing the input so that the network can pick up extra supervision and decide if they want to use them or not.\n\nThere are many ways to implement this.  if you want to mix those with and without annotations, you can add a flag \"is additional annotation available\" as input at training. this is probably the easiest \"human in the loop for training\"\n\nnote: if the kaggle annotation is not enough, add your own annotation, etc. part of lung, etc\n\nin my experiment, I estimated the addition annotation supervision  signals will cause an improvement of about 2.0% to 2.5% for the AUC metric",
      "votes": 169
    },
    {
      "id": 1213323,
      "postDate": "2021-02-22T03:56:17.227Z",
      "content": "<p>check these as well</p>\n<p>Embedding Human Knowledge into Deep Neural Network via Attention Map<br>\n<a href=\"https://arxiv.org/abs/1905.03540\" target=\"_blank\">https://arxiv.org/abs/1905.03540</a></p>\n<p><a href=\"https://github.com/machine-perception-robotics-group/attention_branch_network/\" target=\"_blank\">https://github.com/machine-perception-robotics-group/attention_branch_network/</a></p>\n<p><img src=\"https://storage.googleapis.com/groundai-web-prod/media%2Fusers%2Fuser_233037%2Fproject_360896%2Fimages%2Fx1.png\" alt=\"\"></p>",
      "rawMarkdown": "check these as well\n\nEmbedding Human Knowledge into Deep Neural Network via Attention Map\nhttps://arxiv.org/abs/1905.03540\n\nhttps://github.com/machine-perception-robotics-group/attention_branch_network/\n\n![](https://storage.googleapis.com/groundai-web-prod/media%2Fusers%2Fuser_233037%2Fproject_360896%2Fimages%2Fx1.png)\n",
      "votes": 8,
      "replies": [
        {
          "id": 1213408,
          "postDate": "2021-02-22T05:20:03.200Z",
          "content": "<p>Human-in-the-loop!</p>",
          "rawMarkdown": "Human-in-the-loop!"
        }
      ]
    },
    {
      "id": 1226219,
      "postDate": "2021-03-04T10:49:16.970Z",
      "content": "<p>reference code and model is here!<br>\n<a href=\"https://drive.google.com/drive/folders/1JRqCpV3NMGFeyd5ws6_pyXwQ5x3n8CHF\" target=\"_blank\">https://drive.google.com/drive/folders/1JRqCpV3NMGFeyd5ws6_pyXwQ5x3n8CHF</a><br>\nsee folder \"2021-03-04\" and refer to the readme.pptx</p>\n<p><img src=\"https://i.ibb.co/cvNVfSy/Selection-035.png\" alt=\"\"></p>",
      "rawMarkdown": "reference code and model is here!\nhttps://drive.google.com/drive/folders/1JRqCpV3NMGFeyd5ws6_pyXwQ5x3n8CHF\nsee folder \"2021-03-04\" and refer to the readme.pptx\n\n![](https://i.ibb.co/cvNVfSy/Selection-035.png)",
      "votes": 5,
      "replies": [
        {
          "id": 1229168,
          "postDate": "2021-03-07T05:40:54.820Z",
          "content": "<p>i further compare my implementation and those by public kernels:<br>\n<a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/215910\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/215910</a><br>\n<a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207577\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207577</a></p>\n<p>i confirm a third stage will improve results:<br>\nstage3 : remove teacher consistency loss and finetune with only bce</p>\n<p>this is because teacher consistency loss can limit performance at fine tunning. L2 loss is higher than BCE.</p>\n<p>L2 loss acts as regularsier in the early stage, but hinders improvement at fine tunning.</p>",
          "rawMarkdown": "i further compare my implementation and those by public kernels:\nhttps://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/215910\nhttps://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207577\n\ni confirm a third stage will improve results:\nstage3 : remove teacher consistency loss and finetune with only bce\n\nthis is because teacher consistency loss can limit performance at fine tunning. L2 loss is higher than BCE.\n\n L2 loss acts as regularsier in the early stage, but hinders improvement at fine tunning."
        }
      ]
    },
    {
      "id": 1118777,
      "postDate": "2020-12-19T11:46:47.753Z",
      "content": "<p>if one wants to be fanciful (typo error: should have been x = (1+a)*x</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd5d7d39813aa6095732c9df13900f9fc%2FSelection_092.png?generation=1608378425203989&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "if one wants to be fanciful (typo error: should have been x = (1+a)*x\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd5d7d39813aa6095732c9df13900f9fc%2FSelection_092.png?generation=1608378425203989&alt=media)",
      "votes": 4,
      "replies": [
        {
          "id": 1118784,
          "postDate": "2020-12-19T11:52:07.100Z",
          "content": "<p>even more fanciful<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fdfd1a32b29050b169e0ab8965afd6c02%2FSelection_094.png?generation=1608378725312235&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc7d613c6a38945ecd88249728b1c997a%2FSelection_099.png?generation=1608428351163583&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "even more fanciful\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fdfd1a32b29050b169e0ab8965afd6c02%2FSelection_094.png?generation=1608378725312235&alt=media)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc7d613c6a38945ecd88249728b1c997a%2FSelection_099.png?generation=1608428351163583&alt=media)",
          "votes": 6
        }
      ]
    },
    {
      "id": 1221979,
      "postDate": "2021-03-01T13:01:36.987Z",
      "content": "<blockquote>\n  <p>I estimated the addition annotation supervision  signals will cause an improvement of about 2.0% to 2.5% for the AUC metric</p>\n</blockquote>\n<p>Is your mean your CV increase from <em>x</em> to <em>x+0.02</em>?<br>\nThanks for sharing your ideas. These are really very impressive. </p>",
      "rawMarkdown": "> I estimated the addition annotation supervision  signals will cause an improvement of about 2.0% to 2.5% for the AUC metric\n\nIs your mean your CV increase from *x* to *x+0.02*?\nThanks for sharing your ideas. These are really very impressive. \n\n\n\n\n",
      "votes": 1
    },
    {
      "id": 1218383,
      "postDate": "2021-02-25T19:31:03.053Z",
      "content": "<p>i haven't tried this yet, but it is an idea: contrastive attention for self-supervised learning</p>\n<p>the idea is:</p>\n<ul>\n<li>use all images (annotated and non annotated) for contrastive learning:<br>\nattention_loss (image, augment_image) </li>\n<li>then finetune with annotated images</li>\n</ul>",
      "rawMarkdown": "i haven't tried this yet, but it is an idea: contrastive attention for self-supervised learning\n\nthe idea is:\n- use all images (annotated and non annotated) for contrastive learning:\n   attention\\_loss (image, augment\\_image) \n- then finetune with annotated images",
      "votes": 1,
      "replies": [
        {
          "id": 1218646,
          "postDate": "2021-02-26T04:51:01.890Z",
          "content": "<p>Hello, are we allowed to use both normal images dataset and annotated images dataset? If I understand this  competition correctly? Haha sorry, I joined late.</p>",
          "rawMarkdown": "Hello, are we allowed to use both normal images dataset and annotated images dataset? If I understand this  competition correctly? Haha sorry, I joined late."
        },
        {
          "id": 1220918,
          "postDate": "2021-02-28T13:45:03.440Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/h3309089\" target=\"_blank\">@h3309089</a> can you please give us a intuition on what contrastive learning is , I am very new to this topic , it will help me a lot</p>\n<p>Thanks</p>",
          "rawMarkdown": "Hi @h3309089 can you please give us a intuition on what contrastive learning is , I am very new to this topic , it will help me a lot\n\nThanks",
          "votes": -1
        },
        {
          "id": 1221994,
          "postDate": "2021-03-01T13:10:31.807Z",
          "content": "<p>search for MoCo or SimCLR. Note it requires a lot of computation and sometimes huge batch sizes</p>",
          "rawMarkdown": "search for MoCo or SimCLR. Note it requires a lot of computation and sometimes huge batch sizes"
        },
        {
          "id": 1224362,
          "postDate": "2021-03-02T16:52:56.470Z",
          "content": "<p>added a few interesting ideas<br>\n<a href=\"https://ibb.co/2NhCgKT\" target=\"_blank\">https://ibb.co/2NhCgKT</a><br>\n<a href=\"https://ibb.co/CMc22pW\" target=\"_blank\">https://ibb.co/CMc22pW</a><br>\n<a href=\"https://ibb.co/S09dJ3x\" target=\"_blank\">https://ibb.co/S09dJ3x</a><br>\n<a href=\"https://ibb.co/C7mC5F1\" target=\"_blank\">https://ibb.co/C7mC5F1</a></p>\n<p>extreme augmentation with teacher attention distillation improves my LB from 0.959 to 0.960 for single fold resnet50d 640 input.</p>\n<p>These extreme augmentations failed without teacher distillation</p>",
          "rawMarkdown": "added a few interesting ideas\nhttps://ibb.co/2NhCgKT\nhttps://ibb.co/CMc22pW\nhttps://ibb.co/S09dJ3x\nhttps://ibb.co/C7mC5F1\n\nextreme augmentation with teacher attention distillation improves my LB from 0.959 to 0.960 for single fold resnet50d 640 input.\n\nThese extreme augmentations failed without teacher distillation",
          "votes": 1
        },
        {
          "id": 1224366,
          "postDate": "2021-03-02T16:54:11.363Z",
          "content": "<p>\"search for MoCo or SimCLR. Note it requires a lot of computation and sometimes huge batch sizes\"</p>\n<p>i believe strong supervision reduce batch size. maybe someone can write a paper on it … </p>",
          "rawMarkdown": "\"search for MoCo or SimCLR. Note it requires a lot of computation and sometimes huge batch sizes\"\n\ni believe strong supervision reduce batch size. maybe someone can write a paper on it ... ",
          "votes": -1
        },
        {
          "id": 1225777,
          "postDate": "2021-03-03T23:13:24.487Z",
          "content": "<p>multi task teacher is also giving higher LB score (LB 0.960 for single fold resnet50d).<br>\nit is better to have a separate feature map teacher for each task.</p>",
          "rawMarkdown": "multi task teacher is also giving higher LB score (LB 0.960 for single fold resnet50d).\nit is better to have a separate feature map teacher for each task."
        }
      ]
    },
    {
      "id": 1136710,
      "postDate": "2021-01-03T11:02:31.183Z",
      "content": "<p>Could you please share some work or paper that implements the presented concepts using TF2.0. Thanks</p>",
      "rawMarkdown": "Could you please share some work or paper that implements the presented concepts using TF2.0. Thanks",
      "votes": 1
    },
    {
      "id": 1118743,
      "postDate": "2020-12-19T11:04:21.597Z",
      "content": "<p>Thank you for sharing your idea! I'm not familiar with this concept. So I would like to ask some questions if you don't mind. What do you use for consistency loss? Do you train the teacher first and then distill the teacher to the student?</p>",
      "rawMarkdown": "Thank you for sharing your idea! I'm not familiar with this concept. So I would like to ask some questions if you don't mind. What do you use for consistency loss? Do you train the teacher first and then distill the teacher to the student?\n",
      "votes": 1,
      "replies": [
        {
          "id": 1118765,
          "postDate": "2020-12-19T11:32:40.230Z",
          "content": "<p>\"What do you use for consistency loss?\"<br>\nI think this is not so important. i think normal L2 loss is good enough</p>\n<p>\"Do you train the teacher first and then distill the teacher to the student?\"<br>\nmy plan is to train the teacher first and freeze it.  a model trained on rendered image can achieve about 96% in local CV</p>",
          "rawMarkdown": "\"What do you use for consistency loss?\"\nI think this is not so important. i think normal L2 loss is good enough\n\n\"Do you train the teacher first and then distill the teacher to the student?\"\nmy plan is to train the teacher first and freeze it.  a model trained on rendered image can achieve about 96% in local CV",
          "votes": 5
        },
        {
          "id": 1118770,
          "postDate": "2020-12-19T11:39:37.037Z",
          "content": "<p>Thank you!</p>",
          "rawMarkdown": "Thank you!"
        },
        {
          "id": 1136613,
          "postDate": "2021-01-03T08:49:09.620Z",
          "content": "<p>Got great insights from this discussion. Thank you <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>",
          "rawMarkdown": "Got great insights from this discussion. Thank you @hengck23 "
        },
        {
          "id": 1212293,
          "postDate": "2021-02-21T04:39:34.167Z",
          "content": "<p>Will teacher better than student?</p>",
          "rawMarkdown": "Will teacher better than student?"
        }
      ]
    },
    {
      "id": 1239780,
      "postDate": "2021-03-16T02:25:41.553Z",
      "content": "<p>i find a new home for this method. it think it will shine there<br>\n<a href=\"https://www.kaggle.com/c/bms-molecular-translation\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation</a></p>\n<p><img src=\"https://i.ibb.co/L1S3cW0/Selection-352.png\" alt=\"\"></p>",
      "rawMarkdown": "i find a new home for this method. it think it will shine there\nhttps://www.kaggle.com/c/bms-molecular-translation\n\n![](https://i.ibb.co/L1S3cW0/Selection-352.png)",
      "votes": 2
    },
    {
      "id": 1217634,
      "postDate": "2021-02-25T08:05:45.050Z",
      "content": "<p>some results.</p>\n<p>i use two-stage (instead of three-stage in some public code)</p>\n<p><img src=\"https://i.ibb.co/nmJBJNx/Selection-191.png\" alt=\"\"><br>\n<img src=\"https://i.ibb.co/4YfY3xG/Selection-193.png\" alt=\"\"></p>",
      "rawMarkdown": "some results.\n\ni use two-stage (instead of three-stage in some public code)\n\n![](https://i.ibb.co/nmJBJNx/Selection-191.png)\n![](https://i.ibb.co/4YfY3xG/Selection-193.png)",
      "votes": 1,
      "replies": [
        {
          "id": 1217647,
          "postDate": "2021-02-25T08:09:43.957Z",
          "content": "<pre><code>            batch_size = len(batch['index'])\n            df  = batch['df']\n\n            target = batch['target'].cuda()\n            image  = batch['image'].cuda()\n            render = batch['render'].cuda()\n            is_annotate = batch['is_annotate'].cuda()\n\n            with torch.no_grad():\n                teacher.eval()\n                logit_t, f_t = teacher(render)\n\n            #----\n            net.train()\n            optimizer.zero_grad()\n\n            def loss_consistency(f,truth):\n                f = f[is_annotate==1]\n                t = truth[is_annotate==1]\n                loss = F.mse_loss(f,t)\n                return loss\n\n            if is_mixed_precision: \n                image = image.half()\n                with amp.autocast():\n                    logit,f = net(image)\n                    loss0 = binary_cross_entropy_loss(logit, target)\n                    loss1 = loss_consistency(f, f_t)\n\n                scaler.scale(loss0+loss1).backward()\n                scaler.step(optimizer)\n                scaler.update()\n</code></pre>",
          "rawMarkdown": "```\n            batch_size = len(batch['index'])\n            df  = batch['df']\n\n            target = batch['target'].cuda()\n            image  = batch['image'].cuda()\n            render = batch['render'].cuda()\n            is_annotate = batch['is_annotate'].cuda()\n\n            with torch.no_grad():\n                teacher.eval()\n                logit_t, f_t = teacher(render)\n\n            #----\n            net.train()\n            optimizer.zero_grad()\n\n            def loss_consistency(f,truth):\n                f = f[is_annotate==1]\n                t = truth[is_annotate==1]\n                loss = F.mse_loss(f,t)\n                return loss\n \n            if is_mixed_precision: \n                image = image.half()\n                with amp.autocast():\n                    logit,f = net(image)\n                    loss0 = binary_cross_entropy_loss(logit, target)\n                    loss1 = loss_consistency(f, f_t)\n \n                scaler.scale(loss0+loss1).backward()\n                scaler.step(optimizer)\n                scaler.update()\n\n```",
          "votes": 2
        },
        {
          "id": 1219558,
          "postDate": "2021-02-27T00:23:23.673Z",
          "content": "<p>since the score on the \"my submissions\" page can be ranked, I can see:</p>\n<p>only one fold:<br>\nfor 640x640 input, tta= origina+flip:<br>\nresnet200d + teacher attention : local cv 0.956805, lb 0.959 (best)<br>\nresnet50d + teacher attention : local cv 0.956385, lb 0.959 <br>\nresnet200d : local cv 0.955946, lb 0.959 (worst)<br>\n(<a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/204950\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/204950</a>)</p>\n<p>it is seen that now resnet50d is as good as resnet200d with extra supervision</p>",
          "rawMarkdown": "since the score on the \"my submissions\" page can be ranked, I can see:\n\nonly one fold:\nfor 640x640 input, tta= origina+flip:\nresnet200d + teacher attention : local cv 0.956805, lb 0.959 (best)\nresnet50d + teacher attention : local cv 0.956385, lb 0.959 \nresnet200d : local cv 0.955946, lb 0.959 (worst)\n(https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/204950)\n\n\nit is seen that now resnet50d is as good as resnet200d with extra supervision",
          "votes": 1
        },
        {
          "id": 1220422,
          "postDate": "2021-02-28T01:26:45.613Z",
          "content": "<p>i will not be surprised that more hand-annotation will be the key to winning this challenge.<br>\nwe will see who can do it \"smartly\" </p>\n<p><img src=\"https://i.ibb.co/pd14X6t/Selection-215.png\" alt=\"\"></p>",
          "rawMarkdown": "i will not be surprised that more hand-annotation will be the key to winning this challenge.\nwe will see who can do it \"smartly\" \n\n![](https://i.ibb.co/pd14X6t/Selection-215.png)",
          "votes": 1
        },
        {
          "id": 1220430,
          "postDate": "2021-02-28T01:40:04.663Z",
          "content": "<p>replacing line annotation with endpoint annotation. how is teacher distillation affected?<br>\nNote this is results for CVC only. the others like ETT are not used at all</p>\n<p><img src=\"https://i.ibb.co/Y3j63HP/Selection-221.png\" alt=\"\"></p>",
          "rawMarkdown": "replacing line annotation with endpoint annotation. how is teacher distillation affected?\nNote this is results for CVC only. the others like ETT are not used at all\n\n![](https://i.ibb.co/Y3j63HP/Selection-221.png)",
          "votes": 2
        },
        {
          "id": 1220685,
          "postDate": "2021-02-28T08:50:36.873Z",
          "content": "<p>be creative!</p>\n<p>multi scale annotation, find pretain model that segment lung,heart, etc …. maybe they would work?</p>",
          "rawMarkdown": "be creative!\n\nmulti scale annotation, find pretain model that segment lung,heart, etc .... maybe they would work?",
          "votes": 2
        },
        {
          "id": 1220873,
          "postDate": "2021-02-28T12:39:29.817Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> you tried normfree nets? how does it work for you?</p>",
          "rawMarkdown": "@hengck23 you tried normfree nets? how does it work for you?",
          "votes": 1
        }
      ]
    },
    {
      "id": 1151143,
      "postDate": "2021-01-13T06:53:54.077Z",
      "content": "<p>i would like to combine my idea with this google paper:<br>\nConcept Bottleneck Models<br>\n<a href=\"https://arxiv.org/pdf/2007.04612.pdf\" target=\"_blank\">https://arxiv.org/pdf/2007.04612.pdf</a></p>\n<p>\"By construction, we can intervene on these concept bottleneck models by editing their predicted concept values and propagating these changes to the final prediction\"</p>\n<p>i want to intervene by sketching on the image</p>",
      "rawMarkdown": "i would like to combine my idea with this google paper:\nConcept Bottleneck Models\nhttps://arxiv.org/pdf/2007.04612.pdf\n\n\"By construction, we can intervene on these concept bottleneck models by editing their predicted concept values and propagating these changes to the final prediction\"\n\ni want to intervene by sketching on the image",
      "votes": 1
    },
    {
      "id": 1214113,
      "postDate": "2021-02-22T15:43:44.077Z",
      "content": "<p>Great discussion thread. Can anyone share an example of the Teacher-Student model example kernel?  </p>",
      "rawMarkdown": "Great discussion thread. Can anyone share an example of the Teacher-Student model example kernel?  "
    },
    {
      "id": 1133438,
      "postDate": "2020-12-31T08:08:51.367Z",
      "content": "<p>Sir how you ensemble with annotation without annotations can you explain? And what part of the images you annotated ?  Thanks in advance.  </p>",
      "rawMarkdown": "Sir how you ensemble with annotation without annotations can you explain? And what part of the images you annotated ?  Thanks in advance.  "
    },
    {
      "id": 1121948,
      "postDate": "2020-12-22T03:40:55.427Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1121735,
      "postDate": "2020-12-21T21:34:26.697Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 1118630,
      "postDate": "2020-12-19T08:55:24.363Z",
      "content": "<p>Thanks for insights!</p>",
      "rawMarkdown": "Thanks for insights!"
    }
  ],
  "comments": [
    {
      "id": 1213323,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-02-22T03:56:17.227000",
      "content": "<p>check these as well</p>\n<p>Embedding Human Knowledge into Deep Neural Network via Attention Map<br>\n<a href=\"https://arxiv.org/abs/1905.03540\" target=\"_blank\">https://arxiv.org/abs/1905.03540</a></p>\n<p><a href=\"https://github.com/machine-perception-robotics-group/attention_branch_network/\" target=\"_blank\">https://github.com/machine-perception-robotics-group/attention_branch_network/</a></p>\n<p><img src=\"https://storage.googleapis.com/groundai-web-prod/media%2Fusers%2Fuser_233037%2Fproject_360896%2Fimages%2Fx1.png\" alt=\"\"></p>",
      "votes": 8,
      "replies": [
        {
          "id": 1213408,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2021-02-22T05:20:03.200000",
          "content": "<p>Human-in-the-loop!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1226219,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-03-04T10:49:16.970000",
      "content": "<p>reference code and model is here!<br>\n<a href=\"https://drive.google.com/drive/folders/1JRqCpV3NMGFeyd5ws6_pyXwQ5x3n8CHF\" target=\"_blank\">https://drive.google.com/drive/folders/1JRqCpV3NMGFeyd5ws6_pyXwQ5x3n8CHF</a><br>\nsee folder \"2021-03-04\" and refer to the readme.pptx</p>\n<p><img src=\"https://i.ibb.co/cvNVfSy/Selection-035.png\" alt=\"\"></p>",
      "votes": 5,
      "replies": [
        {
          "id": 1229168,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-03-07T05:40:54.820000",
          "content": "<p>i further compare my implementation and those by public kernels:<br>\n<a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/215910\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/215910</a><br>\n<a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207577\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207577</a></p>\n<p>i confirm a third stage will improve results:<br>\nstage3 : remove teacher consistency loss and finetune with only bce</p>\n<p>this is because teacher consistency loss can limit performance at fine tunning. L2 loss is higher than BCE.</p>\n<p>L2 loss acts as regularsier in the early stage, but hinders improvement at fine tunning.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1118777,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-12-19T11:46:47.753000",
      "content": "<p>if one wants to be fanciful (typo error: should have been x = (1+a)*x</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd5d7d39813aa6095732c9df13900f9fc%2FSelection_092.png?generation=1608378425203989&amp;alt=media\" alt=\"\"></p>",
      "votes": 4,
      "replies": [
        {
          "id": 1118784,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-12-19T11:52:07.100000",
          "content": "<p>even more fanciful<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fdfd1a32b29050b169e0ab8965afd6c02%2FSelection_094.png?generation=1608378725312235&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fc7d613c6a38945ecd88249728b1c997a%2FSelection_099.png?generation=1608428351163583&amp;alt=media\" alt=\"\"></p>",
          "votes": 6,
          "replies": []
        }
      ]
    },
    {
      "id": 1221979,
      "author_name": "Aman Deep Gupta",
      "author_url": "",
      "post_date": "2021-03-01T13:01:36.987000",
      "content": "<blockquote>\n  <p>I estimated the addition annotation supervision  signals will cause an improvement of about 2.0% to 2.5% for the AUC metric</p>\n</blockquote>\n<p>Is your mean your CV increase from <em>x</em> to <em>x+0.02</em>?<br>\nThanks for sharing your ideas. These are really very impressive. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1218383,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-02-25T19:31:03.053000",
      "content": "<p>i haven't tried this yet, but it is an idea: contrastive attention for self-supervised learning</p>\n<p>the idea is:</p>\n<ul>\n<li>use all images (annotated and non annotated) for contrastive learning:<br>\nattention_loss (image, augment_image) </li>\n<li>then finetune with annotated images</li>\n</ul>",
      "votes": 1,
      "replies": [
        {
          "id": 1218646,
          "author_name": "Andy Jian Zhou",
          "author_url": "",
          "post_date": "2021-02-26T04:51:01.890000",
          "content": "<p>Hello, are we allowed to use both normal images dataset and annotated images dataset? If I understand this  competition correctly? Haha sorry, I joined late.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1220918,
          "author_name": "Mr_KnowNothing",
          "author_url": "",
          "post_date": "2021-02-28T13:45:03.440000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/h3309089\" target=\"_blank\">@h3309089</a> can you please give us a intuition on what contrastive learning is , I am very new to this topic , it will help me a lot</p>\n<p>Thanks</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1221994,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2021-03-01T13:10:31.807000",
          "content": "<p>search for MoCo or SimCLR. Note it requires a lot of computation and sometimes huge batch sizes</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1224362,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-03-02T16:52:56.470000",
          "content": "<p>added a few interesting ideas<br>\n<a href=\"https://ibb.co/2NhCgKT\" target=\"_blank\">https://ibb.co/2NhCgKT</a><br>\n<a href=\"https://ibb.co/CMc22pW\" target=\"_blank\">https://ibb.co/CMc22pW</a><br>\n<a href=\"https://ibb.co/S09dJ3x\" target=\"_blank\">https://ibb.co/S09dJ3x</a><br>\n<a href=\"https://ibb.co/C7mC5F1\" target=\"_blank\">https://ibb.co/C7mC5F1</a></p>\n<p>extreme augmentation with teacher attention distillation improves my LB from 0.959 to 0.960 for single fold resnet50d 640 input.</p>\n<p>These extreme augmentations failed without teacher distillation</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1224366,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-03-02T16:54:11.363000",
          "content": "<p>\"search for MoCo or SimCLR. Note it requires a lot of computation and sometimes huge batch sizes\"</p>\n<p>i believe strong supervision reduce batch size. maybe someone can write a paper on it … </p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1225777,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-03-03T23:13:24.487000",
          "content": "<p>multi task teacher is also giving higher LB score (LB 0.960 for single fold resnet50d).<br>\nit is better to have a separate feature map teacher for each task.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1136710,
      "author_name": "konwarsky",
      "author_url": "",
      "post_date": "2021-01-03T11:02:31.183000",
      "content": "<p>Could you please share some work or paper that implements the presented concepts using TF2.0. Thanks</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1118743,
      "author_name": "Tolga",
      "author_url": "",
      "post_date": "2020-12-19T11:04:21.597000",
      "content": "<p>Thank you for sharing your idea! I'm not familiar with this concept. So I would like to ask some questions if you don't mind. What do you use for consistency loss? Do you train the teacher first and then distill the teacher to the student?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1118765,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-12-19T11:32:40.230000",
          "content": "<p>\"What do you use for consistency loss?\"<br>\nI think this is not so important. i think normal L2 loss is good enough</p>\n<p>\"Do you train the teacher first and then distill the teacher to the student?\"<br>\nmy plan is to train the teacher first and freeze it.  a model trained on rendered image can achieve about 96% in local CV</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1118770,
          "author_name": "Tolga",
          "author_url": "",
          "post_date": "2020-12-19T11:39:37.037000",
          "content": "<p>Thank you!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1136613,
          "author_name": "Dr. Amritpal Singh",
          "author_url": "",
          "post_date": "2021-01-03T08:49:09.620000",
          "content": "<p>Got great insights from this discussion. Thank you <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1212293,
          "author_name": "Zekun",
          "author_url": "",
          "post_date": "2021-02-21T04:39:34.167000",
          "content": "<p>Will teacher better than student?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1239780,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-03-16T02:25:41.553000",
      "content": "<p>i find a new home for this method. it think it will shine there<br>\n<a href=\"https://www.kaggle.com/c/bms-molecular-translation\" target=\"_blank\">https://www.kaggle.com/c/bms-molecular-translation</a></p>\n<p><img src=\"https://i.ibb.co/L1S3cW0/Selection-352.png\" alt=\"\"></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1217634,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-02-25T08:05:45.050000",
      "content": "<p>some results.</p>\n<p>i use two-stage (instead of three-stage in some public code)</p>\n<p><img src=\"https://i.ibb.co/nmJBJNx/Selection-191.png\" alt=\"\"><br>\n<img src=\"https://i.ibb.co/4YfY3xG/Selection-193.png\" alt=\"\"></p>",
      "votes": 1,
      "replies": [
        {
          "id": 1217647,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-02-25T08:09:43.957000",
          "content": "<pre><code>            batch_size = len(batch['index'])\n            df  = batch['df']\n\n            target = batch['target'].cuda()\n            image  = batch['image'].cuda()\n            render = batch['render'].cuda()\n            is_annotate = batch['is_annotate'].cuda()\n\n            with torch.no_grad():\n                teacher.eval()\n                logit_t, f_t = teacher(render)\n\n            #----\n            net.train()\n            optimizer.zero_grad()\n\n            def loss_consistency(f,truth):\n                f = f[is_annotate==1]\n                t = truth[is_annotate==1]\n                loss = F.mse_loss(f,t)\n                return loss\n\n            if is_mixed_precision: \n                image = image.half()\n                with amp.autocast():\n                    logit,f = net(image)\n                    loss0 = binary_cross_entropy_loss(logit, target)\n                    loss1 = loss_consistency(f, f_t)\n\n                scaler.scale(loss0+loss1).backward()\n                scaler.step(optimizer)\n                scaler.update()\n</code></pre>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1219558,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-02-27T00:23:23.673000",
          "content": "<p>since the score on the \"my submissions\" page can be ranked, I can see:</p>\n<p>only one fold:<br>\nfor 640x640 input, tta= origina+flip:<br>\nresnet200d + teacher attention : local cv 0.956805, lb 0.959 (best)<br>\nresnet50d + teacher attention : local cv 0.956385, lb 0.959 <br>\nresnet200d : local cv 0.955946, lb 0.959 (worst)<br>\n(<a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/204950\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/204950</a>)</p>\n<p>it is seen that now resnet50d is as good as resnet200d with extra supervision</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1220422,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-02-28T01:26:45.613000",
          "content": "<p>i will not be surprised that more hand-annotation will be the key to winning this challenge.<br>\nwe will see who can do it \"smartly\" </p>\n<p><img src=\"https://i.ibb.co/pd14X6t/Selection-215.png\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1220430,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-02-28T01:40:04.663000",
          "content": "<p>replacing line annotation with endpoint annotation. how is teacher distillation affected?<br>\nNote this is results for CVC only. the others like ETT are not used at all</p>\n<p><img src=\"https://i.ibb.co/Y3j63HP/Selection-221.png\" alt=\"\"></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1220685,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2021-02-28T08:50:36.873000",
          "content": "<p>be creative!</p>\n<p>multi scale annotation, find pretain model that segment lung,heart, etc …. maybe they would work?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1220873,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2021-02-28T12:39:29.817000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> you tried normfree nets? how does it work for you?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1151143,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2021-01-13T06:53:54.077000",
      "content": "<p>i would like to combine my idea with this google paper:<br>\nConcept Bottleneck Models<br>\n<a href=\"https://arxiv.org/pdf/2007.04612.pdf\" target=\"_blank\">https://arxiv.org/pdf/2007.04612.pdf</a></p>\n<p>\"By construction, we can intervene on these concept bottleneck models by editing their predicted concept values and propagating these changes to the final prediction\"</p>\n<p>i want to intervene by sketching on the image</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1214113,
      "author_name": "Md. Masud Rana",
      "author_url": "",
      "post_date": "2021-02-22T15:43:44.077000",
      "content": "<p>Great discussion thread. Can anyone share an example of the Teacher-Student model example kernel?  </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1133438,
      "author_name": "AIFahim",
      "author_url": "",
      "post_date": "2020-12-31T08:08:51.367000",
      "content": "<p>Sir how you ensemble with annotation without annotations can you explain? And what part of the images you annotated ?  Thanks in advance.  </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1121948,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-22T03:40:55.427000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1121735,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-21T21:34:26.697000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1118630,
      "author_name": "Saurabh Shahane",
      "author_url": "",
      "post_date": "2020-12-19T08:55:24.363000",
      "content": "<p>Thanks for insights!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1118515": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd8e62984343f73764b13346d5d16aa71%2FSelection_089.png?generation=1608359259306689&alt=media)\n\nthe key idea is to modify the activation (activation of the feature map in this case) by changing the input so that the network can pick up extra supervision and decide if they want to use them or not.\n\nThere are many ways to implement this.  if you want to mix those with and without annotations, you can add a flag \"is additional annotation available\" as input at training. this is probably the easiest \"human in the loop for training\"\n\nnote: if the kaggle annotation is not enough, add your own annotation, etc. part of lung, etc\n\nin my experiment, I estimated the addition annotation supervision  signals will cause an improvement of about 2.0% to 2.5% for the AUC metric",
    "1213323": "check these as well\n\nEmbedding Human Knowledge into Deep Neural Network via Attention Map\nhttps://arxiv.org/abs/1905.03540\n\nhttps://github.com/machine-perception-robotics-group/attention_branch_network/\n\n![](https://storage.googleapis.com/groundai-web-prod/media%2Fusers%2Fuser_233037%2Fproject_360896%2Fimages%2Fx1.png)\n",
    "1226219": "reference code and model is here!\nhttps://drive.google.com/drive/folders/1JRqCpV3NMGFeyd5ws6_pyXwQ5x3n8CHF\nsee folder \"2021-03-04\" and refer to the readme.pptx\n\n![](https://i.ibb.co/cvNVfSy/Selection-035.png)",
    "1118777": "if one wants to be fanciful (typo error: should have been x = (1+a)*x\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fd5d7d39813aa6095732c9df13900f9fc%2FSelection_092.png?generation=1608378425203989&alt=media)",
    "1221979": "> I estimated the addition annotation supervision  signals will cause an improvement of about 2.0% to 2.5% for the AUC metric\n\nIs your mean your CV increase from *x* to *x+0.02*?\nThanks for sharing your ideas. These are really very impressive. \n\n\n\n\n",
    "1218383": "i haven't tried this yet, but it is an idea: contrastive attention for self-supervised learning\n\nthe idea is:\n- use all images (annotated and non annotated) for contrastive learning:\n   attention\\_loss (image, augment\\_image) \n- then finetune with annotated images",
    "1136710": "Could you please share some work or paper that implements the presented concepts using TF2.0. Thanks",
    "1118743": "Thank you for sharing your idea! I'm not familiar with this concept. So I would like to ask some questions if you don't mind. What do you use for consistency loss? Do you train the teacher first and then distill the teacher to the student?\n",
    "1239780": "i find a new home for this method. it think it will shine there\nhttps://www.kaggle.com/c/bms-molecular-translation\n\n![](https://i.ibb.co/L1S3cW0/Selection-352.png)",
    "1217634": "some results.\n\ni use two-stage (instead of three-stage in some public code)\n\n![](https://i.ibb.co/nmJBJNx/Selection-191.png)\n![](https://i.ibb.co/4YfY3xG/Selection-193.png)",
    "1151143": "i would like to combine my idea with this google paper:\nConcept Bottleneck Models\nhttps://arxiv.org/pdf/2007.04612.pdf\n\n\"By construction, we can intervene on these concept bottleneck models by editing their predicted concept values and propagating these changes to the final prediction\"\n\ni want to intervene by sketching on the image",
    "1214113": "Great discussion thread. Can anyone share an example of the Teacher-Student model example kernel?  ",
    "1133438": "Sir how you ensemble with annotation without annotations can you explain? And what part of the images you annotated ?  Thanks in advance.  ",
    "1121948": "",
    "1121735": "",
    "1118630": "Thanks for insights!"
  }
}