{
  "id": 205208,
  "title": "Multi-Head Approach",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/205208",
  "author_name": "",
  "post_date": "2020-12-19T02:33:59.379211700Z",
  "votes": 44,
  "comment_count": 15,
  "views": 0,
  "content": "<p>We have 11 targets in this competition and they can be divided into 4 groups: <code>ETT</code>, <code>NGT</code>, <code>CVC</code>,  and <code>Swan</code>.  <br>\nMaybe, different groups have different areas in images to focus on.</p>\n<p>One possible way to leverage this idea is a multi-head approach. Groups share CNN backbone but have independent classifier heads. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F473234%2F5a65c8c830ba257b8b8b461c2cdd423a%2Fsingle-head_approach.png?generation=1608349285787683&amp;alt=media\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F473234%2Fb51948275686bd2ee68464439492e21a%2Fmulti-head_approach.png?generation=1608344387543491&amp;alt=media\"></p>\n<p>I haven't tried it yet, but I'm thinking it might work effectively.</p>",
  "messages": [
    {
      "id": "1118369",
      "postDate": "12/19/2020 02:33:59",
      "content": "<p>We have 11 targets in this competition and they can be divided into 4 groups: <code>ETT</code>, <code>NGT</code>, <code>CVC</code>,  and <code>Swan</code>.  <br>\nMaybe, different groups have different areas in images to focus on.</p>\n<p>One possible way to leverage this idea is a multi-head approach. Groups share CNN backbone but have independent classifier heads. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F473234%2F5a65c8c830ba257b8b8b461c2cdd423a%2Fsingle-head_approach.png?generation=1608349285787683&amp;alt=media\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F473234%2Fb51948275686bd2ee68464439492e21a%2Fmulti-head_approach.png?generation=1608344387543491&amp;alt=media\"></p>\n<p>I haven't tried it yet, but I'm thinking it might work effectively.</p>",
      "rawMarkdown": "We have 11 targets in this competition and they can be divided into 4 groups: `ETT`, `NGT`, `CVC`,  and `Swan`.  \nMaybe, different groups have different areas in images to focus on.\n\nOne possible way to leverage this idea is a multi-head approach. Groups share CNN backbone but have independent classifier heads. \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F473234%2F5a65c8c830ba257b8b8b461c2cdd423a%2Fsingle-head_approach.png?generation=1608349285787683&alt=media\" width=85%>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F473234%2Fb51948275686bd2ee68464439492e21a%2Fmulti-head_approach.png?generation=1608344387543491&alt=media\" width=85%>\n\nI haven't tried it yet, but I'm thinking it might work effectively.",
      "votes": null
    },
    {
      "id": "1118385",
      "postDate": "12/19/2020 02:54:00",
      "content": "<p>With the <strong>Functional API</strong> of Keras i see myself doing something like this, but with PyTorch i not able to \"see\" how to code something like that. So. i'm very curious to see how are you going to do that.</p>",
      "rawMarkdown": "With the **Functional API** of Keras i see myself doing something like this, but with PyTorch i not able to \"see\" how to code something like that. So. i'm very curious to see how are you going to do that.",
      "votes": null
    },
    {
      "id": "1118392",
      "postDate": "12/19/2020 03:07:35",
      "content": "<p>I didn't know there was such a method. I  found the similar method in RSNA 2020 solutions.<br>\n<a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/193505\" target=\"_blank\">https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/193505</a></p>\n<p>I am a newbie of deep learning, so I'm sorry if I misunderstood what the method is. <br>\nAnyway, I will try it as soon as possible!</p>",
      "rawMarkdown": "I didn't know there was such a method. I  found the similar method in RSNA 2020 solutions.\nhttps://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/193505\n\nI am a newbie of deep learning, so I'm sorry if I misunderstood what the method is. \nAnyway, I will try it as soon as possible!",
      "votes": null
    },
    {
      "id": "1118403",
      "postDate": "12/19/2020 03:32:15",
      "content": "<p>I have tried Multi-Head approach on first day but no improvement for old metric, I will try it again for new metric.</p>",
      "rawMarkdown": "I have tried Multi-Head approach on first day but no improvement for old metric, I will try it again for new metric.",
      "votes": null
    },
    {
      "id": "1118432",
      "postDate": "12/19/2020 04:19:12",
      "content": "<p>A simple way is replacing the <code>fc</code> layer of pretrained models (e.g. models of torchvision).</p>\n<p>For example:</p>\n<pre><code>import torch\nfrom torch import nn\nfrom torchvision import models\n\nclass CustomMultiHead(nn.Module):\n\n    def __init__(self, out_dims=[3, 4, 3, 1], in_dim=2048, mid_dim=1024):\n        super(CustomMultiHead, self).__init__()\n\n        self.n_heads = len(out_dims)\n        for i in range(self.n_heads):\n            layer = nn.Sequential(\n                nn.Linear(in_dim, mid_dim),\n                nn.ReLU(inplace=True), nn.Dropout(0.5),\n                nn.Linear(mid_dim, out_dims[i]))\n            setattr(self, f\"head{i}\", layer)\n\n    def forward(self, x):\n        out_list = [getattr(self, f\"head{i}\")(x) for i in range(self.n_heads)]\n        for h in out_list:\n            print(h.shape)\n        out = torch.cat(out_list, 1)\n        return out\n\n\n# # test: replace fc with custom head\nm = models.resnet50(pretrained=True)\ndel m.fc\nm.fc = CustomMultiHead()\n</code></pre>",
      "rawMarkdown": "A simple way is replacing the `fc` layer of pretrained models (e.g. models of torchvision).\n\nFor example:\n\n```\nimport torch\nfrom torch import nn\nfrom torchvision import models\n\nclass CustomMultiHead(nn.Module):\n\n    def __init__(self, out_dims=[3, 4, 3, 1], in_dim=2048, mid_dim=1024):\n        super(CustomMultiHead, self).__init__()\n\n        self.n_heads = len(out_dims)\n        for i in range(self.n_heads):\n            layer = nn.Sequential(\n                nn.Linear(in_dim, mid_dim),\n                nn.ReLU(inplace=True), nn.Dropout(0.5),\n                nn.Linear(mid_dim, out_dims[i]))\n            setattr(self, f\"head{i}\", layer)\n\n    def forward(self, x):\n        out_list = [getattr(self, f\"head{i}\")(x) for i in range(self.n_heads)]\n        for h in out_list:\n            print(h.shape)\n        out = torch.cat(out_list, 1)\n        return out\n\n\n# # test: replace fc with custom head\nm = models.resnet50(pretrained=True)\ndel m.fc\nm.fc = CustomMultiHead()\n```",
      "votes": null
    },
    {
      "id": "1118434",
      "postDate": "12/19/2020 04:21:43",
      "content": "<p>Thank you for information. That is a little surprising to me. I'll try later.</p>",
      "rawMarkdown": "Thank you for information. That is a little surprising to me. I'll try later.",
      "votes": null
    },
    {
      "id": "1118523",
      "postDate": "12/19/2020 06:40:13",
      "content": "<p>it actually will not affect results because </p>\n<ul>\n<li>you are using BCE loss</li>\n<li>if you do the maths, the back gradient propagation is exactly the same in values in both case<br>\n(i.e. you end up learning the same parameters)</li>\n</ul>",
      "rawMarkdown": "it actually will not affect results because \n- you are using BCE loss\n- if you do the maths, the back gradient propagation is exactly the same in values in both case\n(i.e. you end up learning the same parameters)",
      "votes": null
    },
    {
      "id": "1118537",
      "postDate": "12/19/2020 06:57:27",
      "content": "<p>Yes, you are right if we use 1 linear layer as heads.</p>\n<p>I think there is still room for consideration on a model's branch point and type of heads.</p>",
      "rawMarkdown": "Yes, you are right if we use 1 linear layer as heads.\n\nI think there is still room for consideration on a model's branch point and type of heads.",
      "votes": null
    },
    {
      "id": "1118539",
      "postDate": "12/19/2020 07:00:38",
      "content": "<p>\" if we use 1 linear layer as heads.\"<br>\noh, I and sorry that I missed that out. </p>\n<p>you are correct. if the heads are non-linear then the dependency of each group becomes decoupled. <br>\nThis does indeed reduce overfitting in some cases. There could be an improvement. </p>\n<p>I am excited to see the experimental results. </p>\n<p>Thanks!</p>",
      "rawMarkdown": "\" if we use 1 linear layer as heads.\"\noh, I and sorry that I missed that out. \n\nyou are correct. if the heads are non-linear then the dependency of each group becomes decoupled. \nThis does indeed reduce overfitting in some cases. There could be an improvement. \n\nI am excited to see the experimental results. \n\nThanks!",
      "votes": null
    },
    {
      "id": "1118541",
      "postDate": "12/19/2020 07:03:38",
      "content": "<p>I guess heads can be used in another fully connected layer for stacking?</p>",
      "rawMarkdown": "I guess heads can be used in another fully connected layer for stacking?",
      "votes": null
    },
    {
      "id": "1118545",
      "postDate": "12/19/2020 07:09:18",
      "content": "<p>i am thinking of a transformer like attention head. let the transformer to decide connections (i.e. dependency among groups) instead. But I am not sure if we have enough train samples for that</p>",
      "rawMarkdown": "i am thinking of a transformer like attention head. let the transformer to decide connections (i.e. dependency among groups) instead. But I am not sure if we have enough train samples for that",
      "votes": null
    },
    {
      "id": "1118573",
      "postDate": "12/19/2020 07:51:57",
      "content": "<p>I tried it again, result is same, no improvement for me. How about you?</p>",
      "rawMarkdown": "I tried it again, result is same, no improvement for me. How about you?",
      "votes": null
    },
    {
      "id": "1119028",
      "postDate": "12/19/2020 16:39:41",
      "content": "<p><a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a>, in my case, CV was 0.006 higher than that of the original model. Many thanks!</p>",
      "rawMarkdown": "ttahara, in my case, CV was 0.006 higher than that of the original model. Many thanks!",
      "votes": null
    },
    {
      "id": "1119338",
      "postDate": "12/20/2020 00:11:44",
      "content": "<p>That is the good news😃 Thanks for sharing your result.</p>\n<p>BTW, <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> says it did not work for him. I'm interested in the difference between two. </p>",
      "rawMarkdown": "That is the good news😃 Thanks for sharing your result.\n\nBTW, @yasufuminakama says it did not work for him. I'm interested in the difference between two.",
      "votes": null
    },
    {
      "id": "1119351",
      "postDate": "12/20/2020 01:25:12",
      "content": "<p>Unfortunately, in public LB, the score was not improved.<br>\nMy CV might not be precise. However, the blend with the original one got slightly better (0.003 higher than the original in LB).</p>",
      "rawMarkdown": "Unfortunately, in public LB, the score was not improved.\nMy CV might not be precise. However, the blend with the original one got slightly better (0.003 higher than the original in LB).",
      "votes": null
    },
    {
      "id": "1130252",
      "postDate": "12/28/2020 21:53:19",
      "content": "<p><a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> <br>\nI'm sorry to keep you waiting so long. I've published a topic about this approach.<br>\n<a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207230\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207230</a></p>\n<p>So far, my approach has not gotten good score. I will continue to improve it ;)</p>",
      "rawMarkdown": "yasufuminakama \nI'm sorry to keep you waiting so long. I've published a topic about this approach.\nhttps://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207230\n\nSo far, my approach has not gotten good score. I will continue to improve it ;)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1118385,
      "author_name": "hiramcho",
      "author_url": "",
      "post_date": "12/19/2020 02:54:00",
      "content": "<p>With the <strong>Functional API</strong> of Keras i see myself doing something like this, but with PyTorch i not able to \"see\" how to code something like that. So. i'm very curious to see how are you going to do that.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1118432,
          "author_name": "ttahara",
          "author_url": "",
          "post_date": "12/19/2020 04:19:12",
          "content": "<p>A simple way is replacing the <code>fc</code> layer of pretrained models (e.g. models of torchvision).</p>\n<p>For example:</p>\n<pre><code>import torch\nfrom torch import nn\nfrom torchvision import models\n\nclass CustomMultiHead(nn.Module):\n\n    def __init__(self, out_dims=[3, 4, 3, 1], in_dim=2048, mid_dim=1024):\n        super(CustomMultiHead, self).__init__()\n\n        self.n_heads = len(out_dims)\n        for i in range(self.n_heads):\n            layer = nn.Sequential(\n                nn.Linear(in_dim, mid_dim),\n                nn.ReLU(inplace=True), nn.Dropout(0.5),\n                nn.Linear(mid_dim, out_dims[i]))\n            setattr(self, f\"head{i}\", layer)\n\n    def forward(self, x):\n        out_list = [getattr(self, f\"head{i}\")(x) for i in range(self.n_heads)]\n        for h in out_list:\n            print(h.shape)\n        out = torch.cat(out_list, 1)\n        return out\n\n\n# # test: replace fc with custom head\nm = models.resnet50(pretrained=True)\ndel m.fc\nm.fc = CustomMultiHead()\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1118392,
      "author_name": "yosukeyama",
      "author_url": "",
      "post_date": "12/19/2020 03:07:35",
      "content": "<p>I didn't know there was such a method. I  found the similar method in RSNA 2020 solutions.<br>\n<a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/193505\" target=\"_blank\">https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/193505</a></p>\n<p>I am a newbie of deep learning, so I'm sorry if I misunderstood what the method is. <br>\nAnyway, I will try it as soon as possible!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1119028,
          "author_name": "yosukeyama",
          "author_url": "",
          "post_date": "12/19/2020 16:39:41",
          "content": "<p><a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a>, in my case, CV was 0.006 higher than that of the original model. Many thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1119338,
          "author_name": "ttahara",
          "author_url": "",
          "post_date": "12/20/2020 00:11:44",
          "content": "<p>That is the good news😃 Thanks for sharing your result.</p>\n<p>BTW, <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> says it did not work for him. I'm interested in the difference between two. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1119351,
          "author_name": "yosukeyama",
          "author_url": "",
          "post_date": "12/20/2020 01:25:12",
          "content": "<p>Unfortunately, in public LB, the score was not improved.<br>\nMy CV might not be precise. However, the blend with the original one got slightly better (0.003 higher than the original in LB).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1118403,
      "author_name": "yasufuminakama",
      "author_url": "",
      "post_date": "12/19/2020 03:32:15",
      "content": "<p>I have tried Multi-Head approach on first day but no improvement for old metric, I will try it again for new metric.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1118434,
          "author_name": "ttahara",
          "author_url": "",
          "post_date": "12/19/2020 04:21:43",
          "content": "<p>Thank you for information. That is a little surprising to me. I'll try later.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1118573,
          "author_name": "yasufuminakama",
          "author_url": "",
          "post_date": "12/19/2020 07:51:57",
          "content": "<p>I tried it again, result is same, no improvement for me. How about you?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1130252,
          "author_name": "ttahara",
          "author_url": "",
          "post_date": "12/28/2020 21:53:19",
          "content": "<p><a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> <br>\nI'm sorry to keep you waiting so long. I've published a topic about this approach.<br>\n<a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207230\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207230</a></p>\n<p>So far, my approach has not gotten good score. I will continue to improve it ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1118523,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "12/19/2020 06:40:13",
      "content": "<p>it actually will not affect results because </p>\n<ul>\n<li>you are using BCE loss</li>\n<li>if you do the maths, the back gradient propagation is exactly the same in values in both case<br>\n(i.e. you end up learning the same parameters)</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1118537,
          "author_name": "ttahara",
          "author_url": "",
          "post_date": "12/19/2020 06:57:27",
          "content": "<p>Yes, you are right if we use 1 linear layer as heads.</p>\n<p>I think there is still room for consideration on a model's branch point and type of heads.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1118539,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "12/19/2020 07:00:38",
          "content": "<p>\" if we use 1 linear layer as heads.\"<br>\noh, I and sorry that I missed that out. </p>\n<p>you are correct. if the heads are non-linear then the dependency of each group becomes decoupled. <br>\nThis does indeed reduce overfitting in some cases. There could be an improvement. </p>\n<p>I am excited to see the experimental results. </p>\n<p>Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1118541,
          "author_name": "gunesevitan",
          "author_url": "",
          "post_date": "12/19/2020 07:03:38",
          "content": "<p>I guess heads can be used in another fully connected layer for stacking?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1118545,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "12/19/2020 07:09:18",
          "content": "<p>i am thinking of a transformer like attention head. let the transformer to decide connections (i.e. dependency among groups) instead. But I am not sure if we have enough train samples for that</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1118369": "We have 11 targets in this competition and they can be divided into 4 groups: `ETT`, `NGT`, `CVC`,  and `Swan`.  \nMaybe, different groups have different areas in images to focus on.\n\nOne possible way to leverage this idea is a multi-head approach. Groups share CNN backbone but have independent classifier heads. \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F473234%2F5a65c8c830ba257b8b8b461c2cdd423a%2Fsingle-head_approach.png?generation=1608349285787683&alt=media\" width=85%>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F473234%2Fb51948275686bd2ee68464439492e21a%2Fmulti-head_approach.png?generation=1608344387543491&alt=media\" width=85%>\n\nI haven't tried it yet, but I'm thinking it might work effectively.",
    "1118385": "With the **Functional API** of Keras i see myself doing something like this, but with PyTorch i not able to \"see\" how to code something like that. So. i'm very curious to see how are you going to do that.",
    "1118392": "I didn't know there was such a method. I  found the similar method in RSNA 2020 solutions.\nhttps://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/193505\n\nI am a newbie of deep learning, so I'm sorry if I misunderstood what the method is. \nAnyway, I will try it as soon as possible!",
    "1118403": "I have tried Multi-Head approach on first day but no improvement for old metric, I will try it again for new metric.",
    "1118432": "A simple way is replacing the `fc` layer of pretrained models (e.g. models of torchvision).\n\nFor example:\n\n```\nimport torch\nfrom torch import nn\nfrom torchvision import models\n\nclass CustomMultiHead(nn.Module):\n\n    def __init__(self, out_dims=[3, 4, 3, 1], in_dim=2048, mid_dim=1024):\n        super(CustomMultiHead, self).__init__()\n\n        self.n_heads = len(out_dims)\n        for i in range(self.n_heads):\n            layer = nn.Sequential(\n                nn.Linear(in_dim, mid_dim),\n                nn.ReLU(inplace=True), nn.Dropout(0.5),\n                nn.Linear(mid_dim, out_dims[i]))\n            setattr(self, f\"head{i}\", layer)\n\n    def forward(self, x):\n        out_list = [getattr(self, f\"head{i}\")(x) for i in range(self.n_heads)]\n        for h in out_list:\n            print(h.shape)\n        out = torch.cat(out_list, 1)\n        return out\n\n\n# # test: replace fc with custom head\nm = models.resnet50(pretrained=True)\ndel m.fc\nm.fc = CustomMultiHead()\n```",
    "1118434": "Thank you for information. That is a little surprising to me. I'll try later.",
    "1118523": "it actually will not affect results because \n- you are using BCE loss\n- if you do the maths, the back gradient propagation is exactly the same in values in both case\n(i.e. you end up learning the same parameters)",
    "1118537": "Yes, you are right if we use 1 linear layer as heads.\n\nI think there is still room for consideration on a model's branch point and type of heads.",
    "1118539": "\" if we use 1 linear layer as heads.\"\noh, I and sorry that I missed that out. \n\nyou are correct. if the heads are non-linear then the dependency of each group becomes decoupled. \nThis does indeed reduce overfitting in some cases. There could be an improvement. \n\nI am excited to see the experimental results. \n\nThanks!",
    "1118541": "I guess heads can be used in another fully connected layer for stacking?",
    "1118545": "i am thinking of a transformer like attention head. let the transformer to decide connections (i.e. dependency among groups) instead. But I am not sure if we have enough train samples for that",
    "1118573": "I tried it again, result is same, no improvement for me. How about you?",
    "1119028": "ttahara, in my case, CV was 0.006 higher than that of the original model. Many thanks!",
    "1119338": "That is the good news😃 Thanks for sharing your result.\n\nBTW, @yasufuminakama says it did not work for him. I'm interested in the difference between two.",
    "1119351": "Unfortunately, in public LB, the score was not improved.\nMy CV might not be precise. However, the blend with the original one got slightly better (0.003 higher than the original in LB).",
    "1130252": "yasufuminakama \nI'm sorry to keep you waiting so long. I've published a topic about this approach.\nhttps://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/207230\n\nSo far, my approach has not gotten good score. I will continue to improve it ;)"
  },
  "source": "meta"
}