{
  "id": 123432,
  "title": "Some experiments with CNN tails ",
  "url": "/competitions/bengaliai-cv19/discussion/123432",
  "author_name": "DrHB",
  "post_date": "2019-12-27T15:45:10.088000",
  "votes": 70,
  "comment_count": 26,
  "views": 0,
  "content": "<p>Just sharing some of quick tests that I have done over last few days. </p>\n\n<p>It seems like there is few different way to design custom tail (Last Layers of your CNN). Below you will find which one worked the best for me:</p>\n\n<p>My experimental setup:</p>\n\n<p>```\nmodel: resnet34(pretrained=True)\nimg: 3 channel \nimg_sz: 128\nsplit: random (80/20) (consistent across all experiments)\noptim: Adam\nepoch: 10\nsched: Cosine decay</p>\n\n<p>```</p>\n\n<p><strong>METHOD 1:</strong>\nI assume very simple method is to cut pertained network after <code>AdaptiveAvgPool2d</code> and add 3 simple <code>nn.Linear</code> please see image below (scissors and dotted line indicate the place of cut) \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2Fba551d5f815a002fbbf536c8792c964d%2FScreen%20Shot%202019-12-27%20at%2010.21.54%20AM.png?generation=1577460762607005&amp;alt=media\" alt=\"\"></p>\n\n<p>After training this network I got: \n<code>\nCV - 0.9650 \nLB - 0.9610\n</code></p>\n\n<p><strong>METHOD 2:</strong>\nHere we cut  above <code>AdaptiveAvgPool2d</code> and add for each of 3 <code>nn.Linear</code> there own <code>AdaptiveAvgPool2d</code> </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F29cfddbd570990893d21175c6f1ef6a6%2FScreen%20Shot%202019-12-28%20at%2012.06.09%20AM.png?generation=1577509638632635&amp;alt=media\" alt=\"\"></p>\n\n<p>Scores:\n<code>\nCV - 0.9645 \nLB - 0.9625\n</code></p>\n\n<p><strong>METHOD 3:</strong>\nHere we cut at end of <code>layer4</code> and add a small block inspired by @Iafoss kernels: \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2Fbf7038f0ccac2171d35856c3ac3572e7%2FScreen%20Shot%202019-12-28%20at%2012.01.34%20AM.png?generation=1577509380771301&amp;alt=media\" alt=\"\"></p>\n\n<p>Scores:\n<code>\nCV - 0.9652 \nLB - 0.9616\n</code></p>\n\n<p>It seems likes <strong>METHOD 2</strong> at least in my hands work the best. It shows a smaller CV and LB gap and yields better result. </p>\n\n<p>I can imagine there are bunch of other ways one can design custom tails. I will keep experimenting and posting my results. Hope this will inspire some creativity and people will come up with better designs =) </p>",
  "messages": [
    {
      "id": 704545,
      "postDate": "2019-12-27T15:45:10.090Z",
      "content": "<p>Just sharing some of quick tests that I have done over last few days. </p>\n\n<p>It seems like there is few different way to design custom tail (Last Layers of your CNN). Below you will find which one worked the best for me:</p>\n\n<p>My experimental setup:</p>\n\n<p>```\nmodel: resnet34(pretrained=True)\nimg: 3 channel \nimg_sz: 128\nsplit: random (80/20) (consistent across all experiments)\noptim: Adam\nepoch: 10\nsched: Cosine decay</p>\n\n<p>```</p>\n\n<p><strong>METHOD 1:</strong>\nI assume very simple method is to cut pertained network after <code>AdaptiveAvgPool2d</code> and add 3 simple <code>nn.Linear</code> please see image below (scissors and dotted line indicate the place of cut) \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2Fba551d5f815a002fbbf536c8792c964d%2FScreen%20Shot%202019-12-27%20at%2010.21.54%20AM.png?generation=1577460762607005&amp;alt=media\" alt=\"\"></p>\n\n<p>After training this network I got: \n<code>\nCV - 0.9650 \nLB - 0.9610\n</code></p>\n\n<p><strong>METHOD 2:</strong>\nHere we cut  above <code>AdaptiveAvgPool2d</code> and add for each of 3 <code>nn.Linear</code> there own <code>AdaptiveAvgPool2d</code> </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F29cfddbd570990893d21175c6f1ef6a6%2FScreen%20Shot%202019-12-28%20at%2012.06.09%20AM.png?generation=1577509638632635&amp;alt=media\" alt=\"\"></p>\n\n<p>Scores:\n<code>\nCV - 0.9645 \nLB - 0.9625\n</code></p>\n\n<p><strong>METHOD 3:</strong>\nHere we cut at end of <code>layer4</code> and add a small block inspired by @Iafoss kernels: \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2Fbf7038f0ccac2171d35856c3ac3572e7%2FScreen%20Shot%202019-12-28%20at%2012.01.34%20AM.png?generation=1577509380771301&amp;alt=media\" alt=\"\"></p>\n\n<p>Scores:\n<code>\nCV - 0.9652 \nLB - 0.9616\n</code></p>\n\n<p>It seems likes <strong>METHOD 2</strong> at least in my hands work the best. It shows a smaller CV and LB gap and yields better result. </p>\n\n<p>I can imagine there are bunch of other ways one can design custom tails. I will keep experimenting and posting my results. Hope this will inspire some creativity and people will come up with better designs =) </p>",
      "rawMarkdown": "Just sharing some of quick tests that I have done over last few days. \n\nIt seems like there is few different way to design custom tail (Last Layers of your CNN). Below you will find which one worked the best for me:\n\n\n\nMy experimental setup:\n\n```\nmodel: resnet34(pretrained=True)\nimg: 3 channel \nimg_sz: 128\nsplit: random (80/20) (consistent across all experiments)\noptim: Adam\nepoch: 10\nsched: Cosine decay\n\n```\n\n**METHOD 1:**\nI assume very simple method is to cut pertained network after `AdaptiveAvgPool2d` and add 3 simple `nn.Linear` please see image below (scissors and dotted line indicate the place of cut) \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2Fba551d5f815a002fbbf536c8792c964d%2FScreen%20Shot%202019-12-27%20at%2010.21.54%20AM.png?generation=1577460762607005&amp;alt=media)\n\nAfter training this network I got: \n```\nCV - 0.9650 \nLB - 0.9610\n```\n\n**METHOD 2:**\nHere we cut  above `AdaptiveAvgPool2d` and add for each of 3 `nn.Linear` there own `AdaptiveAvgPool2d` \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F29cfddbd570990893d21175c6f1ef6a6%2FScreen%20Shot%202019-12-28%20at%2012.06.09%20AM.png?generation=1577509638632635&amp;alt=media)\n\n\nScores:\n```\nCV - 0.9645 \nLB - 0.9625\n```\n\n**METHOD 3:**\nHere we cut at end of `layer4` and add a small block inspired by @Iafoss kernels: \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2Fbf7038f0ccac2171d35856c3ac3572e7%2FScreen%20Shot%202019-12-28%20at%2012.01.34%20AM.png?generation=1577509380771301&amp;alt=media)\n\n\nScores:\n```\nCV - 0.9652 \nLB - 0.9616\n```\n\nIt seems likes **METHOD 2** at least in my hands work the best. It shows a smaller CV and LB gap and yields better result. \n\nI can imagine there are bunch of other ways one can design custom tails. I will keep experimenting and posting my results. Hope this will inspire some creativity and people will come up with better designs =) ",
      "votes": 70
    },
    {
      "id": 708872,
      "postDate": "2020-01-02T19:39:28.817Z",
      "content": "<p>Update:\nif you use in <strong>METHOD 2</strong>  <code>GeM</code> pooling layer you get following score. \n<code>\nCV - 0.9712\nLB - 0.9654\n</code>\nWhat is <code>GeM</code>?\nits a pooling layer, I took it from the 1st place solution in APTOS challenge (<a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/108065\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/108065</a>)</p>\n\n<p>code below:</p>\n\n<p><code>\nfrom torch.nn.parameter import Parameter\ndef gem(x, p=3, eps=1e-6):\n    return F.avg_pool2d(x.clamp(min=eps).pow(p), (x.size(-2), x.size(-1))).pow(1./p)\nclass GeM(nn.Module):\n    def __init__(self, p=3, eps=1e-6):\n        super(GeM,self).__init__()\n        self.p = Parameter(torch.ones(1)*p)\n        self.eps = eps\n    def forward(self, x):\n        return gem(x, p=self.p, eps=self.eps) <br>\n    def __repr__(self):\n        return self.__class__.__name__ + '(' + 'p=' + '{:.4f}'.format(self.p.data.tolist()[0]) + ', ' + 'eps=' + str(self.eps) + ')'\n</code></p>",
      "rawMarkdown": "Update:\nif you use in **METHOD 2**  `GeM` pooling layer you get following score. \n```\nCV - 0.9712\nLB - 0.9654\n```\nWhat is `GeM`?\nits a pooling layer, I took it from the 1st place solution in APTOS challenge (https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/108065)\n\ncode below:\n\n```\nfrom torch.nn.parameter import Parameter\ndef gem(x, p=3, eps=1e-6):\n    return F.avg_pool2d(x.clamp(min=eps).pow(p), (x.size(-2), x.size(-1))).pow(1./p)\nclass GeM(nn.Module):\n    def __init__(self, p=3, eps=1e-6):\n        super(GeM,self).__init__()\n        self.p = Parameter(torch.ones(1)*p)\n        self.eps = eps\n    def forward(self, x):\n        return gem(x, p=self.p, eps=self.eps)       \n    def __repr__(self):\n        return self.__class__.__name__ + '(' + 'p=' + '{:.4f}'.format(self.p.data.tolist()[0]) + ', ' + 'eps=' + str(self.eps) + ')'\n```\n\n",
      "votes": 26,
      "replies": [
        {
          "id": 731046,
          "postDate": "2020-01-28T09:42:09.633Z",
          "content": "<p>Hi all, \nThanks <a href=\"/drhabib\">@drhabib</a> for sharing lot of good information ..</p>\n\n<p>I tried the GeM in place of adaptive avgpool layer. It starts to throw nan as loss within few batches of training in the first epoch itself.</p>\n\n<p>I am using the pipeline similar to <a href=\"/iafoss\">@iafoss</a> </p>\n\n<p>Is there any specific reason. Or is there a way to fix this</p>",
          "rawMarkdown": "Hi all, \nThanks @drhabib for sharing lot of good information ..\n\nI tried the GeM in place of adaptive avgpool layer. It starts to throw nan as loss within few batches of training in the first epoch itself.\n\nI am using the pipeline similar to @iafoss \n\nIs there any specific reason. Or is there a way to fix this",
          "votes": 1
        },
        {
          "id": 731054,
          "postDate": "2020-01-28T09:50:19.643Z",
          "content": "<p>Are you training your model on half precision? In that case, try single precision.</p>",
          "rawMarkdown": "Are you training your model on half precision? In that case, try single precision."
        },
        {
          "id": 731865,
          "postDate": "2020-01-29T06:47:53.477Z",
          "content": "<p>Thanks for your suggestion <a href=\"/atikur\">@atikur</a> , I will try and update you. Is there some kind of fix , that we can come up to make it work on half precision.</p>",
          "rawMarkdown": "Thanks for your suggestion @atikur , I will try and update you. Is there some kind of fix , that we can come up to make it work on half precision."
        },
        {
          "id": 751431,
          "postDate": "2020-02-20T07:36:37.117Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 765132,
          "postDate": "2020-03-06T09:33:01.247Z",
          "content": "<p><a href=\"/dhakshiin1601\">@dhakshiin1601</a> Sorry to disturb you. Have you solve the nan problem? I'm at the same suitation and use the pipline of iafoss too..</p>",
          "rawMarkdown": "@dhakshiin1601 Sorry to disturb you. Have you solve the nan problem? I'm at the same suitation and use the pipline of iafoss too.."
        },
        {
          "id": 765174,
          "postDate": "2020-03-06T10:21:21.640Z",
          "content": "<p>Please check out my Discussion thread named EfficientNet trails. It has the required code for it. All the best </p>",
          "rawMarkdown": "Please check out my Discussion thread named EfficientNet trails. It has the required code for it. All the best \n\n",
          "votes": 1
        },
        {
          "id": 765683,
          "postDate": "2020-03-07T00:43:38.890Z",
          "content": "<p>Thanks a lot !</p>",
          "rawMarkdown": "Thanks a lot !"
        }
      ]
    },
    {
      "id": 704562,
      "postDate": "2019-12-27T16:10:04.027Z",
      "content": "<p>Hi,  I have one question: From method 2,  Do the 3 single AdaptiveAvgPool2d layer use the same feature from backbone? if yes, maybe that's no different with method 1.... but maybe I misunderstood..</p>",
      "rawMarkdown": "Hi,  I have one question: From method 2,  Do the 3 single AdaptiveAvgPool2d layer use the same feature from backbone? if yes, maybe that's no different with method 1.... but maybe I misunderstood..\n\n\n\n",
      "votes": 3,
      "replies": [
        {
          "id": 704564,
          "postDate": "2019-12-27T16:13:08.470Z",
          "content": "<p>In <strong>method 1</strong> you have 1 <code>AdaptiveAvgPool2d</code> and 3 separate <code>nn.Linear</code>. In <strong>method 2</strong>  for each class you have its own  <code>AdaptiveAvgPool2d</code>  (3 in total) which are connected to there own <code>nn.Linear.</code> So I hoped that in method 2  each AdaptiveAvgPool2d will extract there own features corresponding to the class </p>",
          "rawMarkdown": "In **method 1** you have 1 `AdaptiveAvgPool2d` and 3 separate `nn.Linear`. In **method 2**  for each class you have its own  `AdaptiveAvgPool2d`  (3 in total) which are connected to there own `nn.Linear.` So I hoped that in method 2  each AdaptiveAvgPool2d will extract there own features corresponding to the class ",
          "votes": 1
        },
        {
          "id": 704566,
          "postDate": "2019-12-27T16:16:55.513Z",
          "content": "<p>oh, I got that. maybe you set the different parameters for the 3 separate AdaptiveAvgPool2d, did I right?</p>",
          "rawMarkdown": "oh, I got that. maybe you set the different parameters for the 3 separate AdaptiveAvgPool2d, did I right?",
          "votes": 1
        },
        {
          "id": 704567,
          "postDate": "2019-12-27T16:17:46.747Z",
          "content": "<p>yes! =) </p>",
          "rawMarkdown": "yes! =) "
        },
        {
          "id": 704829,
          "postDate": "2019-12-28T03:51:28.670Z",
          "content": "<p>AdaptiveAvgPool2d have no parameter in it. So three separate AveragePools act in a same way. Method 1 and 2 seem to be the same.</p>",
          "rawMarkdown": "AdaptiveAvgPool2d have no parameter in it. So three separate AveragePools act in a same way. Method 1 and 2 seem to be the same.",
          "votes": 8
        },
        {
          "id": 704835,
          "postDate": "2019-12-28T04:06:52.120Z",
          "content": "<p>Ok I see the confusion now with my schemes.  I have BasicBlocks(with conv and activ) on top of each separate AdaptivePool2D. I will update  the image. Thanks for pointing out.</p>",
          "rawMarkdown": "Ok I see the confusion now with my schemes.  I have BasicBlocks(with conv and activ) on top of each separate AdaptivePool2D. I will update  the image. Thanks for pointing out."
        }
      ]
    },
    {
      "id": 740335,
      "postDate": "2020-02-09T09:26:25.927Z",
      "content": "<p>Hi.  For second method, May i ask you why you use activation-&gt;conv-&gt;bn?\nI found the last operation of resnet34 layer4 is relu, so thats to say your head is activation-&gt;activation-&gt;conv-bn.\nIs there any reason for this? I'm confused......Thank you..</p>",
      "rawMarkdown": "Hi.  For second method, May i ask you why you use activation-&gt;conv-&gt;bn?\nI found the last operation of resnet34 layer4 is relu, so thats to say your head is activation-&gt;activation-&gt;conv-bn.\nIs there any reason for this? I'm confused......Thank you..",
      "votes": 1
    },
    {
      "id": 717204,
      "postDate": "2020-01-12T21:24:55.940Z",
      "content": "<p>Why do you use 3 channels when actual image have just one?\nCannot catch your idea...</p>",
      "rawMarkdown": "Why do you use 3 channels when actual image have just one?\nCannot catch your idea...",
      "votes": 1,
      "replies": [
        {
          "id": 721640,
          "postDate": "2020-01-17T14:59:26.440Z",
          "content": "<p>You can convert 1 channel image to 3 channel by cloning. People in this competition has success using both approaches with 1 or 3 channel.</p>",
          "rawMarkdown": "You can convert 1 channel image to 3 channel by cloning. People in this competition has success using both approaches with 1 or 3 channel.",
          "votes": 1
        },
        {
          "id": 721830,
          "postDate": "2020-01-17T18:52:32.837Z",
          "content": "<p>I guess it's good approach, thank you for suggestion.</p>",
          "rawMarkdown": "I guess it's good approach, thank you for suggestion.",
          "votes": 1
        }
      ]
    },
    {
      "id": 704743,
      "postDate": "2019-12-28T00:16:49.593Z",
      "content": "<p>Did you try training three separate models for each target?</p>",
      "rawMarkdown": "Did you try training three separate models for each target?",
      "votes": 1,
      "replies": [
        {
          "id": 704745,
          "postDate": "2019-12-28T00:24:25.840Z",
          "content": "<p>No I haven’t</p>",
          "rawMarkdown": "No I haven’t"
        },
        {
          "id": 704746,
          "postDate": "2019-12-28T00:25:27.180Z",
          "content": "<p>I am currently working on it but my gpu qouta ran out. I wish we could have gcp credit coupons.</p>",
          "rawMarkdown": "I am currently working on it but my gpu qouta ran out. I wish we could have gcp credit coupons.",
          "votes": 1
        }
      ]
    },
    {
      "id": 704548,
      "postDate": "2019-12-27T15:49:33.930Z",
      "content": "<p>The CV of method 2 is 0.945 or is it a typo?</p>",
      "rawMarkdown": "The CV of method 2 is 0.945 or is it a typo?",
      "votes": 1,
      "replies": [
        {
          "id": 704551,
          "postDate": "2019-12-27T15:52:58.437Z",
          "content": "<p>Thank you! Corrected =) </p>",
          "rawMarkdown": "Thank you! Corrected =) ",
          "votes": 1
        }
      ]
    },
    {
      "id": 705696,
      "postDate": "2019-12-29T10:06:56.937Z",
      "content": "<p>What is <code>nn.Mish</code> mentioned here? Did you mean this activation function? <a href=\"https://arxiv.org/abs/1908.08681\">https://arxiv.org/abs/1908.08681</a></p>",
      "rawMarkdown": "What is `nn.Mish` mentioned here? Did you mean this activation function? https://arxiv.org/abs/1908.08681",
      "votes": 2,
      "replies": [
        {
          "id": 705809,
          "postDate": "2019-12-29T14:05:55.997Z",
          "content": "<p>yes:)</p>",
          "rawMarkdown": "yes:)",
          "votes": 1
        }
      ]
    },
    {
      "id": 763019,
      "postDate": "2020-03-04T03:10:06.680Z",
      "content": "<p>A good share ! <br>\nIntead of your method2, I just add 3 linear layers after the feature map and get a worse result... <br>\nCan it be explained to the lack of conv layer of each part? The linear layers(only linear not conv) can't learn the details among each part while using a same feature map?</p>",
      "rawMarkdown": "A good share !  \nIntead of your method2, I just add 3 linear layers after the feature map and get a worse result...  \nCan it be explained to the lack of conv layer of each part? The linear layers(only linear not conv) can't learn the details among each part while using a same feature map?"
    }
  ],
  "comments": [
    {
      "id": 708872,
      "author_name": "DrHB",
      "author_url": "",
      "post_date": "2020-01-02T19:39:28.817000",
      "content": "<p>Update:\nif you use in <strong>METHOD 2</strong>  <code>GeM</code> pooling layer you get following score. \n<code>\nCV - 0.9712\nLB - 0.9654\n</code>\nWhat is <code>GeM</code>?\nits a pooling layer, I took it from the 1st place solution in APTOS challenge (<a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/108065\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/108065</a>)</p>\n\n<p>code below:</p>\n\n<p><code>\nfrom torch.nn.parameter import Parameter\ndef gem(x, p=3, eps=1e-6):\n    return F.avg_pool2d(x.clamp(min=eps).pow(p), (x.size(-2), x.size(-1))).pow(1./p)\nclass GeM(nn.Module):\n    def __init__(self, p=3, eps=1e-6):\n        super(GeM,self).__init__()\n        self.p = Parameter(torch.ones(1)*p)\n        self.eps = eps\n    def forward(self, x):\n        return gem(x, p=self.p, eps=self.eps) <br>\n    def __repr__(self):\n        return self.__class__.__name__ + '(' + 'p=' + '{:.4f}'.format(self.p.data.tolist()[0]) + ', ' + 'eps=' + str(self.eps) + ')'\n</code></p>",
      "votes": 26,
      "replies": [
        {
          "id": 731046,
          "author_name": "Balaji Selvaraj",
          "author_url": "",
          "post_date": "2020-01-28T09:42:09.633000",
          "content": "<p>Hi all, \nThanks <a href=\"/drhabib\">@drhabib</a> for sharing lot of good information ..</p>\n\n<p>I tried the GeM in place of adaptive avgpool layer. It starts to throw nan as loss within few batches of training in the first epoch itself.</p>\n\n<p>I am using the pipeline similar to <a href=\"/iafoss\">@iafoss</a> </p>\n\n<p>Is there any specific reason. Or is there a way to fix this</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 731054,
          "author_name": "Atikur Rahman",
          "author_url": "",
          "post_date": "2020-01-28T09:50:19.643000",
          "content": "<p>Are you training your model on half precision? In that case, try single precision.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 731865,
          "author_name": "Balaji Selvaraj",
          "author_url": "",
          "post_date": "2020-01-29T06:47:53.477000",
          "content": "<p>Thanks for your suggestion <a href=\"/atikur\">@atikur</a> , I will try and update you. Is there some kind of fix , that we can come up to make it work on half precision.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 751431,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-20T07:36:37.117000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 765132,
          "author_name": "Shiyuan Zeng",
          "author_url": "",
          "post_date": "2020-03-06T09:33:01.247000",
          "content": "<p><a href=\"/dhakshiin1601\">@dhakshiin1601</a> Sorry to disturb you. Have you solve the nan problem? I'm at the same suitation and use the pipline of iafoss too..</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 765174,
          "author_name": "Balaji Selvaraj",
          "author_url": "",
          "post_date": "2020-03-06T10:21:21.640000",
          "content": "<p>Please check out my Discussion thread named EfficientNet trails. It has the required code for it. All the best </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 765683,
          "author_name": "Shiyuan Zeng",
          "author_url": "",
          "post_date": "2020-03-07T00:43:38.890000",
          "content": "<p>Thanks a lot !</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 704562,
      "author_name": "Gary",
      "author_url": "",
      "post_date": "2019-12-27T16:10:04.027000",
      "content": "<p>Hi,  I have one question: From method 2,  Do the 3 single AdaptiveAvgPool2d layer use the same feature from backbone? if yes, maybe that's no different with method 1.... but maybe I misunderstood..</p>",
      "votes": 3,
      "replies": [
        {
          "id": 704564,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-12-27T16:13:08.470000",
          "content": "<p>In <strong>method 1</strong> you have 1 <code>AdaptiveAvgPool2d</code> and 3 separate <code>nn.Linear</code>. In <strong>method 2</strong>  for each class you have its own  <code>AdaptiveAvgPool2d</code>  (3 in total) which are connected to there own <code>nn.Linear.</code> So I hoped that in method 2  each AdaptiveAvgPool2d will extract there own features corresponding to the class </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 704566,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2019-12-27T16:16:55.513000",
          "content": "<p>oh, I got that. maybe you set the different parameters for the 3 separate AdaptiveAvgPool2d, did I right?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 704567,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-12-27T16:17:46.747000",
          "content": "<p>yes! =) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 704829,
          "author_name": "Ildoo Kim",
          "author_url": "",
          "post_date": "2019-12-28T03:51:28.670000",
          "content": "<p>AdaptiveAvgPool2d have no parameter in it. So three separate AveragePools act in a same way. Method 1 and 2 seem to be the same.</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 704835,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-12-28T04:06:52.120000",
          "content": "<p>Ok I see the confusion now with my schemes.  I have BasicBlocks(with conv and activ) on top of each separate AdaptivePool2D. I will update  the image. Thanks for pointing out.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 740335,
      "author_name": "shiba",
      "author_url": "",
      "post_date": "2020-02-09T09:26:25.927000",
      "content": "<p>Hi.  For second method, May i ask you why you use activation-&gt;conv-&gt;bn?\nI found the last operation of resnet34 layer4 is relu, so thats to say your head is activation-&gt;activation-&gt;conv-bn.\nIs there any reason for this? I'm confused......Thank you..</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 717204,
      "author_name": "Dmitry Kravchuk",
      "author_url": "",
      "post_date": "2020-01-12T21:24:55.940000",
      "content": "<p>Why do you use 3 channels when actual image have just one?\nCannot catch your idea...</p>",
      "votes": 1,
      "replies": [
        {
          "id": 721640,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-01-17T14:59:26.440000",
          "content": "<p>You can convert 1 channel image to 3 channel by cloning. People in this competition has success using both approaches with 1 or 3 channel.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 721830,
          "author_name": "Dmitry Kravchuk",
          "author_url": "",
          "post_date": "2020-01-17T18:52:32.837000",
          "content": "<p>I guess it's good approach, thank you for suggestion.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 704743,
      "author_name": "Shangqiu Li",
      "author_url": "",
      "post_date": "2019-12-28T00:16:49.593000",
      "content": "<p>Did you try training three separate models for each target?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 704745,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-12-28T00:24:25.840000",
          "content": "<p>No I haven’t</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 704746,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2019-12-28T00:25:27.180000",
          "content": "<p>I am currently working on it but my gpu qouta ran out. I wish we could have gcp credit coupons.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 704548,
      "author_name": "Peter",
      "author_url": "",
      "post_date": "2019-12-27T15:49:33.930000",
      "content": "<p>The CV of method 2 is 0.945 or is it a typo?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 704551,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-12-27T15:52:58.437000",
          "content": "<p>Thank you! Corrected =) </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 705696,
      "author_name": "Siarhei Fedartsou",
      "author_url": "",
      "post_date": "2019-12-29T10:06:56.937000",
      "content": "<p>What is <code>nn.Mish</code> mentioned here? Did you mean this activation function? <a href=\"https://arxiv.org/abs/1908.08681\">https://arxiv.org/abs/1908.08681</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 705809,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-12-29T14:05:55.997000",
          "content": "<p>yes:)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 763019,
      "author_name": "Shiyuan Zeng",
      "author_url": "",
      "post_date": "2020-03-04T03:10:06.680000",
      "content": "<p>A good share ! <br>\nIntead of your method2, I just add 3 linear layers after the feature map and get a worse result... <br>\nCan it be explained to the lack of conv layer of each part? The linear layers(only linear not conv) can't learn the details among each part while using a same feature map?</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "704545": "Just sharing some of quick tests that I have done over last few days. \n\nIt seems like there is few different way to design custom tail (Last Layers of your CNN). Below you will find which one worked the best for me:\n\n\n\nMy experimental setup:\n\n```\nmodel: resnet34(pretrained=True)\nimg: 3 channel \nimg_sz: 128\nsplit: random (80/20) (consistent across all experiments)\noptim: Adam\nepoch: 10\nsched: Cosine decay\n\n```\n\n**METHOD 1:**\nI assume very simple method is to cut pertained network after `AdaptiveAvgPool2d` and add 3 simple `nn.Linear` please see image below (scissors and dotted line indicate the place of cut) \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2Fba551d5f815a002fbbf536c8792c964d%2FScreen%20Shot%202019-12-27%20at%2010.21.54%20AM.png?generation=1577460762607005&amp;alt=media)\n\nAfter training this network I got: \n```\nCV - 0.9650 \nLB - 0.9610\n```\n\n**METHOD 2:**\nHere we cut  above `AdaptiveAvgPool2d` and add for each of 3 `nn.Linear` there own `AdaptiveAvgPool2d` \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F29cfddbd570990893d21175c6f1ef6a6%2FScreen%20Shot%202019-12-28%20at%2012.06.09%20AM.png?generation=1577509638632635&amp;alt=media)\n\n\nScores:\n```\nCV - 0.9645 \nLB - 0.9625\n```\n\n**METHOD 3:**\nHere we cut at end of `layer4` and add a small block inspired by @Iafoss kernels: \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2Fbf7038f0ccac2171d35856c3ac3572e7%2FScreen%20Shot%202019-12-28%20at%2012.01.34%20AM.png?generation=1577509380771301&amp;alt=media)\n\n\nScores:\n```\nCV - 0.9652 \nLB - 0.9616\n```\n\nIt seems likes **METHOD 2** at least in my hands work the best. It shows a smaller CV and LB gap and yields better result. \n\nI can imagine there are bunch of other ways one can design custom tails. I will keep experimenting and posting my results. Hope this will inspire some creativity and people will come up with better designs =) ",
    "708872": "Update:\nif you use in **METHOD 2**  `GeM` pooling layer you get following score. \n```\nCV - 0.9712\nLB - 0.9654\n```\nWhat is `GeM`?\nits a pooling layer, I took it from the 1st place solution in APTOS challenge (https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/108065)\n\ncode below:\n\n```\nfrom torch.nn.parameter import Parameter\ndef gem(x, p=3, eps=1e-6):\n    return F.avg_pool2d(x.clamp(min=eps).pow(p), (x.size(-2), x.size(-1))).pow(1./p)\nclass GeM(nn.Module):\n    def __init__(self, p=3, eps=1e-6):\n        super(GeM,self).__init__()\n        self.p = Parameter(torch.ones(1)*p)\n        self.eps = eps\n    def forward(self, x):\n        return gem(x, p=self.p, eps=self.eps)       \n    def __repr__(self):\n        return self.__class__.__name__ + '(' + 'p=' + '{:.4f}'.format(self.p.data.tolist()[0]) + ', ' + 'eps=' + str(self.eps) + ')'\n```\n\n",
    "704562": "Hi,  I have one question: From method 2,  Do the 3 single AdaptiveAvgPool2d layer use the same feature from backbone? if yes, maybe that's no different with method 1.... but maybe I misunderstood..\n\n\n\n",
    "740335": "Hi.  For second method, May i ask you why you use activation-&gt;conv-&gt;bn?\nI found the last operation of resnet34 layer4 is relu, so thats to say your head is activation-&gt;activation-&gt;conv-bn.\nIs there any reason for this? I'm confused......Thank you..",
    "717204": "Why do you use 3 channels when actual image have just one?\nCannot catch your idea...",
    "704743": "Did you try training three separate models for each target?",
    "704548": "The CV of method 2 is 0.945 or is it a typo?",
    "705696": "What is `nn.Mish` mentioned here? Did you mean this activation function? https://arxiv.org/abs/1908.08681",
    "763019": "A good share !  \nIntead of your method2, I just add 3 linear layers after the feature map and get a worse result...  \nCan it be explained to the lack of conv layer of each part? The linear layers(only linear not conv) can't learn the details among each part while using a same feature map?"
  }
}