{
  "id": 102812,
  "title": "Transfer learning question ",
  "url": "/competitions/aptos2019-blindness-detection/discussion/102812",
  "author_name": "",
  "post_date": "2019-08-05T08:29:47.495913600Z",
  "votes": 6,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello \nI want to get some feedback about the shape of the tail you add to CNN when you do a transfer learning from deep learning network (like ResNet or EfficentNet ).\nThe basic approach is to replace the last classifier layer with a simple linear layer with in=num_of_features , out = num_classes \nThen I see methods of using global max-pooling (2d)  or using a new sequence with batch normalization.\nAny feedback?  useful articles or links to read?  </p>",
  "messages": [
    {
      "id": "592399",
      "postDate": "08/05/2019 08:29:47",
      "content": "<p>Hello \nI want to get some feedback about the shape of the tail you add to CNN when you do a transfer learning from deep learning network (like ResNet or EfficentNet ).\nThe basic approach is to replace the last classifier layer with a simple linear layer with in=num_of_features , out = num_classes \nThen I see methods of using global max-pooling (2d)  or using a new sequence with batch normalization.\nAny feedback?  useful articles or links to read?  </p>",
      "rawMarkdown": "Hello \nI want to get some feedback about the shape of the tail you add to CNN when you do a transfer learning from deep learning network (like ResNet or EfficentNet ).\nThe basic approach is to replace the last classifier layer with a simple linear layer with in=num_of_features , out = num_classes \nThen I see methods of using global max-pooling (2d)  or using a new sequence with batch normalization.\nAny feedback?  useful articles or links to read?",
      "votes": null
    },
    {
      "id": "593578",
      "postDate": "08/06/2019 20:33:34",
      "content": "<p>Hi, \nGenerally speaking there are two ways to do transfer learning.</p>\n\n<p>METHOD A.\nAs you said you modify last linear layer to match your target output.</p>\n\n<p>METHOD B.\nInstead of modifying last linear  layer you find last <code>AdaptiveAvgPool2d</code> and cut the model. So in the case of resnet you left with your resblock (end of  layer usually <code>Conv2D</code> -- <code>BatchNorm</code>). From here you have 3 options.</p>\n\n<p>Option A:\nAdd new  <code>AdaptiveAvgPool2d</code> --&gt;  <code>Flatten()</code> --&gt; <code>Linear</code> (in features,  your target output)</p>\n\n<p>Option B:\nAdd new <code>AdaptiveMaxPool2d</code>  --&gt;  <code>Flatten()</code> --&gt; <code>Linear</code> (in features,  your target output)</p>\n\n<p>Option C:\nUse both \n[ <code>AdaptiveAvgPool2d</code> ,  <code>AdaptiveMaxPool2d</code>]  --&gt; concatenate  output --&gt; <code>Flatten()</code> --&gt; <code>Linear</code> (in features,  your target output).</p>\n\n<p>Hope it helps =) </p>",
      "rawMarkdown": "Hi, \nGenerally speaking there are two ways to do transfer learning.\n\nMETHOD A.\nAs you said you modify last linear layer to match your target output.\n\nMETHOD B.\nInstead of modifying last linear  layer you find last `AdaptiveAvgPool2d` and cut the model. So in the case of resnet you left with your resblock (end of  layer usually `Conv2D` -- `BatchNorm`). From here you have 3 options.\n\nOption A:\nAdd new  `AdaptiveAvgPool2d` --&gt;  `Flatten()` --&gt; `Linear` (in features,  your target output)\n\nOption B:\nAdd new `AdaptiveMaxPool2d`  --&gt;  `Flatten()` --&gt; `Linear` (in features,  your target output)\n\nOption C:\nUse both \n[ `AdaptiveAvgPool2d` ,  `AdaptiveMaxPool2d `]  --&gt; concatenate  output --&gt; `Flatten()` --&gt; `Linear` (in features,  your target output).\n\n\nHope it helps =)",
      "votes": null
    },
    {
      "id": "593771",
      "postDate": "08/07/2019 05:03:24",
      "content": "<p>Thanks for the detailed reply </p>",
      "rawMarkdown": "Thanks for the detailed reply",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 593578,
      "author_name": "drhabib",
      "author_url": "",
      "post_date": "08/06/2019 20:33:34",
      "content": "<p>Hi, \nGenerally speaking there are two ways to do transfer learning.</p>\n\n<p>METHOD A.\nAs you said you modify last linear layer to match your target output.</p>\n\n<p>METHOD B.\nInstead of modifying last linear  layer you find last <code>AdaptiveAvgPool2d</code> and cut the model. So in the case of resnet you left with your resblock (end of  layer usually <code>Conv2D</code> -- <code>BatchNorm</code>). From here you have 3 options.</p>\n\n<p>Option A:\nAdd new  <code>AdaptiveAvgPool2d</code> --&gt;  <code>Flatten()</code> --&gt; <code>Linear</code> (in features,  your target output)</p>\n\n<p>Option B:\nAdd new <code>AdaptiveMaxPool2d</code>  --&gt;  <code>Flatten()</code> --&gt; <code>Linear</code> (in features,  your target output)</p>\n\n<p>Option C:\nUse both \n[ <code>AdaptiveAvgPool2d</code> ,  <code>AdaptiveMaxPool2d</code>]  --&gt; concatenate  output --&gt; <code>Flatten()</code> --&gt; <code>Linear</code> (in features,  your target output).</p>\n\n<p>Hope it helps =) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 593771,
      "author_name": "omershect",
      "author_url": "",
      "post_date": "08/07/2019 05:03:24",
      "content": "<p>Thanks for the detailed reply </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "592399": "Hello \nI want to get some feedback about the shape of the tail you add to CNN when you do a transfer learning from deep learning network (like ResNet or EfficentNet ).\nThe basic approach is to replace the last classifier layer with a simple linear layer with in=num_of_features , out = num_classes \nThen I see methods of using global max-pooling (2d)  or using a new sequence with batch normalization.\nAny feedback?  useful articles or links to read?",
    "593578": "Hi, \nGenerally speaking there are two ways to do transfer learning.\n\nMETHOD A.\nAs you said you modify last linear layer to match your target output.\n\nMETHOD B.\nInstead of modifying last linear  layer you find last `AdaptiveAvgPool2d` and cut the model. So in the case of resnet you left with your resblock (end of  layer usually `Conv2D` -- `BatchNorm`). From here you have 3 options.\n\nOption A:\nAdd new  `AdaptiveAvgPool2d` --&gt;  `Flatten()` --&gt; `Linear` (in features,  your target output)\n\nOption B:\nAdd new `AdaptiveMaxPool2d`  --&gt;  `Flatten()` --&gt; `Linear` (in features,  your target output)\n\nOption C:\nUse both \n[ `AdaptiveAvgPool2d` ,  `AdaptiveMaxPool2d `]  --&gt; concatenate  output --&gt; `Flatten()` --&gt; `Linear` (in features,  your target output).\n\n\nHope it helps =)",
    "593771": "Thanks for the detailed reply"
  },
  "source": "meta"
}