{
  "id": 210638,
  "title": "How to add dropout and dense layer in PyTorch Pretrained model?",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/210638",
  "author_name": "",
  "post_date": "2021-01-11T15:25:51.559359600Z",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Can anyone help me to understand how can I add dropout, activation functions, and dense layers in the pre-trained PyTorch model? I try to search in the forum and also on stack overflow including PyTorch docs but did not get a working example of the same. Let's say we are writing out a custom class like this  </p>\n<pre><code>class CustomClassifierForPretrainedModel(nn.Module):\n  def __init__(self, pretrained_model, n_class, pretrained=False):\n    super().__init__()\n    self.model = timm.create_model(pretrained_model, pretrained=pretrained)\n    n_features = self.model.classifier.in_features           \n    # tweaking the model\n    self.model.classifier = nn.Sequential(\n                                 nn.Linear(n_features, 512),\n                                 nn.Dropout(0.5),\n                                 nn.ReLU(),\n                                 nn.Linear(256, 128),\n                                 nn.Dropout(0.33),\n                                 nn.ReLU(),\n                                 nn.Linear(256, n_class))             \n    def forward(self, x):\n             x = self.model(x)\n             return x\n</code></pre>\n<p>Am I doing it in the right way? However, I tried to run the same but ended with an OOM issue. Besides that, I'm trying to understand is it the right way to implement a custom pre-trained PyTorch model?</p>",
  "messages": [
    {
      "id": "1149089",
      "postDate": "01/11/2021 15:25:51",
      "content": "<p>Can anyone help me to understand how can I add dropout, activation functions, and dense layers in the pre-trained PyTorch model? I try to search in the forum and also on stack overflow including PyTorch docs but did not get a working example of the same. Let's say we are writing out a custom class like this  </p>\n<pre><code>class CustomClassifierForPretrainedModel(nn.Module):\n  def __init__(self, pretrained_model, n_class, pretrained=False):\n    super().__init__()\n    self.model = timm.create_model(pretrained_model, pretrained=pretrained)\n    n_features = self.model.classifier.in_features           \n    # tweaking the model\n    self.model.classifier = nn.Sequential(\n                                 nn.Linear(n_features, 512),\n                                 nn.Dropout(0.5),\n                                 nn.ReLU(),\n                                 nn.Linear(256, 128),\n                                 nn.Dropout(0.33),\n                                 nn.ReLU(),\n                                 nn.Linear(256, n_class))             \n    def forward(self, x):\n             x = self.model(x)\n             return x\n</code></pre>\n<p>Am I doing it in the right way? However, I tried to run the same but ended with an OOM issue. Besides that, I'm trying to understand is it the right way to implement a custom pre-trained PyTorch model?</p>",
      "rawMarkdown": "Can anyone help me to understand how can I add dropout, activation functions, and dense layers in the pre-trained PyTorch model? I try to search in the forum and also on stack overflow including PyTorch docs but did not get a working example of the same. Let's say we are writing out a custom class like this  \n\n\n```\nclass CustomClassifierForPretrainedModel(nn.Module):\n  def __init__(self, pretrained_model, n_class, pretrained=False):\n    super().__init__()\n    self.model = timm.create_model(pretrained_model, pretrained=pretrained)\n    n_features = self.model.classifier.in_features           \n    # tweaking the model\n    self.model.classifier = nn.Sequential(\n                                 nn.Linear(n_features, 512),\n                                 nn.Dropout(0.5),\n                                 nn.ReLU(),\n                                 nn.Linear(256, 128),\n                                 nn.Dropout(0.33),\n                                 nn.ReLU(),\n                                 nn.Linear(256, n_class))             \n    def forward(self, x):\n             x = self.model(x)\n             return x\n\n```\n\nAm I doing it in the right way? However, I tried to run the same but ended with an OOM issue. Besides that, I'm trying to understand is it the right way to implement a custom pre-trained PyTorch model?",
      "votes": null
    },
    {
      "id": "1149235",
      "postDate": "01/11/2021 17:40:33",
      "content": "<p>There is no layer named <code>classifier</code> in the pretrained model .  It need to be <code>self.model.fc</code> </p>\n<p>If you need to add other custom layers  that you have already defined, you have to extract pre-classifier features of the pretrained model first  and add them during the forward pass</p>\n<p>Example : </p>\n<pre><code>def forward(self, x):\n    x = self.model.forward_features(x)\n    x = self.mycustom_pre_pooled_layer(x) #optional\n    x = self.avg_pool(x).flatten(1)\n    x = self.mycustom_post_pooled_layer(x) #optional\n    x = self.dropout(x)\n    x = self.fc(x)\n    return x\n</code></pre>",
      "rawMarkdown": "There is no layer named `classifier ` in the pretrained model .  It need to be `self.model.fc` \n\nIf you need to add other custom layers  that you have already defined, you have to extract pre-classifier features of the pretrained model first  and add them during the forward pass\n\nExample : \n\n```\ndef forward(self, x):\n    x = self.model.forward_features(x)\n    x = self.mycustom_pre_pooled_layer(x) #optional\n    x = self.avg_pool(x).flatten(1)\n    x = self.mycustom_post_pooled_layer(x) #optional\n    x = self.dropout(x)\n    x = self.fc(x)\n    return x\n```",
      "votes": null
    },
    {
      "id": "1149239",
      "postDate": "01/11/2021 17:48:26",
      "content": "<p>For the EfficientNet models the output layer is called <code>classifier</code> (in timm package)</p>",
      "rawMarkdown": "For the EfficientNet models the output layer is called `classifier` (in timm package)",
      "votes": null
    },
    {
      "id": "1149257",
      "postDate": "01/11/2021 18:11:53",
      "content": "<p>Looks like your classifier has some wrong number of input/ouput in the Linear layers. Also adding some batchnorm layers can help. Here is what you should do:</p>\n<pre><code>self.model.classifier = nn.Sequential(\n                                 nn.Linear(n_features, 512),\n                                 nn.ReLU(),\n                                 nn.BatchNorm1d(512)\n                                 nn.Dropout(0.5),\n                                 nn.Linear(512, 256),\n                                 nn.ReLU(),\n                                 nn.BatchNorm1d(256)\n                                 nn.Dropout(0.33),\n                                 nn.Linear(256, n_class))  \n</code></pre>",
      "rawMarkdown": "Looks like your classifier has some wrong number of input/ouput in the Linear layers. Also adding some batchnorm layers can help. Here is what you should do:\n\n```\nself.model.classifier = nn.Sequential(\n                                 nn.Linear(n_features, 512),\n                                 nn.ReLU(),\n                                 nn.BatchNorm1d(512)\n                                 nn.Dropout(0.5),\n                                 nn.Linear(512, 256),\n                                 nn.ReLU(),\n                                 nn.BatchNorm1d(256)\n                                 nn.Dropout(0.33),\n                                 nn.Linear(256, n_class))  \n```",
      "votes": null
    },
    {
      "id": "1149287",
      "postDate": "01/11/2021 18:32:57",
      "content": "<p>You're right.  </p>\n<p>I always use the pre-classifier features (<code>forward_features</code>) and add my own pooled/custom/fc layers. At least it avoid confusion with resnet like fc layer ^^</p>",
      "rawMarkdown": "You're right.  \n\nI always use the pre-classifier features (`forward_features`) and add my own pooled/custom/fc layers. At least it avoid confusion with resnet like fc layer ^^",
      "votes": null
    },
    {
      "id": "1149449",
      "postDate": "01/11/2021 21:02:00",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/yannmajewski\" target=\"_blank\">@yannmajewski</a>. Your response makes sense. I will try out it :)</p>",
      "rawMarkdown": "Thanks @yannmajewski. Your response makes sense. I will try out it :)",
      "votes": null
    },
    {
      "id": "1149453",
      "postDate": "01/11/2021 21:05:37",
      "content": "<p>Ali has rightly pointed out the typical output layer for EfficientNet( for timm packages). However, I got the idea to follow the general pattern to use different layers for the pre-trained models suggested by Sergine. Thank you both.</p>",
      "rawMarkdown": "Ali has rightly pointed out the typical output layer for EfficientNet( for timm packages). However, I got the idea to follow the general pattern to use different layers for the pre-trained models suggested by Sergine. Thank you both.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1149235,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "01/11/2021 17:40:33",
      "content": "<p>There is no layer named <code>classifier</code> in the pretrained model .  It need to be <code>self.model.fc</code> </p>\n<p>If you need to add other custom layers  that you have already defined, you have to extract pre-classifier features of the pretrained model first  and add them during the forward pass</p>\n<p>Example : </p>\n<pre><code>def forward(self, x):\n    x = self.model.forward_features(x)\n    x = self.mycustom_pre_pooled_layer(x) #optional\n    x = self.avg_pool(x).flatten(1)\n    x = self.mycustom_post_pooled_layer(x) #optional\n    x = self.dropout(x)\n    x = self.fc(x)\n    return x\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1149239,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "01/11/2021 17:48:26",
          "content": "<p>For the EfficientNet models the output layer is called <code>classifier</code> (in timm package)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1149287,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "01/11/2021 18:32:57",
          "content": "<p>You're right.  </p>\n<p>I always use the pre-classifier features (<code>forward_features</code>) and add my own pooled/custom/fc layers. At least it avoid confusion with resnet like fc layer ^^</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1149453,
          "author_name": "saurabh2mishra",
          "author_url": "",
          "post_date": "01/11/2021 21:05:37",
          "content": "<p>Ali has rightly pointed out the typical output layer for EfficientNet( for timm packages). However, I got the idea to follow the general pattern to use different layers for the pre-trained models suggested by Sergine. Thank you both.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1149257,
      "author_name": "yannmajewski",
      "author_url": "",
      "post_date": "01/11/2021 18:11:53",
      "content": "<p>Looks like your classifier has some wrong number of input/ouput in the Linear layers. Also adding some batchnorm layers can help. Here is what you should do:</p>\n<pre><code>self.model.classifier = nn.Sequential(\n                                 nn.Linear(n_features, 512),\n                                 nn.ReLU(),\n                                 nn.BatchNorm1d(512)\n                                 nn.Dropout(0.5),\n                                 nn.Linear(512, 256),\n                                 nn.ReLU(),\n                                 nn.BatchNorm1d(256)\n                                 nn.Dropout(0.33),\n                                 nn.Linear(256, n_class))  \n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1149449,
          "author_name": "saurabh2mishra",
          "author_url": "",
          "post_date": "01/11/2021 21:02:00",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/yannmajewski\" target=\"_blank\">@yannmajewski</a>. Your response makes sense. I will try out it :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1149089": "Can anyone help me to understand how can I add dropout, activation functions, and dense layers in the pre-trained PyTorch model? I try to search in the forum and also on stack overflow including PyTorch docs but did not get a working example of the same. Let's say we are writing out a custom class like this  \n\n\n```\nclass CustomClassifierForPretrainedModel(nn.Module):\n  def __init__(self, pretrained_model, n_class, pretrained=False):\n    super().__init__()\n    self.model = timm.create_model(pretrained_model, pretrained=pretrained)\n    n_features = self.model.classifier.in_features           \n    # tweaking the model\n    self.model.classifier = nn.Sequential(\n                                 nn.Linear(n_features, 512),\n                                 nn.Dropout(0.5),\n                                 nn.ReLU(),\n                                 nn.Linear(256, 128),\n                                 nn.Dropout(0.33),\n                                 nn.ReLU(),\n                                 nn.Linear(256, n_class))             \n    def forward(self, x):\n             x = self.model(x)\n             return x\n\n```\n\nAm I doing it in the right way? However, I tried to run the same but ended with an OOM issue. Besides that, I'm trying to understand is it the right way to implement a custom pre-trained PyTorch model?",
    "1149235": "There is no layer named `classifier ` in the pretrained model .  It need to be `self.model.fc` \n\nIf you need to add other custom layers  that you have already defined, you have to extract pre-classifier features of the pretrained model first  and add them during the forward pass\n\nExample : \n\n```\ndef forward(self, x):\n    x = self.model.forward_features(x)\n    x = self.mycustom_pre_pooled_layer(x) #optional\n    x = self.avg_pool(x).flatten(1)\n    x = self.mycustom_post_pooled_layer(x) #optional\n    x = self.dropout(x)\n    x = self.fc(x)\n    return x\n```",
    "1149239": "For the EfficientNet models the output layer is called `classifier` (in timm package)",
    "1149257": "Looks like your classifier has some wrong number of input/ouput in the Linear layers. Also adding some batchnorm layers can help. Here is what you should do:\n\n```\nself.model.classifier = nn.Sequential(\n                                 nn.Linear(n_features, 512),\n                                 nn.ReLU(),\n                                 nn.BatchNorm1d(512)\n                                 nn.Dropout(0.5),\n                                 nn.Linear(512, 256),\n                                 nn.ReLU(),\n                                 nn.BatchNorm1d(256)\n                                 nn.Dropout(0.33),\n                                 nn.Linear(256, n_class))  \n```",
    "1149287": "You're right.  \n\nI always use the pre-classifier features (`forward_features`) and add my own pooled/custom/fc layers. At least it avoid confusion with resnet like fc layer ^^",
    "1149449": "Thanks @yannmajewski. Your response makes sense. I will try out it :)",
    "1149453": "Ali has rightly pointed out the typical output layer for EfficientNet( for timm packages). However, I got the idea to follow the general pattern to use different layers for the pre-trained models suggested by Sergine. Thank you both."
  },
  "source": "meta"
}