{
  "id": 396406,
  "title": "Question on pretraining in pytorch",
  "url": "/competitions/birdclef-2023/discussion/396406",
  "author_name": "",
  "post_date": "2023-03-21T14:03:11.446352200Z",
  "votes": 2,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi!<br>\nI have pretrained an effnet b1 model on 2021-22 data, now to finetune the model I've loaded the my weights pretrained on the 2021-22 data and added another layer that projects the output the 264 classes.<br>\nMy question is if I'm using the timm library and I've put the <code>pretrained</code> argument to False and loaded my weights do I need to freeze the weights to finetune the model? Here is a sample of my code:</p>\n<pre><code>class PretrainedModel(nn.Module):\n    def __init__(self, model_name=Config.model, num_classes = Config.pretrained_classes, pretrained = False):\n        super().__init__()\n        self.num_classes = num_classes\n\n        self.backbone = timm.create_model(model_name, pretrained=pretrained)\n        self.in_features = self.backbone.classifier.in_features\n        self.backbone.classifier = nn.Sequential(\n            nn.Linear(self.in_features, num_classes)\n        )\n\n    def forward(self,images):\n        logits = self.backbone(images)\n        return logits\n</code></pre>\n<pre><code>class BirdClefModel(nn.Module):\n    def __init__(self, model_name=Config.model, num_classes = Config.num_classes, pretrained = False):\n        super().__init__()\n        self.num_classes = num_classes\n\n        self.backbone = PretrainedModel()\n        self.backbone.load_state_dict(torch.load('/kaggle/input/pretrained/model.pth'))\n        print(\"Pretrained model loaded...\")\n\n        self.in_feats = self.backbone.num_classes\n        self.fc = nn.Linear(self.in_feats, num_classes)\n\n\n    def forward(self,images):\n        logits = self.backbone(images)\n        logits = self.fc(logits)\n        return logits\n</code></pre>\n<p>It would be great of anyonw could help me with this question.<br>\nThanks!</p>",
  "messages": [
    {
      "id": "2190804",
      "postDate": "03/21/2023 14:03:11",
      "content": "<p>Hi!<br>\nI have pretrained an effnet b1 model on 2021-22 data, now to finetune the model I've loaded the my weights pretrained on the 2021-22 data and added another layer that projects the output the 264 classes.<br>\nMy question is if I'm using the timm library and I've put the <code>pretrained</code> argument to False and loaded my weights do I need to freeze the weights to finetune the model? Here is a sample of my code:</p>\n<pre><code>class PretrainedModel(nn.Module):\n    def __init__(self, model_name=Config.model, num_classes = Config.pretrained_classes, pretrained = False):\n        super().__init__()\n        self.num_classes = num_classes\n\n        self.backbone = timm.create_model(model_name, pretrained=pretrained)\n        self.in_features = self.backbone.classifier.in_features\n        self.backbone.classifier = nn.Sequential(\n            nn.Linear(self.in_features, num_classes)\n        )\n\n    def forward(self,images):\n        logits = self.backbone(images)\n        return logits\n</code></pre>\n<pre><code>class BirdClefModel(nn.Module):\n    def __init__(self, model_name=Config.model, num_classes = Config.num_classes, pretrained = False):\n        super().__init__()\n        self.num_classes = num_classes\n\n        self.backbone = PretrainedModel()\n        self.backbone.load_state_dict(torch.load('/kaggle/input/pretrained/model.pth'))\n        print(\"Pretrained model loaded...\")\n\n        self.in_feats = self.backbone.num_classes\n        self.fc = nn.Linear(self.in_feats, num_classes)\n\n\n    def forward(self,images):\n        logits = self.backbone(images)\n        logits = self.fc(logits)\n        return logits\n</code></pre>\n<p>It would be great of anyonw could help me with this question.<br>\nThanks!</p>",
      "rawMarkdown": "Hi!\nI have pretrained an effnet b1 model on 2021-22 data, now to finetune the model I've loaded the my weights pretrained on the 2021-22 data and added another layer that projects the output the 264 classes.\nMy question is if I'm using the timm library and I've put the `pretrained` argument to False and loaded my weights do I need to freeze the weights to finetune the model? Here is a sample of my code:\n```\nclass PretrainedModel(nn.Module):\n    def __init__(self, model_name=Config.model, num_classes = Config.pretrained_classes, pretrained = False):\n        super().__init__()\n        self.num_classes = num_classes\n\n        self.backbone = timm.create_model(model_name, pretrained=pretrained)\n        self.in_features = self.backbone.classifier.in_features\n        self.backbone.classifier = nn.Sequential(\n            nn.Linear(self.in_features, num_classes)\n        )\n    \n    def forward(self,images):\n        logits = self.backbone(images)\n        return logits\n```\n\n```\nclass BirdClefModel(nn.Module):\n    def __init__(self, model_name=Config.model, num_classes = Config.num_classes, pretrained = False):\n        super().__init__()\n        self.num_classes = num_classes\n\n        self.backbone = PretrainedModel()\n        self.backbone.load_state_dict(torch.load('/kaggle/input/pretrained/model.pth'))\n        print(\"Pretrained model loaded...\")\n        \n        self.in_feats = self.backbone.num_classes\n        self.fc = nn.Linear(self.in_feats, num_classes)\n        \n    \n    def forward(self,images):\n        logits = self.backbone(images)\n        logits = self.fc(logits)\n        return logits\n```\nIt would be great of anyonw could help me with this question.\nThanks!",
      "votes": null
    },
    {
      "id": "2190821",
      "postDate": "03/21/2023 14:10:30",
      "content": "<p>In that case, there are two ways to fine-tune your model on 2023 data. </p>\n<ol>\n<li>Freeze the backbone layers, &amp; fine-tune just the classifier head (fc layer)</li>\n<li>Finetune the entire backbone + classifier head (lesser epochs)</li>\n</ol>\n<p>That's up to the experiment, I usually prefer the 2nd option, there are new birds in 2023, could be possible that freezed backbone not be able to produce a good feature map without losing much information.</p>",
      "rawMarkdown": "In that case, there are two ways to fine-tune your model on 2023 data. \n\n1. Freeze the backbone layers, & fine-tune just the classifier head (fc layer)\n2. Finetune the entire backbone + classifier head (lesser epochs)\n\nThat's up to the experiment, I usually prefer the 2nd option, there are new birds in 2023, could be possible that freezed backbone not be able to produce a good feature map without losing much information.",
      "votes": null
    },
    {
      "id": "2190862",
      "postDate": "03/21/2023 14:36:59",
      "content": "<p>So what I'm currently doing is the 2nd option, right?</p>",
      "rawMarkdown": "So what I'm currently doing is the 2nd option, right?",
      "votes": null
    },
    {
      "id": "2190873",
      "postDate": "03/21/2023 14:47:15",
      "content": "<p>I don't see any code to freeze layers in your code you have shared, so I assume yes</p>",
      "rawMarkdown": "I don't see any code to freeze layers in your code you have shared, so I assume yes",
      "votes": null
    },
    {
      "id": "2190876",
      "postDate": "03/21/2023 14:48:16",
      "content": "<p>mhh you can also replace the PretrainedModel classifier head with a newly initialized one. <br>\nIn BirdClefModel after loading older weights, do <code>self.backbone.classifier = nn.Identity()</code> and replace in_feats value with self.in_features</p>",
      "rawMarkdown": "mhh you can also replace the PretrainedModel classifier head with a newly initialized one. \nIn BirdClefModel after loading older weights, do `self.backbone.classifier = nn.Identity()` and replace in_feats value with self.in_features",
      "votes": null
    },
    {
      "id": "2190882",
      "postDate": "03/21/2023 14:51:39",
      "content": "<blockquote>\n  <p>In BirdClefModel after loading older weights, do self.backbone.classifier = nn.Identity() and replace in_feats value with self.in_features</p>\n</blockquote>\n<p>I believe that is generally a better way to do it, the layer backbone.classifier has currently learned to be a classifier head and its output currently has just about no business being useful to another head with different classes.</p>",
      "rawMarkdown": ">In BirdClefModel after loading older weights, do self.backbone.classifier = nn.Identity() and replace in_feats value with self.in_features\n\nI believe that is generally a better way to do it, the layer backbone.classifier has currently learned to be a classifier head and its output currently has just about no business being useful to another head with different classes.",
      "votes": null
    },
    {
      "id": "2190913",
      "postDate": "03/21/2023 15:11:13",
      "content": "<p>Thanks a lot for this!</p>",
      "rawMarkdown": "Thanks a lot for this!",
      "votes": null
    },
    {
      "id": "2242284",
      "postDate": "05/02/2023 06:51:30",
      "content": "<p>I am using baseline notebook of <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> . After pretraining my score is still stuck at 0.78. My cv is 0.82(cmap(5)) , 0.80(AP) but didn't get any boost in lb.</p>",
      "rawMarkdown": "I am using baseline notebook of @nischaydnk . After pretraining my score is still stuck at 0.78. My cv is 0.82(cmap(5)) , 0.80(AP) but didn't get any boost in lb.",
      "votes": null
    },
    {
      "id": "2249556",
      "postDate": "05/07/2023 20:53:11",
      "content": "<p>Hi Himanshu,  I've pursued this path as well, although with 21, 22, 23 data for pre-training, but to be honest it hasn't worked at all for me yet, I haven't even got it to converge on a vaguely sensible solution equivalent to my original model.  It's been somewhat discouraging!</p>\n<p>Sticking with Nischay's numbering, I have been trying option 3:  On fine-tune I freeze the entire backbone and pre-trained head, concatenate another randomly initialised layer, and train that.  After a few epochs I unfreeze the rest of the classifier head.  To me this sounds sensible, but I'd love to get another opinion,  I think the frozen part should be a good feature generator, and now I'm just training a single new layer at the start.</p>\n<p>If we replace the head weights like in 1. &amp; 2. aren't we throwing away much of the learning benefit?  Keeping only the backbone finetuning learning from the pre-training.  But is this what everyone does?  I guess the benefit is that we're using a simple fc layer for the classifier head like the model architecture was optimised for.</p>",
      "rawMarkdown": "Hi Himanshu,  I've pursued this path as well, although with 21, 22, 23 data for pre-training, but to be honest it hasn't worked at all for me yet, I haven't even got it to converge on a vaguely sensible solution equivalent to my original model.  It's been somewhat discouraging!\n\nSticking with Nischay's numbering, I have been trying option 3:  On fine-tune I freeze the entire backbone and pre-trained head, concatenate another randomly initialised layer, and train that.  After a few epochs I unfreeze the rest of the classifier head.  To me this sounds sensible, but I'd love to get another opinion,  I think the frozen part should be a good feature generator, and now I'm just training a single new layer at the start.\n\nIf we replace the head weights like in 1. & 2. aren't we throwing away much of the learning benefit?  Keeping only the backbone finetuning learning from the pre-training.  But is this what everyone does?  I guess the benefit is that we're using a simple fc layer for the classifier head like the model architecture was optimised for.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2190821,
      "author_name": "nischaydnk",
      "author_url": "",
      "post_date": "03/21/2023 14:10:30",
      "content": "<p>In that case, there are two ways to fine-tune your model on 2023 data. </p>\n<ol>\n<li>Freeze the backbone layers, &amp; fine-tune just the classifier head (fc layer)</li>\n<li>Finetune the entire backbone + classifier head (lesser epochs)</li>\n</ol>\n<p>That's up to the experiment, I usually prefer the 2nd option, there are new birds in 2023, could be possible that freezed backbone not be able to produce a good feature map without losing much information.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2190862,
          "author_name": "hackingpirate",
          "author_url": "",
          "post_date": "03/21/2023 14:36:59",
          "content": "<p>So what I'm currently doing is the 2nd option, right?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2190873,
              "author_name": "harshitsheoran",
              "author_url": "",
              "post_date": "03/21/2023 14:47:15",
              "content": "<p>I don't see any code to freeze layers in your code you have shared, so I assume yes</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2190876,
              "author_name": "nischaydnk",
              "author_url": "",
              "post_date": "03/21/2023 14:48:16",
              "content": "<p>mhh you can also replace the PretrainedModel classifier head with a newly initialized one. <br>\nIn BirdClefModel after loading older weights, do <code>self.backbone.classifier = nn.Identity()</code> and replace in_feats value with self.in_features</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2190913,
                  "author_name": "hackingpirate",
                  "author_url": "",
                  "post_date": "03/21/2023 15:11:13",
                  "content": "<p>Thanks a lot for this!</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            },
            {
              "id": 2190882,
              "author_name": "harshitsheoran",
              "author_url": "",
              "post_date": "03/21/2023 14:51:39",
              "content": "<blockquote>\n  <p>In BirdClefModel after loading older weights, do self.backbone.classifier = nn.Identity() and replace in_feats value with self.in_features</p>\n</blockquote>\n<p>I believe that is generally a better way to do it, the layer backbone.classifier has currently learned to be a classifier head and its output currently has just about no business being useful to another head with different classes.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2242284,
      "author_name": "himanshunayal",
      "author_url": "",
      "post_date": "05/02/2023 06:51:30",
      "content": "<p>I am using baseline notebook of <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> . After pretraining my score is still stuck at 0.78. My cv is 0.82(cmap(5)) , 0.80(AP) but didn't get any boost in lb.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2249556,
          "author_name": "ollypowell",
          "author_url": "",
          "post_date": "05/07/2023 20:53:11",
          "content": "<p>Hi Himanshu,  I've pursued this path as well, although with 21, 22, 23 data for pre-training, but to be honest it hasn't worked at all for me yet, I haven't even got it to converge on a vaguely sensible solution equivalent to my original model.  It's been somewhat discouraging!</p>\n<p>Sticking with Nischay's numbering, I have been trying option 3:  On fine-tune I freeze the entire backbone and pre-trained head, concatenate another randomly initialised layer, and train that.  After a few epochs I unfreeze the rest of the classifier head.  To me this sounds sensible, but I'd love to get another opinion,  I think the frozen part should be a good feature generator, and now I'm just training a single new layer at the start.</p>\n<p>If we replace the head weights like in 1. &amp; 2. aren't we throwing away much of the learning benefit?  Keeping only the backbone finetuning learning from the pre-training.  But is this what everyone does?  I guess the benefit is that we're using a simple fc layer for the classifier head like the model architecture was optimised for.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2190804": "Hi!\nI have pretrained an effnet b1 model on 2021-22 data, now to finetune the model I've loaded the my weights pretrained on the 2021-22 data and added another layer that projects the output the 264 classes.\nMy question is if I'm using the timm library and I've put the `pretrained` argument to False and loaded my weights do I need to freeze the weights to finetune the model? Here is a sample of my code:\n```\nclass PretrainedModel(nn.Module):\n    def __init__(self, model_name=Config.model, num_classes = Config.pretrained_classes, pretrained = False):\n        super().__init__()\n        self.num_classes = num_classes\n\n        self.backbone = timm.create_model(model_name, pretrained=pretrained)\n        self.in_features = self.backbone.classifier.in_features\n        self.backbone.classifier = nn.Sequential(\n            nn.Linear(self.in_features, num_classes)\n        )\n    \n    def forward(self,images):\n        logits = self.backbone(images)\n        return logits\n```\n\n```\nclass BirdClefModel(nn.Module):\n    def __init__(self, model_name=Config.model, num_classes = Config.num_classes, pretrained = False):\n        super().__init__()\n        self.num_classes = num_classes\n\n        self.backbone = PretrainedModel()\n        self.backbone.load_state_dict(torch.load('/kaggle/input/pretrained/model.pth'))\n        print(\"Pretrained model loaded...\")\n        \n        self.in_feats = self.backbone.num_classes\n        self.fc = nn.Linear(self.in_feats, num_classes)\n        \n    \n    def forward(self,images):\n        logits = self.backbone(images)\n        logits = self.fc(logits)\n        return logits\n```\nIt would be great of anyonw could help me with this question.\nThanks!",
    "2190821": "In that case, there are two ways to fine-tune your model on 2023 data. \n\n1. Freeze the backbone layers, & fine-tune just the classifier head (fc layer)\n2. Finetune the entire backbone + classifier head (lesser epochs)\n\nThat's up to the experiment, I usually prefer the 2nd option, there are new birds in 2023, could be possible that freezed backbone not be able to produce a good feature map without losing much information.",
    "2190862": "So what I'm currently doing is the 2nd option, right?",
    "2190873": "I don't see any code to freeze layers in your code you have shared, so I assume yes",
    "2190876": "mhh you can also replace the PretrainedModel classifier head with a newly initialized one. \nIn BirdClefModel after loading older weights, do `self.backbone.classifier = nn.Identity()` and replace in_feats value with self.in_features",
    "2190882": ">In BirdClefModel after loading older weights, do self.backbone.classifier = nn.Identity() and replace in_feats value with self.in_features\n\nI believe that is generally a better way to do it, the layer backbone.classifier has currently learned to be a classifier head and its output currently has just about no business being useful to another head with different classes.",
    "2190913": "Thanks a lot for this!",
    "2242284": "I am using baseline notebook of @nischaydnk . After pretraining my score is still stuck at 0.78. My cv is 0.82(cmap(5)) , 0.80(AP) but didn't get any boost in lb.",
    "2249556": "Hi Himanshu,  I've pursued this path as well, although with 21, 22, 23 data for pre-training, but to be honest it hasn't worked at all for me yet, I haven't even got it to converge on a vaguely sensible solution equivalent to my original model.  It's been somewhat discouraging!\n\nSticking with Nischay's numbering, I have been trying option 3:  On fine-tune I freeze the entire backbone and pre-trained head, concatenate another randomly initialised layer, and train that.  After a few epochs I unfreeze the rest of the classifier head.  To me this sounds sensible, but I'd love to get another opinion,  I think the frozen part should be a good feature generator, and now I'm just training a single new layer at the start.\n\nIf we replace the head weights like in 1. & 2. aren't we throwing away much of the learning benefit?  Keeping only the backbone finetuning learning from the pre-training.  But is this what everyone does?  I guess the benefit is that we're using a simple fc layer for the classifier head like the model architecture was optimised for."
  },
  "source": "meta"
}