{
  "id": 172591,
  "title": "Why are we not freezing efficientnet's  non-top layers?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/172591",
  "author_name": "Krishna Kishor Kammaje",
  "post_date": "2020-08-05T17:27:21.463000",
  "votes": 3,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I have seen many kernels that use efficientnet pre-trained weights. Normally in transfer learning, I have seen that CNN layers are frozen and the final fully-connected layer/s are trained. (at least initially, before unfreezing all). But in all the kernels I saw here. no one does that. I tried doing that but had a poor result. It will be great if anyone can throw some light on this. </p>",
  "messages": [
    {
      "id": 959748,
      "postDate": "2020-08-05T21:23:49.387Z",
      "content": "<p>In general, the number of layers that should stay frozen when doing transfer learning depends on two things: the difference between the source and target domains and the sample size of the target data set.</p>\n\n<p>ImageNet, which is a source domain for most of the pre-trained models, contains images that are quite different from the ones used in this competition (e.g., cars, dogs, fruits, etc). Features learned to classify such images may be substantially different from the features needed to distinguish malignant and benign lesions. Some basic features from the early layers (like shapes) can still be useful, so it is a good idea to start from pre-trained weights. But given a decent sample size and large differences between the domains, fine-tuning both convolutional and fully-connected layers is more likely to demonstrate better performance.</p>",
      "rawMarkdown": "In general, the number of layers that should stay frozen when doing transfer learning depends on two things: the difference between the source and target domains and the sample size of the target data set.\n\nImageNet, which is a source domain for most of the pre-trained models, contains images that are quite different from the ones used in this competition (e.g., cars, dogs, fruits, etc). Features learned to classify such images may be substantially different from the features needed to distinguish malignant and benign lesions. Some basic features from the early layers (like shapes) can still be useful, so it is a good idea to start from pre-trained weights. But given a decent sample size and large differences between the domains, fine-tuning both convolutional and fully-connected layers is more likely to demonstrate better performance.",
      "votes": 10,
      "replies": [
        {
          "id": 960175,
          "postDate": "2020-08-06T07:35:45.267Z",
          "content": "<p>Very clear answer. Makes complete sense on why I was unable to get a better result by freezing layers. So this is still a transfer learning, but to a limited extent (just starting with the pre-trained weights). I thought that transfer learning is always about training only last or last few layers. Thanks a lot.</p>",
          "rawMarkdown": "Very clear answer. Makes complete sense on why I was unable to get a better result by freezing layers. So this is still a transfer learning, but to a limited extent (just starting with the pre-trained weights). I thought that transfer learning is always about training only last or last few layers. Thanks a lot."
        }
      ]
    },
    {
      "id": 959595,
      "postDate": "2020-08-05T18:23:30.743Z",
      "content": "<p>Try with freezing and try without freezing and you will know the answer.</p>",
      "rawMarkdown": "Try with freezing and try without freezing and you will know the answer.",
      "votes": 6,
      "replies": [
        {
          "id": 960177,
          "postDate": "2020-08-06T07:38:00.080Z",
          "content": "<p>I mentioned in the question that I tried freezing and got a poorer result. But I could not interpret the reason, so wanted to know. Thanks anyway. </p>",
          "rawMarkdown": "I mentioned in the question that I tried freezing and got a poorer result. But I could not interpret the reason, so wanted to know. Thanks anyway. "
        },
        {
          "id": 960500,
          "postDate": "2020-08-06T12:56:03.550Z",
          "content": "<p><a href=\"/krisho007\">@krisho007</a> this is not so simple\nit could help if you try it different ways, for instance you can freeze just few or freeze for a moment or change learning rate for different layers</p>",
          "rawMarkdown": "@krisho007 this is not so simple\nit could help if you try it different ways, for instance you can freeze just few or freeze for a moment or change learning rate for different layers"
        }
      ]
    },
    {
      "id": 959552,
      "postDate": "2020-08-05T17:27:21.463Z",
      "content": "<p>I have seen many kernels that use efficientnet pre-trained weights. Normally in transfer learning, I have seen that CNN layers are frozen and the final fully-connected layer/s are trained. (at least initially, before unfreezing all). But in all the kernels I saw here. no one does that. I tried doing that but had a poor result. It will be great if anyone can throw some light on this. </p>",
      "rawMarkdown": "I have seen many kernels that use efficientnet pre-trained weights. Normally in transfer learning, I have seen that CNN layers are frozen and the final fully-connected layer/s are trained. (at least initially, before unfreezing all). But in all the kernels I saw here. no one does that. I tried doing that but had a poor result. It will be great if anyone can throw some light on this. ",
      "votes": 3
    },
    {
      "id": 959761,
      "postDate": "2020-08-05T21:42:05.780Z",
      "content": "<p>Also, I would like to point out that in those public kernels that you are referring too, they typically start with a very low learning rate. So, for the first few epochs, the pre-trained weights are not allowed to change very much. It is not 100% equivalent but similar to keeping the weights completely frozen in the beginning of the training process.</p>",
      "rawMarkdown": "Also, I would like to point out that in those public kernels that you are referring too, they typically start with a very low learning rate. So, for the first few epochs, the pre-trained weights are not allowed to change very much. It is not 100% equivalent but similar to keeping the weights completely frozen in the beginning of the training process.",
      "votes": 4
    },
    {
      "id": 959651,
      "postDate": "2020-08-05T19:24:31.183Z",
      "content": "<p>Empirically worse results. Might not hold for all situations but I think in general one can forget this option. Same for dense heads imo.</p>",
      "rawMarkdown": "Empirically worse results. Might not hold for all situations but I think in general one can forget this option. Same for dense heads imo.",
      "votes": 1
    },
    {
      "id": 960447,
      "postDate": "2020-08-06T12:11:33.657Z",
      "content": "<p>There are few ways to start the transfer learning process. The most common one is to freeze the CNN layer and train only FC layers. This works for most of the cases where we are dealing to identify the common cases like dogs , houses etc as the model has already seen these images or look alike and there is no need to train the inner layers. But for those cases where we are going into uncharted territory ,above approach has some drawbacks and the general approach is to copy the structure and let the model to train the whole network. So in this approach instead to randomly initializing the parameters we initialize the parameters to be equals to the weight of pretrained models. I have seen doing this had given me better accuracy as well as faster model convergence. I hope this clarifies the ask</p>",
      "rawMarkdown": "There are few ways to start the transfer learning process. The most common one is to freeze the CNN layer and train only FC layers. This works for most of the cases where we are dealing to identify the common cases like dogs , houses etc as the model has already seen these images or look alike and there is no need to train the inner layers. But for those cases where we are going into uncharted territory ,above approach has some drawbacks and the general approach is to copy the structure and let the model to train the whole network. So in this approach instead to randomly initializing the parameters we initialize the parameters to be equals to the weight of pretrained models. I have seen doing this had given me better accuracy as well as faster model convergence. I hope this clarifies the ask",
      "votes": 2
    },
    {
      "id": 959784,
      "postDate": "2020-08-05T22:11:29.153Z",
      "content": "<p>The way to win the competition is to usually do something better than others, if you think this can lead to better solution you should try it then submit results and get the prize and glory. </p>\n\n<p>My personal problem with this competition is that I am not able to test all ideas I have, so I must choose where to go. But I am learning a lot - including things I never expected before.</p>",
      "rawMarkdown": "The way to win the competition is to usually do something better than others, if you think this can lead to better solution you should try it then submit results and get the prize and glory. \n\nMy personal problem with this competition is that I am not able to test all ideas I have, so I must choose where to go. But I am learning a lot - including things I never expected before."
    },
    {
      "id": 959643,
      "postDate": "2020-08-05T19:12:44.773Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 959748,
      "author_name": "Nikita Kozodoi",
      "author_url": "",
      "post_date": "2020-08-05T21:23:49.387000",
      "content": "<p>In general, the number of layers that should stay frozen when doing transfer learning depends on two things: the difference between the source and target domains and the sample size of the target data set.</p>\n\n<p>ImageNet, which is a source domain for most of the pre-trained models, contains images that are quite different from the ones used in this competition (e.g., cars, dogs, fruits, etc). Features learned to classify such images may be substantially different from the features needed to distinguish malignant and benign lesions. Some basic features from the early layers (like shapes) can still be useful, so it is a good idea to start from pre-trained weights. But given a decent sample size and large differences between the domains, fine-tuning both convolutional and fully-connected layers is more likely to demonstrate better performance.</p>",
      "votes": 10,
      "replies": [
        {
          "id": 960175,
          "author_name": "Krishna Kishor Kammaje",
          "author_url": "",
          "post_date": "2020-08-06T07:35:45.267000",
          "content": "<p>Very clear answer. Makes complete sense on why I was unable to get a better result by freezing layers. So this is still a transfer learning, but to a limited extent (just starting with the pre-trained weights). I thought that transfer learning is always about training only last or last few layers. Thanks a lot.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 959595,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2020-08-05T18:23:30.743000",
      "content": "<p>Try with freezing and try without freezing and you will know the answer.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 960177,
          "author_name": "Krishna Kishor Kammaje",
          "author_url": "",
          "post_date": "2020-08-06T07:38:00.080000",
          "content": "<p>I mentioned in the question that I tried freezing and got a poorer result. But I could not interpret the reason, so wanted to know. Thanks anyway. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 960500,
          "author_name": "Jacek Poplawski",
          "author_url": "",
          "post_date": "2020-08-06T12:56:03.550000",
          "content": "<p><a href=\"/krisho007\">@krisho007</a> this is not so simple\nit could help if you try it different ways, for instance you can freeze just few or freeze for a moment or change learning rate for different layers</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 959761,
      "author_name": "Alexey Pronin",
      "author_url": "",
      "post_date": "2020-08-05T21:42:05.780000",
      "content": "<p>Also, I would like to point out that in those public kernels that you are referring too, they typically start with a very low learning rate. So, for the first few epochs, the pre-trained weights are not allowed to change very much. It is not 100% equivalent but similar to keeping the weights completely frozen in the beginning of the training process.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 959651,
      "author_name": "Roman Weilguny",
      "author_url": "",
      "post_date": "2020-08-05T19:24:31.183000",
      "content": "<p>Empirically worse results. Might not hold for all situations but I think in general one can forget this option. Same for dense heads imo.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 960447,
      "author_name": "Navneet",
      "author_url": "",
      "post_date": "2020-08-06T12:11:33.657000",
      "content": "<p>There are few ways to start the transfer learning process. The most common one is to freeze the CNN layer and train only FC layers. This works for most of the cases where we are dealing to identify the common cases like dogs , houses etc as the model has already seen these images or look alike and there is no need to train the inner layers. But for those cases where we are going into uncharted territory ,above approach has some drawbacks and the general approach is to copy the structure and let the model to train the whole network. So in this approach instead to randomly initializing the parameters we initialize the parameters to be equals to the weight of pretrained models. I have seen doing this had given me better accuracy as well as faster model convergence. I hope this clarifies the ask</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 959784,
      "author_name": "Jacek Poplawski",
      "author_url": "",
      "post_date": "2020-08-05T22:11:29.153000",
      "content": "<p>The way to win the competition is to usually do something better than others, if you think this can lead to better solution you should try it then submit results and get the prize and glory. </p>\n\n<p>My personal problem with this competition is that I am not able to test all ideas I have, so I must choose where to go. But I am learning a lot - including things I never expected before.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 959643,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-05T19:12:44.773000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "959748": "In general, the number of layers that should stay frozen when doing transfer learning depends on two things: the difference between the source and target domains and the sample size of the target data set.\n\nImageNet, which is a source domain for most of the pre-trained models, contains images that are quite different from the ones used in this competition (e.g., cars, dogs, fruits, etc). Features learned to classify such images may be substantially different from the features needed to distinguish malignant and benign lesions. Some basic features from the early layers (like shapes) can still be useful, so it is a good idea to start from pre-trained weights. But given a decent sample size and large differences between the domains, fine-tuning both convolutional and fully-connected layers is more likely to demonstrate better performance.",
    "959595": "Try with freezing and try without freezing and you will know the answer.",
    "959552": "I have seen many kernels that use efficientnet pre-trained weights. Normally in transfer learning, I have seen that CNN layers are frozen and the final fully-connected layer/s are trained. (at least initially, before unfreezing all). But in all the kernels I saw here. no one does that. I tried doing that but had a poor result. It will be great if anyone can throw some light on this. ",
    "959761": "Also, I would like to point out that in those public kernels that you are referring too, they typically start with a very low learning rate. So, for the first few epochs, the pre-trained weights are not allowed to change very much. It is not 100% equivalent but similar to keeping the weights completely frozen in the beginning of the training process.",
    "959651": "Empirically worse results. Might not hold for all situations but I think in general one can forget this option. Same for dense heads imo.",
    "960447": "There are few ways to start the transfer learning process. The most common one is to freeze the CNN layer and train only FC layers. This works for most of the cases where we are dealing to identify the common cases like dogs , houses etc as the model has already seen these images or look alike and there is no need to train the inner layers. But for those cases where we are going into uncharted territory ,above approach has some drawbacks and the general approach is to copy the structure and let the model to train the whole network. So in this approach instead to randomly initializing the parameters we initialize the parameters to be equals to the weight of pretrained models. I have seen doing this had given me better accuracy as well as faster model convergence. I hope this clarifies the ask",
    "959784": "The way to win the competition is to usually do something better than others, if you think this can lead to better solution you should try it then submit results and get the prize and glory. \n\nMy personal problem with this competition is that I am not able to test all ideas I have, so I must choose where to go. But I am learning a lot - including things I never expected before.",
    "959643": ""
  }
}