{
  "id": 222377,
  "title": "Trainable Params Way Less Than Non Trainable Params",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/222377",
  "author_name": "",
  "post_date": "2021-02-26T17:48:42.352792100Z",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Total Params: 64,125,851<br>\nTrainable Params: 28,171<br>\nNon-Trainable Params: 64,097,680</p>\n<p>I am using CNN pre-trained on EfficientNetB7. My guess is my model will not be able to make full sense of the data since it can only train on 28,171 params. If my guess is correct, then I will welcome suggestions from you guys on how to make the non-trainable params very small and the trainable params very large.</p>\n<p>Thanks.</p>\n<p>Note:<br>\n(1) I used test_size=20% for train_test_split<br>\n(2) I used ModelCheckPoint and EarlyStopping</p>",
  "messages": [
    {
      "id": "1219331",
      "postDate": "02/26/2021 17:48:42",
      "content": "<p>Total Params: 64,125,851<br>\nTrainable Params: 28,171<br>\nNon-Trainable Params: 64,097,680</p>\n<p>I am using CNN pre-trained on EfficientNetB7. My guess is my model will not be able to make full sense of the data since it can only train on 28,171 params. If my guess is correct, then I will welcome suggestions from you guys on how to make the non-trainable params very small and the trainable params very large.</p>\n<p>Thanks.</p>\n<p>Note:<br>\n(1) I used test_size=20% for train_test_split<br>\n(2) I used ModelCheckPoint and EarlyStopping</p>",
      "rawMarkdown": "Total Params: 64,125,851\nTrainable Params: 28,171\nNon-Trainable Params: 64,097,680\n\nI am using CNN pre-trained on EfficientNetB7. My guess is my model will not be able to make full sense of the data since it can only train on 28,171 params. If my guess is correct, then I will welcome suggestions from you guys on how to make the non-trainable params very small and the trainable params very large.\n\nThanks.\n\nNote:\n(1) I used test_size=20% for train_test_split\n(2) I used ModelCheckPoint and EarlyStopping",
      "votes": null
    },
    {
      "id": "1219355",
      "postDate": "02/26/2021 18:29:45",
      "content": "<p>I think we need a bit more context here. It sounds a bit like you are doing transfer learning with a pre-trained model (I'm asssuming that it's a convolutional neural network) and have only the final layer(s) of the model unfrozen? Is that correct? </p>\n<p>If so and if the pre-training was an e.g. ImageNet, then I'd guess that you'd indeed fail to get a competitive performance on the leaderboard. This may look different, if you either first optimize the final few layers, then unfreeze the remainder of the network and then fine-tune that, too. Another case, in which this may be less of a problem, is if your frozen layers already have the right representation of these types of images (that might be the case, if the model has been pre-trained on a dataset of chest x-rays) - it's not like 28,171 is a small number of parameters in a sense… However, even then fine-tuning at the end will presumably help a little.</p>",
      "rawMarkdown": "I think we need a bit more context here. It sounds a bit like you are doing transfer learning with a pre-trained model (I'm asssuming that it's a convolutional neural network) and have only the final layer(s) of the model unfrozen? Is that correct? \n\nIf so and if the pre-training was an e.g. ImageNet, then I'd guess that you'd indeed fail to get a competitive performance on the leaderboard. This may look different, if you either first optimize the final few layers, then unfreeze the remainder of the network and then fine-tune that, too. Another case, in which this may be less of a problem, is if your frozen layers already have the right representation of these types of images (that might be the case, if the model has been pre-trained on a dataset of chest x-rays) - it's not like 28,171 is a small number of parameters in a sense... However, even then fine-tuning at the end will presumably help a little.",
      "votes": null
    },
    {
      "id": "1219923",
      "postDate": "02/27/2021 11:12:23",
      "content": "<p>Thanks a lot.</p>\n<p>I am using CNN pre-trained on EfficientNetB7. However, I don't understand what you meant by freezing and unfreezing the final (or other) layers of the model. How can I unfreeze all layers (I am assuming this is the right thing to do)? Also, how can I know if some layers are frozen?</p>\n<p>This is my code:</p>\n<p>`def create_model(input_img, img_weights):<br>\n    base_model = efn.EfficientNetB7(<br>\n        include_top=False,<br>\n        weights=image_weights,<br>\n        input_shape=input_img<br>\n    )<br>\n    base_model.trainable=False</p>\n<pre><code>model = tf.keras.Sequential([\n    base_model,\n    layers.GlobalAveragePooling2D(),\n    layers.Dense(11, activation='sigmoid')\n])\nreturn model`\n</code></pre>\n<p>Thanks once again.</p>",
      "rawMarkdown": "Thanks a lot.\n\nI am using CNN pre-trained on EfficientNetB7. However, I don't understand what you meant by freezing and unfreezing the final (or other) layers of the model. How can I unfreeze all layers (I am assuming this is the right thing to do)? Also, how can I know if some layers are frozen?\n\nThis is my code:\n\n`def create_model(input_img, img_weights):\n    base_model = efn.EfficientNetB7(\n        include_top=False,\n        weights=image_weights,\n        input_shape=input_img\n    )\n    base_model.trainable=False\n\n    model = tf.keras.Sequential([\n        base_model,\n        layers.GlobalAveragePooling2D(),\n        layers.Dense(11, activation='sigmoid')\n    ])\n    return model`\n\nThanks once again.",
      "votes": null
    },
    {
      "id": "1220308",
      "postDate": "02/27/2021 20:42:17",
      "content": "<p>You should have more trainable params. With that many params frozen you likely cant get the model to do much other than what it was already trained for. Should be as simple as model.trainable = True or iterating through the layers and setting each layer to be trainable. </p>",
      "rawMarkdown": "You should have more trainable params. With that many params frozen you likely cant get the model to do much other than what it was already trained for. Should be as simple as model.trainable = True or iterating through the layers and setting each layer to be trainable.",
      "votes": null
    },
    {
      "id": "1220377",
      "postDate": "02/27/2021 22:43:17",
      "content": "<p><code>base_model.trainable=False</code> seems to specify that you only train the top layers you've added on top of a model pre-trained on ImageNet, but not most of those pre-trained layers (these are especially the layers with all the convolutional filters). A typical strategy is to first keep the lower layers frozen, as you have, then after some training, you unfreeze those and train all parameters for some more epochs. A great discussion of that can be found e.g. in the book by Jeremy Howard and Sylvain Gugger (you can take a peek via their online repository for the book e.g. <a href=\"https://github.com/fastai/fastbook/blob/master/01_intro.ipynb\" target=\"_blank\">chapter 1</a> discusses some of that, I love my paper copy of that book - if you prefer videos, their great course is <a href=\"https://course.fast.ai/\" target=\"_blank\">online</a> - you get some of this stuff around 1 hour in the first lesson). Of course, François Chollet's <a href=\"https://www.manning.com/books/deep-learning-with-python-second-edition\" target=\"_blank\">book</a> is also fantastic and may be a better fit for you, if you think you'd like to stick with keras/TensorFlow.</p>",
      "rawMarkdown": "`base_model.trainable=False` seems to specify that you only train the top layers you've added on top of a model pre-trained on ImageNet, but not most of those pre-trained layers (these are especially the layers with all the convolutional filters). A typical strategy is to first keep the lower layers frozen, as you have, then after some training, you unfreeze those and train all parameters for some more epochs. A great discussion of that can be found e.g. in the book by Jeremy Howard and Sylvain Gugger (you can take a peek via their online repository for the book e.g. [chapter 1](https://github.com/fastai/fastbook/blob/master/01_intro.ipynb) discusses some of that, I love my paper copy of that book - if you prefer videos, their great course is [online](https://course.fast.ai/) - you get some of this stuff around 1 hour in the first lesson). Of course, François Chollet's [book](https://www.manning.com/books/deep-learning-with-python-second-edition) is also fantastic and may be a better fit for you, if you think you'd like to stick with keras/TensorFlow.",
      "votes": null
    },
    {
      "id": "1220967",
      "postDate": "02/28/2021 14:56:37",
      "content": "<p>Thanks a lot.</p>",
      "rawMarkdown": "Thanks a lot.",
      "votes": null
    },
    {
      "id": "1220968",
      "postDate": "02/28/2021 14:56:55",
      "content": "<p>Thanks a lot.</p>",
      "rawMarkdown": "Thanks a lot.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1219355,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "02/26/2021 18:29:45",
      "content": "<p>I think we need a bit more context here. It sounds a bit like you are doing transfer learning with a pre-trained model (I'm asssuming that it's a convolutional neural network) and have only the final layer(s) of the model unfrozen? Is that correct? </p>\n<p>If so and if the pre-training was an e.g. ImageNet, then I'd guess that you'd indeed fail to get a competitive performance on the leaderboard. This may look different, if you either first optimize the final few layers, then unfreeze the remainder of the network and then fine-tune that, too. Another case, in which this may be less of a problem, is if your frozen layers already have the right representation of these types of images (that might be the case, if the model has been pre-trained on a dataset of chest x-rays) - it's not like 28,171 is a small number of parameters in a sense… However, even then fine-tuning at the end will presumably help a little.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1219923,
          "author_name": "zammie",
          "author_url": "",
          "post_date": "02/27/2021 11:12:23",
          "content": "<p>Thanks a lot.</p>\n<p>I am using CNN pre-trained on EfficientNetB7. However, I don't understand what you meant by freezing and unfreezing the final (or other) layers of the model. How can I unfreeze all layers (I am assuming this is the right thing to do)? Also, how can I know if some layers are frozen?</p>\n<p>This is my code:</p>\n<p>`def create_model(input_img, img_weights):<br>\n    base_model = efn.EfficientNetB7(<br>\n        include_top=False,<br>\n        weights=image_weights,<br>\n        input_shape=input_img<br>\n    )<br>\n    base_model.trainable=False</p>\n<pre><code>model = tf.keras.Sequential([\n    base_model,\n    layers.GlobalAveragePooling2D(),\n    layers.Dense(11, activation='sigmoid')\n])\nreturn model`\n</code></pre>\n<p>Thanks once again.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1220377,
          "author_name": "bjoernholzhauer",
          "author_url": "",
          "post_date": "02/27/2021 22:43:17",
          "content": "<p><code>base_model.trainable=False</code> seems to specify that you only train the top layers you've added on top of a model pre-trained on ImageNet, but not most of those pre-trained layers (these are especially the layers with all the convolutional filters). A typical strategy is to first keep the lower layers frozen, as you have, then after some training, you unfreeze those and train all parameters for some more epochs. A great discussion of that can be found e.g. in the book by Jeremy Howard and Sylvain Gugger (you can take a peek via their online repository for the book e.g. <a href=\"https://github.com/fastai/fastbook/blob/master/01_intro.ipynb\" target=\"_blank\">chapter 1</a> discusses some of that, I love my paper copy of that book - if you prefer videos, their great course is <a href=\"https://course.fast.ai/\" target=\"_blank\">online</a> - you get some of this stuff around 1 hour in the first lesson). Of course, François Chollet's <a href=\"https://www.manning.com/books/deep-learning-with-python-second-edition\" target=\"_blank\">book</a> is also fantastic and may be a better fit for you, if you think you'd like to stick with keras/TensorFlow.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1220968,
          "author_name": "zammie",
          "author_url": "",
          "post_date": "02/28/2021 14:56:55",
          "content": "<p>Thanks a lot.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1220308,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "02/27/2021 20:42:17",
      "content": "<p>You should have more trainable params. With that many params frozen you likely cant get the model to do much other than what it was already trained for. Should be as simple as model.trainable = True or iterating through the layers and setting each layer to be trainable. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1220967,
          "author_name": "zammie",
          "author_url": "",
          "post_date": "02/28/2021 14:56:37",
          "content": "<p>Thanks a lot.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1219331": "Total Params: 64,125,851\nTrainable Params: 28,171\nNon-Trainable Params: 64,097,680\n\nI am using CNN pre-trained on EfficientNetB7. My guess is my model will not be able to make full sense of the data since it can only train on 28,171 params. If my guess is correct, then I will welcome suggestions from you guys on how to make the non-trainable params very small and the trainable params very large.\n\nThanks.\n\nNote:\n(1) I used test_size=20% for train_test_split\n(2) I used ModelCheckPoint and EarlyStopping",
    "1219355": "I think we need a bit more context here. It sounds a bit like you are doing transfer learning with a pre-trained model (I'm asssuming that it's a convolutional neural network) and have only the final layer(s) of the model unfrozen? Is that correct? \n\nIf so and if the pre-training was an e.g. ImageNet, then I'd guess that you'd indeed fail to get a competitive performance on the leaderboard. This may look different, if you either first optimize the final few layers, then unfreeze the remainder of the network and then fine-tune that, too. Another case, in which this may be less of a problem, is if your frozen layers already have the right representation of these types of images (that might be the case, if the model has been pre-trained on a dataset of chest x-rays) - it's not like 28,171 is a small number of parameters in a sense... However, even then fine-tuning at the end will presumably help a little.",
    "1219923": "Thanks a lot.\n\nI am using CNN pre-trained on EfficientNetB7. However, I don't understand what you meant by freezing and unfreezing the final (or other) layers of the model. How can I unfreeze all layers (I am assuming this is the right thing to do)? Also, how can I know if some layers are frozen?\n\nThis is my code:\n\n`def create_model(input_img, img_weights):\n    base_model = efn.EfficientNetB7(\n        include_top=False,\n        weights=image_weights,\n        input_shape=input_img\n    )\n    base_model.trainable=False\n\n    model = tf.keras.Sequential([\n        base_model,\n        layers.GlobalAveragePooling2D(),\n        layers.Dense(11, activation='sigmoid')\n    ])\n    return model`\n\nThanks once again.",
    "1220308": "You should have more trainable params. With that many params frozen you likely cant get the model to do much other than what it was already trained for. Should be as simple as model.trainable = True or iterating through the layers and setting each layer to be trainable.",
    "1220377": "`base_model.trainable=False` seems to specify that you only train the top layers you've added on top of a model pre-trained on ImageNet, but not most of those pre-trained layers (these are especially the layers with all the convolutional filters). A typical strategy is to first keep the lower layers frozen, as you have, then after some training, you unfreeze those and train all parameters for some more epochs. A great discussion of that can be found e.g. in the book by Jeremy Howard and Sylvain Gugger (you can take a peek via their online repository for the book e.g. [chapter 1](https://github.com/fastai/fastbook/blob/master/01_intro.ipynb) discusses some of that, I love my paper copy of that book - if you prefer videos, their great course is [online](https://course.fast.ai/) - you get some of this stuff around 1 hour in the first lesson). Of course, François Chollet's [book](https://www.manning.com/books/deep-learning-with-python-second-edition) is also fantastic and may be a better fit for you, if you think you'd like to stick with keras/TensorFlow.",
    "1220967": "Thanks a lot.",
    "1220968": "Thanks a lot."
  },
  "source": "meta"
}