{
  "id": 217601,
  "title": "Adding layers after pretrained model",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/217601",
  "author_name": "",
  "post_date": "2021-02-07T14:00:36.839226700Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I have two doubts regarding adding layers above pretrained model.</p>\n<p>1: Will adding the dropout layer directly below Average pooling layer increase the accuracy (in general)<br>\n2: Will adding a FC layer, followed by a dropout layer after Average pooling layer increase the accuracy (in general)</p>\n<p>I know these are basic doubts, But help me with this.<br>\nThanks in advance.</p>",
  "messages": [
    {
      "id": "1190122",
      "postDate": "02/07/2021 14:00:36",
      "content": "<p>I have two doubts regarding adding layers above pretrained model.</p>\n<p>1: Will adding the dropout layer directly below Average pooling layer increase the accuracy (in general)<br>\n2: Will adding a FC layer, followed by a dropout layer after Average pooling layer increase the accuracy (in general)</p>\n<p>I know these are basic doubts, But help me with this.<br>\nThanks in advance.</p>",
      "rawMarkdown": "I have two doubts regarding adding layers above pretrained model.\n\n1: Will adding the dropout layer directly below Average pooling layer increase the accuracy (in general)\n2: Will adding a FC layer, followed by a dropout layer after Average pooling layer increase the accuracy (in general)\n\nI know these are basic doubts, But help me with this.\nThanks in advance.",
      "votes": null
    },
    {
      "id": "1190166",
      "postDate": "02/07/2021 14:32:03",
      "content": "<p>A lot of these things are not general (i.e. you can always find cases where this may turn out to be different) and you need to experiment. On the other hand, others have done a decent amount of experimentation on this and noted what typically works. E.g. Jeremy Howard and Sylvain Gugger in <a href=\"https://www.amazon.com/Deep-Learning-Coders-fastai-PyTorch/dp/1492045527\" target=\"_blank\">their book</a> (see also the <a href=\"https://github.com/fastai/fastbook/blob/master/15_arch_details.ipynb\" target=\"_blank\">repository for the book</a>, the lectures are of course also great to watch, but I'm not sure where it gets discussed) recommend this head that is in the fastai library as a default: <br>\nAdaptiveConcatPool2d - Flatten - BatchNorm1d - Dropout - FC - ReLU - BatchNorm1d - Dropout - FC (output). The AdaptiveConcatPool is a particular alternative to average pooling.</p>\n<p>Obviously, you'd need some kind of pooling and flattening the tensor, and at the end at least one fully connected layer at the end to output your predicted probabilities (or non-softmaxed outputs). What happens inbetween is somewhat optional and something you can experiment with.</p>",
      "rawMarkdown": "A lot of these things are not general (i.e. you can always find cases where this may turn out to be different) and you need to experiment. On the other hand, others have done a decent amount of experimentation on this and noted what typically works. E.g. Jeremy Howard and Sylvain Gugger in [their book](https://www.amazon.com/Deep-Learning-Coders-fastai-PyTorch/dp/1492045527) (see also the [repository for the book](https://github.com/fastai/fastbook/blob/master/15_arch_details.ipynb), the lectures are of course also great to watch, but I'm not sure where it gets discussed) recommend this head that is in the fastai library as a default: \nAdaptiveConcatPool2d - Flatten - BatchNorm1d - Dropout - FC - ReLU - BatchNorm1d - Dropout - FC (output). The AdaptiveConcatPool is a particular alternative to average pooling.\n\nObviously, you'd need some kind of pooling and flattening the tensor, and at the end at least one fully connected layer at the end to output your predicted probabilities (or non-softmaxed outputs). What happens inbetween is somewhat optional and something you can experiment with.",
      "votes": null
    },
    {
      "id": "1191033",
      "postDate": "02/08/2021 07:55:56",
      "content": "<p>Well everything is still experimental in Deep learning, I can't be certain that adding more layers will add something to the model and increase its performance, and I can't be sure when changing the learning rate that the model will converge on an optimum minima, nevertheless we still have rules for everything. For instance, we use the CNNs for image classification and lately there are papers for transformers in computer vision task. We follow the conv layer by an activation function and a normalization step if needed, we put the average pooling layer to reduce the feature maps and reduce the heavy computations, we add the dropout layer to avoid overfitting. The rules are always their as your guide but there is no one rule for performance, yeah sometimes adding the conv layers extracting more features help the model finds a better solution, but sometimes the model performance decreases when adding more layers, so it's experimental in the end, happy learning and good luck converging your model</p>",
      "rawMarkdown": "Well everything is still experimental in Deep learning, I can't be certain that adding more layers will add something to the model and increase its performance, and I can't be sure when changing the learning rate that the model will converge on an optimum minima, nevertheless we still have rules for everything. For instance, we use the CNNs for image classification and lately there are papers for transformers in computer vision task. We follow the conv layer by an activation function and a normalization step if needed, we put the average pooling layer to reduce the feature maps and reduce the heavy computations, we add the dropout layer to avoid overfitting. The rules are always their as your guide but there is no one rule for performance, yeah sometimes adding the conv layers extracting more features help the model finds a better solution, but sometimes the model performance decreases when adding more layers, so it's experimental in the end, happy learning and good luck converging your model",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1190166,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "02/07/2021 14:32:03",
      "content": "<p>A lot of these things are not general (i.e. you can always find cases where this may turn out to be different) and you need to experiment. On the other hand, others have done a decent amount of experimentation on this and noted what typically works. E.g. Jeremy Howard and Sylvain Gugger in <a href=\"https://www.amazon.com/Deep-Learning-Coders-fastai-PyTorch/dp/1492045527\" target=\"_blank\">their book</a> (see also the <a href=\"https://github.com/fastai/fastbook/blob/master/15_arch_details.ipynb\" target=\"_blank\">repository for the book</a>, the lectures are of course also great to watch, but I'm not sure where it gets discussed) recommend this head that is in the fastai library as a default: <br>\nAdaptiveConcatPool2d - Flatten - BatchNorm1d - Dropout - FC - ReLU - BatchNorm1d - Dropout - FC (output). The AdaptiveConcatPool is a particular alternative to average pooling.</p>\n<p>Obviously, you'd need some kind of pooling and flattening the tensor, and at the end at least one fully connected layer at the end to output your predicted probabilities (or non-softmaxed outputs). What happens inbetween is somewhat optional and something you can experiment with.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1191033,
      "author_name": "omarmohamed22",
      "author_url": "",
      "post_date": "02/08/2021 07:55:56",
      "content": "<p>Well everything is still experimental in Deep learning, I can't be certain that adding more layers will add something to the model and increase its performance, and I can't be sure when changing the learning rate that the model will converge on an optimum minima, nevertheless we still have rules for everything. For instance, we use the CNNs for image classification and lately there are papers for transformers in computer vision task. We follow the conv layer by an activation function and a normalization step if needed, we put the average pooling layer to reduce the feature maps and reduce the heavy computations, we add the dropout layer to avoid overfitting. The rules are always their as your guide but there is no one rule for performance, yeah sometimes adding the conv layers extracting more features help the model finds a better solution, but sometimes the model performance decreases when adding more layers, so it's experimental in the end, happy learning and good luck converging your model</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1190122": "I have two doubts regarding adding layers above pretrained model.\n\n1: Will adding the dropout layer directly below Average pooling layer increase the accuracy (in general)\n2: Will adding a FC layer, followed by a dropout layer after Average pooling layer increase the accuracy (in general)\n\nI know these are basic doubts, But help me with this.\nThanks in advance.",
    "1190166": "A lot of these things are not general (i.e. you can always find cases where this may turn out to be different) and you need to experiment. On the other hand, others have done a decent amount of experimentation on this and noted what typically works. E.g. Jeremy Howard and Sylvain Gugger in [their book](https://www.amazon.com/Deep-Learning-Coders-fastai-PyTorch/dp/1492045527) (see also the [repository for the book](https://github.com/fastai/fastbook/blob/master/15_arch_details.ipynb), the lectures are of course also great to watch, but I'm not sure where it gets discussed) recommend this head that is in the fastai library as a default: \nAdaptiveConcatPool2d - Flatten - BatchNorm1d - Dropout - FC - ReLU - BatchNorm1d - Dropout - FC (output). The AdaptiveConcatPool is a particular alternative to average pooling.\n\nObviously, you'd need some kind of pooling and flattening the tensor, and at the end at least one fully connected layer at the end to output your predicted probabilities (or non-softmaxed outputs). What happens inbetween is somewhat optional and something you can experiment with.",
    "1191033": "Well everything is still experimental in Deep learning, I can't be certain that adding more layers will add something to the model and increase its performance, and I can't be sure when changing the learning rate that the model will converge on an optimum minima, nevertheless we still have rules for everything. For instance, we use the CNNs for image classification and lately there are papers for transformers in computer vision task. We follow the conv layer by an activation function and a normalization step if needed, we put the average pooling layer to reduce the feature maps and reduce the heavy computations, we add the dropout layer to avoid overfitting. The rules are always their as your guide but there is no one rule for performance, yeah sometimes adding the conv layers extracting more features help the model finds a better solution, but sometimes the model performance decreases when adding more layers, so it's experimental in the end, happy learning and good luck converging your model"
  },
  "source": "meta"
}