{
  "id": 221556,
  "title": "Deciding number of Trainable parameters 😬",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/221556",
  "author_name": "",
  "post_date": "2021-02-23T06:43:29.364389200Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>First of all, I would like to thank you to Kaggle for organizing this competition and the great Kaggle community also my teammates <a href=\"https://www.kaggle.com/harshwardhanbhangale\" target=\"_blank\">@harshwardhanbhangale</a>  &amp; <a href=\"https://www.kaggle.com/mohit13gidwani\" target=\"_blank\">@mohit13gidwani</a> </p>\n<p>We learned a lot during the competition starting from different backbones to the different loss function LR scheduling etc but would like to share one mistake done by us to which we were stuck for a while also</p>\n<p>Generally, people were connecting outputs of the last layer directly to number of classes but we were adding an additional linear layer of 256 neurons </p>\n<p>Due to which our output probabilities were fluctuating a lot leading to poor scores.<br>\nnumber of parameters with additional layer - 460k<br>\nnumber of parameters without additional layer - 9k (better results)</p>\n<p>We're not sure but due to the less amount of data, the architecture with additional layer is not able to train properly. We shared this for all the newbies which might have faced this problem or can avoid such issues in future</p>\n<p>You can also mention your thoughts in the comments </p>",
  "messages": [
    {
      "id": "1214849",
      "postDate": "02/23/2021 06:43:29",
      "content": "<p>First of all, I would like to thank you to Kaggle for organizing this competition and the great Kaggle community also my teammates <a href=\"https://www.kaggle.com/harshwardhanbhangale\" target=\"_blank\">@harshwardhanbhangale</a>  &amp; <a href=\"https://www.kaggle.com/mohit13gidwani\" target=\"_blank\">@mohit13gidwani</a> </p>\n<p>We learned a lot during the competition starting from different backbones to the different loss function LR scheduling etc but would like to share one mistake done by us to which we were stuck for a while also</p>\n<p>Generally, people were connecting outputs of the last layer directly to number of classes but we were adding an additional linear layer of 256 neurons </p>\n<p>Due to which our output probabilities were fluctuating a lot leading to poor scores.<br>\nnumber of parameters with additional layer - 460k<br>\nnumber of parameters without additional layer - 9k (better results)</p>\n<p>We're not sure but due to the less amount of data, the architecture with additional layer is not able to train properly. We shared this for all the newbies which might have faced this problem or can avoid such issues in future</p>\n<p>You can also mention your thoughts in the comments </p>",
      "rawMarkdown": "First of all, I would like to thank you to Kaggle for organizing this competition and the great Kaggle community also my teammates @harshwardhanbhangale  & @mohit13gidwani \n\nWe learned a lot during the competition starting from different backbones to the different loss function LR scheduling etc but would like to share one mistake done by us to which we were stuck for a while also\n\nGenerally, people were connecting outputs of the last layer directly to number of classes but we were adding an additional linear layer of 256 neurons \n\nDue to which our output probabilities were fluctuating a lot leading to poor scores.\nnumber of parameters with additional layer - 460k\nnumber of parameters without additional layer - 9k (better results)\n\nWe're not sure but due to the less amount of data, the architecture with additional layer is not able to train properly. We shared this for all the newbies which might have faced this problem or can avoid such issues in future\n\nYou can also mention your thoughts in the comments",
      "votes": null
    },
    {
      "id": "1215009",
      "postDate": "02/23/2021 09:04:32",
      "content": "<p>That's an interesting point and having more parameters does indeed more risk of overfitting - if one does not regularize properly (which may require extra experimentation). However, note that another observation is that very often two extra layers <strong>does</strong> help. E.g. one of the options of the custom head of the <code>fastai</code> library (based on their experiments on what seems to help in many cases) is to have two layers. It seems to be mostly a matter of the right size and, of course, the right regularization (e.g. dropout). It did not make too much of a difference either way for me in this competition (I used dropout on those extra layers; perhaps I could have tweaked those layers a bit more…).</p>",
      "rawMarkdown": "That's an interesting point and having more parameters does indeed more risk of overfitting - if one does not regularize properly (which may require extra experimentation). However, note that another observation is that very often two extra layers **does** help. E.g. one of the options of the custom head of the `fastai` library (based on their experiments on what seems to help in many cases) is to have two layers. It seems to be mostly a matter of the right size and, of course, the right regularization (e.g. dropout). It did not make too much of a difference either way for me in this competition (I used dropout on those extra layers; perhaps I could have tweaked those layers a bit more...).",
      "votes": null
    },
    {
      "id": "1215092",
      "postDate": "02/23/2021 10:49:08",
      "content": "<p>I did something like this in the efficient net b4, maybe I did not set the number of neurons properly, which led to a drop in the generalization capability on the final test set. However on the validation data I had very good results…</p>\n<p><a href=\"https://ibb.co/khwc8Jp\"><img src=\"https://i.ibb.co/RCsH0N1/te.png\" alt=\"te\"></a></p>",
      "rawMarkdown": "I did something like this in the efficient net b4, maybe I did not set the number of neurons properly, which led to a drop in the generalization capability on the final test set. However on the validation data I had very good results...\n\n<a href=\"https://ibb.co/khwc8Jp\"><img src=\"https://i.ibb.co/RCsH0N1/te.png\" alt=\"te\" border=\"0\"></a>",
      "votes": null
    },
    {
      "id": "1215442",
      "postDate": "02/23/2021 15:56:48",
      "content": "<p>Perhaps a BatchNorm and a dropout before the first fully connected layer of head would help? E.g. <strong>some kind of pooling</strong> - <strong>flattening</strong> - <strong>BN</strong> - <strong>Dropout</strong> - <strong>Linear</strong> - <strong>ReLU/SiLU/whatever</strong> - <strong>BN</strong> - <strong>Dropout</strong> - <strong>Linear (output = number of classes)</strong>? E.g. the book chapter 15 of the book by the fast.ai team (see the <a href=\"https://github.com/fastai/fastbook/blob/master/15_arch_details.ipynb\" target=\"_blank\">github for the book</a>), which I already mentioned below, describes that order.</p>\n<p>However, it looks to me like e.g. the EfficientNet-B0 seems to just do <strong>Pooling</strong> - <strong>FC</strong> (see <a href=\"https://arxiv.org/abs/1905.11946\" target=\"_blank\">arXiv paper</a>).</p>",
      "rawMarkdown": "Perhaps a BatchNorm and a dropout before the first fully connected layer of head would help? E.g. **some kind of pooling** - **flattening** - **BN** - **Dropout** - **Linear** - **ReLU/SiLU/whatever** - **BN** - **Dropout** - **Linear (output = number of classes)**? E.g. the book chapter 15 of the book by the fast.ai team (see the [github for the book](https://github.com/fastai/fastbook/blob/master/15_arch_details.ipynb)), which I already mentioned below, describes that order.\n\nHowever, it looks to me like e.g. the EfficientNet-B0 seems to just do **Pooling** - **FC** (see [arXiv paper](https://arxiv.org/abs/1905.11946)).",
      "votes": null
    },
    {
      "id": "1215462",
      "postDate": "02/23/2021 16:20:37",
      "content": "<p>Nice, ill try that idea for sure. Tnx man, i will read that book chapter, coz their layer order might help in other future competitions. </p>",
      "rawMarkdown": "Nice, ill try that idea for sure. Tnx man, i will read that book chapter, coz their layer order might help in other future competitions.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1215009,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "02/23/2021 09:04:32",
      "content": "<p>That's an interesting point and having more parameters does indeed more risk of overfitting - if one does not regularize properly (which may require extra experimentation). However, note that another observation is that very often two extra layers <strong>does</strong> help. E.g. one of the options of the custom head of the <code>fastai</code> library (based on their experiments on what seems to help in many cases) is to have two layers. It seems to be mostly a matter of the right size and, of course, the right regularization (e.g. dropout). It did not make too much of a difference either way for me in this competition (I used dropout on those extra layers; perhaps I could have tweaked those layers a bit more…).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1215092,
      "author_name": "marjan1111",
      "author_url": "",
      "post_date": "02/23/2021 10:49:08",
      "content": "<p>I did something like this in the efficient net b4, maybe I did not set the number of neurons properly, which led to a drop in the generalization capability on the final test set. However on the validation data I had very good results…</p>\n<p><a href=\"https://ibb.co/khwc8Jp\"><img src=\"https://i.ibb.co/RCsH0N1/te.png\" alt=\"te\"></a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1215442,
          "author_name": "bjoernholzhauer",
          "author_url": "",
          "post_date": "02/23/2021 15:56:48",
          "content": "<p>Perhaps a BatchNorm and a dropout before the first fully connected layer of head would help? E.g. <strong>some kind of pooling</strong> - <strong>flattening</strong> - <strong>BN</strong> - <strong>Dropout</strong> - <strong>Linear</strong> - <strong>ReLU/SiLU/whatever</strong> - <strong>BN</strong> - <strong>Dropout</strong> - <strong>Linear (output = number of classes)</strong>? E.g. the book chapter 15 of the book by the fast.ai team (see the <a href=\"https://github.com/fastai/fastbook/blob/master/15_arch_details.ipynb\" target=\"_blank\">github for the book</a>), which I already mentioned below, describes that order.</p>\n<p>However, it looks to me like e.g. the EfficientNet-B0 seems to just do <strong>Pooling</strong> - <strong>FC</strong> (see <a href=\"https://arxiv.org/abs/1905.11946\" target=\"_blank\">arXiv paper</a>).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1215462,
          "author_name": "marjan1111",
          "author_url": "",
          "post_date": "02/23/2021 16:20:37",
          "content": "<p>Nice, ill try that idea for sure. Tnx man, i will read that book chapter, coz their layer order might help in other future competitions. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1214849": "First of all, I would like to thank you to Kaggle for organizing this competition and the great Kaggle community also my teammates @harshwardhanbhangale  & @mohit13gidwani \n\nWe learned a lot during the competition starting from different backbones to the different loss function LR scheduling etc but would like to share one mistake done by us to which we were stuck for a while also\n\nGenerally, people were connecting outputs of the last layer directly to number of classes but we were adding an additional linear layer of 256 neurons \n\nDue to which our output probabilities were fluctuating a lot leading to poor scores.\nnumber of parameters with additional layer - 460k\nnumber of parameters without additional layer - 9k (better results)\n\nWe're not sure but due to the less amount of data, the architecture with additional layer is not able to train properly. We shared this for all the newbies which might have faced this problem or can avoid such issues in future\n\nYou can also mention your thoughts in the comments",
    "1215009": "That's an interesting point and having more parameters does indeed more risk of overfitting - if one does not regularize properly (which may require extra experimentation). However, note that another observation is that very often two extra layers **does** help. E.g. one of the options of the custom head of the `fastai` library (based on their experiments on what seems to help in many cases) is to have two layers. It seems to be mostly a matter of the right size and, of course, the right regularization (e.g. dropout). It did not make too much of a difference either way for me in this competition (I used dropout on those extra layers; perhaps I could have tweaked those layers a bit more...).",
    "1215092": "I did something like this in the efficient net b4, maybe I did not set the number of neurons properly, which led to a drop in the generalization capability on the final test set. However on the validation data I had very good results...\n\n<a href=\"https://ibb.co/khwc8Jp\"><img src=\"https://i.ibb.co/RCsH0N1/te.png\" alt=\"te\" border=\"0\"></a>",
    "1215442": "Perhaps a BatchNorm and a dropout before the first fully connected layer of head would help? E.g. **some kind of pooling** - **flattening** - **BN** - **Dropout** - **Linear** - **ReLU/SiLU/whatever** - **BN** - **Dropout** - **Linear (output = number of classes)**? E.g. the book chapter 15 of the book by the fast.ai team (see the [github for the book](https://github.com/fastai/fastbook/blob/master/15_arch_details.ipynb)), which I already mentioned below, describes that order.\n\nHowever, it looks to me like e.g. the EfficientNet-B0 seems to just do **Pooling** - **FC** (see [arXiv paper](https://arxiv.org/abs/1905.11946)).",
    "1215462": "Nice, ill try that idea for sure. Tnx man, i will read that book chapter, coz their layer order might help in other future competitions."
  },
  "source": "meta"
}