{
  "id": 214018,
  "title": "Model Overfitting",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/214018",
  "author_name": "",
  "post_date": "2021-01-25T02:58:43.399065900Z",
  "votes": 5,
  "comment_count": 15,
  "views": 0,
  "content": "<p>While finetuning model by making last 20% of layers trainable, model is <strong>overfitting</strong>. Any suggestion on proper way of doing finetuning? <a href=\"https://www.kaggle.com/shubham219/model-ensembling-with-k-fold\" target=\"_blank\">My Training Model</a></p>",
  "messages": [
    {
      "id": "1168526",
      "postDate": "01/25/2021 02:58:43",
      "content": "<p>While finetuning model by making last 20% of layers trainable, model is <strong>overfitting</strong>. Any suggestion on proper way of doing finetuning? <a href=\"https://www.kaggle.com/shubham219/model-ensembling-with-k-fold\" target=\"_blank\">My Training Model</a></p>",
      "rawMarkdown": "While finetuning model by making last 20% of layers trainable, model is **overfitting**. Any suggestion on proper way of doing finetuning? [My Training Model](https://www.kaggle.com/shubham219/model-ensembling-with-k-fold)",
      "votes": null
    },
    {
      "id": "1168553",
      "postDate": "01/25/2021 03:48:59",
      "content": "<p>Will fork your kernel on my local PC and try a few things - will report back if I find something.</p>",
      "rawMarkdown": "Will fork your kernel on my local PC and try a few things - will report back if I find something.",
      "votes": null
    },
    {
      "id": "1168691",
      "postDate": "01/25/2021 05:36:25",
      "content": "<p>Thanks..and please suggest me something to improve. This is my first time in computer Vision.<br>\nI am trying to do Regularization now.</p>",
      "rawMarkdown": "Thanks..and please suggest me something to improve. This is my first time in computer Vision.\nI am trying to do Regularization now.",
      "votes": null
    },
    {
      "id": "1168828",
      "postDate": "01/25/2021 07:34:55",
      "content": "<p>Have the kernel forked and running on one of my machines.  Often bringing shared kernels to local machine is a pain because folder assignments - your code nicely written and was pretty easy to get running.  I got a little confused by the references to three different forms of the tf records - but figured those were the residue of earlier experiments. </p>\n<p>I did add a learning rate scheduler before I started the run.  Getting the learning rate right is one of the early keys (IMO) to finding a model that works.  I will be using a different schedule for frozen vs not frozen fits.  I did not feel I had a lot of success on this competition with my efforts to run a frozen and not frozen fit.  With a decent learning rate scheduler I have been doing most of my experiments with a single fit - it takes twice as long to do a frozen/unfrozen and I was not getting a big bang for the compute time.  So my current approach is to not fine tune until the last week or two of a competition when I have finished experiments with most of the things I know how to do and have a model I like.</p>\n<p>Added a bit to the early stopping patience and set epochs to high number.  When using a learning rate scheduler you want a robust change in the learning rate to have occurred before you early stop.  That's pretty hard to do on kaggle with the 9 hour time limit.</p>\n<p>Every model I have created in this competition are over fit when looking at the train vs validation plots.  </p>\n<p>Just looking at your kernel and your results I think its too early for you to be worrying about over - fit.   Your training accuracy looks like it needs to fixed first - I may think different tomorrow when the training is done - but until your training accuracy is above 0.90 efforts on reducing over-train may be the wrong path.  Until my training accuracy is higher than the current LB score I like to focus on doing the things that increase train accuracy/loss with only very minor attempts to fix the gap between train and validation plots.   But that's just my method and folks may disagree.</p>",
      "rawMarkdown": "Have the kernel forked and running on one of my machines.  Often bringing shared kernels to local machine is a pain because folder assignments - your code nicely written and was pretty easy to get running.  I got a little confused by the references to three different forms of the tf records - but figured those were the residue of earlier experiments. \n\nI did add a learning rate scheduler before I started the run.  Getting the learning rate right is one of the early keys (IMO) to finding a model that works.  I will be using a different schedule for frozen vs not frozen fits.  I did not feel I had a lot of success on this competition with my efforts to run a frozen and not frozen fit.  With a decent learning rate scheduler I have been doing most of my experiments with a single fit - it takes twice as long to do a frozen/unfrozen and I was not getting a big bang for the compute time.  So my current approach is to not fine tune until the last week or two of a competition when I have finished experiments with most of the things I know how to do and have a model I like.\n\nAdded a bit to the early stopping patience and set epochs to high number.  When using a learning rate scheduler you want a robust change in the learning rate to have occurred before you early stop.  That's pretty hard to do on kaggle with the 9 hour time limit.\n\nEvery model I have created in this competition are over fit when looking at the train vs validation plots.  \n\nJust looking at your kernel and your results I think its too early for you to be worrying about over - fit.   Your training accuracy looks like it needs to fixed first - I may think different tomorrow when the training is done - but until your training accuracy is above 0.90 efforts on reducing over-train may be the wrong path.  Until my training accuracy is higher than the current LB score I like to focus on doing the things that increase train accuracy/loss with only very minor attempts to fix the gap between train and validation plots.   But that's just my method and folks may disagree.",
      "votes": null
    },
    {
      "id": "1168870",
      "postDate": "01/25/2021 08:04:12",
      "content": "<p>Many Thanks! for the observations. While training on the 90% frozen layers My training accuracy is getting increased to ~91% but same time validation accuracy is getting down. So I am trying to fix that.</p>\n<p>Please share the code if you get any success in improving the accuracy just wanna see new changes. </p>\n<p>! Happy Learning</p>",
      "rawMarkdown": "Many Thanks! for the observations. While training on the 90% frozen layers My training accuracy is getting increased to ~91% but same time validation accuracy is getting down. So I am trying to fix that.\n\nPlease share the code if you get any success in improving the accuracy just wanna see new changes. \n\n! Happy Learning",
      "votes": null
    },
    {
      "id": "1169598",
      "postDate": "01/25/2021 15:48:39",
      "content": "<p>Hello! <br>\nThe thing is even a small convnet I've tried out first (containing only 90K parameters) starts to overfit after ~10 epochs. So it seems overfitting would always be an issue in this competition, but you can consider doing these:</p>\n<ul>\n<li>enlarge your data by adding 2019 competition data: TFRecords are available <a href=\"https://www.kaggle.com/nickuzmenkov/cassava-leaf-disease-merged\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-50-tfrecords-external-512x512\" target=\"_blank\">here</a></li>\n<li>data augmentation: a good starting point on TPU <a href=\"https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96\" target=\"_blank\">is here</a></li>\n<li>as mentioned by <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> learning rate schedule would help to improve your score. I also suggest adding warmup steps and unfreeze the entire model</li>\n<li><code>Dropout</code> layers, as well as <code>kernel_regularizer</code> (<code>l1</code>, <code>l2</code>, and <code>l1-l2</code>) do not seem to help here much, but you can also try them out as long as it is a common regularization technique</li>\n</ul>",
      "rawMarkdown": "Hello! \nThe thing is even a small convnet I've tried out first (containing only 90K parameters) starts to overfit after ~10 epochs. So it seems overfitting would always be an issue in this competition, but you can consider doing these:\n* enlarge your data by adding 2019 competition data: TFRecords are available [here](https://www.kaggle.com/nickuzmenkov/cassava-leaf-disease-merged) and [here](https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-50-tfrecords-external-512x512)\n* data augmentation: a good starting point on TPU [is here](https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96)\n* as mentioned by @pcjimmmy learning rate schedule would help to improve your score. I also suggest adding warmup steps and unfreeze the entire model\n* `Dropout` layers, as well as `kernel_regularizer` (`l1`, `l2`, and `l1-l2`) do not seem to help here much, but you can also try them out as long as it is a common regularization technique",
      "votes": null
    },
    {
      "id": "1169693",
      "postDate": "01/25/2021 17:18:23",
      "content": "<p>Kernel still running on my local machine - near the end of the first fold of fine tuning.</p>\n<p>Couple of observations/suggestions for change.</p>\n<p>The distribution of classes very unbalanced - so I would change the KFold to a StratifiedFold.  Little harder to write the <strong>for loop</strong> but IMO you should almost always use StratifiedFold unless the training set is very nicely balanced for classes.</p>\n<p>As you have written the code - your not performing the 5 fold in traditional manner.  Normally you would save the best epoch from each fold (5 total models) and than use those 5 in the prediction.  Your code only saving the best overall epoch and therefore losing some of the power of doing folds.  Combined with use of Kfold this can easily be a source of the part of the over fit you have seen.</p>",
      "rawMarkdown": "Kernel still running on my local machine - near the end of the first fold of fine tuning.\n\nCouple of observations/suggestions for change.\n\nThe distribution of classes very unbalanced - so I would change the KFold to a StratifiedFold.  Little harder to write the **for loop** but IMO you should almost always use StratifiedFold unless the training set is very nicely balanced for classes.\n\nAs you have written the code - your not performing the 5 fold in traditional manner.  Normally you would save the best epoch from each fold (5 total models) and than use those 5 in the prediction.  Your code only saving the best overall epoch and therefore losing some of the power of doing folds.  Combined with use of Kfold this can easily be a source of the part of the over fit you have seen.",
      "votes": null
    },
    {
      "id": "1171289",
      "postDate": "01/26/2021 18:53:38",
      "content": "<p>Kernel completed on my local machine - the LR scheduler and extended number of epochs did help but still a 6% gap between training accuracy and validation accuracy.  If I get a extra available submission over the next day I will submit but I have a local evaluation that I make that suggests a poor LB score.</p>\n<p>I see nothing wrong with how you unfreeze - it's text book - but my experiments with various unfreeze gave me no real help in this competition.</p>\n<p>My suggestions</p>\n<ol>\n<li>StratifiedFold</li>\n<li>Save each fold model and use the 5 models in prediction.</li>\n<li>Add some type of learning rate scheduler - I will post the one I used (with two different settings) in a post shortly.</li>\n<li>Heavy augmentation </li>\n<li>Keep making small changes for dropout and regularization.</li>\n<li>2019 data - I am playing with that now myself but it's only breakeven or slight loss for local CV for me so far.</li>\n</ol>\n<p>A nice summary of other things to try in the post<br>\n<a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/203594\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/203594</a></p>\n<p><strong>Basic question - does a single model with three base pre-trained branches (like your kernel) perform better than three separate models that are than used together to make the predictions?</strong></p>\n<p>I used the three-in-one approach in a past competition with good success.  My approach in this competition is to combine the predictions of multiple models.  Since I am using 4 local machines it's more convenient for me to do the separate models and than make the prediction on Kaggle combining.  </p>",
      "rawMarkdown": "Kernel completed on my local machine - the LR scheduler and extended number of epochs did help but still a 6% gap between training accuracy and validation accuracy.  If I get a extra available submission over the next day I will submit but I have a local evaluation that I make that suggests a poor LB score.\n\nI see nothing wrong with how you unfreeze - it's text book - but my experiments with various unfreeze gave me no real help in this competition.\n\nMy suggestions\n1.  StratifiedFold\n2.  Save each fold model and use the 5 models in prediction.\n3.  Add some type of learning rate scheduler - I will post the one I used (with two different settings) in a post shortly.\n4. Heavy augmentation \n5.  Keep making small changes for dropout and regularization.\n6.  2019 data - I am playing with that now myself but it's only breakeven or slight loss for local CV for me so far.\n\nA nice summary of other things to try in the post\nhttps://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/203594\n\n**Basic question - does a single model with three base pre-trained branches (like your kernel) perform better than three separate models that are than used together to make the predictions?**\n\nI used the three-in-one approach in a past competition with good success.  My approach in this competition is to combine the predictions of multiple models.  Since I am using 4 local machines it's more convenient for me to do the separate models and than make the prediction on Kaggle combining.",
      "votes": null
    },
    {
      "id": "1171291",
      "postDate": "01/26/2021 18:55:50",
      "content": "<p>Used this for LR when frozen. (pasting into these posts makes a mess of the correct format of the code)  This a a very common method used on a number of shared kernels in this and other competitions - not sure who was the original author of the code.</p>\n<p>`LR_START = 0.000005#0.00001<br>\nLR_MAX = 0.001<br>\nLR_MIN = 0.000005<br>\nLR_RAMPUP_EPOCHS = 0<br>\nLR_SUSTAIN_EPOCHS = 0<br>\nLR_EXP_DECAY = .94</p>\n<p>def lrfn(epoch):<br>\n    if epoch &lt; LR_RAMPUP_EPOCHS:<br>\n        lr = (LR_MAX - LR_START) / LR_RAMPUP_EPOCHS * epoch + LR_START<br>\n    elif epoch &lt; LR_RAMPUP_EPOCHS + LR_SUSTAIN_EPOCHS:<br>\n        lr = LR_MAX<br>\n    else:<br>\n        lr = (LR_MAX - LR_MIN) * LR_EXP_DECAY**(epoch - LR_RAMPUP_EPOCHS - LR_SUSTAIN_EPOCHS) + LR_MIN<br>\n    return lr</p>\n<p>lr_callback = tf.keras.callbacks.LearningRateScheduler(lrfn, verbose = True)</p>\n<p>rng = [i for i in range(EPOCHS)]<br>\ny = [lrfn(x) for x in rng]<br>\nplt.plot(rng, y)<br>\nprint(\"Learning rate schedule: {:.3g} to {:.3g} to {:.3g}\".format(y[0], max(y), y[-1]))`</p>",
      "rawMarkdown": "Used this for LR when frozen. (pasting into these posts makes a mess of the correct format of the code)  This a a very common method used on a number of shared kernels in this and other competitions - not sure who was the original author of the code.\n\n`LR_START = 0.000005#0.00001\nLR_MAX = 0.001\nLR_MIN = 0.000005\nLR_RAMPUP_EPOCHS = 0\nLR_SUSTAIN_EPOCHS = 0\nLR_EXP_DECAY = .94\n\ndef lrfn(epoch):\n    if epoch < LR_RAMPUP_EPOCHS:\n        lr = (LR_MAX - LR_START) / LR_RAMPUP_EPOCHS * epoch + LR_START\n    elif epoch < LR_RAMPUP_EPOCHS + LR_SUSTAIN_EPOCHS:\n        lr = LR_MAX\n    else:\n        lr = (LR_MAX - LR_MIN) * LR_EXP_DECAY**(epoch - LR_RAMPUP_EPOCHS - LR_SUSTAIN_EPOCHS) + LR_MIN\n    return lr\n    \nlr_callback = tf.keras.callbacks.LearningRateScheduler(lrfn, verbose = True)\n\nrng = [i for i in range(EPOCHS)]\ny = [lrfn(x) for x in rng]\nplt.plot(rng, y)\nprint(\"Learning rate schedule: {:.3g} to {:.3g} to {:.3g}\".format(y[0], max(y), y[-1]))`",
      "votes": null
    },
    {
      "id": "1171293",
      "postDate": "01/26/2021 18:56:49",
      "content": "<p>Used this changed version for unfrozen</p>\n<p>`LR_START = 0.000005#0.00001<br>\nLR_MAX = 0.0001<br>\nLR_MIN = 0.000005<br>\nLR_RAMPUP_EPOCHS = 4<br>\nLR_SUSTAIN_EPOCHS = 0<br>\nLR_EXP_DECAY = .94</p>\n<p>def lrfn(epoch):<br>\n    if epoch &lt; LR_RAMPUP_EPOCHS:<br>\n        lr = (LR_MAX - LR_START) / LR_RAMPUP_EPOCHS * epoch + LR_START<br>\n    elif epoch &lt; LR_RAMPUP_EPOCHS + LR_SUSTAIN_EPOCHS:<br>\n        lr = LR_MAX<br>\n    else:<br>\n        lr = (LR_MAX - LR_MIN) * LR_EXP_DECAY**(epoch - LR_RAMPUP_EPOCHS - LR_SUSTAIN_EPOCHS) + LR_MIN<br>\n    return lr</p>\n<p>lr_callback = tf.keras.callbacks.LearningRateScheduler(lrfn, verbose = True)</p>\n<p>rng = [i for i in range(EPOCHS)]<br>\ny = [lrfn(x) for x in rng]<br>\nplt.plot(rng, y)<br>\nprint(\"Learning rate schedule: {:.3g} to {:.3g} to {:.3g}\".format(y[0], max(y), y[-1]))`</p>",
      "rawMarkdown": "Used this changed version for unfrozen\n\n`LR_START = 0.000005#0.00001\nLR_MAX = 0.0001\nLR_MIN = 0.000005\nLR_RAMPUP_EPOCHS = 4\nLR_SUSTAIN_EPOCHS = 0\nLR_EXP_DECAY = .94\n\ndef lrfn(epoch):\n    if epoch < LR_RAMPUP_EPOCHS:\n        lr = (LR_MAX - LR_START) / LR_RAMPUP_EPOCHS * epoch + LR_START\n    elif epoch < LR_RAMPUP_EPOCHS + LR_SUSTAIN_EPOCHS:\n        lr = LR_MAX\n    else:\n        lr = (LR_MAX - LR_MIN) * LR_EXP_DECAY**(epoch - LR_RAMPUP_EPOCHS - LR_SUSTAIN_EPOCHS) + LR_MIN\n    return lr\n    \nlr_callback = tf.keras.callbacks.LearningRateScheduler(lrfn, verbose = True)\n\nrng = [i for i in range(EPOCHS)]\ny = [lrfn(x) for x in rng]\nplt.plot(rng, y)\nprint(\"Learning rate schedule: {:.3g} to {:.3g} to {:.3g}\".format(y[0], max(y), y[-1]))`",
      "votes": null
    },
    {
      "id": "1171311",
      "postDate": "01/26/2021 19:11:18",
      "content": "<p>I think it all depends on your training strategie. </p>\n<p>I have a heavy pretrained model which were overfitting (CV = 0.897 and LB 0.989) .  But after changing my training strategie, both CV and LB were improved : CV = 0.901 and LB 0.902).  And I train for larger number of epochs than 10 and don't freeze any layer. </p>",
      "rawMarkdown": "I think it all depends on your training strategie. \n\nI have a heavy pretrained model which were overfitting (CV = 0.897 and LB 0.989) .  But after changing my training strategie, both CV and LB were improved : CV = 0.901 and LB 0.902).  And I train for larger number of epochs than 10 and don't freeze any layer.",
      "votes": null
    },
    {
      "id": "1171373",
      "postDate": "01/26/2021 19:41:36",
      "content": "<p>is that 0.901 with TTA?  and 2019 data?</p>",
      "rawMarkdown": "is that 0.901 with TTA?  and 2019 data?",
      "votes": null
    },
    {
      "id": "1171388",
      "postDate": "01/26/2021 19:52:11",
      "content": "<p>Do you have any suggestions for better regularization?</p>",
      "rawMarkdown": "Do you have any suggestions for better regularization?",
      "votes": null
    },
    {
      "id": "1171457",
      "postDate": "01/26/2021 20:45:41",
      "content": "<p>For LB = 0.902 there is no TTA. </p>\n<p>A part from one Dropout layer (that is present in all my classification pipeline), I don't use any model regularization. All my focus is on the training process.  But I can't disclose the details before the end. </p>",
      "rawMarkdown": "For LB = 0.902 there is no TTA. \n\nA part from one Dropout layer (that is present in all my classification pipeline), I don't use any model regularization. All my focus is on the training process.  But I can't disclose the details before the end.",
      "votes": null
    },
    {
      "id": "1175254",
      "postDate": "01/29/2021 03:40:24",
      "content": "<p>Thank you for sharing, <a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a></p>",
      "rawMarkdown": "Thank you for sharing, @serigne",
      "votes": null
    },
    {
      "id": "1175323",
      "postDate": "01/29/2021 04:53:51",
      "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a>, if you don't mind sharing, what is your training accuracy compared to validation accuracy compared to test accuracy?</p>",
      "rawMarkdown": "serigne, if you don't mind sharing, what is your training accuracy compared to validation accuracy compared to test accuracy?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1168553,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "01/25/2021 03:48:59",
      "content": "<p>Will fork your kernel on my local PC and try a few things - will report back if I find something.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1168691,
          "author_name": "shubham219",
          "author_url": "",
          "post_date": "01/25/2021 05:36:25",
          "content": "<p>Thanks..and please suggest me something to improve. This is my first time in computer Vision.<br>\nI am trying to do Regularization now.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1168828,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "01/25/2021 07:34:55",
          "content": "<p>Have the kernel forked and running on one of my machines.  Often bringing shared kernels to local machine is a pain because folder assignments - your code nicely written and was pretty easy to get running.  I got a little confused by the references to three different forms of the tf records - but figured those were the residue of earlier experiments. </p>\n<p>I did add a learning rate scheduler before I started the run.  Getting the learning rate right is one of the early keys (IMO) to finding a model that works.  I will be using a different schedule for frozen vs not frozen fits.  I did not feel I had a lot of success on this competition with my efforts to run a frozen and not frozen fit.  With a decent learning rate scheduler I have been doing most of my experiments with a single fit - it takes twice as long to do a frozen/unfrozen and I was not getting a big bang for the compute time.  So my current approach is to not fine tune until the last week or two of a competition when I have finished experiments with most of the things I know how to do and have a model I like.</p>\n<p>Added a bit to the early stopping patience and set epochs to high number.  When using a learning rate scheduler you want a robust change in the learning rate to have occurred before you early stop.  That's pretty hard to do on kaggle with the 9 hour time limit.</p>\n<p>Every model I have created in this competition are over fit when looking at the train vs validation plots.  </p>\n<p>Just looking at your kernel and your results I think its too early for you to be worrying about over - fit.   Your training accuracy looks like it needs to fixed first - I may think different tomorrow when the training is done - but until your training accuracy is above 0.90 efforts on reducing over-train may be the wrong path.  Until my training accuracy is higher than the current LB score I like to focus on doing the things that increase train accuracy/loss with only very minor attempts to fix the gap between train and validation plots.   But that's just my method and folks may disagree.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1168870,
          "author_name": "shubham219",
          "author_url": "",
          "post_date": "01/25/2021 08:04:12",
          "content": "<p>Many Thanks! for the observations. While training on the 90% frozen layers My training accuracy is getting increased to ~91% but same time validation accuracy is getting down. So I am trying to fix that.</p>\n<p>Please share the code if you get any success in improving the accuracy just wanna see new changes. </p>\n<p>! Happy Learning</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1169693,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "01/25/2021 17:18:23",
          "content": "<p>Kernel still running on my local machine - near the end of the first fold of fine tuning.</p>\n<p>Couple of observations/suggestions for change.</p>\n<p>The distribution of classes very unbalanced - so I would change the KFold to a StratifiedFold.  Little harder to write the <strong>for loop</strong> but IMO you should almost always use StratifiedFold unless the training set is very nicely balanced for classes.</p>\n<p>As you have written the code - your not performing the 5 fold in traditional manner.  Normally you would save the best epoch from each fold (5 total models) and than use those 5 in the prediction.  Your code only saving the best overall epoch and therefore losing some of the power of doing folds.  Combined with use of Kfold this can easily be a source of the part of the over fit you have seen.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1171289,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "01/26/2021 18:53:38",
          "content": "<p>Kernel completed on my local machine - the LR scheduler and extended number of epochs did help but still a 6% gap between training accuracy and validation accuracy.  If I get a extra available submission over the next day I will submit but I have a local evaluation that I make that suggests a poor LB score.</p>\n<p>I see nothing wrong with how you unfreeze - it's text book - but my experiments with various unfreeze gave me no real help in this competition.</p>\n<p>My suggestions</p>\n<ol>\n<li>StratifiedFold</li>\n<li>Save each fold model and use the 5 models in prediction.</li>\n<li>Add some type of learning rate scheduler - I will post the one I used (with two different settings) in a post shortly.</li>\n<li>Heavy augmentation </li>\n<li>Keep making small changes for dropout and regularization.</li>\n<li>2019 data - I am playing with that now myself but it's only breakeven or slight loss for local CV for me so far.</li>\n</ol>\n<p>A nice summary of other things to try in the post<br>\n<a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/203594\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/203594</a></p>\n<p><strong>Basic question - does a single model with three base pre-trained branches (like your kernel) perform better than three separate models that are than used together to make the predictions?</strong></p>\n<p>I used the three-in-one approach in a past competition with good success.  My approach in this competition is to combine the predictions of multiple models.  Since I am using 4 local machines it's more convenient for me to do the separate models and than make the prediction on Kaggle combining.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1171291,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "01/26/2021 18:55:50",
          "content": "<p>Used this for LR when frozen. (pasting into these posts makes a mess of the correct format of the code)  This a a very common method used on a number of shared kernels in this and other competitions - not sure who was the original author of the code.</p>\n<p>`LR_START = 0.000005#0.00001<br>\nLR_MAX = 0.001<br>\nLR_MIN = 0.000005<br>\nLR_RAMPUP_EPOCHS = 0<br>\nLR_SUSTAIN_EPOCHS = 0<br>\nLR_EXP_DECAY = .94</p>\n<p>def lrfn(epoch):<br>\n    if epoch &lt; LR_RAMPUP_EPOCHS:<br>\n        lr = (LR_MAX - LR_START) / LR_RAMPUP_EPOCHS * epoch + LR_START<br>\n    elif epoch &lt; LR_RAMPUP_EPOCHS + LR_SUSTAIN_EPOCHS:<br>\n        lr = LR_MAX<br>\n    else:<br>\n        lr = (LR_MAX - LR_MIN) * LR_EXP_DECAY**(epoch - LR_RAMPUP_EPOCHS - LR_SUSTAIN_EPOCHS) + LR_MIN<br>\n    return lr</p>\n<p>lr_callback = tf.keras.callbacks.LearningRateScheduler(lrfn, verbose = True)</p>\n<p>rng = [i for i in range(EPOCHS)]<br>\ny = [lrfn(x) for x in rng]<br>\nplt.plot(rng, y)<br>\nprint(\"Learning rate schedule: {:.3g} to {:.3g} to {:.3g}\".format(y[0], max(y), y[-1]))`</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1171293,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "01/26/2021 18:56:49",
          "content": "<p>Used this changed version for unfrozen</p>\n<p>`LR_START = 0.000005#0.00001<br>\nLR_MAX = 0.0001<br>\nLR_MIN = 0.000005<br>\nLR_RAMPUP_EPOCHS = 4<br>\nLR_SUSTAIN_EPOCHS = 0<br>\nLR_EXP_DECAY = .94</p>\n<p>def lrfn(epoch):<br>\n    if epoch &lt; LR_RAMPUP_EPOCHS:<br>\n        lr = (LR_MAX - LR_START) / LR_RAMPUP_EPOCHS * epoch + LR_START<br>\n    elif epoch &lt; LR_RAMPUP_EPOCHS + LR_SUSTAIN_EPOCHS:<br>\n        lr = LR_MAX<br>\n    else:<br>\n        lr = (LR_MAX - LR_MIN) * LR_EXP_DECAY**(epoch - LR_RAMPUP_EPOCHS - LR_SUSTAIN_EPOCHS) + LR_MIN<br>\n    return lr</p>\n<p>lr_callback = tf.keras.callbacks.LearningRateScheduler(lrfn, verbose = True)</p>\n<p>rng = [i for i in range(EPOCHS)]<br>\ny = [lrfn(x) for x in rng]<br>\nplt.plot(rng, y)<br>\nprint(\"Learning rate schedule: {:.3g} to {:.3g} to {:.3g}\".format(y[0], max(y), y[-1]))`</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1169598,
      "author_name": "nickuzmenkov",
      "author_url": "",
      "post_date": "01/25/2021 15:48:39",
      "content": "<p>Hello! <br>\nThe thing is even a small convnet I've tried out first (containing only 90K parameters) starts to overfit after ~10 epochs. So it seems overfitting would always be an issue in this competition, but you can consider doing these:</p>\n<ul>\n<li>enlarge your data by adding 2019 competition data: TFRecords are available <a href=\"https://www.kaggle.com/nickuzmenkov/cassava-leaf-disease-merged\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-50-tfrecords-external-512x512\" target=\"_blank\">here</a></li>\n<li>data augmentation: a good starting point on TPU <a href=\"https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96\" target=\"_blank\">is here</a></li>\n<li>as mentioned by <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> learning rate schedule would help to improve your score. I also suggest adding warmup steps and unfreeze the entire model</li>\n<li><code>Dropout</code> layers, as well as <code>kernel_regularizer</code> (<code>l1</code>, <code>l2</code>, and <code>l1-l2</code>) do not seem to help here much, but you can also try them out as long as it is a common regularization technique</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1171311,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "01/26/2021 19:11:18",
          "content": "<p>I think it all depends on your training strategie. </p>\n<p>I have a heavy pretrained model which were overfitting (CV = 0.897 and LB 0.989) .  But after changing my training strategie, both CV and LB were improved : CV = 0.901 and LB 0.902).  And I train for larger number of epochs than 10 and don't freeze any layer. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1171373,
          "author_name": "proletheus",
          "author_url": "",
          "post_date": "01/26/2021 19:41:36",
          "content": "<p>is that 0.901 with TTA?  and 2019 data?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1171388,
          "author_name": "ayu055",
          "author_url": "",
          "post_date": "01/26/2021 19:52:11",
          "content": "<p>Do you have any suggestions for better regularization?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1171457,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "01/26/2021 20:45:41",
          "content": "<p>For LB = 0.902 there is no TTA. </p>\n<p>A part from one Dropout layer (that is present in all my classification pipeline), I don't use any model regularization. All my focus is on the training process.  But I can't disclose the details before the end. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1175254,
          "author_name": "nickuzmenkov",
          "author_url": "",
          "post_date": "01/29/2021 03:40:24",
          "content": "<p>Thank you for sharing, <a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1175323,
          "author_name": "ayu055",
          "author_url": "",
          "post_date": "01/29/2021 04:53:51",
          "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a>, if you don't mind sharing, what is your training accuracy compared to validation accuracy compared to test accuracy?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1168526": "While finetuning model by making last 20% of layers trainable, model is **overfitting**. Any suggestion on proper way of doing finetuning? [My Training Model](https://www.kaggle.com/shubham219/model-ensembling-with-k-fold)",
    "1168553": "Will fork your kernel on my local PC and try a few things - will report back if I find something.",
    "1168691": "Thanks..and please suggest me something to improve. This is my first time in computer Vision.\nI am trying to do Regularization now.",
    "1168828": "Have the kernel forked and running on one of my machines.  Often bringing shared kernels to local machine is a pain because folder assignments - your code nicely written and was pretty easy to get running.  I got a little confused by the references to three different forms of the tf records - but figured those were the residue of earlier experiments. \n\nI did add a learning rate scheduler before I started the run.  Getting the learning rate right is one of the early keys (IMO) to finding a model that works.  I will be using a different schedule for frozen vs not frozen fits.  I did not feel I had a lot of success on this competition with my efforts to run a frozen and not frozen fit.  With a decent learning rate scheduler I have been doing most of my experiments with a single fit - it takes twice as long to do a frozen/unfrozen and I was not getting a big bang for the compute time.  So my current approach is to not fine tune until the last week or two of a competition when I have finished experiments with most of the things I know how to do and have a model I like.\n\nAdded a bit to the early stopping patience and set epochs to high number.  When using a learning rate scheduler you want a robust change in the learning rate to have occurred before you early stop.  That's pretty hard to do on kaggle with the 9 hour time limit.\n\nEvery model I have created in this competition are over fit when looking at the train vs validation plots.  \n\nJust looking at your kernel and your results I think its too early for you to be worrying about over - fit.   Your training accuracy looks like it needs to fixed first - I may think different tomorrow when the training is done - but until your training accuracy is above 0.90 efforts on reducing over-train may be the wrong path.  Until my training accuracy is higher than the current LB score I like to focus on doing the things that increase train accuracy/loss with only very minor attempts to fix the gap between train and validation plots.   But that's just my method and folks may disagree.",
    "1168870": "Many Thanks! for the observations. While training on the 90% frozen layers My training accuracy is getting increased to ~91% but same time validation accuracy is getting down. So I am trying to fix that.\n\nPlease share the code if you get any success in improving the accuracy just wanna see new changes. \n\n! Happy Learning",
    "1169598": "Hello! \nThe thing is even a small convnet I've tried out first (containing only 90K parameters) starts to overfit after ~10 epochs. So it seems overfitting would always be an issue in this competition, but you can consider doing these:\n* enlarge your data by adding 2019 competition data: TFRecords are available [here](https://www.kaggle.com/nickuzmenkov/cassava-leaf-disease-merged) and [here](https://www.kaggle.com/dimitreoliveira/cassava-leaf-disease-50-tfrecords-external-512x512)\n* data augmentation: a good starting point on TPU [is here](https://www.kaggle.com/cdeotte/rotation-augmentation-gpu-tpu-0-96)\n* as mentioned by @pcjimmmy learning rate schedule would help to improve your score. I also suggest adding warmup steps and unfreeze the entire model\n* `Dropout` layers, as well as `kernel_regularizer` (`l1`, `l2`, and `l1-l2`) do not seem to help here much, but you can also try them out as long as it is a common regularization technique",
    "1169693": "Kernel still running on my local machine - near the end of the first fold of fine tuning.\n\nCouple of observations/suggestions for change.\n\nThe distribution of classes very unbalanced - so I would change the KFold to a StratifiedFold.  Little harder to write the **for loop** but IMO you should almost always use StratifiedFold unless the training set is very nicely balanced for classes.\n\nAs you have written the code - your not performing the 5 fold in traditional manner.  Normally you would save the best epoch from each fold (5 total models) and than use those 5 in the prediction.  Your code only saving the best overall epoch and therefore losing some of the power of doing folds.  Combined with use of Kfold this can easily be a source of the part of the over fit you have seen.",
    "1171289": "Kernel completed on my local machine - the LR scheduler and extended number of epochs did help but still a 6% gap between training accuracy and validation accuracy.  If I get a extra available submission over the next day I will submit but I have a local evaluation that I make that suggests a poor LB score.\n\nI see nothing wrong with how you unfreeze - it's text book - but my experiments with various unfreeze gave me no real help in this competition.\n\nMy suggestions\n1.  StratifiedFold\n2.  Save each fold model and use the 5 models in prediction.\n3.  Add some type of learning rate scheduler - I will post the one I used (with two different settings) in a post shortly.\n4. Heavy augmentation \n5.  Keep making small changes for dropout and regularization.\n6.  2019 data - I am playing with that now myself but it's only breakeven or slight loss for local CV for me so far.\n\nA nice summary of other things to try in the post\nhttps://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/203594\n\n**Basic question - does a single model with three base pre-trained branches (like your kernel) perform better than three separate models that are than used together to make the predictions?**\n\nI used the three-in-one approach in a past competition with good success.  My approach in this competition is to combine the predictions of multiple models.  Since I am using 4 local machines it's more convenient for me to do the separate models and than make the prediction on Kaggle combining.",
    "1171291": "Used this for LR when frozen. (pasting into these posts makes a mess of the correct format of the code)  This a a very common method used on a number of shared kernels in this and other competitions - not sure who was the original author of the code.\n\n`LR_START = 0.000005#0.00001\nLR_MAX = 0.001\nLR_MIN = 0.000005\nLR_RAMPUP_EPOCHS = 0\nLR_SUSTAIN_EPOCHS = 0\nLR_EXP_DECAY = .94\n\ndef lrfn(epoch):\n    if epoch < LR_RAMPUP_EPOCHS:\n        lr = (LR_MAX - LR_START) / LR_RAMPUP_EPOCHS * epoch + LR_START\n    elif epoch < LR_RAMPUP_EPOCHS + LR_SUSTAIN_EPOCHS:\n        lr = LR_MAX\n    else:\n        lr = (LR_MAX - LR_MIN) * LR_EXP_DECAY**(epoch - LR_RAMPUP_EPOCHS - LR_SUSTAIN_EPOCHS) + LR_MIN\n    return lr\n    \nlr_callback = tf.keras.callbacks.LearningRateScheduler(lrfn, verbose = True)\n\nrng = [i for i in range(EPOCHS)]\ny = [lrfn(x) for x in rng]\nplt.plot(rng, y)\nprint(\"Learning rate schedule: {:.3g} to {:.3g} to {:.3g}\".format(y[0], max(y), y[-1]))`",
    "1171293": "Used this changed version for unfrozen\n\n`LR_START = 0.000005#0.00001\nLR_MAX = 0.0001\nLR_MIN = 0.000005\nLR_RAMPUP_EPOCHS = 4\nLR_SUSTAIN_EPOCHS = 0\nLR_EXP_DECAY = .94\n\ndef lrfn(epoch):\n    if epoch < LR_RAMPUP_EPOCHS:\n        lr = (LR_MAX - LR_START) / LR_RAMPUP_EPOCHS * epoch + LR_START\n    elif epoch < LR_RAMPUP_EPOCHS + LR_SUSTAIN_EPOCHS:\n        lr = LR_MAX\n    else:\n        lr = (LR_MAX - LR_MIN) * LR_EXP_DECAY**(epoch - LR_RAMPUP_EPOCHS - LR_SUSTAIN_EPOCHS) + LR_MIN\n    return lr\n    \nlr_callback = tf.keras.callbacks.LearningRateScheduler(lrfn, verbose = True)\n\nrng = [i for i in range(EPOCHS)]\ny = [lrfn(x) for x in rng]\nplt.plot(rng, y)\nprint(\"Learning rate schedule: {:.3g} to {:.3g} to {:.3g}\".format(y[0], max(y), y[-1]))`",
    "1171311": "I think it all depends on your training strategie. \n\nI have a heavy pretrained model which were overfitting (CV = 0.897 and LB 0.989) .  But after changing my training strategie, both CV and LB were improved : CV = 0.901 and LB 0.902).  And I train for larger number of epochs than 10 and don't freeze any layer.",
    "1171373": "is that 0.901 with TTA?  and 2019 data?",
    "1171388": "Do you have any suggestions for better regularization?",
    "1171457": "For LB = 0.902 there is no TTA. \n\nA part from one Dropout layer (that is present in all my classification pipeline), I don't use any model regularization. All my focus is on the training process.  But I can't disclose the details before the end.",
    "1175254": "Thank you for sharing, @serigne",
    "1175323": "serigne, if you don't mind sharing, what is your training accuracy compared to validation accuracy compared to test accuracy?"
  },
  "source": "meta"
}