{
  "id": 205879,
  "title": "Extremely low validation accuracy ",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/205879",
  "author_name": "",
  "post_date": "2020-12-22T10:09:36.502044800Z",
  "votes": 2,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Hello everyone, <br>\nI am getting abnormally low validation accuracy in the range of 0.1-0.12 with some epochs shooting to 0.60. </p>\n<p>I have tried efficientnet b0 and b3. Also, I have tried various image sizes and replicated settings of some of the kernels available here but the accuracy remains the same. </p>\n<p>I have not been able to debug it. This is the kernel. <br>\n<a href=\"https://www.kaggle.com/zainahmedsharif/cassanava22/edit/run/49953329\" target=\"_blank\">https://www.kaggle.com/zainahmedsharif/cassanava22/edit/run/49953329</a></p>",
  "messages": [
    {
      "id": "1122264",
      "postDate": "12/22/2020 10:09:36",
      "content": "<p>Hello everyone, <br>\nI am getting abnormally low validation accuracy in the range of 0.1-0.12 with some epochs shooting to 0.60. </p>\n<p>I have tried efficientnet b0 and b3. Also, I have tried various image sizes and replicated settings of some of the kernels available here but the accuracy remains the same. </p>\n<p>I have not been able to debug it. This is the kernel. <br>\n<a href=\"https://www.kaggle.com/zainahmedsharif/cassanava22/edit/run/49953329\" target=\"_blank\">https://www.kaggle.com/zainahmedsharif/cassanava22/edit/run/49953329</a></p>",
      "rawMarkdown": "Hello everyone, \nI am getting abnormally low validation accuracy in the range of 0.1-0.12 with some epochs shooting to 0.60. \n\nI have tried efficientnet b0 and b3. Also, I have tried various image sizes and replicated settings of some of the kernels available here but the accuracy remains the same. \n\nI have not been able to debug it. This is the kernel. \nhttps://www.kaggle.com/zainahmedsharif/cassanava22/edit/run/49953329",
      "votes": null
    },
    {
      "id": "1122429",
      "postDate": "12/22/2020 12:39:10",
      "content": "<p>Folks who have shared kernels that start with a long page showing the directory contents have huge risk of killing any interest in my looking at the code.  My suggestion - comment out the directory listing before you ask someone to look over the code.</p>",
      "rawMarkdown": "Folks who have shared kernels that start with a long page showing the directory contents have huge risk of killing any interest in my looking at the code.  My suggestion - comment out the directory listing before you ask someone to look over the code.",
      "votes": null
    },
    {
      "id": "1122859",
      "postDate": "12/22/2020 18:44:57",
      "content": "<p>Have you normalized the images before training? What is the loss value?</p>",
      "rawMarkdown": "Have you normalized the images before training? What is the loss value?",
      "votes": null
    },
    {
      "id": "1123188",
      "postDate": "12/23/2020 02:53:08",
      "content": "<p>Sure,<br>\nI have removed the listing. Added more comments to make it more readable. <br>\nThanks for the tip. If you could have a look, that would be great. </p>",
      "rawMarkdown": "Sure,\nI have removed the listing. Added more comments to make it more readable. \nThanks for the tip. If you could have a look, that would be great.",
      "votes": null
    },
    {
      "id": "1123193",
      "postDate": "12/23/2020 02:58:20",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5608364%2F7d5face2c31792ace9573adb45734f4f%2Fefficientb3.jpg?generation=1608692453629578&amp;alt=media\" alt=\"![\">]</p>\n<p>These were the results. <br>\nwith only 4 epochs in, early stopping kicks in. <br>\nimage size = 300 <br>\nbatch size = 32 <br>\nmodel = efficientnetb3 </p>",
      "rawMarkdown": "![![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5608364%2F7d5face2c31792ace9573adb45734f4f%2Fefficientb3.jpg?generation=1608692453629578&alt=media)]\n\n\nThese were the results. \nwith only 4 epochs in, early stopping kicks in. \nimage size = 300 \nbatch size = 32 \nmodel = efficientnetb3",
      "votes": null
    },
    {
      "id": "1123275",
      "postDate": "12/23/2020 05:13:54",
      "content": "<p>The model is seemingly overfitting in your case.I suggest you to add some Dropout layers and again try to train it.</p>",
      "rawMarkdown": "The model is seemingly overfitting in your case.I suggest you to add some Dropout layers and again try to train it.",
      "votes": null
    },
    {
      "id": "1123292",
      "postDate": "12/23/2020 05:26:46",
      "content": "<p>Thanks for the response. <br>\nI am using efficientnetb0 with a output layer with 5 neurons.<br>\nHow do you suggest I proceed? <br>\nI have added a dense layer with dropout for now. </p>",
      "rawMarkdown": "Thanks for the response. \nI am using efficientnetb0 with a output layer with 5 neurons.\nHow do you suggest I proceed? \nI have added a dense layer with dropout for now.",
      "votes": null
    },
    {
      "id": "1123480",
      "postDate": "12/23/2020 08:58:35",
      "content": "<p>Looks more readable for sure.  If I was sharing to ask for advice I would also delete any cells that are 100% commented - it's hard to jump into someone's code and find suggestions - you want to make it as clean as practical so they can focus on the critical code.  </p>\n<p>When getting the huge swings in values your model is early stopping while the model is still trying to learn.  With patience=3 you need to really be close to the right learning rate to avoid early stopping before the model has really stopped learning.  Unless your script is pushing the hours limit be more generous with patience - I like a value that lets  any learning curve algorithm make at least 3 changes to the learning rate before it early stops.  Make it high enough that the model runs for the full 50 epochs the first time.  So I would set patience to at least 15 - maybe even 25.</p>",
      "rawMarkdown": "Looks more readable for sure.  If I was sharing to ask for advice I would also delete any cells that are 100% commented - it's hard to jump into someone's code and find suggestions - you want to make it as clean as practical so they can focus on the critical code.  \n\nWhen getting the huge swings in values your model is early stopping while the model is still trying to learn.  With patience=3 you need to really be close to the right learning rate to avoid early stopping before the model has really stopped learning.  Unless your script is pushing the hours limit be more generous with patience - I like a value that lets  any learning curve algorithm make at least 3 changes to the learning rate before it early stops.  Make it high enough that the model runs for the full 50 epochs the first time.  So I would set patience to at least 15 - maybe even 25.",
      "votes": null
    },
    {
      "id": "1123530",
      "postDate": "12/23/2020 09:57:21",
      "content": "<p>Great. Thanks for the response. I had a similar feeling. I will try setting a more generous value for patience in early stopping. </p>",
      "rawMarkdown": "Great. Thanks for the response. I had a similar feeling. I will try setting a more generous value for patience in early stopping.",
      "votes": null
    },
    {
      "id": "1124177",
      "postDate": "12/23/2020 17:58:19",
      "content": "<p>Ran your script with patience of 25 - looks much better.  Fix your model number in the last bit where you generate the submission (model2 should be model) and its good to go,  Next you need to add a learning curve algorithm.</p>",
      "rawMarkdown": "Ran your script with patience of 25 - looks much better.  Fix your model number in the last bit where you generate the submission (model2 should be model) and its good to go,  Next you need to add a learning curve algorithm.",
      "votes": null
    },
    {
      "id": "1124285",
      "postDate": "12/23/2020 19:46:13",
      "content": "<p>Hey try to reduce your learning rate around 1e-4 and make sure to use a lr scheduler that will decay it over time. Im not too familiar with keras but i dont think i saw it in your notebook</p>",
      "rawMarkdown": "Hey try to reduce your learning rate around 1e-4 and make sure to use a lr scheduler that will decay it over time. Im not too familiar with keras but i dont think i saw it in your notebook",
      "votes": null
    },
    {
      "id": "1124681",
      "postDate": "12/24/2020 06:14:17",
      "content": "<p>Thanks for the tip. Yes, I did not use learning rata decay </p>",
      "rawMarkdown": "Thanks for the tip. Yes, I did not use learning rata decay",
      "votes": null
    },
    {
      "id": "1124688",
      "postDate": "12/24/2020 06:17:04",
      "content": "<p>Yes I am getting much better result as well. I am plotting the learning curve at the end or do you mean something else? </p>",
      "rawMarkdown": "Yes I am getting much better result as well. I am plotting the learning curve at the end or do you mean something else?",
      "votes": null
    },
    {
      "id": "1124872",
      "postDate": "12/24/2020 08:44:46",
      "content": "<p>I mean something else.  In your code I do not see any type of learning rate scheduler.</p>\n<p>There are several that are commonly found in shared kernels.  The idea being that the learning rate should be smaller as the epochs continue and the model is not getting a better val_loss.    They are often a bit like the early stopping - but rather than stopping the model after the patience value they lower the learning rate.  Yaan mentioned this same need in his recent response.  </p>\n<p>Here is a simple one - create a callback and than reference that callback in your fit.  I don't get correct formatting when I paste into these posts, so you need to clean up the formatting.</p>\n<p><code>callbacks = [ReduceLROnPlateau(monitor='val_loss', patience=11, verbose=1, factor=0.2),\n             EarlyStopping(monitor='val_loss', patience=39),\n             ModelCheckpoint(filepath=pc + 'best_model.h5', monitor='val_loss', save_best_only=True)]</code></p>\n<p><code>history = model.fit(train_generator,\n                  epochs = EPOCHS, \n                  validation_data=validation_generator,\n                  callbacks=callbacks)</code></p>\n<p>I run 4 machines so my filepath for the checkpoint adds the name of the machine (PC) being used to the model save.</p>\n<p>For ReduceLROnPlateau:<br>\nThe patience number of epochs and the factor are two things to play with to tune up.  After you try the above code you might want to play with the factor.  I like 0.8 to 0.95 when I want a nice slow (long running) model fit.  But start with 0.2 because of the hours limit on kaggle.  Also should increase your epochs from 50 to high number.  You want the model to stop the fit from early stopping rather than reaching the epochs - I use 5000 epochs on my local machines since I don't have to worry about kaggle timeout.</p>",
      "rawMarkdown": "I mean something else.  In your code I do not see any type of learning rate scheduler.\n\nThere are several that are commonly found in shared kernels.  The idea being that the learning rate should be smaller as the epochs continue and the model is not getting a better val_loss.    They are often a bit like the early stopping - but rather than stopping the model after the patience value they lower the learning rate.  Yaan mentioned this same need in his recent response.  \n\nHere is a simple one - create a callback and than reference that callback in your fit.  I don't get correct formatting when I paste into these posts, so you need to clean up the formatting.\n\n`callbacks = [ReduceLROnPlateau(monitor='val_loss', patience=11, verbose=1, factor=0.2),\n             EarlyStopping(monitor='val_loss', patience=39),\n             ModelCheckpoint(filepath=pc + 'best_model.h5', monitor='val_loss', save_best_only=True)]`\n\n`history = model.fit(train_generator,\n                  epochs = EPOCHS, \n                  validation_data=validation_generator,\n                  callbacks=callbacks)`\n\nI run 4 machines so my filepath for the checkpoint adds the name of the machine (PC) being used to the model save.\n\nFor ReduceLROnPlateau:\nThe patience number of epochs and the factor are two things to play with to tune up.  After you try the above code you might want to play with the factor.  I like 0.8 to 0.95 when I want a nice slow (long running) model fit.  But start with 0.2 because of the hours limit on kaggle.  Also should increase your epochs from 50 to high number.  You want the model to stop the fit from early stopping rather than reaching the epochs - I use 5000 epochs on my local machines since I don't have to worry about kaggle timeout.",
      "votes": null
    },
    {
      "id": "1124934",
      "postDate": "12/24/2020 09:29:07",
      "content": "<p>I get it. Thanks. <br>\nand I suppose the patience value in <strong>ReduceLROnPlateau</strong> would be 3 to 4 times less than patience value in early stopping. </p>",
      "rawMarkdown": "I get it. Thanks. \nand I suppose the patience value in **ReduceLROnPlateau** would be 3 to 4 times less than patience value in early stopping.",
      "votes": null
    },
    {
      "id": "1124963",
      "postDate": "12/24/2020 09:49:30",
      "content": "<p>You got the idea - depends on the factor.   If I use a high factor (0.9) than I might put early stopping at 6 or more times - since a 0.9 is a slow change to the learning rate.   </p>\n<p>After you look at the plots of accuracy and loss you can get a feel for what to change-but that takes a lot of experience or a very long post :)</p>",
      "rawMarkdown": "You got the idea - depends on the factor.   If I use a high factor (0.9) than I might put early stopping at 6 or more times - since a 0.9 is a slow change to the learning rate.   \n\nAfter you look at the plots of accuracy and loss you can get a feel for what to change-but that takes a lot of experience or a very long post :)",
      "votes": null
    },
    {
      "id": "1124998",
      "postDate": "12/24/2020 10:14:55",
      "content": "<p>Thanks for your help. <br>\nIt was extremely helpful. <br>\nGood luck!!! </p>",
      "rawMarkdown": "Thanks for your help. \nIt was extremely helpful. \nGood luck!!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1122429,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "12/22/2020 12:39:10",
      "content": "<p>Folks who have shared kernels that start with a long page showing the directory contents have huge risk of killing any interest in my looking at the code.  My suggestion - comment out the directory listing before you ask someone to look over the code.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1123188,
          "author_name": "zainahmedsharif",
          "author_url": "",
          "post_date": "12/23/2020 02:53:08",
          "content": "<p>Sure,<br>\nI have removed the listing. Added more comments to make it more readable. <br>\nThanks for the tip. If you could have a look, that would be great. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1123480,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "12/23/2020 08:58:35",
          "content": "<p>Looks more readable for sure.  If I was sharing to ask for advice I would also delete any cells that are 100% commented - it's hard to jump into someone's code and find suggestions - you want to make it as clean as practical so they can focus on the critical code.  </p>\n<p>When getting the huge swings in values your model is early stopping while the model is still trying to learn.  With patience=3 you need to really be close to the right learning rate to avoid early stopping before the model has really stopped learning.  Unless your script is pushing the hours limit be more generous with patience - I like a value that lets  any learning curve algorithm make at least 3 changes to the learning rate before it early stops.  Make it high enough that the model runs for the full 50 epochs the first time.  So I would set patience to at least 15 - maybe even 25.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1123530,
          "author_name": "zainahmedsharif",
          "author_url": "",
          "post_date": "12/23/2020 09:57:21",
          "content": "<p>Great. Thanks for the response. I had a similar feeling. I will try setting a more generous value for patience in early stopping. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1124177,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "12/23/2020 17:58:19",
          "content": "<p>Ran your script with patience of 25 - looks much better.  Fix your model number in the last bit where you generate the submission (model2 should be model) and its good to go,  Next you need to add a learning curve algorithm.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1124688,
          "author_name": "zainahmedsharif",
          "author_url": "",
          "post_date": "12/24/2020 06:17:04",
          "content": "<p>Yes I am getting much better result as well. I am plotting the learning curve at the end or do you mean something else? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1124872,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "12/24/2020 08:44:46",
          "content": "<p>I mean something else.  In your code I do not see any type of learning rate scheduler.</p>\n<p>There are several that are commonly found in shared kernels.  The idea being that the learning rate should be smaller as the epochs continue and the model is not getting a better val_loss.    They are often a bit like the early stopping - but rather than stopping the model after the patience value they lower the learning rate.  Yaan mentioned this same need in his recent response.  </p>\n<p>Here is a simple one - create a callback and than reference that callback in your fit.  I don't get correct formatting when I paste into these posts, so you need to clean up the formatting.</p>\n<p><code>callbacks = [ReduceLROnPlateau(monitor='val_loss', patience=11, verbose=1, factor=0.2),\n             EarlyStopping(monitor='val_loss', patience=39),\n             ModelCheckpoint(filepath=pc + 'best_model.h5', monitor='val_loss', save_best_only=True)]</code></p>\n<p><code>history = model.fit(train_generator,\n                  epochs = EPOCHS, \n                  validation_data=validation_generator,\n                  callbacks=callbacks)</code></p>\n<p>I run 4 machines so my filepath for the checkpoint adds the name of the machine (PC) being used to the model save.</p>\n<p>For ReduceLROnPlateau:<br>\nThe patience number of epochs and the factor are two things to play with to tune up.  After you try the above code you might want to play with the factor.  I like 0.8 to 0.95 when I want a nice slow (long running) model fit.  But start with 0.2 because of the hours limit on kaggle.  Also should increase your epochs from 50 to high number.  You want the model to stop the fit from early stopping rather than reaching the epochs - I use 5000 epochs on my local machines since I don't have to worry about kaggle timeout.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1124934,
          "author_name": "zainahmedsharif",
          "author_url": "",
          "post_date": "12/24/2020 09:29:07",
          "content": "<p>I get it. Thanks. <br>\nand I suppose the patience value in <strong>ReduceLROnPlateau</strong> would be 3 to 4 times less than patience value in early stopping. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1124963,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "12/24/2020 09:49:30",
          "content": "<p>You got the idea - depends on the factor.   If I use a high factor (0.9) than I might put early stopping at 6 or more times - since a 0.9 is a slow change to the learning rate.   </p>\n<p>After you look at the plots of accuracy and loss you can get a feel for what to change-but that takes a lot of experience or a very long post :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1124998,
          "author_name": "zainahmedsharif",
          "author_url": "",
          "post_date": "12/24/2020 10:14:55",
          "content": "<p>Thanks for your help. <br>\nIt was extremely helpful. <br>\nGood luck!!! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1122859,
      "author_name": "shivamjohri",
      "author_url": "",
      "post_date": "12/22/2020 18:44:57",
      "content": "<p>Have you normalized the images before training? What is the loss value?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1123193,
          "author_name": "zainahmedsharif",
          "author_url": "",
          "post_date": "12/23/2020 02:58:20",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5608364%2F7d5face2c31792ace9573adb45734f4f%2Fefficientb3.jpg?generation=1608692453629578&amp;alt=media\" alt=\"![\">]</p>\n<p>These were the results. <br>\nwith only 4 epochs in, early stopping kicks in. <br>\nimage size = 300 <br>\nbatch size = 32 <br>\nmodel = efficientnetb3 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1123275,
      "author_name": "shivamjohri",
      "author_url": "",
      "post_date": "12/23/2020 05:13:54",
      "content": "<p>The model is seemingly overfitting in your case.I suggest you to add some Dropout layers and again try to train it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1123292,
          "author_name": "zainahmedsharif",
          "author_url": "",
          "post_date": "12/23/2020 05:26:46",
          "content": "<p>Thanks for the response. <br>\nI am using efficientnetb0 with a output layer with 5 neurons.<br>\nHow do you suggest I proceed? <br>\nI have added a dense layer with dropout for now. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1124285,
      "author_name": "yannmajewski",
      "author_url": "",
      "post_date": "12/23/2020 19:46:13",
      "content": "<p>Hey try to reduce your learning rate around 1e-4 and make sure to use a lr scheduler that will decay it over time. Im not too familiar with keras but i dont think i saw it in your notebook</p>",
      "votes": null,
      "replies": [
        {
          "id": 1124681,
          "author_name": "zainahmedsharif",
          "author_url": "",
          "post_date": "12/24/2020 06:14:17",
          "content": "<p>Thanks for the tip. Yes, I did not use learning rata decay </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1122264": "Hello everyone, \nI am getting abnormally low validation accuracy in the range of 0.1-0.12 with some epochs shooting to 0.60. \n\nI have tried efficientnet b0 and b3. Also, I have tried various image sizes and replicated settings of some of the kernels available here but the accuracy remains the same. \n\nI have not been able to debug it. This is the kernel. \nhttps://www.kaggle.com/zainahmedsharif/cassanava22/edit/run/49953329",
    "1122429": "Folks who have shared kernels that start with a long page showing the directory contents have huge risk of killing any interest in my looking at the code.  My suggestion - comment out the directory listing before you ask someone to look over the code.",
    "1122859": "Have you normalized the images before training? What is the loss value?",
    "1123188": "Sure,\nI have removed the listing. Added more comments to make it more readable. \nThanks for the tip. If you could have a look, that would be great.",
    "1123193": "![![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5608364%2F7d5face2c31792ace9573adb45734f4f%2Fefficientb3.jpg?generation=1608692453629578&alt=media)]\n\n\nThese were the results. \nwith only 4 epochs in, early stopping kicks in. \nimage size = 300 \nbatch size = 32 \nmodel = efficientnetb3",
    "1123275": "The model is seemingly overfitting in your case.I suggest you to add some Dropout layers and again try to train it.",
    "1123292": "Thanks for the response. \nI am using efficientnetb0 with a output layer with 5 neurons.\nHow do you suggest I proceed? \nI have added a dense layer with dropout for now.",
    "1123480": "Looks more readable for sure.  If I was sharing to ask for advice I would also delete any cells that are 100% commented - it's hard to jump into someone's code and find suggestions - you want to make it as clean as practical so they can focus on the critical code.  \n\nWhen getting the huge swings in values your model is early stopping while the model is still trying to learn.  With patience=3 you need to really be close to the right learning rate to avoid early stopping before the model has really stopped learning.  Unless your script is pushing the hours limit be more generous with patience - I like a value that lets  any learning curve algorithm make at least 3 changes to the learning rate before it early stops.  Make it high enough that the model runs for the full 50 epochs the first time.  So I would set patience to at least 15 - maybe even 25.",
    "1123530": "Great. Thanks for the response. I had a similar feeling. I will try setting a more generous value for patience in early stopping.",
    "1124177": "Ran your script with patience of 25 - looks much better.  Fix your model number in the last bit where you generate the submission (model2 should be model) and its good to go,  Next you need to add a learning curve algorithm.",
    "1124285": "Hey try to reduce your learning rate around 1e-4 and make sure to use a lr scheduler that will decay it over time. Im not too familiar with keras but i dont think i saw it in your notebook",
    "1124681": "Thanks for the tip. Yes, I did not use learning rata decay",
    "1124688": "Yes I am getting much better result as well. I am plotting the learning curve at the end or do you mean something else?",
    "1124872": "I mean something else.  In your code I do not see any type of learning rate scheduler.\n\nThere are several that are commonly found in shared kernels.  The idea being that the learning rate should be smaller as the epochs continue and the model is not getting a better val_loss.    They are often a bit like the early stopping - but rather than stopping the model after the patience value they lower the learning rate.  Yaan mentioned this same need in his recent response.  \n\nHere is a simple one - create a callback and than reference that callback in your fit.  I don't get correct formatting when I paste into these posts, so you need to clean up the formatting.\n\n`callbacks = [ReduceLROnPlateau(monitor='val_loss', patience=11, verbose=1, factor=0.2),\n             EarlyStopping(monitor='val_loss', patience=39),\n             ModelCheckpoint(filepath=pc + 'best_model.h5', monitor='val_loss', save_best_only=True)]`\n\n`history = model.fit(train_generator,\n                  epochs = EPOCHS, \n                  validation_data=validation_generator,\n                  callbacks=callbacks)`\n\nI run 4 machines so my filepath for the checkpoint adds the name of the machine (PC) being used to the model save.\n\nFor ReduceLROnPlateau:\nThe patience number of epochs and the factor are two things to play with to tune up.  After you try the above code you might want to play with the factor.  I like 0.8 to 0.95 when I want a nice slow (long running) model fit.  But start with 0.2 because of the hours limit on kaggle.  Also should increase your epochs from 50 to high number.  You want the model to stop the fit from early stopping rather than reaching the epochs - I use 5000 epochs on my local machines since I don't have to worry about kaggle timeout.",
    "1124934": "I get it. Thanks. \nand I suppose the patience value in **ReduceLROnPlateau** would be 3 to 4 times less than patience value in early stopping.",
    "1124963": "You got the idea - depends on the factor.   If I use a high factor (0.9) than I might put early stopping at 6 or more times - since a 0.9 is a slow change to the learning rate.   \n\nAfter you look at the plots of accuracy and loss you can get a feel for what to change-but that takes a lot of experience or a very long post :)",
    "1124998": "Thanks for your help. \nIt was extremely helpful. \nGood luck!!!"
  },
  "source": "meta"
}