{
  "id": 70253,
  "title": "comparison of optimizers for keras",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/70253",
  "author_name": "",
  "post_date": "2018-11-01T10:46:08.188823600Z",
  "votes": 37,
  "comment_count": 19,
  "views": 0,
  "content": "<p>I am wondering how much we can improve the results by using different optimizers, so I run all but Nadam from keras, using ResNet50 with pre-trained weights of imagenet. The input shape is (224, 224, 3).</p>\n\n<p>the parameters used as following:\nSGD, lr=0.01, momentum=True, nesterov=True, decay=1e-06,</p>\n\n<p>all others used the default setting with decay=1e-06. </p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/inbox/642483/73906ba9af381a9e4a5bc0e7ca764896/pre_trained.png\" alt=\"Image\"></p>\n\n<p>overall, the AMSGrad variant of Adam using learning rate 0.0001 has the fastest convergence speed for training, and also shows very good validation accuracy. RMSprop is the worst optimizer according to the figure, which is in line with the observation from here <a href=\"https://shaoanlu.wordpress.com/2017/05/29/sgd-all-which-one-is-the-best-optimizer-dogs-vs-cats-toy-experiment/\">https://shaoanlu.wordpress.com/2017/05/29/sgd-all-which-one-is-the-best-optimizer-dogs-vs-cats-toy-experiment/</a></p>\n\n<p>Next, I will evaluate the optimizers based on training from scratch for Resnet50, let's see if the results change or not. </p>\n\n<p><strong>Update</strong>:\nbelow is the evaluation of different optimizers on ResNet50 training from scratch, the input size is (512, 512, 4).\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/inbox/642483/7152f51a584227b2dec0468885c35ec4/from_scratch.png\" alt=\"Image\"></p>\n\n<p>as we can see from the figure, now Adamax is the best one, which implies that we might need different optimizers for pre-trained and training-from-scratch model.</p>",
  "messages": [
    {
      "id": "413673",
      "postDate": "11/01/2018 10:46:08",
      "content": "<p>I am wondering how much we can improve the results by using different optimizers, so I run all but Nadam from keras, using ResNet50 with pre-trained weights of imagenet. The input shape is (224, 224, 3).</p>\n\n<p>the parameters used as following:\nSGD, lr=0.01, momentum=True, nesterov=True, decay=1e-06,</p>\n\n<p>all others used the default setting with decay=1e-06. </p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/inbox/642483/73906ba9af381a9e4a5bc0e7ca764896/pre_trained.png\" alt=\"Image\"></p>\n\n<p>overall, the AMSGrad variant of Adam using learning rate 0.0001 has the fastest convergence speed for training, and also shows very good validation accuracy. RMSprop is the worst optimizer according to the figure, which is in line with the observation from here <a href=\"https://shaoanlu.wordpress.com/2017/05/29/sgd-all-which-one-is-the-best-optimizer-dogs-vs-cats-toy-experiment/\">https://shaoanlu.wordpress.com/2017/05/29/sgd-all-which-one-is-the-best-optimizer-dogs-vs-cats-toy-experiment/</a></p>\n\n<p>Next, I will evaluate the optimizers based on training from scratch for Resnet50, let's see if the results change or not. </p>\n\n<p><strong>Update</strong>:\nbelow is the evaluation of different optimizers on ResNet50 training from scratch, the input size is (512, 512, 4).\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/inbox/642483/7152f51a584227b2dec0468885c35ec4/from_scratch.png\" alt=\"Image\"></p>\n\n<p>as we can see from the figure, now Adamax is the best one, which implies that we might need different optimizers for pre-trained and training-from-scratch model.</p>",
      "rawMarkdown": "I am wondering how much we can improve the results by using different optimizers, so I run all but Nadam from keras, using ResNet50 with pre-trained weights of imagenet. The input shape is (224, 224, 3).\n\nthe parameters used as following:\nSGD, lr=0.01, momentum=True, nesterov=True, decay=1e-06,\n\nall others used the default setting with decay=1e-06. \n\n![Image](https://storage.googleapis.com/kaggle-forum-message-attachments/inbox/642483/73906ba9af381a9e4a5bc0e7ca764896/pre_trained.png)\n\n\noverall, the AMSGrad variant of Adam using learning rate 0.0001 has the fastest convergence speed for training, and also shows very good validation accuracy. RMSprop is the worst optimizer according to the figure, which is in line with the observation from here https://shaoanlu.wordpress.com/2017/05/29/sgd-all-which-one-is-the-best-optimizer-dogs-vs-cats-toy-experiment/\n\nNext, I will evaluate the optimizers based on training from scratch for Resnet50, let's see if the results change or not. \n\n\n**Update**:\nbelow is the evaluation of different optimizers on ResNet50 training from scratch, the input size is (512, 512, 4).\n![Image](https://storage.googleapis.com/kaggle-forum-message-attachments/inbox/642483/7152f51a584227b2dec0468885c35ec4/from_scratch.png)\n\nas we can see from the figure, now Adamax is the best one, which implies that we might need different optimizers for pre-trained and training-from-scratch model.",
      "votes": null
    },
    {
      "id": "413829",
      "postDate": "11/01/2018 16:05:07",
      "content": "<p>Thanks for the very interesting study!\nI have a few questions: <br>\n - which loss function are you using? BCE? <br>\n - did you try looking at the macro F1 metric? (rather than accuracy) <br>\n - how long did it take you to train from scratch ResNet50?  </p>",
      "rawMarkdown": "Thanks for the very interesting study!\nI have a few questions:  \n - which loss function are you using? BCE?  \n - did you try looking at the macro F1 metric? (rather than accuracy)  \n - how long did it take you to train from scratch ResNet50?",
      "votes": null
    },
    {
      "id": "413839",
      "postDate": "11/01/2018 16:20:40",
      "content": "<p>Hi, </p>\n\n<p>Q1: yes, I used BCE as loss function</p>\n\n<p>Q2: not yet, because macro F1 metric requires fine-tune the thresholds, so far, I only used accuracy.</p>\n\n<p>Q3: ~1000s for one epoch for training from scratch, I splited the training set as 8-folds, and trainied the model using 7-folds and validated on 1 fold.</p>",
      "rawMarkdown": "Hi, \n\nQ1: yes, I used BCE as loss function\n\nQ2: not yet, because macro F1 metric requires fine-tune the thresholds, so far, I only used accuracy.\n\nQ3: ~1000s for one epoch for training from scratch, I splited the training set as 8-folds, and trainied the model using 7-folds and validated on 1 fold.",
      "votes": null
    },
    {
      "id": "413867",
      "postDate": "11/01/2018 17:06:09",
      "content": "<p>many thanks!</p>",
      "rawMarkdown": "many thanks!",
      "votes": null
    },
    {
      "id": "413945",
      "postDate": "11/01/2018 20:50:56",
      "content": "<p>Great information here. These are inline with my own observations. I've been using Adadelta as it had performed the best out of all I had tried. I have yet to try Adam with amsgrad or Adamax, but from your data I might have some room to improve there.</p>",
      "rawMarkdown": "Great information here. These are inline with my own observations. I've been using Adadelta as it had performed the best out of all I had tried. I have yet to try Adam with amsgrad or Adamax, but from your data I might have some room to improve there.",
      "votes": null
    },
    {
      "id": "414069",
      "postDate": "11/02/2018 03:32:36",
      "content": "<p>Thank you so much for posting this analysis. I just did a quick check with my kernel, and it looks adamax indeed outperforms adam. Here is that I got for identical setup (the same as <a href=\"https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb\">hear</a> but with 8-fold TTA and the same train and val sets in both runs):</p>\n\n<pre><code>Adam\nF1 macro:  0.645\nF1 macro (th = 0.5):  0.618\nF1 micro:  0.770\n\nAdamax\nF1 macro:  0.664\nF1 macro (th = 0.5):  0.631\nF1 micro:  0.785\n</code></pre>",
      "rawMarkdown": "Thank you so much for posting this analysis. I just did a quick check with my kernel, and it looks adamax indeed outperforms adam. Here is that I got for identical setup (the same as [hear][1] but with 8-fold TTA and the same train and val sets in both runs):\n\n    Adam\n    F1 macro:  0.645\n    F1 macro (th = 0.5):  0.618\n    F1 micro:  0.770\n    \n    Adamax\n    F1 macro:  0.664\n    F1 macro (th = 0.5):  0.631\n    F1 micro:  0.785\n\n\n  [1]: https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb",
      "votes": null
    },
    {
      "id": "414075",
      "postDate": "11/02/2018 04:01:23",
      "content": "<p>Did you have to re-train everything from scratch or just finetune the Adam models with new optimizer?</p>",
      "rawMarkdown": "Did you have to re-train everything from scratch or just finetune the Adam models with new optimizer?",
      "votes": null
    },
    {
      "id": "414080",
      "postDate": "11/02/2018 04:20:47",
      "content": "<p>Update on further testing on my side: I see the same results using Adamax and Adam+AMSGrad here locally. I've only made it through 10 epochs but the initial performance is doing well. </p>\n\n<p>In previous testing I even see the same large validation spikes occasionally with Adadelta. The only difference is that RMSProp worked reasonably well for me.</p>",
      "rawMarkdown": "Update on further testing on my side: I see the same results using Adamax and Adam+AMSGrad here locally. I've only made it through 10 epochs but the initial performance is doing well. \n\nIn previous testing I even see the same large validation spikes occasionally with Adadelta. The only difference is that RMSProp worked reasonably well for me.",
      "votes": null
    },
    {
      "id": "414081",
      "postDate": "11/02/2018 04:33:48",
      "content": "<p>I just ran training from scratch for adamax and adam with above mentioned difference from <a href=\"https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb\">this kernel</a>: 8-fold TTA and the same train and val sets in both runs</p>\n\n<p>I'm going to start this competition in 2 weeks, after ship detection one is over, so no real models just several kaggle kernels that I run for testing my ideas...</p>",
      "rawMarkdown": "I just ran training from scratch for adamax and adam with above mentioned difference from [this kernel][1]: 8-fold TTA and the same train and val sets in both runs\n\nI'm going to start this competition in 2 weeks, after ship detection one is over, so no real models just several kaggle kernels that I run for testing my ideas...\n\n\n  [1]: https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb",
      "votes": null
    },
    {
      "id": "414162",
      "postDate": "11/02/2018 08:40:22",
      "content": "<p>If you're interested in teaming up later just sent me a DM. I think the room for improvement is huge in this competition.</p>",
      "rawMarkdown": "If you're interested in teaming up later just sent me a DM. I think the room for improvement is huge in this competition.",
      "votes": null
    },
    {
      "id": "414175",
      "postDate": "11/02/2018 09:06:05",
      "content": "<p>Thank you for updating.</p>\n\n<p>maybe I used RMSProp incorrectly,\n but given the observations of mine and here:\n<a href=\"https://shaoanlu.wordpress.com/2017/05/29/sgd-all-which-one-is-the-best-optimizer-dogs-vs-cats-toy-experiment/\">https://shaoanlu.wordpress.com/2017/05/29/sgd-all-which-one-is-the-best-optimizer-dogs-vs-cats-toy-experiment/</a></p>\n\n<p>I would not use it to train my model.</p>",
      "rawMarkdown": "Thank you for updating.\n\nmaybe I used RMSProp incorrectly,\n but given the observations of mine and here:\nhttps://shaoanlu.wordpress.com/2017/05/29/sgd-all-which-one-is-the-best-optimizer-dogs-vs-cats-toy-experiment/\n\n\nI would not use it to train my model.",
      "votes": null
    },
    {
      "id": "414199",
      "postDate": "11/02/2018 09:58:28",
      "content": "<p>Why didn't you try Nadam？</p>",
      "rawMarkdown": "Why didn't you try Nadam？",
      "votes": null
    },
    {
      "id": "414220",
      "postDate": "11/02/2018 10:48:44",
      "content": "<p>because there are some memory errors when I used Nadam.</p>",
      "rawMarkdown": "because there are some memory errors when I used Nadam.",
      "votes": null
    },
    {
      "id": "414341",
      "postDate": "11/02/2018 15:22:17",
      "content": "<p>I've tried Nadam, it performed similar to Adam, Adadelta was better.</p>",
      "rawMarkdown": "I've tried Nadam, it performed similar to Adam, Adadelta was better.",
      "votes": null
    },
    {
      "id": "414346",
      "postDate": "11/02/2018 15:29:12",
      "content": "<p>For me RMSProp worked well at the start of training. Initially it is faster than Adam, but I definitely wont use it for fine tuning.</p>\n\n<p>With the size of the model and training time, I've looked towards something that would converge quickly at the start, and then fine tune with SGD + cosine learning schedule + restarts to get better generalization at the end. For now Adadelta has been the best to start with. Still testing on the others but with nearly an hour per epoch it may take some more time.</p>",
      "rawMarkdown": "For me RMSProp worked well at the start of training. Initially it is faster than Adam, but I definitely wont use it for fine tuning.\n\nWith the size of the model and training time, I've looked towards something that would converge quickly at the start, and then fine tune with SGD + cosine learning schedule + restarts to get better generalization at the end. For now Adadelta has been the best to start with. Still testing on the others but with nearly an hour per epoch it may take some more time.",
      "votes": null
    },
    {
      "id": "414368",
      "postDate": "11/02/2018 16:03:55",
      "content": "<p>one hour per epoch seems pretty slow, \njust curious, what's your input size?</p>",
      "rawMarkdown": "one hour per epoch seems pretty slow, \njust curious, what's your input size?",
      "votes": null
    },
    {
      "id": "414880",
      "postDate": "11/03/2018 19:49:57",
      "content": "<p>1 hour = 1024x1024x3, 90,000 images. Batch size 20, GTX Titan X. Scales down pretty much linearly with the number of images. With 12GB video memory, I can do very small batches at 2048x2048, but it also takes significantly longer. For now going to put some time in at 1024x1024. My 512x512 model is set for now with 0.508 public LB using the same threshold across all classes.</p>\n\n<p>I ran a test overnight, using only 30k images. Ran 50 epochs with Adam + amsgrad and it overfit really badly. With an 80/20 split, the train overfit to higher than 0.9 f1.  I've added some dropout and maxnorm constraitint as well as back to the 90k imageset.  To get to the same 50 epochs though means I'll have to wait a few days. Until then, smaller tests on the GTX 970 I also have in this pieced together machine.</p>",
      "rawMarkdown": "1 hour = 1024x1024x3, 90,000 images. Batch size 20, GTX Titan X. Scales down pretty much linearly with the number of images. With 12GB video memory, I can do very small batches at 2048x2048, but it also takes significantly longer. For now going to put some time in at 1024x1024. My 512x512 model is set for now with 0.508 public LB using the same threshold across all classes.\n\nI ran a test overnight, using only 30k images. Ran 50 epochs with Adam + amsgrad and it overfit really badly. With an 80/20 split, the train overfit to higher than 0.9 f1.  I've added some dropout and maxnorm constraitint as well as back to the 90k imageset.  To get to the same 50 epochs though means I'll have to wait a few days. Until then, smaller tests on the GTX 970 I also have in this pieced together machine.",
      "votes": null
    },
    {
      "id": "419035",
      "postDate": "11/11/2018 05:47:52",
      "content": "<p>Somewhat unrelated question, but how did you manage to add another color channel to the premade Keras model? I've tried to figure this out but haven't had any luck so far. Any good resources that you know of? Thanks!</p>",
      "rawMarkdown": "Somewhat unrelated question, but how did you manage to add another color channel to the premade Keras model? I've tried to figure this out but haven't had any luck so far. Any good resources that you know of? Thanks!",
      "votes": null
    },
    {
      "id": "902308",
      "postDate": "06/26/2020 04:16:41",
      "content": "<p>Great info, Adamax seems to work the best.</p>",
      "rawMarkdown": "Great info, Adamax seems to work the best.",
      "votes": null
    },
    {
      "id": "964672",
      "postDate": "08/10/2020 05:01:59",
      "content": "<p>Great work on the comparison.</p>",
      "rawMarkdown": "Great work on the comparison.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 413829,
      "author_name": "stecasasso",
      "author_url": "",
      "post_date": "11/01/2018 16:05:07",
      "content": "<p>Thanks for the very interesting study!\nI have a few questions: <br>\n - which loss function are you using? BCE? <br>\n - did you try looking at the macro F1 metric? (rather than accuracy) <br>\n - how long did it take you to train from scratch ResNet50?  </p>",
      "votes": null,
      "replies": [
        {
          "id": 413839,
          "author_name": "zhijianli",
          "author_url": "",
          "post_date": "11/01/2018 16:20:40",
          "content": "<p>Hi, </p>\n\n<p>Q1: yes, I used BCE as loss function</p>\n\n<p>Q2: not yet, because macro F1 metric requires fine-tune the thresholds, so far, I only used accuracy.</p>\n\n<p>Q3: ~1000s for one epoch for training from scratch, I splited the training set as 8-folds, and trainied the model using 7-folds and validated on 1 fold.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 413867,
          "author_name": "stecasasso",
          "author_url": "",
          "post_date": "11/01/2018 17:06:09",
          "content": "<p>many thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 413945,
      "author_name": "ldm314",
      "author_url": "",
      "post_date": "11/01/2018 20:50:56",
      "content": "<p>Great information here. These are inline with my own observations. I've been using Adadelta as it had performed the best out of all I had tried. I have yet to try Adam with amsgrad or Adamax, but from your data I might have some room to improve there.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 414069,
      "author_name": "iafoss",
      "author_url": "",
      "post_date": "11/02/2018 03:32:36",
      "content": "<p>Thank you so much for posting this analysis. I just did a quick check with my kernel, and it looks adamax indeed outperforms adam. Here is that I got for identical setup (the same as <a href=\"https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb\">hear</a> but with 8-fold TTA and the same train and val sets in both runs):</p>\n\n<pre><code>Adam\nF1 macro:  0.645\nF1 macro (th = 0.5):  0.618\nF1 micro:  0.770\n\nAdamax\nF1 macro:  0.664\nF1 macro (th = 0.5):  0.631\nF1 micro:  0.785\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 414075,
          "author_name": "suicaokhoailang",
          "author_url": "",
          "post_date": "11/02/2018 04:01:23",
          "content": "<p>Did you have to re-train everything from scratch or just finetune the Adam models with new optimizer?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 414081,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "11/02/2018 04:33:48",
          "content": "<p>I just ran training from scratch for adamax and adam with above mentioned difference from <a href=\"https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb\">this kernel</a>: 8-fold TTA and the same train and val sets in both runs</p>\n\n<p>I'm going to start this competition in 2 weeks, after ship detection one is over, so no real models just several kaggle kernels that I run for testing my ideas...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 414162,
          "author_name": "suicaokhoailang",
          "author_url": "",
          "post_date": "11/02/2018 08:40:22",
          "content": "<p>If you're interested in teaming up later just sent me a DM. I think the room for improvement is huge in this competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 414080,
      "author_name": "ldm314",
      "author_url": "",
      "post_date": "11/02/2018 04:20:47",
      "content": "<p>Update on further testing on my side: I see the same results using Adamax and Adam+AMSGrad here locally. I've only made it through 10 epochs but the initial performance is doing well. </p>\n\n<p>In previous testing I even see the same large validation spikes occasionally with Adadelta. The only difference is that RMSProp worked reasonably well for me.</p>",
      "votes": null,
      "replies": [
        {
          "id": 414175,
          "author_name": "zhijianli",
          "author_url": "",
          "post_date": "11/02/2018 09:06:05",
          "content": "<p>Thank you for updating.</p>\n\n<p>maybe I used RMSProp incorrectly,\n but given the observations of mine and here:\n<a href=\"https://shaoanlu.wordpress.com/2017/05/29/sgd-all-which-one-is-the-best-optimizer-dogs-vs-cats-toy-experiment/\">https://shaoanlu.wordpress.com/2017/05/29/sgd-all-which-one-is-the-best-optimizer-dogs-vs-cats-toy-experiment/</a></p>\n\n<p>I would not use it to train my model.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 414346,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "11/02/2018 15:29:12",
          "content": "<p>For me RMSProp worked well at the start of training. Initially it is faster than Adam, but I definitely wont use it for fine tuning.</p>\n\n<p>With the size of the model and training time, I've looked towards something that would converge quickly at the start, and then fine tune with SGD + cosine learning schedule + restarts to get better generalization at the end. For now Adadelta has been the best to start with. Still testing on the others but with nearly an hour per epoch it may take some more time.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 414368,
          "author_name": "zhijianli",
          "author_url": "",
          "post_date": "11/02/2018 16:03:55",
          "content": "<p>one hour per epoch seems pretty slow, \njust curious, what's your input size?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 414880,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "11/03/2018 19:49:57",
          "content": "<p>1 hour = 1024x1024x3, 90,000 images. Batch size 20, GTX Titan X. Scales down pretty much linearly with the number of images. With 12GB video memory, I can do very small batches at 2048x2048, but it also takes significantly longer. For now going to put some time in at 1024x1024. My 512x512 model is set for now with 0.508 public LB using the same threshold across all classes.</p>\n\n<p>I ran a test overnight, using only 30k images. Ran 50 epochs with Adam + amsgrad and it overfit really badly. With an 80/20 split, the train overfit to higher than 0.9 f1.  I've added some dropout and maxnorm constraitint as well as back to the 90k imageset.  To get to the same 50 epochs though means I'll have to wait a few days. Until then, smaller tests on the GTX 970 I also have in this pieced together machine.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 414199,
      "author_name": "vikeezhou",
      "author_url": "",
      "post_date": "11/02/2018 09:58:28",
      "content": "<p>Why didn't you try Nadam？</p>",
      "votes": null,
      "replies": [
        {
          "id": 414220,
          "author_name": "zhijianli",
          "author_url": "",
          "post_date": "11/02/2018 10:48:44",
          "content": "<p>because there are some memory errors when I used Nadam.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 414341,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "11/02/2018 15:22:17",
          "content": "<p>I've tried Nadam, it performed similar to Adam, Adadelta was better.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 419035,
      "author_name": "mrsnappingturtle",
      "author_url": "",
      "post_date": "11/11/2018 05:47:52",
      "content": "<p>Somewhat unrelated question, but how did you manage to add another color channel to the premade Keras model? I've tried to figure this out but haven't had any luck so far. Any good resources that you know of? Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 902308,
      "author_name": "chethanhebbar",
      "author_url": "",
      "post_date": "06/26/2020 04:16:41",
      "content": "<p>Great info, Adamax seems to work the best.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 964672,
      "author_name": "partha1189",
      "author_url": "",
      "post_date": "08/10/2020 05:01:59",
      "content": "<p>Great work on the comparison.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "413673": "I am wondering how much we can improve the results by using different optimizers, so I run all but Nadam from keras, using ResNet50 with pre-trained weights of imagenet. The input shape is (224, 224, 3).\n\nthe parameters used as following:\nSGD, lr=0.01, momentum=True, nesterov=True, decay=1e-06,\n\nall others used the default setting with decay=1e-06. \n\n![Image](https://storage.googleapis.com/kaggle-forum-message-attachments/inbox/642483/73906ba9af381a9e4a5bc0e7ca764896/pre_trained.png)\n\n\noverall, the AMSGrad variant of Adam using learning rate 0.0001 has the fastest convergence speed for training, and also shows very good validation accuracy. RMSprop is the worst optimizer according to the figure, which is in line with the observation from here https://shaoanlu.wordpress.com/2017/05/29/sgd-all-which-one-is-the-best-optimizer-dogs-vs-cats-toy-experiment/\n\nNext, I will evaluate the optimizers based on training from scratch for Resnet50, let's see if the results change or not. \n\n\n**Update**:\nbelow is the evaluation of different optimizers on ResNet50 training from scratch, the input size is (512, 512, 4).\n![Image](https://storage.googleapis.com/kaggle-forum-message-attachments/inbox/642483/7152f51a584227b2dec0468885c35ec4/from_scratch.png)\n\nas we can see from the figure, now Adamax is the best one, which implies that we might need different optimizers for pre-trained and training-from-scratch model.",
    "413829": "Thanks for the very interesting study!\nI have a few questions:  \n - which loss function are you using? BCE?  \n - did you try looking at the macro F1 metric? (rather than accuracy)  \n - how long did it take you to train from scratch ResNet50?",
    "413839": "Hi, \n\nQ1: yes, I used BCE as loss function\n\nQ2: not yet, because macro F1 metric requires fine-tune the thresholds, so far, I only used accuracy.\n\nQ3: ~1000s for one epoch for training from scratch, I splited the training set as 8-folds, and trainied the model using 7-folds and validated on 1 fold.",
    "413867": "many thanks!",
    "413945": "Great information here. These are inline with my own observations. I've been using Adadelta as it had performed the best out of all I had tried. I have yet to try Adam with amsgrad or Adamax, but from your data I might have some room to improve there.",
    "414069": "Thank you so much for posting this analysis. I just did a quick check with my kernel, and it looks adamax indeed outperforms adam. Here is that I got for identical setup (the same as [hear][1] but with 8-fold TTA and the same train and val sets in both runs):\n\n    Adam\n    F1 macro:  0.645\n    F1 macro (th = 0.5):  0.618\n    F1 micro:  0.770\n    \n    Adamax\n    F1 macro:  0.664\n    F1 macro (th = 0.5):  0.631\n    F1 micro:  0.785\n\n\n  [1]: https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb",
    "414075": "Did you have to re-train everything from scratch or just finetune the Adam models with new optimizer?",
    "414080": "Update on further testing on my side: I see the same results using Adamax and Adam+AMSGrad here locally. I've only made it through 10 epochs but the initial performance is doing well. \n\nIn previous testing I even see the same large validation spikes occasionally with Adadelta. The only difference is that RMSProp worked reasonably well for me.",
    "414081": "I just ran training from scratch for adamax and adam with above mentioned difference from [this kernel][1]: 8-fold TTA and the same train and val sets in both runs\n\nI'm going to start this competition in 2 weeks, after ship detection one is over, so no real models just several kaggle kernels that I run for testing my ideas...\n\n\n  [1]: https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb",
    "414162": "If you're interested in teaming up later just sent me a DM. I think the room for improvement is huge in this competition.",
    "414175": "Thank you for updating.\n\nmaybe I used RMSProp incorrectly,\n but given the observations of mine and here:\nhttps://shaoanlu.wordpress.com/2017/05/29/sgd-all-which-one-is-the-best-optimizer-dogs-vs-cats-toy-experiment/\n\n\nI would not use it to train my model.",
    "414199": "Why didn't you try Nadam？",
    "414220": "because there are some memory errors when I used Nadam.",
    "414341": "I've tried Nadam, it performed similar to Adam, Adadelta was better.",
    "414346": "For me RMSProp worked well at the start of training. Initially it is faster than Adam, but I definitely wont use it for fine tuning.\n\nWith the size of the model and training time, I've looked towards something that would converge quickly at the start, and then fine tune with SGD + cosine learning schedule + restarts to get better generalization at the end. For now Adadelta has been the best to start with. Still testing on the others but with nearly an hour per epoch it may take some more time.",
    "414368": "one hour per epoch seems pretty slow, \njust curious, what's your input size?",
    "414880": "1 hour = 1024x1024x3, 90,000 images. Batch size 20, GTX Titan X. Scales down pretty much linearly with the number of images. With 12GB video memory, I can do very small batches at 2048x2048, but it also takes significantly longer. For now going to put some time in at 1024x1024. My 512x512 model is set for now with 0.508 public LB using the same threshold across all classes.\n\nI ran a test overnight, using only 30k images. Ran 50 epochs with Adam + amsgrad and it overfit really badly. With an 80/20 split, the train overfit to higher than 0.9 f1.  I've added some dropout and maxnorm constraitint as well as back to the 90k imageset.  To get to the same 50 epochs though means I'll have to wait a few days. Until then, smaller tests on the GTX 970 I also have in this pieced together machine.",
    "419035": "Somewhat unrelated question, but how did you manage to add another color channel to the premade Keras model? I've tried to figure this out but haven't had any luck so far. Any good resources that you know of? Thanks!",
    "902308": "Great info, Adamax seems to work the best.",
    "964672": "Great work on the comparison."
  },
  "source": "meta"
}