{
  "id": 214592,
  "title": "The performance of TaylorCrossEntropyLossn under Adam and SGD is quit different",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/214592",
  "author_name": "Hanson0910",
  "post_date": "2021-01-27T06:21:43.859000",
  "votes": 11,
  "comment_count": 14,
  "views": 0,
  "content": "<p><strong>My experimental results show that there is a great difference between Adam and SGD under TaylorCrossEntropyLossn train with efficinetnet4，I didn't do other experiments to prove this conclusion. The accuracy of the verification set is shown in the following figure，adam is better than sgd</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4848897%2F6b5fe983c2780f69fc256913952b5763%2F1.png?generation=1611728294714195&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 1171854,
      "postDate": "2021-01-27T06:21:43.860Z",
      "content": "<p><strong>My experimental results show that there is a great difference between Adam and SGD under TaylorCrossEntropyLossn train with efficinetnet4，I didn't do other experiments to prove this conclusion. The accuracy of the verification set is shown in the following figure，adam is better than sgd</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4848897%2F6b5fe983c2780f69fc256913952b5763%2F1.png?generation=1611728294714195&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "**My experimental results show that there is a great difference between Adam and SGD under TaylorCrossEntropyLossn train with efficinetnet4，I didn't do other experiments to prove this conclusion. The accuracy of the verification set is shown in the following figure，adam is better than sgd**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4848897%2F6b5fe983c2780f69fc256913952b5763%2F1.png?generation=1611728294714195&alt=media)\n",
      "votes": 11
    },
    {
      "id": 1172754,
      "postDate": "2021-01-27T14:37:53.213Z",
      "content": "<p>I've never heard about TaylorCrossEntropyLossn yet. I will try out later. Thanks for sharing <a href=\"https://www.kaggle.com/hanson0910\" target=\"_blank\">@hanson0910</a> !! +upvoted:) And in my experience it was little bit better when I use lookahead with Adam. The final result is almost same, but I could save some time because it needed less epochs. I want to try out RAdam, too.</p>",
      "rawMarkdown": "I've never heard about TaylorCrossEntropyLossn yet. I will try out later. Thanks for sharing @hanson0910 !! +upvoted:) And in my experience it was little bit better when I use lookahead with Adam. The final result is almost same, but I could save some time because it needed less epochs. I want to try out RAdam, too.",
      "replies": [
        {
          "id": 1172822,
          "postDate": "2021-01-27T15:04:09.030Z",
          "content": "<p>agree with you</p>",
          "rawMarkdown": "agree with you",
          "votes": 1
        }
      ]
    },
    {
      "id": 1172116,
      "postDate": "2021-01-27T09:12:13.117Z",
      "content": "<p>if you want to squeeze every juices of score, you may use SGD. However, you need to allot more resources to train a model longer than Adam. SGD is highly recommended for fine-tuning. </p>",
      "rawMarkdown": "if you want to squeeze every juices of score, you may use SGD. However, you need to allot more resources to train a model longer than Adam. SGD is highly recommended for fine-tuning. "
    },
    {
      "id": 1171986,
      "postDate": "2021-01-27T07:54:41.720Z",
      "content": "<p>I don't use TaylorCE loss but it may depend on your setting and SGD needs in general much larger number of epoches than Adam</p>",
      "rawMarkdown": "I don't use TaylorCE loss but it may depend on your setting and SGD needs in general much larger number of epoches than Adam\n",
      "replies": [
        {
          "id": 1172006,
          "postDate": "2021-01-27T08:00:40.237Z",
          "content": "<p>you can have a try</p>",
          "rawMarkdown": "you can have a try"
        },
        {
          "id": 1172830,
          "postDate": "2021-01-27T15:06:55.367Z",
          "content": "<p>I think I will stick with my current loss function ^^</p>",
          "rawMarkdown": "I think I will stick with my current loss function ^^",
          "votes": 1
        },
        {
          "id": 1172863,
          "postDate": "2021-01-27T15:20:44.037Z",
          "content": "<p>You're smart👀</p>",
          "rawMarkdown": "You're smart👀"
        }
      ]
    },
    {
      "id": 1171971,
      "postDate": "2021-01-27T07:50:53.453Z",
      "content": "<p>This Taylor loss function is very unstable and often crashes, so adam is advantageous. Of course, for me, Adam also occasionally crashes.</p>",
      "rawMarkdown": "This Taylor loss function is very unstable and often crashes, so adam is advantageous. Of course, for me, Adam also occasionally crashes.",
      "replies": [
        {
          "id": 1172000,
          "postDate": "2021-01-27T07:58:34.507Z",
          "content": "<p>for me SGD is better now.</p>",
          "rawMarkdown": "for me SGD is better now."
        },
        {
          "id": 1176043,
          "postDate": "2021-01-29T13:25:50.660Z",
          "content": "<p>Can you share your learning rate? I use Adam optimizer for 1e-4 learning rate,but I want to try other optimizers such as SGD, so if you can share your optimization method ,it will help me a lot.Thanks.</p>",
          "rawMarkdown": "Can you share your learning rate? I use Adam optimizer for 1e-4 learning rate,but I want to try other optimizers such as SGD, so if you can share your optimization method ,it will help me a lot.Thanks."
        },
        {
          "id": 1176184,
          "postDate": "2021-01-29T14:30:56.367Z",
          "content": "<p>adam is 1e-4 and sgd is 1e-2</p>",
          "rawMarkdown": "adam is 1e-4 and sgd is 1e-2"
        }
      ]
    },
    {
      "id": 1171885,
      "postDate": "2021-01-27T06:53:32Z",
      "content": "<p>Adam's approach is usually better when there are fewer iterations</p>",
      "rawMarkdown": "Adam's approach is usually better when there are fewer iterations",
      "replies": [
        {
          "id": 1171921,
          "postDate": "2021-01-27T07:10:55.780Z",
          "content": "<p>Sorry,the coordinate represent epoch not iters.</p>",
          "rawMarkdown": "Sorry,the coordinate represent epoch not iters."
        }
      ]
    },
    {
      "id": 1171972,
      "postDate": "2021-01-27T07:50:53.453Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1172754,
      "author_name": "Dongkyu Kim",
      "author_url": "",
      "post_date": "2021-01-27T14:37:53.213000",
      "content": "<p>I've never heard about TaylorCrossEntropyLossn yet. I will try out later. Thanks for sharing <a href=\"https://www.kaggle.com/hanson0910\" target=\"_blank\">@hanson0910</a> !! +upvoted:) And in my experience it was little bit better when I use lookahead with Adam. The final result is almost same, but I could save some time because it needed less epochs. I want to try out RAdam, too.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1172822,
          "author_name": "Hanson0910",
          "author_url": "",
          "post_date": "2021-01-27T15:04:09.030000",
          "content": "<p>agree with you</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1172116,
      "author_name": "FGPC",
      "author_url": "",
      "post_date": "2021-01-27T09:12:13.117000",
      "content": "<p>if you want to squeeze every juices of score, you may use SGD. However, you need to allot more resources to train a model longer than Adam. SGD is highly recommended for fine-tuning. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1171986,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2021-01-27T07:54:41.720000",
      "content": "<p>I don't use TaylorCE loss but it may depend on your setting and SGD needs in general much larger number of epoches than Adam</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1172006,
          "author_name": "Hanson0910",
          "author_url": "",
          "post_date": "2021-01-27T08:00:40.237000",
          "content": "<p>you can have a try</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1172830,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2021-01-27T15:06:55.367000",
          "content": "<p>I think I will stick with my current loss function ^^</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1172863,
          "author_name": "Hanson0910",
          "author_url": "",
          "post_date": "2021-01-27T15:20:44.037000",
          "content": "<p>You're smart👀</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1171971,
      "author_name": "Leon",
      "author_url": "",
      "post_date": "2021-01-27T07:50:53.453000",
      "content": "<p>This Taylor loss function is very unstable and often crashes, so adam is advantageous. Of course, for me, Adam also occasionally crashes.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1172000,
          "author_name": "Hanson0910",
          "author_url": "",
          "post_date": "2021-01-27T07:58:34.507000",
          "content": "<p>for me SGD is better now.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1176043,
          "author_name": "hujianxin",
          "author_url": "",
          "post_date": "2021-01-29T13:25:50.660000",
          "content": "<p>Can you share your learning rate? I use Adam optimizer for 1e-4 learning rate,but I want to try other optimizers such as SGD, so if you can share your optimization method ,it will help me a lot.Thanks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1176184,
          "author_name": "Hanson0910",
          "author_url": "",
          "post_date": "2021-01-29T14:30:56.367000",
          "content": "<p>adam is 1e-4 and sgd is 1e-2</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1171885,
      "author_name": "DarknessZX",
      "author_url": "",
      "post_date": "2021-01-27T06:53:32",
      "content": "<p>Adam's approach is usually better when there are fewer iterations</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1171921,
          "author_name": "Hanson0910",
          "author_url": "",
          "post_date": "2021-01-27T07:10:55.780000",
          "content": "<p>Sorry,the coordinate represent epoch not iters.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1171972,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-01-27T07:50:53.453000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1171854": "**My experimental results show that there is a great difference between Adam and SGD under TaylorCrossEntropyLossn train with efficinetnet4，I didn't do other experiments to prove this conclusion. The accuracy of the verification set is shown in the following figure，adam is better than sgd**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4848897%2F6b5fe983c2780f69fc256913952b5763%2F1.png?generation=1611728294714195&alt=media)\n",
    "1172754": "I've never heard about TaylorCrossEntropyLossn yet. I will try out later. Thanks for sharing @hanson0910 !! +upvoted:) And in my experience it was little bit better when I use lookahead with Adam. The final result is almost same, but I could save some time because it needed less epochs. I want to try out RAdam, too.",
    "1172116": "if you want to squeeze every juices of score, you may use SGD. However, you need to allot more resources to train a model longer than Adam. SGD is highly recommended for fine-tuning. ",
    "1171986": "I don't use TaylorCE loss but it may depend on your setting and SGD needs in general much larger number of epoches than Adam\n",
    "1171971": "This Taylor loss function is very unstable and often crashes, so adam is advantageous. Of course, for me, Adam also occasionally crashes.",
    "1171885": "Adam's approach is usually better when there are fewer iterations",
    "1171972": ""
  }
}