{
  "id": 91412,
  "title": "AdaBound doesn't help ",
  "url": "/competitions/imet-2019-fgvc6/discussion/91412",
  "author_name": "",
  "post_date": "2019-05-04T11:43:47.401115100Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>AdaBound: <a href=\"https://github.com/Luolc/AdaBound\">https://github.com/Luolc/AdaBound</a>\nIt's an optimizer that trains as fast as Adam and as good as SGD. It works well at most of tasks.\nAnd I don't know why it makes the f2 score drop at this task, maybe you can skip this experiment.</p>\n\n<p>the log like that\n<code>\n Epoch 1, lr 0.0001: 100% 107873/107904 [1:12:33&lt;00:01, 24.78it/s, loss=16.637]\nvalid_loss 19.526 | valid_f2_th_0.10 0.338 | valid_f2_th_0.09 0.338 | valid_f2_th_0.11 0.338 | valid_f2_th_0.15 0.336 | valid_f2_th_0.12 0.336 | valid_f2_th_0.14 0.336 | valid_f2_th_0.13 0.336 | valid_f2_th_0.08 0.335 | valid_f2_th_0.07 0.334 | valid_f2_th_0.20 0.334 | valid_f2_th_0.06 0.333 | valid_f2_th_0.05 0.332\nLoaded model from epoch 2, step 3,372\nEpoch 2, lr 0.0001: 100% 107873/107904 [1:15:36&lt;00:01, 23.78it/s, loss=13.476]\nvalid_loss 16.457 | valid_f2_th_0.05 0.183 | valid_f2_th_0.06 0.179 | valid_f2_th_0.07 0.173 | valid_f2_th_0.08 0.159 | valid_f2_th_0.09 0.147 | valid_f2_th_0.10 0.142 | valid_f2_th_0.11 0.134 | valid_f2_th_0.12 0.127 | valid_f2_th_0.13 0.119 | valid_f2_th_0.14 0.105 | valid_f2_th_0.15 0.095 | valid_f2_th_0.20 0.047\n</code></p>",
  "messages": [
    {
      "id": "527022",
      "postDate": "05/04/2019 11:43:47",
      "content": "<p>AdaBound: <a href=\"https://github.com/Luolc/AdaBound\">https://github.com/Luolc/AdaBound</a>\nIt's an optimizer that trains as fast as Adam and as good as SGD. It works well at most of tasks.\nAnd I don't know why it makes the f2 score drop at this task, maybe you can skip this experiment.</p>\n\n<p>the log like that\n<code>\n Epoch 1, lr 0.0001: 100% 107873/107904 [1:12:33&lt;00:01, 24.78it/s, loss=16.637]\nvalid_loss 19.526 | valid_f2_th_0.10 0.338 | valid_f2_th_0.09 0.338 | valid_f2_th_0.11 0.338 | valid_f2_th_0.15 0.336 | valid_f2_th_0.12 0.336 | valid_f2_th_0.14 0.336 | valid_f2_th_0.13 0.336 | valid_f2_th_0.08 0.335 | valid_f2_th_0.07 0.334 | valid_f2_th_0.20 0.334 | valid_f2_th_0.06 0.333 | valid_f2_th_0.05 0.332\nLoaded model from epoch 2, step 3,372\nEpoch 2, lr 0.0001: 100% 107873/107904 [1:15:36&lt;00:01, 23.78it/s, loss=13.476]\nvalid_loss 16.457 | valid_f2_th_0.05 0.183 | valid_f2_th_0.06 0.179 | valid_f2_th_0.07 0.173 | valid_f2_th_0.08 0.159 | valid_f2_th_0.09 0.147 | valid_f2_th_0.10 0.142 | valid_f2_th_0.11 0.134 | valid_f2_th_0.12 0.127 | valid_f2_th_0.13 0.119 | valid_f2_th_0.14 0.105 | valid_f2_th_0.15 0.095 | valid_f2_th_0.20 0.047\n</code></p>",
      "rawMarkdown": "AdaBound: [https://github.com/Luolc/AdaBound](https://github.com/Luolc/AdaBound)\nIt's an optimizer that trains as fast as Adam and as good as SGD. It works well at most of tasks.\nAnd I don't know why it makes the f2 score drop at this task, maybe you can skip this experiment.\n \nthe log like that\n```\n Epoch 1, lr 0.0001: 100% 107873/107904 [1:12:33&lt;00:01, 24.78it/s, loss=16.637]\nvalid_loss 19.526 | valid_f2_th_0.10 0.338 | valid_f2_th_0.09 0.338 | valid_f2_th_0.11 0.338 | valid_f2_th_0.15 0.336 | valid_f2_th_0.12 0.336 | valid_f2_th_0.14 0.336 | valid_f2_th_0.13 0.336 | valid_f2_th_0.08 0.335 | valid_f2_th_0.07 0.334 | valid_f2_th_0.20 0.334 | valid_f2_th_0.06 0.333 | valid_f2_th_0.05 0.332\nLoaded model from epoch 2, step 3,372\nEpoch 2, lr 0.0001: 100% 107873/107904 [1:15:36&lt;00:01, 23.78it/s, loss=13.476]\nvalid_loss 16.457 | valid_f2_th_0.05 0.183 | valid_f2_th_0.06 0.179 | valid_f2_th_0.07 0.173 | valid_f2_th_0.08 0.159 | valid_f2_th_0.09 0.147 | valid_f2_th_0.10 0.142 | valid_f2_th_0.11 0.134 | valid_f2_th_0.12 0.127 | valid_f2_th_0.13 0.119 | valid_f2_th_0.14 0.105 | valid_f2_th_0.15 0.095 | valid_f2_th_0.20 0.047\n```",
      "votes": null
    },
    {
      "id": "527126",
      "postDate": "05/04/2019 16:41:45",
      "content": "<p>I guess you haven't properly tuned hyperparameters of proposed method. What final_lr, gamma did you use?\nThe method was experimented on CIFAR-10 which is simpler problem than imet, thus we need to tune hyperparameters to attend to taht. For example gamma controls the transformation speed and need to be adjusted based on total steps.</p>",
      "rawMarkdown": "I guess you haven't properly tuned hyperparameters of proposed method. What final_lr, gamma did you use?\nThe method was experimented on CIFAR-10 which is simpler problem than imet, thus we need to tune hyperparameters to attend to taht. For example gamma controls the transformation speed and need to be adjusted based on total steps.",
      "votes": null
    },
    {
      "id": "527277",
      "postDate": "05/05/2019 02:00:51",
      "content": "<p>I tried final_lr 1e-1 and 1e-2, 1e-1 will increase much loss in first epoch, 1e-2 just like the log </p>",
      "rawMarkdown": "I tried final_lr 1e-1 and 1e-2, 1e-1 will increase much loss in first epoch, 1e-2 just like the log",
      "votes": null
    },
    {
      "id": "527491",
      "postDate": "05/05/2019 16:45:51",
      "content": "<p>This challenge is more about architectures, losses, etc than optimizers. From my observation (although I have already given up on this competition and I do not have a strong baseline), AdaBound is better than Adam, but not enough to make a huge difference.</p>",
      "rawMarkdown": "This challenge is more about architectures, losses, etc than optimizers. From my observation (although I have already given up on this competition and I do not have a strong baseline), AdaBound is better than Adam, but not enough to make a huge difference.",
      "votes": null
    },
    {
      "id": "527677",
      "postDate": "05/06/2019 04:32:09",
      "content": "<p>Can you tell me your parameters for Adabound in your experience? </p>",
      "rawMarkdown": "Can you tell me your parameters for Adabound in your experience?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 527126,
      "author_name": "appian",
      "author_url": "",
      "post_date": "05/04/2019 16:41:45",
      "content": "<p>I guess you haven't properly tuned hyperparameters of proposed method. What final_lr, gamma did you use?\nThe method was experimented on CIFAR-10 which is simpler problem than imet, thus we need to tune hyperparameters to attend to taht. For example gamma controls the transformation speed and need to be adjusted based on total steps.</p>",
      "votes": null,
      "replies": [
        {
          "id": 527277,
          "author_name": "",
          "author_url": "",
          "post_date": "05/05/2019 02:00:51",
          "content": "<p>I tried final_lr 1e-1 and 1e-2, 1e-1 will increase much loss in first epoch, 1e-2 just like the log </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 527491,
      "author_name": "alexanderliao",
      "author_url": "",
      "post_date": "05/05/2019 16:45:51",
      "content": "<p>This challenge is more about architectures, losses, etc than optimizers. From my observation (although I have already given up on this competition and I do not have a strong baseline), AdaBound is better than Adam, but not enough to make a huge difference.</p>",
      "votes": null,
      "replies": [
        {
          "id": 527677,
          "author_name": "",
          "author_url": "",
          "post_date": "05/06/2019 04:32:09",
          "content": "<p>Can you tell me your parameters for Adabound in your experience? </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "527022": "AdaBound: [https://github.com/Luolc/AdaBound](https://github.com/Luolc/AdaBound)\nIt's an optimizer that trains as fast as Adam and as good as SGD. It works well at most of tasks.\nAnd I don't know why it makes the f2 score drop at this task, maybe you can skip this experiment.\n \nthe log like that\n```\n Epoch 1, lr 0.0001: 100% 107873/107904 [1:12:33&lt;00:01, 24.78it/s, loss=16.637]\nvalid_loss 19.526 | valid_f2_th_0.10 0.338 | valid_f2_th_0.09 0.338 | valid_f2_th_0.11 0.338 | valid_f2_th_0.15 0.336 | valid_f2_th_0.12 0.336 | valid_f2_th_0.14 0.336 | valid_f2_th_0.13 0.336 | valid_f2_th_0.08 0.335 | valid_f2_th_0.07 0.334 | valid_f2_th_0.20 0.334 | valid_f2_th_0.06 0.333 | valid_f2_th_0.05 0.332\nLoaded model from epoch 2, step 3,372\nEpoch 2, lr 0.0001: 100% 107873/107904 [1:15:36&lt;00:01, 23.78it/s, loss=13.476]\nvalid_loss 16.457 | valid_f2_th_0.05 0.183 | valid_f2_th_0.06 0.179 | valid_f2_th_0.07 0.173 | valid_f2_th_0.08 0.159 | valid_f2_th_0.09 0.147 | valid_f2_th_0.10 0.142 | valid_f2_th_0.11 0.134 | valid_f2_th_0.12 0.127 | valid_f2_th_0.13 0.119 | valid_f2_th_0.14 0.105 | valid_f2_th_0.15 0.095 | valid_f2_th_0.20 0.047\n```",
    "527126": "I guess you haven't properly tuned hyperparameters of proposed method. What final_lr, gamma did you use?\nThe method was experimented on CIFAR-10 which is simpler problem than imet, thus we need to tune hyperparameters to attend to taht. For example gamma controls the transformation speed and need to be adjusted based on total steps.",
    "527277": "I tried final_lr 1e-1 and 1e-2, 1e-1 will increase much loss in first epoch, 1e-2 just like the log",
    "527491": "This challenge is more about architectures, losses, etc than optimizers. From my observation (although I have already given up on this competition and I do not have a strong baseline), AdaBound is better than Adam, but not enough to make a huge difference.",
    "527677": "Can you tell me your parameters for Adabound in your experience?"
  },
  "source": "meta"
}