{
  "id": 132244,
  "title": "Training with fp16",
  "url": "/competitions/bengaliai-cv19/discussion/132244",
  "author_name": "",
  "post_date": "2020-02-25T03:41:49.321096900Z",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I've seen that on the forum people were suggesting to turn back to fp32 and finetune model after training with fp16. </p>\n\n<p>I have a few questions about it:</p>\n\n<p>1). Are 3-5 epochs enough? Or maybe even one?</p>\n\n<p>2). If finetuning stage has a small number of epochs than the score might degrade a bit, as it usually degrades on first epochs after you've uploaded a checkpoint. That means we probably want to finetune with a very low lr in order to tackle these issues, don't we?</p>\n\n<p>3). Did anyone notice difference in the score when training with O1 optimization level vs O2? \nApex docs page says:\nO1 and O2 are different implementations of mixed precision. Try both, and see what gives the best speedup and accuracy for your model.</p>",
  "messages": [
    {
      "id": "755708",
      "postDate": "02/25/2020 03:41:49",
      "content": "<p>I've seen that on the forum people were suggesting to turn back to fp32 and finetune model after training with fp16. </p>\n\n<p>I have a few questions about it:</p>\n\n<p>1). Are 3-5 epochs enough? Or maybe even one?</p>\n\n<p>2). If finetuning stage has a small number of epochs than the score might degrade a bit, as it usually degrades on first epochs after you've uploaded a checkpoint. That means we probably want to finetune with a very low lr in order to tackle these issues, don't we?</p>\n\n<p>3). Did anyone notice difference in the score when training with O1 optimization level vs O2? \nApex docs page says:\nO1 and O2 are different implementations of mixed precision. Try both, and see what gives the best speedup and accuracy for your model.</p>",
      "rawMarkdown": "I've seen that on the forum people were suggesting to turn back to fp32 and finetune model after training with fp16. \n\nI have a few questions about it:\n\n1). Are 3-5 epochs enough? Or maybe even one?\n\n2). If finetuning stage has a small number of epochs than the score might degrade a bit, as it usually degrades on first epochs after you've uploaded a checkpoint. That means we probably want to finetune with a very low lr in order to tackle these issues, don't we?\n\n3). Did anyone notice difference in the score when training with O1 optimization level vs O2? \nApex docs page says:\nO1 and O2 are different implementations of mixed precision. Try both, and see what gives the best speedup and accuracy for your model.",
      "votes": null
    },
    {
      "id": "756098",
      "postDate": "02/25/2020 12:05:01",
      "content": "<p>I'm curious about how to turn back to fp32 ?</p>",
      "rawMarkdown": "I'm curious about how to turn back to fp32 ?",
      "votes": null
    },
    {
      "id": "756732",
      "postDate": "02/26/2020 02:44:08",
      "content": "<p>I guess you can change your optimization level to <code>O0</code>.</p>\n\n<p>Check this out: <a href=\"https://nvidia.github.io/apex/amp.html\">https://nvidia.github.io/apex/amp.html</a></p>",
      "rawMarkdown": "I guess you can change your optimization level to `O0`.\n\nCheck this out: https://nvidia.github.io/apex/amp.html",
      "votes": null
    },
    {
      "id": "756752",
      "postDate": "02/26/2020 03:17:53",
      "content": "<p>&gt; 3). Did anyone notice difference in the score when training with O1 optimization level vs O2?</p>\n\n<p>You can try it.</p>\n\n<p>I've tried in another task and my experience is just like Apex docs said \"O2 is slightly faster, but could be harder to converge/stabilize, or may not converge to FP32 results. In O2, all the ops are in FP16, so generally not recommended.\"</p>\n\n<p>just FYR.</p>",
      "rawMarkdown": "&gt; 3). Did anyone notice difference in the score when training with O1 optimization level vs O2?\n\nYou can try it.\n\nI've tried in another task and my experience is just like Apex docs said \"O2 is slightly faster, but could be harder to converge/stabilize, or may not converge to FP32 results. In O2, all the ops are in FP16, so generally not recommended.\"\n\njust FYR.",
      "votes": null
    },
    {
      "id": "757056",
      "postDate": "02/26/2020 11:35:23",
      "content": "<p>Update:</p>\n\n<p>In terms of speed O2 is faster than both O1 and O0.</p>",
      "rawMarkdown": "Update:\n\nIn terms of speed O2 is faster than both O1 and O0.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 756098,
      "author_name": "thefatcat",
      "author_url": "",
      "post_date": "02/25/2020 12:05:01",
      "content": "<p>I'm curious about how to turn back to fp32 ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 756732,
          "author_name": "lightnezzofbeing",
          "author_url": "",
          "post_date": "02/26/2020 02:44:08",
          "content": "<p>I guess you can change your optimization level to <code>O0</code>.</p>\n\n<p>Check this out: <a href=\"https://nvidia.github.io/apex/amp.html\">https://nvidia.github.io/apex/amp.html</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 756752,
      "author_name": "radream",
      "author_url": "",
      "post_date": "02/26/2020 03:17:53",
      "content": "<p>&gt; 3). Did anyone notice difference in the score when training with O1 optimization level vs O2?</p>\n\n<p>You can try it.</p>\n\n<p>I've tried in another task and my experience is just like Apex docs said \"O2 is slightly faster, but could be harder to converge/stabilize, or may not converge to FP32 results. In O2, all the ops are in FP16, so generally not recommended.\"</p>\n\n<p>just FYR.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 757056,
      "author_name": "lightnezzofbeing",
      "author_url": "",
      "post_date": "02/26/2020 11:35:23",
      "content": "<p>Update:</p>\n\n<p>In terms of speed O2 is faster than both O1 and O0.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "755708": "I've seen that on the forum people were suggesting to turn back to fp32 and finetune model after training with fp16. \n\nI have a few questions about it:\n\n1). Are 3-5 epochs enough? Or maybe even one?\n\n2). If finetuning stage has a small number of epochs than the score might degrade a bit, as it usually degrades on first epochs after you've uploaded a checkpoint. That means we probably want to finetune with a very low lr in order to tackle these issues, don't we?\n\n3). Did anyone notice difference in the score when training with O1 optimization level vs O2? \nApex docs page says:\nO1 and O2 are different implementations of mixed precision. Try both, and see what gives the best speedup and accuracy for your model.",
    "756098": "I'm curious about how to turn back to fp32 ?",
    "756732": "I guess you can change your optimization level to `O0`.\n\nCheck this out: https://nvidia.github.io/apex/amp.html",
    "756752": "&gt; 3). Did anyone notice difference in the score when training with O1 optimization level vs O2?\n\nYou can try it.\n\nI've tried in another task and my experience is just like Apex docs said \"O2 is slightly faster, but could be harder to converge/stabilize, or may not converge to FP32 results. In O2, all the ops are in FP16, so generally not recommended.\"\n\njust FYR.",
    "757056": "Update:\n\nIn terms of speed O2 is faster than both O1 and O0."
  },
  "source": "meta"
}