{
  "id": 139455,
  "title": "TPU-optimized custom training loop",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/139455",
  "author_name": "",
  "post_date": "2020-03-28T22:04:24.070124500Z",
  "votes": 12,
  "comment_count": 7,
  "views": 0,
  "content": "<p>As most of us are using TPUs for training that are a few tweaks that we can do to improve training time, they were better described <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443\">here</a> and show <a href=\"https://www.kaggle.com/mgornergoogle/custom-training-loop-with-100-flowers-on-tpu\">here</a> by @mgornergoogle , the most important is to use a custom training loop, I have created a kernel using those tricks, <a href=\"https://www.kaggle.com/dimitreoliveira/jigsaw-tpu-optimized-training-loops\">check it out here</a></p>\n\n<p>Also, feel free to post here any further improvement or tips.</p>",
  "messages": [
    {
      "id": "789680",
      "postDate": "03/28/2020 22:04:24",
      "content": "<p>As most of us are using TPUs for training that are a few tweaks that we can do to improve training time, they were better described <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443\">here</a> and show <a href=\"https://www.kaggle.com/mgornergoogle/custom-training-loop-with-100-flowers-on-tpu\">here</a> by @mgornergoogle , the most important is to use a custom training loop, I have created a kernel using those tricks, <a href=\"https://www.kaggle.com/dimitreoliveira/jigsaw-tpu-optimized-training-loops\">check it out here</a></p>\n\n<p>Also, feel free to post here any further improvement or tips.</p>",
      "rawMarkdown": "As most of us are using TPUs for training that are a few tweaks that we can do to improve training time, they were better described [here](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443) and show [here](https://www.kaggle.com/mgornergoogle/custom-training-loop-with-100-flowers-on-tpu) by @mgornergoogle , the most important is to use a custom training loop, I have created a kernel using those tricks, [check it out here](https://www.kaggle.com/dimitreoliveira/jigsaw-tpu-optimized-training-loops)\n\nAlso, feel free to post here any further improvement or tips.",
      "votes": null
    },
    {
      "id": "791963",
      "postDate": "03/30/2020 20:05:43",
      "content": "<p>Very nice notebook. Thank you.\nWhat kind of speedup did you get from this optimized version, compared to model.fit() ?</p>",
      "rawMarkdown": "Very nice notebook. Thank you.\nWhat kind of speedup did you get from this optimized version, compared to model.fit() ?",
      "votes": null
    },
    {
      "id": "792092",
      "postDate": "03/30/2020 22:31:05",
      "content": "<p>Hey <a href=\"/mgornergoogle\">@mgornergoogle</a> this is a good point,</p>\n\n<p>Here is the log of <code>model.fit()</code>:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1182060%2F5d4c3730523a57c612591b70be45554a%2Fmodedel_fit.png?generation=1585607249794547&amp;alt=media\" alt=\"\"></p>\n\n<p>And here the optimized loop:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1182060%2F2f2738b2effd01425668427c8f7ecbd8%2Fmodedel_optmized.png?generation=1585607297214983&amp;alt=media\" alt=\"\"></p>\n\n<p>as you can see the optimized version is faster, but only a bit, so this left me thinking that the code can be further improved, what do you think?</p>",
      "rawMarkdown": "Hey @mgornergoogle this is a good point,\n\nHere is the log of `model.fit()`:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1182060%2F5d4c3730523a57c612591b70be45554a%2Fmodedel_fit.png?generation=1585607249794547&amp;alt=media)\n\nAnd here the optimized loop:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1182060%2F2f2738b2effd01425668427c8f7ecbd8%2Fmodedel_optmized.png?generation=1585607297214983&amp;alt=media)\n \nas you can see the optimized version is faster, but only a bit, so this left me thinking that the code can be further improved, what do you think?",
      "votes": null
    },
    {
      "id": "792118",
      "postDate": "03/30/2020 22:49:24",
      "content": "<p>good kernel. 1 vote from me </p>",
      "rawMarkdown": "good kernel. 1 vote from me",
      "votes": null
    },
    {
      "id": "792146",
      "postDate": "03/30/2020 23:43:54",
      "content": "<p>thanks <a href=\"/podsyp\">@podsyp</a> </p>",
      "rawMarkdown": "thanks @podsyp",
      "votes": null
    },
    {
      "id": "792167",
      "postDate": "03/31/2020 00:11:17",
      "content": "<p>I think BERT is a very large model and most of the time is already spent in useful computations rather than overheads, even in model.fit(). When I run your notebook, I see 0% idle time consistently during training. I think it's as good as can be. <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4003597%2Fc39f7211852e41aec2100c827a9e7180%2FScreen%20Shot%202020-03-30%20at%2017.07.06.png?generation=1585613536489187&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I think BERT is a very large model and most of the time is already spent in useful computations rather than overheads, even in model.fit(). When I run your notebook, I see 0% idle time consistently during training. I think it's as good as can be. ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4003597%2Fc39f7211852e41aec2100c827a9e7180%2FScreen%20Shot%202020-03-30%20at%2017.07.06.png?generation=1585613536489187&amp;alt=media)",
      "votes": null
    },
    {
      "id": "792180",
      "postDate": "03/31/2020 00:33:28",
      "content": "<p>this explains a lot, thanks <a href=\"/mgornergoogle\">@mgornergoogle</a> </p>",
      "rawMarkdown": "this explains a lot, thanks @mgornergoogle",
      "votes": null
    },
    {
      "id": "794734",
      "postDate": "04/02/2020 03:34:40",
      "content": "<p>Good notebook. thanks for sharing..</p>",
      "rawMarkdown": "Good notebook. thanks for sharing..",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 791963,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "03/30/2020 20:05:43",
      "content": "<p>Very nice notebook. Thank you.\nWhat kind of speedup did you get from this optimized version, compared to model.fit() ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 792092,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "03/30/2020 22:31:05",
          "content": "<p>Hey <a href=\"/mgornergoogle\">@mgornergoogle</a> this is a good point,</p>\n\n<p>Here is the log of <code>model.fit()</code>:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1182060%2F5d4c3730523a57c612591b70be45554a%2Fmodedel_fit.png?generation=1585607249794547&amp;alt=media\" alt=\"\"></p>\n\n<p>And here the optimized loop:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1182060%2F2f2738b2effd01425668427c8f7ecbd8%2Fmodedel_optmized.png?generation=1585607297214983&amp;alt=media\" alt=\"\"></p>\n\n<p>as you can see the optimized version is faster, but only a bit, so this left me thinking that the code can be further improved, what do you think?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 792167,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "03/31/2020 00:11:17",
          "content": "<p>I think BERT is a very large model and most of the time is already spent in useful computations rather than overheads, even in model.fit(). When I run your notebook, I see 0% idle time consistently during training. I think it's as good as can be. <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4003597%2Fc39f7211852e41aec2100c827a9e7180%2FScreen%20Shot%202020-03-30%20at%2017.07.06.png?generation=1585613536489187&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 792180,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "03/31/2020 00:33:28",
          "content": "<p>this explains a lot, thanks <a href=\"/mgornergoogle\">@mgornergoogle</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 792118,
      "author_name": "podsyp",
      "author_url": "",
      "post_date": "03/30/2020 22:49:24",
      "content": "<p>good kernel. 1 vote from me </p>",
      "votes": null,
      "replies": [
        {
          "id": 792146,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "03/30/2020 23:43:54",
          "content": "<p>thanks <a href=\"/podsyp\">@podsyp</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 794734,
      "author_name": "usam9048",
      "author_url": "",
      "post_date": "04/02/2020 03:34:40",
      "content": "<p>Good notebook. thanks for sharing..</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "789680": "As most of us are using TPUs for training that are a few tweaks that we can do to improve training time, they were better described [here](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/135443) and show [here](https://www.kaggle.com/mgornergoogle/custom-training-loop-with-100-flowers-on-tpu) by @mgornergoogle , the most important is to use a custom training loop, I have created a kernel using those tricks, [check it out here](https://www.kaggle.com/dimitreoliveira/jigsaw-tpu-optimized-training-loops)\n\nAlso, feel free to post here any further improvement or tips.",
    "791963": "Very nice notebook. Thank you.\nWhat kind of speedup did you get from this optimized version, compared to model.fit() ?",
    "792092": "Hey @mgornergoogle this is a good point,\n\nHere is the log of `model.fit()`:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1182060%2F5d4c3730523a57c612591b70be45554a%2Fmodedel_fit.png?generation=1585607249794547&amp;alt=media)\n\nAnd here the optimized loop:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1182060%2F2f2738b2effd01425668427c8f7ecbd8%2Fmodedel_optmized.png?generation=1585607297214983&amp;alt=media)\n \nas you can see the optimized version is faster, but only a bit, so this left me thinking that the code can be further improved, what do you think?",
    "792118": "good kernel. 1 vote from me",
    "792146": "thanks @podsyp",
    "792167": "I think BERT is a very large model and most of the time is already spent in useful computations rather than overheads, even in model.fit(). When I run your notebook, I see 0% idle time consistently during training. I think it's as good as can be. ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4003597%2Fc39f7211852e41aec2100c827a9e7180%2FScreen%20Shot%202020-03-30%20at%2017.07.06.png?generation=1585613536489187&amp;alt=media)",
    "792180": "this explains a lot, thanks @mgornergoogle",
    "794734": "Good notebook. thanks for sharing.."
  },
  "source": "meta"
}